The model was trained on OpenAI’s largest run to date, using more than 100,000 GPUs at the company’s Stargate site in Texas, and is the first Astra release where earlier models played a significant supervisory role in training. Rollout starts with enterprise customers in OpenAI’s Daybreak program, with ChatGPT Plus, Pro, Business and Enterprise access, plus the API, following in the coming days.
The Benchmark Claims
OpenAI’s own released figures show Astra leading most of the field it compared itself against, though these are the company’s reported numbers rather than independently reproduced results:
| Evaluation | GPT-6 Astra | GPT-5.6 Sol | Claude Opus 5 |
|---|---|---|---|
| ARC-AGI-3 | 98.6% | 7.8% | 30.2% (high) |
| FrontierMath Tier 4 (v2) | 97.6% | 83.0% | 73.2% |
| DeepSWE v1.1 | 74.1% | 70.8% | 68.8% |
| GPQA Diamond | 96.0% | 94.6% | 93.2% |
| ExploitBench | 100.0% | 78.5% | 70% |
| SRE-Bench (four attempts) | 99.2% | 68.7% | – |
The ARC-AGI-3 jump from Sol’s 7.8% to Astra’s 98.6% is the single largest gap in the table, and it lines up with OpenAI naming Astra its first model to cross a “critical” cybersecurity threshold under its own preparedness framework, meaning it can find and exploit unknown vulnerabilities in well-protected systems without step-by-step human guidance. That is also why OpenAI is limiting Astra’s most advanced cybersecurity capabilities to a smaller group of trusted testers through a program called Daybreak Blue rather than general release.
Pricing Still Unclear
OpenAI has not published Astra’s API pricing in its own materials as of this writing. Early accounts from developers with preview access put it at roughly double the rate of GPT-5.6 Sol, OpenAI’s current flagship, which lists at $5 per million input tokens and $30 per million output tokens. If accurate, that would put Astra in the same range as Anthropic’s higher-tier models, though OpenAI should be treated as the source once it confirms official numbers.
What the Announcement Skipped Over
Brockman framed AGI as a “mission concept or spiritual concept” rather than a technical milestone with agreed criteria, which lets OpenAI claim the moment without committing to a definition anyone could hold it to later. Meanwhile, the more concrete news in the release is arguably the cybersecurity threshold, not the AGI framing. A model capable enough to autonomously find zero-day exploits is a bigger operational fact for security teams than a philosophical debate about general intelligence, and it is the reason access is being staged rather than opened to everyone at once.
Follow us on Bluesky, LinkedIn, X, and Telegram to Get Instant Updates



