One API ID gpt-6-astra, 1.05M context, April 2026 cutoff, and Enterprise admins have to switch it on
Pricing is 10 USD in and 50 USD out per 1M tokens, 2.5x GPT-5.6 Sol on both sides
Terminal-Bench 4.0 at 57.9 percent and 96.3 percent recall at 1M context are the real wins
Astra loses Humanity's Last Exam to all three Claude models and ties Sol on Artificial Analysis
ARC-AGI-3 scores 62.7 percent on a neutral harness and 99.9 percent on OpenAI's own adapter
OpenAI shipped GPT-6 Astra on September 3. Greg Brockman closed the press briefing with "welcome to the AGI era," which is the line every outlet ran with. I spent the morning in the pricing page, the model card and the benchmark table instead, because the interesting part of a launch like this is never the quote. It is the row in the table that nobody screenshots.
What Ships and Who Can Actually Get It
There is one API model ID: gpt-6-astra. No mini, no dated snapshot, no separate pro ID. The single knob is reasoning.effort, which now accepts five values instead of four: low, medium, high, xhigh, and max.
The context window is 1,050,000 tokens, with a 128,000 token maximum output and roughly 922,000 usable for input. Knowledge cutoff is April 30, 2026. Text and images go in, only text comes out. There is no audio and no video, which is worth knowing before you plan around it.
Rollout is staged and the staging is messier than the headlines suggest. Day one was the API plus a limited set of organizations in OpenAI's cyber program. ChatGPT Plus, Pro, Business and Enterprise follow over what OpenAI calls "the coming days," along with AWS Bedrock. Enterprise is the trap: the model is off by default and a workspace admin has to enable it. If your Enterprise account does not show Astra, that is probably not a rollout delay.
Naming is a mess and it will cost somebody an afternoon. OpenAI's launch post says "GPT-6 Astra Pro." The Help Center calls the same thing "GPT-6 Pro." The launch blog says Astra reaches all ChatGPT Plus users; the Help Center says GPT-6 Pro is not included with Plus. Both are probably true, describing base Astra and Pro-effort Astra separately, but nothing on either page says so.
In Codex it appears as three power levels: Astra Light, Astra Medium and Astra Extra High. Codex also gets a genuinely new context mechanism that replaces compaction. Instead of compressing old turns into a summary, Astra keeps searchable notes and reads back into earlier messages and tool output. It is off by default, lives in config.toml, and at launch it does not work with Business, Enterprise or API-key sign-in. Only ChatGPT Plus and Pro sign-in.
The Price Table
Per 1M tokens, standard tier, for requests under 272K input:
Tier
Input
Cached input
Output
Standard
10.00 USD
1.00 USD
50.00 USD
Batch
5.00 USD
0.50 USD
25.00 USD
Flex
5.00 USD
0.50 USD
25.00 USD
Fast mode
20.00 USD
2.00 USD
100.00 USD
GPT-5.6 Sol runs 4.00 USD in and 20.00 USD out. Astra is exactly 2.5x on both sides.
Two multipliers are easy to miss. Cross 272K input tokens and the whole request reprices: input and cache go to 2x, output goes to 1.5x, so a long-context standard call lands at 20.00 USD in and 75.00 USD out. Not the portion over the threshold. The full request. Second, data residency endpoints add 10 percent, and fast mode is simply unavailable for Astra under EU data residency.
If you are on ChatGPT rather than the API, the number that actually constrains you is the message allowance, and it is tighter than people expect. Pro at 200 USD a month gets 200 GPT-6 Pro messages per week. Pro at 100 USD gets 50 per week, shared with GPT-5.6 Sol Pro. Business Standard gets 15 per month. Hit the weekly ceiling on the 200 USD plan and you silently fall back to GPT-5.6 Thinking at medium effort, which is the kind of downgrade you notice in the output before you notice in the UI. Codex and Work carry separate allowances from Chat.
Brockman's own framing is the honest one here. "Pricing tokens doesn't make any sense," he told VentureBeat. "What you actually want is the price per task." On that measure Astra looks better than 2.5x suggests: OpenAI's own DeepSWE figures put cost per completed task around 57 percent below Sol's highest configuration, because Astra needs fewer turns. That is a vendor number and nobody has reproduced it yet. If you want the comparison baseline on the other side, Claude API pricing breaks down the same math for Anthropic's tiers.
Where Astra Genuinely Wins
Terminal-Bench 4.0 is the cleanest result on the sheet. Astra scores 57.9 percent against Sol's 37.3 percent, Claude Fable 5.1 at 55.8 percent and Gemini 3.8 Flash at 19.1 percent. A 20 point jump over the previous OpenAI model on agentic terminal work is not a rounding artifact.
Long context is the second real win, and it is underrated. On OpenAI's MRCR 8-needle test in the 512K to 1M band, Astra hits 96.3 percent against Sol's 73.8 percent. Most models advertise a huge window and quietly fall apart in the top half of it. This one does not, and if you run long agent sessions that matters more than any reasoning score.
Computer use moved too. ScreenSpot-Pro goes from 76.9 to 92.7 percent with no tools. OSWorld 2.0 lands at 72.6 percent against Sol's 65.7, and OpenAI's latency simulation has it finishing those tasks in roughly 40 minutes where Sol took 75.
Then there is cybersecurity, where the numbers stop being normal. ExploitBench 100 percent. SRE-Bench 88 percent on a single attempt, 99.2 percent across four. On an internal port of 20 high-severity V8 CVEs, Astra scored 39 percent against Sol's 11.5, and found two previously unknown zero-days along the way. Every one of those cyber evals was run with production safeguards disabled, which is the caveat OpenAI states plainly and most coverage dropped.
Where Astra Loses
This is the section the launch posts skipped, and it is short but real.
Humanity's Last Exam with tools: Astra 57.2 percent. Claude Fable 5.1 scores 65.0, Fable 5 scores 63.8, Opus 5 scores 63.6. Astra loses to all three, on OpenAI's own comparison table.
Artificial Analysis, which is independent, puts Astra's Intelligence Index at 61.2, essentially tied with Sol at 60.9, behind Fable 5.1 at 65.7 and Opus 5 at 63.1. Their Coding Agent Index has Astra at 67.0 and Fable 5 at 68.1.
FrontierCode 1.1 Main: Astra 53.3, Fable 5 53.5, Opus 5 53.4. That is a three-way tie inside the noise floor. DeepSWE v1.1 has Astra at 74.1 against Opus 5 at 73.7 and Gemini 3.8 Flash at 73.8, which is a 0.4 point lead over a cheaper model. GPQA Diamond moves 1.4 points and is effectively saturated.
Also worth noting: no SWE-bench Verified number was published at all. For a launch leaning this hard on software engineering, that omission is loud. The earlier three-way Opus 5 vs GPT-5.6 Sol vs Kimi K3 comparison holds up better than I expected against this table.
The ARC-AGI Number That Has Two Answers
The 99.9 percent on ARC-AGI-3 is the stat driving the AGI headlines, and it needs an asterisk that OpenAI did not print.
ARC Prize, who own the benchmark, published results split by harness. On their standard provider-neutral harness at max effort, Astra scores 62.7 percent and burns 26,098 USD. On a provider adapter that preserves OpenAI's opaque reasoning state between turns and uses custom compaction, it scores 99.9 percent for 18,817 USD. The human baseline is about 12.78 USD per attempted game.
Both numbers are honest. They measure different things. The 62.7 is what you get through a normal API integration; the 99.9 needs a harness built around OpenAI's own state handling. ARC Prize will now report both on the leaderboard, labeled, which tells you how much the distinction bothered them. Press coverage has separately quoted 66 percent and 98.6 percent for the same benchmark, so if you see a lone ARC-AGI-3 figure with no harness attached, it is not telling you much.
The pattern repeats across the sheet. FrontierMath Tier 4 shows 97.6 percent in the table and "98%" in the prose, on a benchmark OpenAI funded and holds partial exclusive access to. The cyber scores ran unguarded. Every table entry is footnoted "maximum at any effort."
Bottom Line
Astra is a real step up on agentic work, long context and offensive security, and a lateral move on raw reasoning. If your workload is long-horizon terminal tasks, browser automation or anything living past 500K tokens, the 2.5x price is probably worth it and the per-task cost may well be lower. If you are doing single-shot reasoning, Claude Fable 5.1 currently beats it on two independent measures and costs less.
What I would not do is repeat the 99.9 percent without saying which harness produced it. The most useful habit with a launch this loud is to read the benchmark footnotes before the quotes, because on this sheet the footnotes carry most of the information. I keep every model comparison I have run in the RAXXO Lab overview if you want the longer trail.
Top comments (0)