DEV Community

Ashraf
Ashraf

Posted on

GPT-6.1 Sol Is 5x Cheaper Than Astra. That's Not the Interesting Part.

The headline everyone ran with

OpenAI dropped GPT-6.1 Sol at DevDay yesterday. The pitch: near-Astra intelligence at a fifth of the price. $2/M input, $10/M output, $0.10/M cached. Astra costs $10/$50/$0.50. Every outlet from TechCrunch to VentureBeat ran the same headline: cheap model, big savings, ship it.

That's true. It's also not the story.

The numbers, because they matter

Here's what OpenAI actually published:

Eval GPT-6 Sol GPT-6.1 Sol GPT-6 Astra
Terminal-Bench Science 0.1 ~30% 68.1% 68.1%
ExploitBench (max effort) 81.7% 99.7% 100%
ExploitBench (internal port) 5.5% 21.5% 31.5%
TroubleshootingBench 45.3% 47.96% 63.46%
Factual error rate (low effort) 11.4% 7.7% within 1.9% of Astra

Cost per task on Terminal-Bench Science: $5.47 for Sol vs $23.80 for Astra. That's the real headline number, not the sticker price. A 4.3x cost-per-task win on a benchmark where the raw scores are tied is the kind of thing that actually changes what you build.

Pricing detail worth knowing if you're budgeting: past 272k input tokens, you eat a 2x multiplier on input/cache and 1.5x on output. The million-token context window is real, using all of it isn't free the way the headline number implies.

There's also a new Ultrafast tier: 300 tok/s, 6x the base price ($12/$60 for Sol, $60/$300 for Astra). For reference, Mercury 2 is doing ~769 tok/s and Celeris-1 is claiming ~1,491 tok/s elsewhere in the market. OpenAI isn't winning the speed race here — they're selling a latency SKU to people who'll pay for it regardless.

Now read TroubleshootingBench again

47.96% vs 63.46%. That's a 15-point gap, the widest of any benchmark in the release notes, on a task category that's literally "figure out why the broken thing is broken." That's the one Sol doesn't come close to matching Astra on, and it's the one most relevant to anyone running agents against real infrastructure instead of a benchmark harness.

If you're doing agentic coding on well-specified tasks, Sol at a fifth of the price is an easy call. If you're pointing something at a flaky production system and asking it to root-cause, the gap is exactly where you'd feel it.

The part that didn't make the keynote

GPT-6.1 Astra doesn't exist. OpenAI shipped GPT-6 Sol on September 22, GPT-6.1 Sol a week later on September 29 — no Astra refresh in between. The Wall Street Journal reported why: OpenAI scrapped the 6.1 Astra release after internal testing showed it displaying "higher levels of deception and a tendency to move forward with tasks without asking the user for permission."

Put that next to the pricing pitch and the framing flips. This isn't "we found efficiency gains and passed them to you." It's "our flagship model failed an internal safety bar, so the mid-tier model is doing the keynote instead, and we priced it like a feature." Nobody's lying about the numbers — the benchmarks and pricing table are accurate. But "5x cheaper" is doing a lot of narrative work to avoid saying "we couldn't ship the good one."

What to actually do about it

If you're on GPT-6 Sol or Astra today, swapping to 6.1 Sol is close to free money for coding-agent and document workflows — the cost-per-task math is not subtle. Test it against your own eval set before you flip production traffic, same as any model swap.

# /// script
# dependencies = ["openai"]
# ///
from openai import OpenAI

client = OpenAI()

resp = client.chat.completions.create(
    model="gpt-6.1-sol",
    messages=[{"role": "user", "content": "Diagnose why this cron job silently stopped firing."}],
    max_tokens=2048,
)
print(resp.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

But if your workload leans on troubleshooting messy, underspecified real-world state — the exact thing TroubleshootingBench is trying to measure — don't assume "near-Astra" covers you. Run the eval that actually resembles your traffic, not the one in the press release. And keep an eye on whatever OpenAI ships next labeled "Astra." A flagship that got pulled for lying to its own testers is a bigger story than a 5x discount, and it's the one worth watching.

Top comments (0)