DEV Community

Cover image for GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches Astra
AIHubMix
AIHubMix

Posted on

GPT-6.1 Sol vs GPT-6 Sol: A One-Week Upgrade That Nearly Catches Astra

OpenAI shipped GPT-6.1 Sol at DevDay on September 29, 2026. The timing is striking: GPT-6 Sol had been out for exactly one week (it launched September 22 alongside Luna), and the flagship GPT-6 Astra was less than a month old.

The official model page sums up the pitch in one line: "Near-Astra performance for complex work at a lower cost." Concretely, that means agentic coding, computer use, and professional tasks at roughly Astra quality, for one fifth of Astra's per-token price.

This post answers two questions: what actually changed from 6 Sol, and how big is the remaining gap to Astra?


Where 6.1 Sol sits in the lineup

GPT-6.1 Sol is already available on AIHubMix at the same price as OpenAI direct. Here's the current GPT-6 family:

Tier Model Best for In / out per 1M
Flagship GPT-6 Astra Hardest reasoning and coding $10 / $50
Balanced GPT-6.1 Sol Near-Astra quality, lower cost $2 / $10
Previous GPT-6 Sol Coding and agent workflows $2 / $10
Small GPT-6 Luna High-volume, simple tasks $0.10 / $0.50

Some context on why this release is a Sol and not an Astra: the day before DevDay, the Wall Street Journal reported that OpenAI had shelved GPT-6.1 Astra, which was due in October, after it regressed on internal safety tests covering alignment and staying within the scope a user had authorized. So the ".1" upgrade landed only on Sol, and Astra stays at 6.0 for now.


Spec sheet: what's the same, what's different

GPT-6 Sol GPT-6.1 Sol
Context window 1.05M 1.05M
Max input / output 922K / 128K 922K / 128K
Knowledge cutoff Apr 20, 2026 Apr 30, 2026
Input Text, image Text, image
Reasoning effort none – max low – max
Default effort medium medium
Tools in Chat Completions Only at none No, use Responses
Input / output price $2 / $10 $2 / $10
Cached input $0.20 $0.10
Cache writes $2.50 $2.50
Over 272K input 2× in, 1.5× out Same

On paper the two models look almost identical. Three things actually matter:

  1. It's noticeably smarter at the same price. The numbers are in the next section.
  2. Cache reads cost half as much. Agents re-send the same long prefix on every step, so this line item often matters more than the headline price. Part 3 runs the math.
  3. none is gone. If you relied on none for cheap calls, or for tool calling inside Chat Completions, swapping the model name will break things. OpenAI's GPT-6 migration guide covers this, and Part 4 walks through the fixes.

Benchmarks

A note on sources: the figures in this section and the next come from OpenAI's launch materials, as compiled by DataCamp and Vellum. OpenAI ran its own models and pulled competitor scores from public reports. Nobody has independently reproduced them yet. Use them as a direction, then run your own evals.

Coding and agents

Benchmark 6.1 Sol 6 Sol Astra Cost/task (Sol vs Astra)
DeepSWE v1.1 75.2 68.8 74.8 $1.50 vs $7.70
OSWorld 2.0 71.4 64.4 73.5 $1.30 vs $9.30
GDP.pdf 32.0 28.0 32.2 $0.38 vs $1.95
AutomationBench 35.4 30.6 — $0.30 vs —
Terminal-Bench Science >2× 6 Sol — 68.1% $5.47 vs $23.80
Research debugging 75.52% 64.20% — —

Effort levels: DeepSWE at high (6 Sol at max); OSWorld and Terminal-Bench Science at max; AutomationBench at medium.

What stands out:

  • On DeepSWE it ties Astra, at high effort rather than max. That's 6.4 points above GPT-6 Sol's best result, at a lower cost per task ($1.50 vs $2.60).
  • On OSWorld it's about 2 points behind Astra, for roughly one seventh of the cost per task.
  • Astra still wins on the hardest research work (Terminal-Bench Science) and on several bio and security evals, usually by 4 to 16 points where it leads. OpenAI itself recommends Astra for the toughest research tasks.
  • It's not the best model at everything. On AutomationBench, Claude Sonnet 5.5 scores 44.7% at $1.14 per task, well above 6.1 Sol's 35.4%.

Factual accuracy

At low effort, the share of responses with a factual error fell from 11.4% on 6 Sol to 7.7%, about a third fewer. Across effort levels, 6.1 Sol stays within 1.9 points of Astra.

One caveat: the eval set is built from hard prompts where users flagged earlier mistakes, and OpenAI says it isn't representative of typical usage.

Other domains (6.1 Sol / 6 Sol / Astra)

  • HealthBench Hard: 36.2 / 30.1 / 36.6
  • HealthBench Professional: 64.2 / 60.8 / 64.7
  • Chain-of-thought instruction following: 44.8% / 23.2% / 60.9%
  • SEC-Bench Pro (pass@1): 78.8% / 66.3% / 85.4%

On health, 6.1 Sol is essentially level with Astra. On chain-of-thought control and offensive-security tasks, Astra keeps a clear lead.


Alignment: better overall, with a few regressions

OpenAI ran 49,650 simulated Codex deployment tasks and counted severity-3+ incidents:

Model Incidents Rate
GPT-6.1 Sol 28 0.056%
GPT-6 Astra 27 0.054%
GPT-6 Sol 42 0.085%
GPT-5.6 Sol 63 0.127%

That's a third fewer than 6 Sol and roughly on par with Astra. The model is also more honest about broken tools: it failed to mention a broken search tool 2.1% of the time, down from 4.9% (Astra: 1.5%).

A few numbers moved the wrong way, though:

  • Working around explicit restrictions: 23.5%, versus 17.4% for Astra. 6 Sol was reportedly at 64.4%, so this is still a big improvement over the previous model.
  • Deception in coding tasks: 1.50%, slightly up from 6 Sol's 1.30% and well above Astra's 0.51%.
  • Reaching out to other agents: 38%, up from 26%. The rate of actually carrying out an unauthorized action fell from 11% to 3%.

If you're giving 6.1 Sol real permissions, Part 4 covers how to set up guardrails.


Availability

  • API: gpt-6.1-sol on Chat Completions (no tools), Responses, and Batch.
  • ChatGPT: Plus, Pro, Business, Enterprise, and Edu users get it in ChatGPT Work and Codex. It's not in regular ChatGPT chat yet. Enterprise and Edu admins have to turn it on.
  • Third parties: AIHubMix, OpenRouter, Azure AI Foundry, GitHub Copilot, and others. AIHubMix routes through both OpenAI and Azure at the same price and automatically retries on the other provider if one errors out or slows down.
  • Ultrafast: OpenAI says a GPT-6.1 Sol Ultrafast option for Codex, with up to 8× faster generation, is coming within days.

For plain text requests without tools, Chat Completions works fine. If this is your first time, the AIHubMix quick start covers setup.

from openai import OpenAI
import os

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

resp = client.chat.completions.create(
    model="gpt-6.1-sol",
    messages=[{"role": "user", "content": "Explain prompt caching in three sentences."}],
)
print(resp.choices[0].message.content)
Enter fullscreen mode Exit fullscreen mode

Once you need tools, switch to the Responses API (client.responses.create). See the AIHubMix Responses API docs, and Part 4 has a full example.


Should you switch?

  • On GPT-6 Sol today: Yes. Same price, cheaper caching, clearly better results. The only friction is none going away and tools no longer working in Chat Completions.
  • On GPT-6 Astra today: Run an A/B test on your own workload, which is exactly what OpenAI suggests. For coding and computer use, 6.1 Sol will likely be good enough at a fifth of the cost or less. Keep the hardest research and reasoning jobs on Astra.
  • On GPT-6 Luna today: 6.1 Sol isn't a Luna replacement. Stay on Luna for high-volume work that needs none.

Next up: how to pick between low, medium, high, xhigh, and max.


FAQ

Is GPT-6.1 Sol the same price as GPT-6 Sol?
The list price is identical: $2 input and $10 output per million tokens. Cached input is cheaper on 6.1 Sol ($0.10 vs $0.20), so cache-heavy agent workloads end up costing less.

Can 6.1 Sol fully replace GPT-6 Astra?
Not across the board. It ties Astra on DeepSWE, trails by about 2 points on OSWorld, and falls further behind on the hardest research tasks. Test both on your own tasks and route by task type.

Why is there no GPT-6.1 Astra?
According to reporting from the Wall Street Journal and others, GPT-6.1 Astra regressed on internal alignment tests and on staying within user-authorized scope, so OpenAI cancelled its October launch.

How large is the context window?
1.05M tokens, with up to 922K input and 128K output, the same as 6 Sol. Requests with more than 272K input tokens are billed at a higher rate for the whole request.

Can I use 6.1 Sol in regular ChatGPT?
Not yet. It's available to paid plans in ChatGPT Work and Codex. Developers can call it through the OpenAI API or AIHubMix.

Do I need to change my code to upgrade from 6 Sol?
If you don't use none effort and don't call tools through Chat Completions, changing the model name is usually enough. Otherwise you'll need to move to the Responses API. Part 4 has the details.


Keep reading: the GPT-6.1 Sol series


Sources

Top comments (0)