GPT-6.1 Sol and Claude Sonnet 5.5 both list at $2 per million input tokens and $10 per million output tokens. They launched a day apart: Sonnet 5.5 on September 28, 2026, and GPT-6.1 Sol on September 29. Neither vendor has benchmarked one directly against the other. OpenAI compared GPT-6.1 Sol with Claude Opus 5.5 and Fable 5.1, while Anthropic compared Sonnet 5.5 with GPT-6 Sol, the previous Sol model. To choose between them, run both on your own tasks and compare cost per passing task, not just price per token.
This guide covers the specs and prices side by side, the vendor-reported results and their comparison baselines, available third-party numbers (all for GPT-6 Sol, not 6.1), and a repeatable way to test both APIs in Apidog. For model-specific details, see what is GPT-6.1 Sol and what is Claude Sonnet 5.5.
GPT-6.1 Sol vs Claude Sonnet 5.5: specs and pricing
Sources: OpenAI’s pricing page, Anthropic’s Sonnet 5.5 overview and pricing, plus the GPT-6.1 Sol model page linked below.
| GPT-6.1 Sol | Claude Sonnet 5.5 | |
|---|---|---|
| Model ID | gpt-6.1-sol |
claude-sonnet-5-5 (Bedrock: anthropic.claude-sonnet-5-5) |
| Released | Sep 29, 2026 | Sep 28, 2026 |
| Input / output per 1M | $2 / $10 | $2 / $10 |
| Cache reads per 1M | $0.10 | $0.20 |
| Cache writes per 1M | $2.50 | $2.50 (5-minute), $4 (1-hour) |
| Batch per 1M | $1 / $5 | $1 / $5 |
| Prompts over 272K input tokens | $4 input, $0.20 cached, $15 output for the whole request | No premium across the 1M window |
| Faster tier | Fast at $4 / $20; Ultrafast “coming soon” | Fast mode not available |
| Context window | 1,050,000 (922,000 max input) | 1M |
| Max output | 128,000 | 128K (300K on Batch with a beta header) |
| Knowledge cutoff | Apr 30, 2026 | June 2026 |
| Effort levels |
low, medium (default), high, xhigh, max; no none
|
low, medium, high (API default), xhigh, max; thinking can’t be turned off (lowest setting: between_tools) |
| API shape | Responses (with tools), Chat Completions (no tools), Batch | Messages, Batch |
| Other platforms | OpenRouter (openai/gpt-6.1-sol) |
Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS |
| Free access | None. Plus, Pro, Business, Enterprise and Edu get it in ChatGPT Work and Codex, not Chat; Enterprise and Edu require admin enablement, per the models docs. No free API tier. | Anyone can chat with it on Claude.ai; the API bills per token. |
Two pricing details have the biggest impact on implementation:
- Cache-heavy prompts under 272K input tokens: GPT-6.1 Sol cache reads cost half as much as Sonnet 5.5 cache reads. Reusing a large system prompt can therefore cost less on OpenAI.
- Large prompts over 272K input tokens: the GPT-6.1 Sol model page applies premium pricing to the full request: 2x input and cache rates, plus 1.5x output pricing. Sonnet 5.5 remains at $2 input and $10 output per million tokens.
For example, an uncached 400,000-token prompt with a 5,000-token response costs approximately:
- GPT-6.1 Sol: $1.68
- Claude Sonnet 5.5: $0.85
Do not use this example as a final decision metric. Actual task cost also depends on reasoning-token consumption, cache behavior, retries, and how many outputs pass your checks.
What each vendor claims, and against which model
| OpenAI on GPT-6.1 Sol | Anthropic on Claude Sonnet 5.5 | |
|---|---|---|
| Compared against | GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 | Sonnet 5, Opus 5.5, GPT-6 Sol (GPT-5.6 Sol in two charts) |
| Includes the other model | No | No; GPT-6.1 Sol did not exist yet |
| Headline | “Near-Astra intelligence” at one-fifth of Astra’s standard token prices | 30%+ faster than Sonnet 5, up to 30% less per task |
| Who ran the numbers | OpenAI; competitor results came from public reports, per its footnote | Anthropic, plus Cognition, Cursor, Artificial Analysis, and Surge AI for named benchmarks |
OpenAI’s GPT-6.1 Sol launch post reports:
- AutomationBench 1.0.6, medium effort: +2.2 percentage points over Opus 5.5 at roughly one-third of the cost, and +4.8 percentage points over GPT-6 Sol.
- Terminal-Bench Science 0.1, max effort: $5.47 per task, versus $23.21 for Opus 5.5. GPT-6 Astra still scores highest at 68.1%.
- OSWorld 2.0 offline, max effort: +7 percentage points over GPT-6 Sol for less than half the cost.
Anthropic’s Sonnet 5.5 launch post includes a GPT-6 Sol column:
- FrontierCode 1.1, run by Cognition: Sonnet 5.5 scores 52.1% at xhigh effort and 46.2% at max, versus 49.3% for GPT-6 Sol. Anthropic says Sonnet 5.5 at high effort matches GPT-6 Sol’s best score for about one-fifth the cost.
- GDPval-AA v2.1 and AA-Briefcase v1.1, run by Artificial Analysis: 1844 versus 1487, and 1811 versus 1483.
See the Sonnet 5.5 benchmarks breakdown for the rest of Anthropic’s table.
Do not bridge benchmark charts without matching settings
Both vendors include Opus 5.5 in their charts and report AutomationBench and Terminal-Bench Science. That does not make the published values directly comparable.
The test conditions differ:
- OpenAI reports AutomationBench 1.0.6 at medium effort in its own harness.
- Anthropic’s system card summary runs Claude models at max effort: Sonnet 5.5 at 44.7, Opus 5.5 at 42.5, and GPT-6 Sol at 32.0.
- Anthropic reports Sonnet 5.5 at 59.9% on Terminal-Bench Science, while OpenAI does not publish a raw GPT-6.1 Sol score in its launch-post text.
- OpenAI tested computer use on OSWorld 2.0 offline, while Anthropic tested OSWorld 2.1.
The GPT-6 Sol vs Claude Opus 5.5 comparison runs into the same issue.
What third parties measured: GPT-6 Sol, not GPT-6.1
Current third-party comparisons pair Sonnet 5.5 with GPT-6 Sol rather than GPT-6.1 Sol. The most complete source is Artificial Analysis, whose Intelligence Index v4.3.2 aggregates 10 evaluations.
| Effort | Sonnet 5.5 index | Sonnet 5.5 cost per task | GPT-6 Sol index | GPT-6 Sol cost per task |
|---|---|---|---|---|
| medium | 41 | $0.59 | 40 | $0.25 |
| high | 47 | $1.08 | 43 | $0.38 |
| xhigh | 52 | $2.74 | 44 | $0.52 |
| max | 56 | $7.60 | 48 | $1.05 |
At max effort, the same source measures:
- Sonnet 5.5: 138 output tokens per second
- GPT-6 Sol: 76 output tokens per second
Sonnet 5.5 scores higher and costs more per task at every effort shown, despite equal list pricing. The difference is token consumption. For example, Sonnet 5.5 at high effort scores 47 at $1.08 per task, while GPT-6 Sol at max scores 48 at $1.05 per task.
Anthropic’s own chart data also shows task-specific tradeoffs:
- On AA-Briefcase, GPT-6 Sol costs less per task at every effort, such as $0.34 versus $1.64 at medium, while Sonnet 5.5 scores higher.
- On FrontierCode, Sonnet 5.5 at high reaches 49.4% for $0.42 per task, compared with GPT-6 Sol at max at 49.3% for $2.07.
These results are for the older Sol model. OpenAI says GPT-6.1 Sol beats GPT-6 Sol at the same or lower effort on several benchmarks, so expect third-party comparisons to change when GPT-6.1 Sol is independently tested.
Where each is likely stronger, based on vendor framing
GPT-6.1 Sol
OpenAI positions GPT-6.1 Sol for complex coding, computer use, and professional work where near-Astra performance is needed at lower cost.
Its stated gains focus on:
- Agentic automation
- Computer use
- Science tasks
- Fewer factual errors at low effort
From a cost and deployment perspective, it is more favorable for cache-heavy workloads under 272K input tokens. It is also the only model in this pair with a paid speed tier: Fast.
Claude Sonnet 5.5
Anthropic positions Sonnet 5.5 for well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets. Anthropic says Opus 5.5 remains stronger for complex, open-ended work; see Sonnet 5.5 vs Opus 5.5.
Its reported strengths include:
- Agentic coding, including 70.6% on Terminal-Bench 4.0 at max effort in Anthropic-run testing
- Knowledge work
- Computer use, including 80.1% partial on OSWorld 2.1
For pricing, Sonnet 5.5 is more favorable when prompts exceed 272K tokens. It is also available through Bedrock, Google Cloud, and Microsoft Foundry.
Control for effort and sampling differences
A naive A/B test can produce misleading results for two reasons:
-
Different default effort levels
- GPT-6.1 Sol defaults to
medium. - Sonnet 5.5 defaults to
highin the API.
- GPT-6.1 Sol defaults to
Always set matching named effort levels in both requests. Then test at least medium and high.
-
Sampling controls are restricted
- Sonnet 5.5 returns HTTP 400 for non-default
temperature,top_p, ortop_k. - OpenAI’s GPT-6 guidance says to omit
temperatureandtop_pwhen effort is notnone.
- Sonnet 5.5 returns HTTP 400 for non-default
Use repeated runs to measure variance instead of trying to normalize results with unsupported sampling parameters.
Run the comparison yourself in Apidog
Use 20 to 50 tasks from production traffic or your backlog. Each task should have a checkable result, such as:
- An expected JSON field
- A classification label
- A unit-test result
- A schema-valid output
- A deterministic text match where appropriate
1. Create an environment
In Apidog, create an environment with these secret variables:
OPENAI_API_KEY
ANTHROPIC_API_KEY
EFFORT
Also define a shared prompt variable for your test data.
2. Create one request per provider
The APIs use different request formats. See the OpenAI Responses reference, the Anthropic Messages reference, and this API formats comparison.
OpenAI Responses API
POST https://api.openai.com/v1/responses
Authorization: Bearer {{OPENAI_API_KEY}}
Content-Type: application/json
{"model": "gpt-6.1-sol", "reasoning": {"effort": "{{EFFORT}}"},
"max_output_tokens": 25000, "input": "{{prompt}}"}
Anthropic Messages API
POST https://api.anthropic.com/v1/messages
x-api-key: {{ANTHROPIC_API_KEY}}
anthropic-version: 2023-06-01
content-type: application/json
{"model": "claude-sonnet-5-5", "max_tokens": 25000,
"output_config": {"effort": "{{EFFORT}}"},
"messages": [{"role": "user", "content": "{{prompt}}"}]}
3. Add response assertions
First, assert that each request completed normally. Then add an assertion against an expected column in your test dataset.
| Check | OpenAI Responses | Anthropic Messages |
|---|---|---|
| Finished normally |
$.status equals completed
|
$.stop_reason equals end_turn
|
| Answer present |
$.output[*].type contains message
|
$.content[*].type contains text
|
| Output tokens, reasoning included | $.usage.output_tokens |
$.usage.output_tokens |
| Cache reads | $.usage.input_tokens_details.cached_tokens |
$.usage.cache_read_input_tokens |
| Cache writes | $.usage.input_tokens_details.cache_write_tokens |
$.usage.cache_creation_input_tokens |
For structured outputs, validate the actual payload rather than only checking whether a text block exists. For example, assert a required JSON field, an expected enum value, or a schema validation result.
4. Calculate cost per response
Calculate request cost in a post-processor script.
Important accounting differences:
- OpenAI’s
input_tokensincludes cached and cache-write tokens. Subtract those values before applying the standard $2 input rate. See OpenAI prompt caching. - Anthropic’s
input_tokenscounts only tokens after the last cache breakpoint. See Anthropic prompt caching. - On OpenAI, apply the 272K input-token pricing rule when applicable.
For each request, record at least:
provider
model
effort
input_tokens
output_tokens
cache_read_tokens
cache_write_tokens
request_cost
passed
latency
5. Run both requests from one scenario
Put the OpenAI and Anthropic requests in the same test scenario.
Use a CSV dataset with one prompt per row. Keep prompts on a single line and avoid double quotes.
apidog run --access-token "$APIDOG_ACCESS_TOKEN" -t "$SCENARIO_ID" -e "$ENV_ID" \
-d prompts.csv --env-var "EFFORT=medium" -r cli,junit
Repeat the run for each effort level you want to evaluate:
# Example effort runs
EFFORT=medium
EFFORT=high
Run both models at least at medium and high. The effort names match, but their internal calibration is not guaranteed to match across vendors.
6. Compare cost per passing task
Use this calculation:
cost_per_passing_task = total_spend / number_of_tasks_that_passed
Also report:
pass_rate = passed_tasks / total_tasks
average_latency = total_latency / completed_tasks
This gives you a practical decision table:
| Model | Effort | Pass rate | Total spend | Cost per passing task | Average latency |
|---|---|---|---|---|---|
| GPT-6.1 Sol | medium | ||||
| Claude Sonnet 5.5 | medium | ||||
| GPT-6.1 Sol | high | ||||
| Claude Sonnet 5.5 | high |
For request-level implementation details, see the GPT-6.1 Sol API guide and how to use the Claude Sonnet 5.5 API.
FAQ
Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?
They have the same list price: $2 per million input tokens and $10 per million output tokens. GPT-6.1 Sol has cheaper cache reads, while Sonnet 5.5 is cheaper for prompts above 272K input tokens. Actual cost depends on token use per task.
Which is better at coding?
There is no shared GPT-6.1 Sol versus Sonnet 5.5 coding benchmark. Anthropic reports Sonnet 5.5 above GPT-6 Sol on FrontierCode. OpenAI reports GPT-6.1 Sol beating GPT-6 Sol’s best DeepSWE score by 6.4 percentage points. Test both on your own repository and validation suite.
Which is faster?
Artificial Analysis measured Sonnet 5.5 at 138 tokens per second and GPT-6 Sol at 76 tokens per second, both at max effort. There is no third-party throughput measurement for GPT-6.1 Sol yet.
Is either model free?
GPT-6.1 Sol has no free tier. Sonnet 5.5 is available in Claude.ai chat apps, but API use is billed per token. See how to use Claude Sonnet 5.5 for free.
Pick by running both
Start with ten representative tasks from your backlog. Run both models at medium and high, validate each response automatically, and compare cost per passing task.
Use Apidog to keep both provider requests in one reusable scenario. Rerun it when Artificial Analysis publishes GPT-6.1 Sol measurements or when either vendor releases a new snapshot.
Top comments (0)