DEV Community

Cover image for GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark
Hassann
Hassann

Posted on Originally published at apidog.com

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol and Claude Sonnet 5.5 both list at $2 per million input tokens and $10 per million output tokens. They launched a day apart: Sonnet 5.5 on September 28, 2026, and GPT-6.1 Sol on September 29. Neither vendor has benchmarked one directly against the other. OpenAI compared GPT-6.1 Sol with Claude Opus 5.5 and Fable 5.1, while Anthropic compared Sonnet 5.5 with GPT-6 Sol, the previous Sol model. To choose between them, run both on your own tasks and compare cost per passing task, not just price per token.

Try Apidog today

This guide covers the specs and prices side by side, the vendor-reported results and their comparison baselines, available third-party numbers (all for GPT-6 Sol, not 6.1), and a repeatable way to test both APIs in Apidog. For model-specific details, see what is GPT-6.1 Sol and what is Claude Sonnet 5.5.

GPT-6.1 Sol vs Claude Sonnet 5.5: specs and pricing

Sources: OpenAI’s pricing page, Anthropic’s Sonnet 5.5 overview and pricing, plus the GPT-6.1 Sol model page linked below.

GPT-6.1 Sol Claude Sonnet 5.5
Model ID gpt-6.1-sol claude-sonnet-5-5 (Bedrock: anthropic.claude-sonnet-5-5)
Released Sep 29, 2026 Sep 28, 2026
Input / output per 1M $2 / $10 $2 / $10
Cache reads per 1M $0.10 $0.20
Cache writes per 1M $2.50 $2.50 (5-minute), $4 (1-hour)
Batch per 1M $1 / $5 $1 / $5
Prompts over 272K input tokens $4 input, $0.20 cached, $15 output for the whole request No premium across the 1M window
Faster tier Fast at $4 / $20; Ultrafast “coming soon” Fast mode not available
Context window 1,050,000 (922,000 max input) 1M
Max output 128,000 128K (300K on Batch with a beta header)
Knowledge cutoff Apr 30, 2026 June 2026
Effort levels low, medium (default), high, xhigh, max; no none low, medium, high (API default), xhigh, max; thinking can’t be turned off (lowest setting: between_tools)
API shape Responses (with tools), Chat Completions (no tools), Batch Messages, Batch
Other platforms OpenRouter (openai/gpt-6.1-sol) Bedrock, Google Cloud, Microsoft Foundry, Claude Platform on AWS
Free access None. Plus, Pro, Business, Enterprise and Edu get it in ChatGPT Work and Codex, not Chat; Enterprise and Edu require admin enablement, per the models docs. No free API tier. Anyone can chat with it on Claude.ai; the API bills per token.

Two pricing details have the biggest impact on implementation:

  1. Cache-heavy prompts under 272K input tokens: GPT-6.1 Sol cache reads cost half as much as Sonnet 5.5 cache reads. Reusing a large system prompt can therefore cost less on OpenAI.
  2. Large prompts over 272K input tokens: the GPT-6.1 Sol model page applies premium pricing to the full request: 2x input and cache rates, plus 1.5x output pricing. Sonnet 5.5 remains at $2 input and $10 output per million tokens.

For example, an uncached 400,000-token prompt with a 5,000-token response costs approximately:

  • GPT-6.1 Sol: $1.68
  • Claude Sonnet 5.5: $0.85

Do not use this example as a final decision metric. Actual task cost also depends on reasoning-token consumption, cache behavior, retries, and how many outputs pass your checks.

What each vendor claims, and against which model

OpenAI on GPT-6.1 Sol Anthropic on Claude Sonnet 5.5
Compared against GPT-6 Sol, GPT-6 Astra, Claude Opus 5.5, Claude Fable 5.1 Sonnet 5, Opus 5.5, GPT-6 Sol (GPT-5.6 Sol in two charts)
Includes the other model No No; GPT-6.1 Sol did not exist yet
Headline “Near-Astra intelligence” at one-fifth of Astra’s standard token prices 30%+ faster than Sonnet 5, up to 30% less per task
Who ran the numbers OpenAI; competitor results came from public reports, per its footnote Anthropic, plus Cognition, Cursor, Artificial Analysis, and Surge AI for named benchmarks

OpenAI’s GPT-6.1 Sol launch post reports:

  • AutomationBench 1.0.6, medium effort: +2.2 percentage points over Opus 5.5 at roughly one-third of the cost, and +4.8 percentage points over GPT-6 Sol.
  • Terminal-Bench Science 0.1, max effort: $5.47 per task, versus $23.21 for Opus 5.5. GPT-6 Astra still scores highest at 68.1%.
  • OSWorld 2.0 offline, max effort: +7 percentage points over GPT-6 Sol for less than half the cost.

Anthropic’s Sonnet 5.5 launch post includes a GPT-6 Sol column:

  • FrontierCode 1.1, run by Cognition: Sonnet 5.5 scores 52.1% at xhigh effort and 46.2% at max, versus 49.3% for GPT-6 Sol. Anthropic says Sonnet 5.5 at high effort matches GPT-6 Sol’s best score for about one-fifth the cost.
  • GDPval-AA v2.1 and AA-Briefcase v1.1, run by Artificial Analysis: 1844 versus 1487, and 1811 versus 1483.

See the Sonnet 5.5 benchmarks breakdown for the rest of Anthropic’s table.

Do not bridge benchmark charts without matching settings

Both vendors include Opus 5.5 in their charts and report AutomationBench and Terminal-Bench Science. That does not make the published values directly comparable.

The test conditions differ:

  • OpenAI reports AutomationBench 1.0.6 at medium effort in its own harness.
  • Anthropic’s system card summary runs Claude models at max effort: Sonnet 5.5 at 44.7, Opus 5.5 at 42.5, and GPT-6 Sol at 32.0.
  • Anthropic reports Sonnet 5.5 at 59.9% on Terminal-Bench Science, while OpenAI does not publish a raw GPT-6.1 Sol score in its launch-post text.
  • OpenAI tested computer use on OSWorld 2.0 offline, while Anthropic tested OSWorld 2.1.

The GPT-6 Sol vs Claude Opus 5.5 comparison runs into the same issue.

What third parties measured: GPT-6 Sol, not GPT-6.1

Current third-party comparisons pair Sonnet 5.5 with GPT-6 Sol rather than GPT-6.1 Sol. The most complete source is Artificial Analysis, whose Intelligence Index v4.3.2 aggregates 10 evaluations.

Effort Sonnet 5.5 index Sonnet 5.5 cost per task GPT-6 Sol index GPT-6 Sol cost per task
medium 41 $0.59 40 $0.25
high 47 $1.08 43 $0.38
xhigh 52 $2.74 44 $0.52
max 56 $7.60 48 $1.05

At max effort, the same source measures:

  • Sonnet 5.5: 138 output tokens per second
  • GPT-6 Sol: 76 output tokens per second

Sonnet 5.5 scores higher and costs more per task at every effort shown, despite equal list pricing. The difference is token consumption. For example, Sonnet 5.5 at high effort scores 47 at $1.08 per task, while GPT-6 Sol at max scores 48 at $1.05 per task.

Anthropic’s own chart data also shows task-specific tradeoffs:

  • On AA-Briefcase, GPT-6 Sol costs less per task at every effort, such as $0.34 versus $1.64 at medium, while Sonnet 5.5 scores higher.
  • On FrontierCode, Sonnet 5.5 at high reaches 49.4% for $0.42 per task, compared with GPT-6 Sol at max at 49.3% for $2.07.

These results are for the older Sol model. OpenAI says GPT-6.1 Sol beats GPT-6 Sol at the same or lower effort on several benchmarks, so expect third-party comparisons to change when GPT-6.1 Sol is independently tested.

Where each is likely stronger, based on vendor framing

GPT-6.1 Sol

OpenAI positions GPT-6.1 Sol for complex coding, computer use, and professional work where near-Astra performance is needed at lower cost.

Its stated gains focus on:

  • Agentic automation
  • Computer use
  • Science tasks
  • Fewer factual errors at low effort

From a cost and deployment perspective, it is more favorable for cache-heavy workloads under 272K input tokens. It is also the only model in this pair with a paid speed tier: Fast.

Claude Sonnet 5.5

Anthropic positions Sonnet 5.5 for well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets. Anthropic says Opus 5.5 remains stronger for complex, open-ended work; see Sonnet 5.5 vs Opus 5.5.

Its reported strengths include:

  • Agentic coding, including 70.6% on Terminal-Bench 4.0 at max effort in Anthropic-run testing
  • Knowledge work
  • Computer use, including 80.1% partial on OSWorld 2.1

For pricing, Sonnet 5.5 is more favorable when prompts exceed 272K tokens. It is also available through Bedrock, Google Cloud, and Microsoft Foundry.

Control for effort and sampling differences

A naive A/B test can produce misleading results for two reasons:

  1. Different default effort levels
    • GPT-6.1 Sol defaults to medium.
    • Sonnet 5.5 defaults to high in the API.

Always set matching named effort levels in both requests. Then test at least medium and high.

  1. Sampling controls are restricted
    • Sonnet 5.5 returns HTTP 400 for non-default temperature, top_p, or top_k.
    • OpenAI’s GPT-6 guidance says to omit temperature and top_p when effort is not none.

Use repeated runs to measure variance instead of trying to normalize results with unsupported sampling parameters.

Run the comparison yourself in Apidog

Use 20 to 50 tasks from production traffic or your backlog. Each task should have a checkable result, such as:

  • An expected JSON field
  • A classification label
  • A unit-test result
  • A schema-valid output
  • A deterministic text match where appropriate

1. Create an environment

In Apidog, create an environment with these secret variables:

OPENAI_API_KEY
ANTHROPIC_API_KEY
EFFORT
Enter fullscreen mode Exit fullscreen mode

Also define a shared prompt variable for your test data.

2. Create one request per provider

The APIs use different request formats. See the OpenAI Responses reference, the Anthropic Messages reference, and this API formats comparison.

OpenAI Responses API

POST https://api.openai.com/v1/responses
Authorization: Bearer {{OPENAI_API_KEY}}
Content-Type: application/json

{"model": "gpt-6.1-sol", "reasoning": {"effort": "{{EFFORT}}"},
 "max_output_tokens": 25000, "input": "{{prompt}}"}
Enter fullscreen mode Exit fullscreen mode

Anthropic Messages API

POST https://api.anthropic.com/v1/messages
x-api-key: {{ANTHROPIC_API_KEY}}
anthropic-version: 2023-06-01
content-type: application/json

{"model": "claude-sonnet-5-5", "max_tokens": 25000,
 "output_config": {"effort": "{{EFFORT}}"},
 "messages": [{"role": "user", "content": "{{prompt}}"}]}
Enter fullscreen mode Exit fullscreen mode

3. Add response assertions

First, assert that each request completed normally. Then add an assertion against an expected column in your test dataset.

Check OpenAI Responses Anthropic Messages
Finished normally $.status equals completed $.stop_reason equals end_turn
Answer present $.output[*].type contains message $.content[*].type contains text
Output tokens, reasoning included $.usage.output_tokens $.usage.output_tokens
Cache reads $.usage.input_tokens_details.cached_tokens $.usage.cache_read_input_tokens
Cache writes $.usage.input_tokens_details.cache_write_tokens $.usage.cache_creation_input_tokens

For structured outputs, validate the actual payload rather than only checking whether a text block exists. For example, assert a required JSON field, an expected enum value, or a schema validation result.

4. Calculate cost per response

Calculate request cost in a post-processor script.

Important accounting differences:

  • OpenAI’s input_tokens includes cached and cache-write tokens. Subtract those values before applying the standard $2 input rate. See OpenAI prompt caching.
  • Anthropic’s input_tokens counts only tokens after the last cache breakpoint. See Anthropic prompt caching.
  • On OpenAI, apply the 272K input-token pricing rule when applicable.

For each request, record at least:

provider
model
effort
input_tokens
output_tokens
cache_read_tokens
cache_write_tokens
request_cost
passed
latency
Enter fullscreen mode Exit fullscreen mode

5. Run both requests from one scenario

Put the OpenAI and Anthropic requests in the same test scenario.

Use a CSV dataset with one prompt per row. Keep prompts on a single line and avoid double quotes.

apidog run --access-token "$APIDOG_ACCESS_TOKEN" -t "$SCENARIO_ID" -e "$ENV_ID" \
  -d prompts.csv --env-var "EFFORT=medium" -r cli,junit
Enter fullscreen mode Exit fullscreen mode

Repeat the run for each effort level you want to evaluate:

# Example effort runs
EFFORT=medium
EFFORT=high
Enter fullscreen mode Exit fullscreen mode

Run both models at least at medium and high. The effort names match, but their internal calibration is not guaranteed to match across vendors.

6. Compare cost per passing task

Use this calculation:

cost_per_passing_task = total_spend / number_of_tasks_that_passed
Enter fullscreen mode Exit fullscreen mode

Also report:

pass_rate = passed_tasks / total_tasks
average_latency = total_latency / completed_tasks
Enter fullscreen mode Exit fullscreen mode

This gives you a practical decision table:

Model Effort Pass rate Total spend Cost per passing task Average latency
GPT-6.1 Sol medium
Claude Sonnet 5.5 medium
GPT-6.1 Sol high
Claude Sonnet 5.5 high

For request-level implementation details, see the GPT-6.1 Sol API guide and how to use the Claude Sonnet 5.5 API.

FAQ

Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?

They have the same list price: $2 per million input tokens and $10 per million output tokens. GPT-6.1 Sol has cheaper cache reads, while Sonnet 5.5 is cheaper for prompts above 272K input tokens. Actual cost depends on token use per task.

Which is better at coding?

There is no shared GPT-6.1 Sol versus Sonnet 5.5 coding benchmark. Anthropic reports Sonnet 5.5 above GPT-6 Sol on FrontierCode. OpenAI reports GPT-6.1 Sol beating GPT-6 Sol’s best DeepSWE score by 6.4 percentage points. Test both on your own repository and validation suite.

Which is faster?

Artificial Analysis measured Sonnet 5.5 at 138 tokens per second and GPT-6 Sol at 76 tokens per second, both at max effort. There is no third-party throughput measurement for GPT-6.1 Sol yet.

Is either model free?

GPT-6.1 Sol has no free tier. Sonnet 5.5 is available in Claude.ai chat apps, but API use is billed per token. See how to use Claude Sonnet 5.5 for free.

Pick by running both

Start with ten representative tasks from your backlog. Run both models at medium and high, validate each response automatically, and compare cost per passing task.

Use Apidog to keep both provider requests in one reusable scenario. Rerun it when Artificial Analysis publishes GPT-6.1 Sol measurements or when either vendor releases a new snapshot.

Top comments (0)