DEV Community

Cover image for GPT-6.1 Sol vs Claude Sonnet 5.5: Same $2/$10, Half the Cost per Task
synthorai
synthorai

Posted on Originally published at synthorai.io

GPT-6.1 Sol vs Claude Sonnet 5.5: Same $2/$10, Half the Cost per Task

GPT-6.1 Sol and Claude Sonnet 5.5 have the same list price, $2 per million input tokens and $10 per million output, and very different costs per task. With each model at its API default effort (the setting for how much it reasons before answering), GPT-6.1 Sol cost about half as much (0.49x), and at a matched score on an independent index about a quarter. Against Claude Opus 5.5, the rival OpenAI named, it cost 0.29x. What Sonnet 5.5 offers for the extra cost is a top score 4 to 6 points higher on independent indexes, faster responses, and no wrong answers where GPT-6.1 Sol repeatedly missed one task.

TL;DR

  • At the API defaults, GPT-6.1 Sol cost 0.49x Sonnet 5.5 and 0.29x Opus 5.5 per task on 13 tasks.
  • Artificial Analysis scores GPT-6.1 Sol at max and Sonnet 5.5 at xhigh both 52, at $0.72 and $2.74 per task.
  • GPT-6.1 Sol returned a bare answer (only the value asked for) on 234 of 234 prompts; Sonnet 5.5 on 22 of 39 at its default.
  • Sonnet 5.5 was correct on 234 of 234 calls; GPT-6.1 Sol missed one task on 7 of 18 calls.

What is GPT-6.1 Sol, and what does OpenAI claim for it?

GPT-6.1 Sol is OpenAI's update to its mid-tier model, released on 2026-09-29, seven days after GPT-6 Sol at the same price. In the week between, developers criticized GPT-6 Sol for short answers, Anthropic shipped Sonnet 5.5 at the same $2 / $10, and OpenAI cancelled its planned GPT-6.1 Astra flagship after internal tests found deception and unauthorized actions, as CBC reported.

Token prices, the 1,050,000-token context window and the 128,000-token output limit are unchanged; cached input now costs $0.10 per million tokens, down from $0.20. The change that breaks code is in reasoning.effort, the Responses API parameter that sets how much hidden reasoning (billed as output) the model does before it answers: it runs from low to max with medium as the default, and none is gone. Requests that turned reasoning off on GPT-6 Sol need low, and tool calls, which Chat Completions allowed only at none, now need the Responses API.

OpenAI positions the model against two flagships, its own GPT-6 Astra and Anthropic's Claude Opus 5.5: level with Astra on DeepSWE v1.1 (agentic coding) at about a fifth of the cost, 2.1 points short of Astra on OSWorld 2.0 (computer use) at about a seventh of the cost, 2.2 points above Opus 5.5 on AutomationBench (business workflows) at medium, and $5.47 per Terminal-Bench Science task at max against Opus 5.5's $23.21. These are OpenAI's own runs, and Sonnet 5.5 does not appear in them. The launch also announced an Ultrafast tier, up to 6x faster at 6x the price.

Which model should it be compared with: Opus 5.5, Sonnet 5.5 or GPT-6 Sol?

Sonnet 5.5 is the right comparison, because it is the alternative at the same list price, and the comparison shows how little a shared price list says about cost. The table puts the three candidates next to GPT-6.1 Sol, with independent benchmark results read from the publishers' own pages.

GPT-6.1 Sol Opus 5.5 Sonnet 5.5 GPT-6 Sol
List price, input / output per 1M tokens $2 / $10 $4 / $20 $2 / $10 $2 / $10
Released 2026-09-29 2026-09-22 2026-09-28 2026-09-22
Its launch compared it with GPT-6 Astra, Opus 5.5 Fable 5.1, Opus 5, GPT-6 Astra Opus 5.5, GPT-6 Sol GPT-6 Astra, Opus 5, Fable 5
Artificial Analysis Intelligence Index at max (reasoning, knowledge and coding evals) 52 58 56 48
Cost per index task at max $0.72 $5.98 $7.60 $1.04
Vals Index (coding, legal, tax, medical and other professional tasks) 61.2% 67.0% 67.0% 57.5%
Cost per Vals test $3.24 $32.14 $21.34 $7.58
AutomationBench 1.0.6 public leaderboard at max (business workflows across apps) not listed 42.47%, $1.44 per task 44.75%, $1.14 per task not listed
  • Opus 5.5 is the rival OpenAI named, but it lists at twice the price and scores about 6 points higher on both indexes.
  • GPT-6 Sol is the same-price predecessor, which GPT-6.1 Sol beats on both indexes for less per task. That comparison only decides whether to upgrade.
  • Sonnet 5.5 has the same list price, launched a day earlier and was pitched the same way, against Opus 5.5. It scores within 2 points of Opus 5.5 and 4 to 6 points above GPT-6.1 Sol.

The shared price list does not make the two equivalent. GPT-6.1 Sol is ahead of both Claude models on cost for the score it delivers:

  • At a matched score, it costs a fifth to a quarter of what Sonnet 5.5 does: Artificial Analysis has Sonnet 5.5 at xhigh and GPT-6.1 Sol at max both on 52, at $2.74 against $0.72 per task. One step down, Sonnet 5.5 at high scores 47 at $1.08 and GPT-6.1 Sol at medium 48 at $0.21.
  • In the headline index results, Sonnet 5.5's cost per task is in Opus 5.5's range ($7.60 against $5.98 on Artificial Analysis at max, $21.34 against $32.14 on Vals), and each costs 7 to 11 times as much as GPT-6.1 Sol.

In return, Sonnet 5.5 and Opus 5.5 reach Artificial Analysis scores of 56 and 58, which no GPT-6.1 Sol setting does.

How did we test it?

We ran three sets of tasks, all graded in code, on each model's native API (Responses for GPT, Messages for Claude):

Test What it contains What we measured
Single-shot tasks 13 tasks with known answers (for example a 10-item knapsack: find the best set of items under a weight limit), each ending "Reply with a single integer, nothing else"; 3 repeats per effort setting right answer, bare answer (the reply is only the value), cost, time
Business documents 80 goal-only prompts ("Write the Q3 regional review") and 80 that name the required parts: quarterly reviews, incident postmortems, vendor comparisons, support replies every expected part present, words, cost
Tool loop 4 code-reading questions over a small synthetic codebase, 3 tools, 3 repeats questions solved, tool calls, cost, time

Every prompt carried a unique random string so no response came from a cache, and cost uses list prices. The Sonnet 5.5 and Opus 5.5 single-shot and tool-loop numbers come from our Sonnet 5.5 measurements a day earlier, on the same tasks and graders. A gap is called significant only if it survives a Holm correction, which tightens the threshold when several comparisons are tested at once; bare-answer comparisons are counted per task, because repeats of one task are not independent.

Is GPT-6.1 Sol cheaper than Sonnet 5.5 per task?

GPT-6.1 Sol cost about half as much: at each model's API default, it cost $0.0020 per single-shot task and Sonnet 5.5 $0.0042, a ratio of 0.49 (95% interval 0.40 to 0.59, from resampling the 13 tasks). Sonnet 5.5 wrote 1.7 to 2.5 times the output tokens at the settings below max and 3.1 times at max. Opus 5.5, at twice the list price, cost $0.0070 at its default, so GPT-6.1 Sol cost 0.29x as much.

The ratio held at every effort setting, at 0.33x to 0.58x of Sonnet 5.5's cost and 0.20x to 0.31x of Opus 5.5's. With no effort sent, GPT-6.1 Sol and Opus 5.5 run at medium and Sonnet 5.5 at high. The other two tests gave the same ratio against Sonnet 5.5, 0.46x: $0.0103 against $0.0223 per goal-only business document and $0.0052 against $0.0112 per tool-loop run.

Grouped bar chart of cost per 1,000 single-shot tasks by effort setting for GPT-6.1 Sol, Claude Sonnet 5.5 and Claude Opus 5.5. At the defaults: $2.0, $4.1 and $7.0. At max: $4.7, $14.6 and $23.4. GPT-6.1 Sol is the lowest bar at every setting. Bare-answer replies are 39 of 39 at every setting for GPT-6.1 Sol and Opus 5.5, and 12 to 39 of 39 for Sonnet 5.5

Effort is one request field, and the billed tokens come back in usage:

from openai import OpenAI

client = OpenAI()
resp = client.responses.create(
    model="gpt-6.1-sol",
    reasoning={"effort": "low"},  # low, medium (default), high, xhigh, max; "none" is not listed for 6.1
    input="Compute 7 raised to the power 222, modulo 1000. Reply with a single integer, nothing else.",
)
print(resp.output_text)
u = resp.usage
print(u.output_tokens, u.output_tokens_details.reasoning_tokens, u.input_tokens_details.cached_tokens)
Enter fullscreen mode Exit fullscreen mode

output_tokens includes the reasoning tokens, and cached_tokens is the part of the input billed at $0.10 per million.

Which one follows "answer only", and which gets the answer right?

GPT-6.1 Sol followed the format on every prompt, and Sonnet 5.5 made no errors where GPT-6.1 Sol repeatedly missed one task. GPT-6.1 Sol replied with the bare value on all 234 prompts, as did Opus 5.5. Sonnet 5.5 did so on 22 of 39 at its default, 12 at low, 24 at medium and 30 at high; below xhigh it sometimes skips its thinking block (the separate reasoning part of a Claude response) and writes the working into the reply, as our Sonnet 5.5 post describes. Counted per task, the gap survives the correction at low (GPT-6.1 Sol better on 9 of 13 tasks, worse on none) and not at the defaults (6 and none). Code that parses a single value can read GPT-6.1 Sol's reply as-is; with Sonnet 5.5, take the last line or use xhigh.

Sonnet 5.5 answered all 234 calls correctly and Opus 5.5 missed 2. GPT-6.1 Sol answered the knapsack question with 72 on 7 of its 18 calls (the correct value, checked by trying every subset, is 62): three of three at low, one of three at its default and at medium, two of three at xhigh, none at high or max. One task out of 13 does not establish an accuracy gap, but it shows the error is silent: the reply is a well-formed integer, and raising effort did not reliably prevent it.

Did GPT-6.1 Sol fix GPT-6 Sol's short answers?

GPT-6.1 Sol writes full-length answers again. On goal-only prompts, which name the document and nothing about its parts, it wrote 565 words per document against 325 for GPT-6 Sol, longer on all 80 items and level with GPT-5.6 Sol at 592. The complaint that GPT-6 Sol dropped content did not reproduce: it was complete (every region, every incident event, a reply to every ticket) on 58 of 60 graded documents, the same as GPT-5.6 Sol (the 20 vendor comparisons are not graded, because the right vendor is a judgment call).

With the required parts named in the prompt, all three OpenAI models were complete on 80 of 80. Sonnet 5.5 wrote the longest goal-only documents (747 words) and was also complete on every graded one, but missed 12 of 80 with named parts, each a memo or recommendation 1 to 28 words under the requested range. That gap survives the correction, but 10 of the 12 are quarterly reviews, so it says little about other documents.

Longer answers raised the cost per goal-only document from $0.0078 on GPT-6 Sol to $0.0103. That is still 59% below GPT-5.6 Sol, which lists at $4 / $20 (a promotional price through 2026-11-21); with GPT-5.6 Sol's tokens repriced at $2 / $10, GPT-6.1 Sol is 19% cheaper, from writing fewer tokens for the same length.

Is GPT-6.1 Sol slower?

GPT-6.1 Sol is slower than its predecessors and than Sonnet 5.5, most of all on long output. On the business documents it generated about 43 output tokens per second of total request time, against 100 for GPT-6 Sol, 80 for GPT-5.6 Sol and 127 for Sonnet 5.5; a goal-only document took a median of 22.4 seconds against 7.2, 13.9 and 17.2. On the short single-shot tasks the gap is smaller: 5.7 seconds at the defaults against 3.9 for Sonnet 5.5 and 5.5 for Opus 5.5. Absolute speeds depend on the host and the time of day; Artificial Analysis also has it at less than half Sonnet 5.5's speed, 64 tokens per second against 139.

What happens in a tool loop?

GPT-6.1 Sol and Sonnet 5.5 both solved all 12 runs at every setting; GPT-6.1 Sol did it for less than half the cost and took almost twice as long. In a tool loop the model asks for a tool, the calling code runs it and returns the result, and this repeats until the model answers; ours has three tools (list files, read a file, search) over a small synthetic codebase. At the defaults a run cost $0.0052 on GPT-6.1 Sol, $0.0112 on Sonnet 5.5 and $0.0332 on Opus 5.5, with median times of 12.2, 6.7 and 25.3 seconds. At max, GPT-6.1 Sol rose to $0.0086 and Sonnet 5.5 to $0.0422, with 6.4 and 14.7 tool calls per run.

Do our results match other published tests?

Our results agree with the published tests on direction. The sizes differ because our tasks are short and mostly run at the API defaults, and the public indexes run long tasks at max.

Question Published elsewhere Our measurement Verdict
Cost of GPT-6.1 Sol against Sonnet 5.5 Artificial Analysis: $0.72 against $7.60 per index task at max, about 11x 0.49x at the defaults, 0.33x at max Same direction; at max the gap is wider on longer tasks: 0.33x single-shot, 0.20x in our tool loop, 0.09x on the index
Speed Artificial Analysis: 64 against 139 output tokens per second; Vals AI: agentic tasks run 2 to 3 times longer than on GPT-6 Sol 43 against 127 tokens per second; 22.4 s against GPT-6 Sol's 7.2 s per document Agrees
Quality against Sonnet 5.5 Artificial Analysis 52 against 56; Vals 61.2% against 67.0% Sonnet 5.5 correct on 234 of 234; GPT-6.1 Sol wrong on one task 7 of 18 times Same direction; ours rests on one task
GPT-6 Sol answers too short or missing content Developer complaints, for example in the Hacker News launch thread 325 words per document against 592 for GPT-5.6 Sol; complete on 58 of 60, the same as GPT-5.6 Sol Brevity confirmed; dropped content did not reproduce
Sonnet 5.5 ignoring "answer only" No published report found 22 of 39 bare answers at its default, 12 of 39 at low Ours only

What did we observe, and which should you use?

GPT-6.1 Sol is the better value against both Sonnet 5.5 and Opus 5.5; pay for Sonnet 5.5 where latency or the last 4 to 6 points of quality matter. Four observations lead there:

  1. GPT-6.1 Sol leads both Claude models on cost for the quality delivered. Against Sonnet 5.5 it costs about half per short task, about a quarter at a matched index score, and a seventh to a tenth in the headline index results. Against Opus 5.5 it costs 0.29x per short task and an eighth to a tenth on the indexes.
  2. It follows the output format more closely. It returned what was asked on every answer-only prompt and every document with named parts.
  3. The fix for short answers cost speed. Answers are back to GPT-5.6 Sol's length, and each document takes three times as long as on GPT-6 Sol.
  4. Its errors are hard to spot. The 7 wrong answers came at low, medium and xhigh, and each was a well-formed integer.
Workload Pick Numbers
Output parsed by code: extraction, classification, single values GPT-6.1 Sol at medium or high bare answer 234 / 234; $0.0021 per task at medium
Multi-step math or logic where a wrong value is costly Sonnet 5.5, or GPT-6.1 Sol with a check on the result Sonnet 5.5 correct on 234 / 234; GPT-6.1 Sol wrong on 7 of 18 calls for one task
Chat and interactive tools, where latency matters Sonnet 5.5 at its default 3.9 s against 5.7 s per short task; 17.2 s against 22.4 s per document
Agent loops where cost per run matters more than latency GPT-6.1 Sol at its default $0.0052 per run against $0.0112; 12.2 s against 6.7 s
Documents with required sections GPT-6.1 Sol, or Sonnet 5.5 with the lengths stated generously complete on 80 / 80 against 68 / 80, the misses 1 to 28 words short
Hard, open-ended work where the top score matters Sonnet 5.5 or Opus 5.5 at max public index scores of 56 and 58 against 52, at 7 to 11 times the cost per task

With either model, validate any value that a program acts on.

FAQ

Is GPT-6.1 Sol cheaper than Claude Sonnet 5.5?
GPT-6.1 Sol costs less per task than Sonnet 5.5 at the same $2 / $10 list price: 0.49x per single-shot task at each model's default and 0.46x per tool-loop run in our tests, and $0.72 against $2.74 at a matched Artificial Analysis score of 52. Sonnet 5.5 was faster and scores higher at max.

Is GPT-6.1 Sol as good as Claude Opus 5.5?
GPT-6.1 Sol scores about 6 points below Opus 5.5 on the Artificial Analysis Intelligence Index (52 against 58) and on the Vals Index (61.2% against 67.0%), at an eighth to a tenth of the cost per task; in our tests it cost 0.29x per short task. OpenAI's own runs put it 2.2 points ahead on AutomationBench.

Can I still turn reasoning off on GPT-6.1 Sol?
GPT-6.1 Sol lists low as its lowest effort; none, which GPT-6 Sol accepted, is gone. Use low, which cost $0.0018 per single-shot task in our runs.

Related measurements: Claude Sonnet 5.5 vs Sonnet 5, the GPT-6 Astra effort ladder, and GPT-5.6 cost levers.

Measured 2026-09-30, one day after GPT-6.1 Sol's release, through a gateway to the OpenAI Responses API and the Anthropic Messages API: 234 single-shot calls on GPT-6.1 Sol (13 tasks, 3 repeats, 6 settings), 800 business-document calls (80 goal-only items on 6 model and effort settings, 80 named-part items on 4) and 36 tool-loop runs. Sonnet 5.5 and Opus 5.5 single-shot and tool-loop numbers are from 2026-09-29 on the same tasks. Public benchmark figures were read from the publishers' pages on 2026-10-01.

Top comments (0)