Claude Opus 5.5 and GPT-6.1 Sol are the two models most teams are choosing between this week. Opus 5.5 is Anthropic's everyday flagship, released on 22 September 2026. GPT-6.1 Sol is OpenAI's upgraded mid-tier model, released at DevDay on 29 September 2026. The short answer: pick GPT-6.1 Sol for high-volume and cost-sensitive work, and Claude Opus 5.5 for long agentic coding and tasks where a wrong answer costs more than the tokens. The longer answer depends on reasoning effort, and you can check it on your own prompts in about ten minutes.
Disclosure up front: I work on apimodels.app, a multi-model API gateway that serves both models. Prices for it below are its public prices on 1 October 2026. This post was drafted with AI assistance and checked against the vendors' pages.
How I compared them
- List price per 1M tokens, from the vendors' pages on 1 October 2026.
- Specs that change your code: context window, max output, reasoning controls.
- Benchmarks both models were run on in the same table. The labs mostly publish different suites, so I only use rows where both appear, and I name who ran them.
- Cost per call for a realistic agent turn, computed from the prices above.
At a glance
| Claude Opus 5.5 | GPT-6.1 Sol | |
|---|---|---|
| Released | 22 Sep 2026 (Anthropic) | 29 Sep 2026 (OpenAI DevDay) |
| List price, input / output | $4 / $20 | $2 / $10 |
| Cached input (list) | $0.20 read, $5 write | $0.10 read, $2.50 write |
| Context / max output | 1M / 128K | 1.05M / 128K |
| Knowledge cutoff | June 2026 | 30 April 2026 |
| Reasoning control | effort low to max, default medium, always thinks | reasoning_effort low to max, default medium, no none
|
| Long context surcharge | — | above 272K input: input and cache 2×, output 1.5× |
| Best for | long agent sessions, review, careful writing | volume, extraction, cost-capped agents |
| Not for | high-volume bulk classification | code that relies on reasoning_effort: none
|
The benchmarks they share
OpenAI's GPT-6.1 Sol launch included Opus 5.5 in several rows. These are OpenAI's runs and OpenAI's choice of benchmarks, so read them as a vendor's claim:
| Benchmark (OpenAI launch) | Effort | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|---|
| AutomationBench | medium | 31.7% | 29.5% |
| AutomationBench | max | 36.1% ($0.30/task) | 42.5% ($1.44/task) |
| GDP.pdf | medium | 30.0% ($0.34/task) | 25.6% ($0.80/task) |
| Terminal-Bench Science 0.1 | max | 57.0% ($5.47/task) | 63.3% ($23.21/task) |
The pattern is consistent: at medium effort GPT-6.1 Sol edges ahead, at max effort Opus 5.5 pulls ahead, and Sol costs less per task in every row. Independent aggregates point the same way on capability. Reported max-effort figures put Opus 5.5 at 58 and GPT-6.1 Sol at 52 on the Artificial Analysis Intelligence Index, with cost per index task of $5.98 and $0.72.
Anthropic's own launch numbers for Opus 5.5 (66.4% on Terminal-Bench 4.0, 57.8% on CursorBench 4.0) have no GPT-6.1 Sol row, so they don't settle the head-to-head.
What one agent turn costs
Take a coding-agent turn that re-sends 50,000 input tokens, 40,000 of them cached, and returns 2,000 output tokens:
| Claude Opus 5.5 | GPT-6.1 Sol | |
|---|---|---|
| At list price | $0.088 | $0.044 |
| On apimodels.app | $0.053 | $0.022 |
Two things move these numbers in practice. Opus 5.5 always thinks, so its reasoning tokens land in the output bill even on short answers. GPT-6.1 Sol at low effort spends far fewer reasoning tokens, which is where most of its cost-per-task advantage comes from.
Measure it on your own prompts
Benchmarks are someone else's tasks. The fastest way to decide is to send the same twenty prompts to both models and compare answers and cost. Both are OpenAI-compatible on apimodels.app, so one client covers both, and every response carries the billed amount in an x-apimodels-cost header.
Tested with Python 3.11 and openai 1.x:
import os, time
from openai import OpenAI
client = OpenAI(
api_key=os.environ["APIMODELS_API_KEY"],
base_url="https://api.apimodels.app/v1",
)
MODELS = ["claude-opus-5-5", "gpt-6.1-sol"]
def ask(model: str, prompt: str) -> dict:
t0 = time.time()
raw = client.chat.completions.with_raw_response.create(
model=model,
messages=[{"role": "user", "content": prompt}],
max_tokens=8000,
)
r = raw.parse()
return {
"model": model,
"seconds": round(time.time() - t0, 1),
"in": r.usage.prompt_tokens,
"out": r.usage.completion_tokens,
"billed_usd": raw.headers.get("x-apimodels-cost"), # what you were actually charged
"answer": r.choices[0].message.content,
}
prompts = [line for line in open("prompts.txt") if line.strip()]
for p in prompts:
for m in MODELS:
res = ask(m, p)
print(f'{res["model"]:16} {res["seconds"]:>6}s in={res["in"]:>6} out={res["out"]:>6} ${res["billed_usd"]}')
Score the answers yourself, or with a third model as judge, then divide the total billed cost by the number of answers you would ship. That number is the one that matters.
Gotchas when switching
-
reasoning_effort: "none"is not on GPT-6.1 Sol's list. OpenAI's model page lists low, medium, high, xhigh and max. Code written for GPT-6 Sol that sendsnoneshould sendlow. -
Small
max_tokenscan return an empty answer. Reasoning counts against the budget on both models. Iffinish_reasonislengthand the content is empty, raisemax_tokensor lower the effort. - Long prompts change Sol's price. Above 272K input tokens the whole request is billed at 2× input and 1.5× output.
- Opus 5.5 can't be made to cost less by turning thinking off. If a task is simple enough that thinking is waste, it is probably a Sol task.
Which one should you pick
- Classification, extraction, routing, bulk summarisation → GPT-6.1 Sol at low effort.
- Agents with a hard cost cap per task → GPT-6.1 Sol at medium effort.
- Long agentic coding sessions, code review, careful long-form writing → Claude Opus 5.5.
- The hardest reasoning tasks → run both at max effort on twenty real examples before you commit.
- You need a direct contract with OpenAI or Anthropic, such as zero data retention or a regulated-industry agreement → call the vendor directly, not a gateway.
About apimodels.app
apimodels.app is a multi-model API gateway: one key and one OpenAI-compatible endpoint for 140+ image, video, audio and language models. When an upstream channel fails, requests switch to another channel automatically, failed calls are not charged, and long-running jobs can report back through async callbacks. On price, on 1 October 2026 it lists Claude Opus 5.5 at $2.40 / $12 per 1M tokens, 40% below Anthropic's list, and GPT-6.1 Sol at $1 / $5, half of OpenAI's list, billed per token. It is not the right choice if you only ever call one model and already have the vendor's SDK and contract in place.
Which prompts in your own workload flip between the two models? I'd like to hear where your results disagree with the launch tables.
Top comments (0)