DEV Community

Super Lewis
Super Lewis

Posted on

Claude Opus 5.5 vs GPT-6.1 Sol: which one to call, and how to measure it on your own prompts

Claude Opus 5.5 and GPT-6.1 Sol are the two models most teams are choosing between this week. Opus 5.5 is Anthropic's everyday flagship, released on 22 September 2026. GPT-6.1 Sol is OpenAI's upgraded mid-tier model, released at DevDay on 29 September 2026. The short answer: pick GPT-6.1 Sol for high-volume and cost-sensitive work, and Claude Opus 5.5 for long agentic coding and tasks where a wrong answer costs more than the tokens. The longer answer depends on reasoning effort, and you can check it on your own prompts in about ten minutes.

Disclosure up front: I work on apimodels.app, a multi-model API gateway that serves both models. Prices for it below are its public prices on 1 October 2026. This post was drafted with AI assistance and checked against the vendors' pages.

How I compared them

  • List price per 1M tokens, from the vendors' pages on 1 October 2026.
  • Specs that change your code: context window, max output, reasoning controls.
  • Benchmarks both models were run on in the same table. The labs mostly publish different suites, so I only use rows where both appear, and I name who ran them.
  • Cost per call for a realistic agent turn, computed from the prices above.

At a glance

Claude Opus 5.5 GPT-6.1 Sol
Released 22 Sep 2026 (Anthropic) 29 Sep 2026 (OpenAI DevDay)
List price, input / output $4 / $20 $2 / $10
Cached input (list) $0.20 read, $5 write $0.10 read, $2.50 write
Context / max output 1M / 128K 1.05M / 128K
Knowledge cutoff June 2026 30 April 2026
Reasoning control effort low to max, default medium, always thinks reasoning_effort low to max, default medium, no none
Long context surcharge — above 272K input: input and cache 2×, output 1.5×
Best for long agent sessions, review, careful writing volume, extraction, cost-capped agents
Not for high-volume bulk classification code that relies on reasoning_effort: none

The benchmarks they share

OpenAI's GPT-6.1 Sol launch included Opus 5.5 in several rows. These are OpenAI's runs and OpenAI's choice of benchmarks, so read them as a vendor's claim:

Benchmark (OpenAI launch) Effort GPT-6.1 Sol Claude Opus 5.5
AutomationBench medium 31.7% 29.5%
AutomationBench max 36.1% ($0.30/task) 42.5% ($1.44/task)
GDP.pdf medium 30.0% ($0.34/task) 25.6% ($0.80/task)
Terminal-Bench Science 0.1 max 57.0% ($5.47/task) 63.3% ($23.21/task)

The pattern is consistent: at medium effort GPT-6.1 Sol edges ahead, at max effort Opus 5.5 pulls ahead, and Sol costs less per task in every row. Independent aggregates point the same way on capability. Reported max-effort figures put Opus 5.5 at 58 and GPT-6.1 Sol at 52 on the Artificial Analysis Intelligence Index, with cost per index task of $5.98 and $0.72.

Anthropic's own launch numbers for Opus 5.5 (66.4% on Terminal-Bench 4.0, 57.8% on CursorBench 4.0) have no GPT-6.1 Sol row, so they don't settle the head-to-head.

What one agent turn costs

Take a coding-agent turn that re-sends 50,000 input tokens, 40,000 of them cached, and returns 2,000 output tokens:

Claude Opus 5.5 GPT-6.1 Sol
At list price $0.088 $0.044
On apimodels.app $0.053 $0.022

Two things move these numbers in practice. Opus 5.5 always thinks, so its reasoning tokens land in the output bill even on short answers. GPT-6.1 Sol at low effort spends far fewer reasoning tokens, which is where most of its cost-per-task advantage comes from.

Measure it on your own prompts

Benchmarks are someone else's tasks. The fastest way to decide is to send the same twenty prompts to both models and compare answers and cost. Both are OpenAI-compatible on apimodels.app, so one client covers both, and every response carries the billed amount in an x-apimodels-cost header.

Tested with Python 3.11 and openai 1.x:

import os, time
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["APIMODELS_API_KEY"],
    base_url="https://api.apimodels.app/v1",
)

MODELS = ["claude-opus-5-5", "gpt-6.1-sol"]

def ask(model: str, prompt: str) -> dict:
    t0 = time.time()
    raw = client.chat.completions.with_raw_response.create(
        model=model,
        messages=[{"role": "user", "content": prompt}],
        max_tokens=8000,
    )
    r = raw.parse()
    return {
        "model": model,
        "seconds": round(time.time() - t0, 1),
        "in": r.usage.prompt_tokens,
        "out": r.usage.completion_tokens,
        "billed_usd": raw.headers.get("x-apimodels-cost"),  # what you were actually charged
        "answer": r.choices[0].message.content,
    }

prompts = [line for line in open("prompts.txt") if line.strip()]
for p in prompts:
    for m in MODELS:
        res = ask(m, p)
        print(f'{res["model"]:16} {res["seconds"]:>6}s  in={res["in"]:>6} out={res["out"]:>6}  ${res["billed_usd"]}')
Enter fullscreen mode Exit fullscreen mode

Score the answers yourself, or with a third model as judge, then divide the total billed cost by the number of answers you would ship. That number is the one that matters.

Gotchas when switching

  • reasoning_effort: "none" is not on GPT-6.1 Sol's list. OpenAI's model page lists low, medium, high, xhigh and max. Code written for GPT-6 Sol that sends none should send low.
  • Small max_tokens can return an empty answer. Reasoning counts against the budget on both models. If finish_reason is length and the content is empty, raise max_tokens or lower the effort.
  • Long prompts change Sol's price. Above 272K input tokens the whole request is billed at 2× input and 1.5× output.
  • Opus 5.5 can't be made to cost less by turning thinking off. If a task is simple enough that thinking is waste, it is probably a Sol task.

Which one should you pick

  • Classification, extraction, routing, bulk summarisation → GPT-6.1 Sol at low effort.
  • Agents with a hard cost cap per task → GPT-6.1 Sol at medium effort.
  • Long agentic coding sessions, code review, careful long-form writing → Claude Opus 5.5.
  • The hardest reasoning tasks → run both at max effort on twenty real examples before you commit.
  • You need a direct contract with OpenAI or Anthropic, such as zero data retention or a regulated-industry agreement → call the vendor directly, not a gateway.

About apimodels.app

apimodels.app is a multi-model API gateway: one key and one OpenAI-compatible endpoint for 140+ image, video, audio and language models. When an upstream channel fails, requests switch to another channel automatically, failed calls are not charged, and long-running jobs can report back through async callbacks. On price, on 1 October 2026 it lists Claude Opus 5.5 at $2.40 / $12 per 1M tokens, 40% below Anthropic's list, and GPT-6.1 Sol at $1 / $5, half of OpenAI's list, billed per token. It is not the right choice if you only ever call one model and already have the vendor's SDK and contract in place.

Which prompts in your own workload flip between the two models? I'd like to hear where your results disagree with the launch tables.

Top comments (0)