DEV Community

ServBay
ServBay

Posted on

What the Haiku 5.5 vs GPT-6.1 Sol numbers changed about how I route my own tasks

Haiku 5.5 and GPT-6.1 Sol shipped three weeks apart, and the headline number makes the choice look settled. For prompts under 100K tokens, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. GPT-6.1 Sol costs $2 and $10 for the same units, 20 times more.

That gap is real. It also turned out to be the least useful part of the comparison. I went through the pricing tiers, the public benchmark tables, the migration guides and the community threads, mostly to answer a narrower question. If most of my own traffic is classification, field extraction and document summarization, how much does the model choice actually change, and what does it cost to keep both endpoints alive.

The price gap sits inside a specific range

Specs at a glance

Item Claude Haiku 5.5 GPT-6.1 Sol
Release date 2026-10-07 2026-09-29
Context window 1M tokens 1.05M tokens
Max output 128K tokens 128K tokens

Claude Haiku 5.5 is the model identifier claude-haiku-5-5. GPT-6.1 Sol uses gpt-6.1-sol. Both vendors default the model to medium reasoning effort.

Published prices per million tokens

Tier Input Cache read Cache write Output
Haiku 5.5, prompt ≤100K 0.10 0.01 0.125 0.50
Haiku 5.5, prompt >100K 0.50 0.05 0.625 2.50
Sol 6.1, prompt ≤272K 2.00 0.10 2.50 10.00
Sol 6.1, prompt >272K 4.00 0.20 5.00 15.00

Tier thresholds apply to the whole request

Both vendors price the tier against the entire request. When a Haiku prompt moves from 100,000 tokens to 100,001 tokens, the higher rate applies to the whole request, not just to the extra token. Long system prompts, tool definitions and accumulated conversation history all count toward that line. Sol's threshold sits much further out, at 272K tokens.

There is a second thing that affects the bill. Anthropic notes that Haiku 5.5 uses a new tokenizer, and the same plain text produces roughly 30% more tokens than it does with Haiku 4.5. Budgets built on old token counts need to be recounted with the new model.

What a single request costs at different sizes

Fixed token counts, both models

The table below uses fixed input and output token counts plus the pricing above. It covers input and output only, with no tool fees, retries or taxes.

Input / Output Haiku 5.5 Sol 6.1 Sol as a multiple of Haiku
5K / 1K $0.001 $0.020 20x
100K / 2K $0.011 $0.220 20x
150K / 5K $0.0875 $0.350 4x
300K / 5K $0.1625 $1.275 ~7.8x
900K / 10K $0.475 $3.750 ~7.9x

Run a request with 5K input and 1K output 100,000 times and Haiku 5.5 comes to about $100, while Sol comes to about $2,000.

Where the two prices converge

Once the prompt lands between 100K and 272K, the gap narrows to roughly 4x, because Haiku's rate has already moved up a tier while Sol is still on its first tier. Above 272K, Sol's own rate rises as well, and the multiple settles around 7.8x to 7.9x.

This is still a comparison of token unit prices. Two models consume different numbers of tokens to finish the same task, and the FrontierCode data further down shows how far apart that can get.

Where each model actually leads

A note on provenance before the tables. Most Haiku 5.5 scores come from Anthropic's release materials and third-party evaluation by Artificial Analysis, while Sol's scores come from AA and various public leaderboards. Reasoning effort, harness and scoring method differ from test to test, so none of this is a like-for-like run under identical conditions.

Knowledge work

Test (Elo) Haiku 5.5 (max) Sol 6.1 (max) Haiku 5.5 (medium)
GDPval-AA v2.1 1620 1575 1277
AA-Briefcase v1.1 1578 1564 1372

At the highest reasoning effort Haiku 5.5 leads by 45 and 14 Elo points. Elo is a relative score, and 14 points is not enough to show a meaningful difference in practice.

At the default medium effort, those scores fall to 1277 and 1372. AA reports a weighted task cost of about $0.047 at medium and about $0.213 at max, so max costs roughly 4.5 times medium. The 1620 figure belongs to a high-cost configuration, not to how the model performs in default usage.

Terminal coding

Test Haiku 5.5 Sol 6.1
Terminal-Bench 4.0 (public leaderboard and Anthropic's own testing) 39.2% 58.2%
Terminal-Bench 4.0 (AA, max) 32.8% 56.1%

Both datasets point the same way, with Sol ahead by 19 to 23 percentage points. The Haiku 39.2% comes from Anthropic running the model in Claude Code, while the Sol 58.2% comes from the public Codex leaderboard, so the harnesses differ. Anthropic itself states that Sonnet 5.5 and Opus 5.5 suit complex agentic coding better, and that Haiku 5.5 fits narrowly scoped tasks.

Repository tasks on FrontierCode 1.1 Main

Item Haiku 5.5 Sol 6.1
Score 46.4% 50.2%
Pass rate 51.6% 55.8%
Best effort max (Claude Code) medium (Codex)
Average cost per run $1.33 $0.36
Average output tokens 181.4K 11.4K

This is the table I keep coming back to. Haiku 5.5 has the lower unit price, but at max effort it emits roughly 180K tokens, which puts its cost per run at about 3.7 times Sol running at medium. A cheap token does not automatically produce a cheap finished task.

Computer use on OSWorld 2.1

Metric Haiku 5.5 Sol 6.1
Partial score 72.4% 76.6%
Full pass rate 37.1% 39.8%

These come from Anthropic, measured on the same batch of 82 tasks with each task run 5 times, so they are vendor self-reported. The gap is small, and neither model reaches a 40% full pass rate. A high partial score does not mean the task was completed, and offline tasks say nothing about logins, network errors or content that changes on live pages.

Composite index, reasoning and generation speed

Item Haiku 5.5 Sol 6.1
Intelligence Index v4.3.2 43.4 51.8
Humanity's Last Exam (no tools) 44.4% 52.9%
GDP.pdf (all passed) 20.8% 31.0%
AA-LCR v1.1 (long context) 82.7% 83.0%
Output speed (tokens/sec) 241.9 ~55
Weighted task cost $0.213 $0.724

Sol leads on the composite index and on the reasoning tests. Long-context retrieval is nearly even. Haiku generates about 4.4 times faster and its weighted task cost is about 29% of Sol's.

Generation speed is only one part of latency. Thinking time, tool execution and retries all shape how long a user actually waits, so the honest way to settle this is to measure the target tasks at the effort level you plan to run.

Migration cost before switching

GPT-6.1 Sol requires the Responses API

OpenAI removed tool calling from the standard Chat Completions API for Sol. External function calls, web search and structured data operations all go through the Responses API. Parameter accuracy improves for multi-stage function calls, but any application built on Chat Completions has to change its API layer at the same time as the model.

Moving from Haiku 4.5 to Haiku 5.5

That migration touches several areas. Anthropic's guide adjusts the thinking configuration, removes assistant prefill, restricts some sampling parameters and replaces the toolset for computer use. The safety classifier can return a refusal outright, with no server-side fallback. Haiku 5.5 does not yet support Priority Tier.

For an application running stably, switching models for a marginal gain may cost more in retesting than the token spend it saves. That calculation belongs next to the pricing table, not after it.

Where a cc-switch alternative runs out

Task difficulty varies more than the model choice does

Looking at these numbers, the practical problem is not picking a winner. It is that task difficulty inside one codebase varies enormously. Cleaning meeting notes, triaging tickets and extracting fields with the most expensive model is waste. Refactoring a core module or chasing a dependency problem with the cheapest model fails again and again.

The community discussion reflects the same split. One view holds that not every task needs the most expensive model, and that a limited budget is poorly spent on top-tier models running simple work. Another view holds that asking a person to judge each task by hand is unrealistic, and that dropping to a lower tier produces a visible drop in quality, so automatic routing is the only workable answer.

Both positions hold up. A more practical setup is to fix the division of labor by project and scenario first, letting lightweight tasks default to a low-cost model and important tasks go to a high-performance one, then put a single entry point in charge of keys, channels and switching.

The switching decision is not the routing decision

Tools built around switching give a way to change which provider is active. That solves one problem and leaves another one open, because the decision of where each request should go is a per-project, per-scenario question rather than a global toggle. Once a project has three tasks with three different cost profiles, the configuration has to carry that logic. This is the point where a cc-switch alternative stops covering the whole job.

Managing model sources and keys behind one entry point

This is the part where I use ServBay, an AI development management tool with an AI Gateway built in, so that several model sources sit behind one entry point.

Sources, scheduling and priority

The gateway accepts official AI APIs from the major vendors, subscription accounts, and relay providers. Traffic can be distributed automatically, channels can be hot switched, and a channel that becomes unavailable fails over to a backup. Usage statistics come out of the same view. Channel priority can be set, so lower-cost channels take the first pass and high-performance channels fill in behind them.

Model mapping and virtual keys

Model mapping lets a model name hardcoded in an application point at a different model underneath. Mapping claude-opus-5 to glm-5.2 is one example. In the same way, a project whose tasks are mostly summarization and classification can map the name it calls to Claude Haiku 5.5 without touching application code.

Local virtual keys are created per project and assigned separately. Project A's key can map to Haiku 5.5, Project B's key can map to GPT-6.1 Sol, and each project's usage can be tracked on its own.

Protocol conversion

An upper-layer application can speak the OpenAI, Anthropic or Gemini protocol, and the gateway converts it into whatever the underlying channel requires. A tool that only supports the OpenAI protocol can still reach Anthropic models, which removes the need to adapt each integration individually.

A configuration that fits these two models

Everyday summarization, classification and subagent work point at Haiku 5.5. Complex repository and terminal work points at GPT-6.1 Sol. Since Haiku 5.5 raises its unit price above 100K tokens, that project gets an extra backup channel. When a model version changes, the mapping changes and the code in each project stays as it is.

How I would choose between them

By use case

Use case Recommendation Basis
Classification, routing, fixed-format extraction Try Haiku 5.5 first Lower unit price, faster
Summarization and document processing under 100K Try Haiku 5.5 first Knowledge-work score is close to Sol
Difficult terminal and repo tasks GPT-6.1 Sol Clear lead on Terminal-Bench
Screenshot-driven computer use Test both Small gap, and the full pass rate is low for both
Repeated long prompts from 100K to 272K Compare Luna and Sol Haiku's unit price goes up 5x in this range

Subscription plans compare differently

Everything above is token-based API pricing, and subscriptions do not follow the same curve. SemiAnalysis tested several subscriptions with agent workloads and converted the quotas into API-equivalent prices. Its conclusion was that a Claude subscription running Opus 5.5 is worth roughly 5x an OpenAI plan of the same tier running GPT-6.1 Sol. The direct reason is that OpenAI halved the token quotas for models at every tier of its $200 plan.

About 96.6% of the tokens in that test came from cache reads, which matches how agents reuse context repeatedly. Ordinary chat users may not see the same converted value. SemiAnalysis also estimates that subscriptions make up only about 10% of Anthropic's revenue while potentially consuming more than 40% of its inference compute. How heavily subscription quotas are subsidized shifts with each vendor's strategy, so leaning on a single subscription over the long term carries uncertainty.

Questions that came up while working through this

Which model is better overall?

There is no single answer. Haiku 5.5 is cheaper and faster, and its knowledge work scores are close to Sol's. Sol leads clearly in terminal coding and on the composite reasoning index. The task type decides.

How much does Claude Haiku 5.5 cost?

For prompts up to 100K tokens, $0.10 per million input tokens and $0.50 per million output tokens. Above 100K, those rates become $0.50 and $2.50.

How much does GPT-6.1 Sol cost?

For prompts up to 272K tokens, $2 per million input tokens and $10 per million output tokens. Above 272K, those rates become $4 and $15.

Can Haiku 5.5 replace Sol for coding?

For simple, well-scoped editing tasks, yes. For complex terminal and repository tasks, the available data supports Sol. Anthropic recommends Sonnet 5.5 or Opus 5.5 for complex agentic coding.

Do the two models share the same context window?

They are close. Haiku 5.5 has 1 million tokens and Sol 6.1 has 1.05 million tokens, and both cap output at 128K tokens. Tokenizers differ between vendors, so the same text converts to different token counts.

Can one project use several models at once?

Yes. With multiple channels and model mappings configured through an AI gateway, the application connects to one entry point and the gateway decides which model is called underneath.

Closing

The three differences that matter are a 20x price gap, Sol's clear lead on terminal coding, and a near tie on knowledge work and computer operation. Total cost to finish a task also depends on reasoning effort and token consumption, so unit price alone makes for a poor decision rule.

Splitting work by task type holds up better than settling on a single model. Teams that have to manage several models and channels at the same time can route them through an AI gateway such as ServBay, with keys per project, priorities and model mappings, which moves the model decision into the configuration layer while application code stays unchanged.

Top comments (0)