DEV Community

Cover image for Opus 5.5 Shipped With a 20% Price Cut. The Money Moved.
Max Quimby
Max Quimby

Posted on Originally published at computeleap.com

Opus 5.5 Shipped With a 20% Price Cut. The Money Moved.

Anthropic shipped Claude Opus 5.5 on September 22, 2026 -- model ID claude-opus-5-5 -- with input tokens at $4/MTok and output at $20/MTok, a flat 20% cut from Opus 5's $5/$25 pricing. Cache reads dropped 60%, from $0.50 to $0.20 per million tokens. On typical agent workloads at default settings, the total cost reduction lands around 40%.

📖 Read the full version with charts and embedded sources on ComputeLeap →

That is the product news. Here is the story nobody is writing about.

@AnthropicAI — Claude Opus 5.5 is available today.

View original post on X →

Within 24 hours of the launch, three separate Polymarket prediction markets shifted double digits toward Anthropic: LiveBench Mathematics surged 22.5 percentage points to 64%, Text Arena Math climbed 12.6 points to 58%, and LiveBench Coding hit 77% (up 17.3% on the week). While the crowd's attention was on Elon Musk's Grok 4.7 tweet -- 21 million views -- and OpenAI's coordinated GPT-6 Astra demo blitz, the betting money was quietly crowning Anthropic the benchmark king.

When hype and money point at different labs, follow the money.

The Pricing Math That Matters

The headline is 20% cheaper per token. The reality for builders is more aggressive than that.

Opus 5.5's default effort level is medium, one notch below Opus 5's high default. This means the model thinks less by default -- burning fewer thinking tokens per request. Combined with 30%+ faster output generation, the real-world cost of running an agent workflow drops closer to 40% when you factor in reduced compute time and lower thinking overhead.

Here is the full pricing comparison:

Metric Opus 5 Opus 5.5 Change
Input tokens $5/MTok $4/MTok -20%
Output tokens $25/MTok $20/MTok -20%
Cache reads $0.50/MTok $0.20/MTok -60%
Fast mode $10/$50 $8/$40 -20%
Default effort high medium Lower token burn
Output speed Baseline 30%+ faster Less compute time

The cache read reduction is the sleeper hit. For long-running agent sessions -- the workload Anthropic is clearly targeting -- cached context dominates the bill. A 60% cut on cache reads makes multi-hour coding sessions dramatically cheaper.

@VaibhavSisinty — Anthropic dropped Claude Opus 5.5 and the price-to-performance ratio is the best they have ever shipped. Fable 5.1 level performance. 40% cheaper than Opus 5. 30% faster output. Cache reads 60% cheaper.

View original post on X →

For API users migrating from Opus 5: The default effort dropped from high to medium. If your application relies on deep reasoning, explicitly set output_config: {effort: "high"} or higher. The model also rejects thinking: {type: "disabled"} and tool_choice: {type: "any"} -- both return 400 errors. Test your integration before routing production traffic.

The Benchmark Picture

Anthropic is not being shy about the numbers. From the official announcement:

  • Terminal-Bench 4.0: 66.4% (vs. Opus 5's 52.3%) -- a 27% relative improvement in agentic coding
  • OSWorld 2.0 (computer use): 81.8% (vs. Opus 5's 74.0%)
  • GDPval-AA v2.1: 1846 Elo, surpassing Fable 5.1's 1735
  • Output speed: 30%+ faster than Opus 5

The claim that keeps surfacing from reviewers: Opus 5.5 performs at Fable 5.1 level for most tasks while costing 60% less ($4/$20 vs. $10/$50). If that holds, it collapses the price-performance gap between Anthropic's own model tiers.

One concrete data point from the announcement: Opus 5.5 completed a 680,000-line code migration in under one day and improved web app load times in 39 of 40 attempts. These are the kinds of agentic workloads where the cache pricing cut compounds.

The Prediction Market Divergence

This is where the story gets interesting.

On launch day, four model releases competed for attention: Jev (TypeSafe's classification-only model), GPT-6 Astra, Opus 5.5, and Grok 4.7. The attention economy crowned Grok -- Elon's tweet pulled 21 million views. But the Polymarket prediction markets told a completely different story.

Three AI benchmark markets moved double digits toward Anthropic in a single 24-hour window:

  1. LiveBench Mathematics (end of Oct): Anthropic surged to 64%, up 22.5 percentage points. OpenAI sits at 35%.
  2. Text Arena Math (end of Nov): Anthropic reached 58%, up 12.6 points. Google at 16%, OpenAI at 14%.
  3. LiveBench Coding (end of Sept): Anthropic at 77%, up 17.3% on the week.

This is not normal market behavior. A double-digit swing across three separate markets in 24 hours means traders are processing information the attention economy has not caught up with yet.

We have documented this pattern before. When we tracked Polymarket's AI markets earlier this year, Anthropic was sitting at 92% on our prediction market telemetry. The trend line has not broken.

The contrarian tell: On the Code Arena WebDev market, OpenAI leads at 70% while Anthropic sits at 30%. But on LiveBench Coding, Anthropic is at 77%. The prediction markets cannot agree on which benchmark matters -- and that disagreement is itself a signal. Different benchmarks measure different things, and traders are placing bets based on which benchmark they think will prove decisive.

"Pacing the Frontier" -- Branding Judo or Genuine Restraint?

Opus 5.5 is Anthropic's first model release since CEO Dario Amodei published "We Must Pace the Frontier" -- an essay arguing that AI labs should deliberately slow capability advancement to keep safety ahead of capabilities. The Opus 5.5 announcement's opening line references this framing directly.

The Hacker News community, with 1,313 points and 861 comments, immediately identified the tension. The top comment thread called it "branding judo: release a better model while claiming to be the responsible one." Others pointed out that "pacing the frontier" is "so open for interpretation that it is meaningless" -- impossible to verify whether Anthropic actually practices restraint versus just claims it.

Hacker News discussion on Claude Opus 5.5 with 1313 points and 861 comments, debating the pacing the frontier framing

View on Hacker News →

Naval Ravikant weighed in on X: "The best way to pace the frontier is to hold the labs fully liable for the behavior of their models." This reframes the debate entirely -- from voluntary restraint to legal accountability.

The ZeroHedge analysis puts the contradiction in financial terms: flagship AI token prices have collapsed 73% over 13 months (from Opus 4.1's $15/$75 in August 2025 to Opus 5.5's $4/$20 today). OpenAI launched GPT-6 Sol and Luna within hours of the Opus 5.5 announcement, further halving prices. "Pacing, it turns out, is a team sport," the piece notes -- while both labs race on price and capability simultaneously.

@QuintinPope5 — Anthropic should prove their commitment to pacing the frontier by calling it Opus 5.0.5

View original post on X →

The contrarian take: Shipping a model that matches Fable 5.1's performance at 60% of the cost, just 10 days after calling for the industry to slow down, is not "pacing." It is accelerating while holding a sign that says "pace." Whether the safety testing genuinely justifies the release speed is a question the benchmarks cannot answer -- it requires trusting the external evaluation process (METR, Frontier Design) that Anthropic cites in its system card. The community is right to be skeptical.

The Safety Card

Anthropic is leaning into the safety narrative with specific claims:

  • External evaluation by METR and Frontier Design before release
  • "Best-performing model" on Anthropic's automated behavioral audit
  • 85% fewer boundary circumvention attempts than Opus 5
  • Prompt injection resistance matching or exceeding Opus 5
  • Cybersecurity tasks routed to Opus 4.8 for non-verified users
  • Biology work restricted to vetted organizations via the Life Sciences Verification Program
  • Preserved thinking (anti-distillation) enabled by default

The 230-page system card calls Opus 5.5's cyber capabilities "the strongest of any model we have released" -- which is precisely why they restrict those capabilities to verified users. This is the specific, verifiable kind of safety claim that matters more than vague "pacing" rhetoric.

@claudeai — Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

View original post on X →

What This Means for Builders

Migration Checklist

If you are running production workloads on Opus 5, here is what breaks:

  1. Thinking cannot be disabled. thinking: {type: "disabled"} returns a 400 at every effort level. Use output_config: {effort: "low"} instead to reduce thinking overhead.

  2. Forced tool_choice is gone. tool_choice: {type: "any"} and {type: "tool", name: "..."} both return 400. Use {type: "auto"} with strict: true on the tool definition and steer from the prompt.

  3. Computer use requires the new toolset. computer_20251124 returns a 400. Use computer_toolset_20260801.

  4. Default effort is medium, not high. If your workload needs deep reasoning, explicitly set effort to high or xhigh. The quality difference is real.

  5. Preserved thinking is enforced. Thinking blocks are tied to the model and the conversation. A fallback to Opus 5 runs without them. Accounts created on or after August 31, 2026 get stricter enforcement on the history-editing check.

For a complete migration guide, see our Claude Code and Opus development guide.

The Cost Optimization Play

For agent-heavy workloads, Opus 5.5 is the obvious upgrade:

  • Cache-heavy sessions save 60% on reads alone
  • Agent loops with many tool calls benefit from 30%+ faster output
  • Budget-conscious teams can run at medium effort (the new default) for routine work and escalate to high only when the task demands it
  • Fast mode at $8/$40 per MTok is 20% cheaper than Opus 5's fast mode

The Fable 5.1 comparison is the real strategic move. If Opus 5.5 genuinely matches Fable-level performance at $4/$20 versus Fable's $10/$50, there is limited reason to pay for Fable outside of tasks that specifically require its capabilities. Anthropic is cannibalizing its own premium tier to hold the value-for-money position -- a pattern we analyzed when GPT-5.6 won the headlines but the money bet on Anthropic.

The Competitive Landscape

The four-model pile-up on September 22 produced a natural experiment in market attention versus market conviction:

Model Attention Money
Grok 4.7 21M views on Elon's tweet Five reviewers: "a shrug, not a leap"
GPT-6 Astra Coordinated demo blitz (Figma, Notion, Box) LiveBench Coding: OpenAI at 23% vs Anthropic 77%
Opus 5.5 Moderate coverage Three Polymarket markets: double-digit swings
Jev (TypeSafe) AI YouTube's "actual story of the week" No prediction market yet

The inverse correlation between launch volume and market substance is the meta-story. Grok got the most attention and the thinnest substance. The Chinese open-weights ecosystem (Xiaomi's MiMo-V2.6, Alibaba's Qwen 4) got the broadest cross-source convergence at six sources. And Opus 5.5's prediction market movement was the sharpest single-day divergence -- a replay of the pattern we documented during the 48-hour frontier release war.

The Bigger Picture

The 73% price collapse in frontier model tokens over 13 months is the structural story underneath all four launches. Anthropic is executing what ZeroHedge calls a "Jevons bet" -- cutting unit prices to drive volume growth that sustains revenue despite deflation. With 64% of revenue reportedly flowing through Vercel's gateway, Anthropic needs developer adoption to hold that position.

Meanwhile, open-weight models from Chinese labs now capture 56% of token traffic but only 14% of spending. The commercial moat for proprietary models is not the model itself -- it is the pricing, the tooling, the safety compliance, and the enterprise distribution. Opus 5.5's price cut is a move in that game, not just a product launch.

Sonnet 5.5 and Haiku 5.5 are coming in the next few weeks. If the pricing pattern holds, every tier gets cheaper while the capability floor rises. For builders, the practical implication is straightforward: the cost of intelligence is dropping faster than most financial models predicted. Build accordingly.


Follow the prediction markets, not the headlines. The money usually knows something the crowd does not -- until it does not.


Originally published at ComputeLeap

Top comments (1)

Collapse
 
kaziava profile image
Hardcore Engineer

the sep 19 / sep 23 pair reads as a clean natural experiment: capability news
didn't move the money, capability plus a price cut did. we see the same
asymmetry one level down in production LLMOps — benchmark headlines never
change our routing, price changes always tempt it, and neither should move
anything without an eval gate.

our rule since the february embedder incident: any change to a slot, embedder
or generator, is a swap event, and swap events re-run the golden set before
traffic moves. a 20% cut makes the swap tempting; the gate makes it safe. and
the cut that actually matters for agentic bills isn't the sticker 20%: Opus
5.5 cache reads at $0.20/MTok against Opus 5's $0.50 is a 60% cut on the line
item that dominates long agent loops, since agents spend most of their tokens
re-reading context they already paid for — a trace posted here last week put
the share near total. on cache-heavy workloads the real discount is the one
nobody puts in the headline.

so we tag every eval verdict with model + date, and "Opus 5.5 at launch" and
"Opus 5.5 in three months" are different vendors to our CI. price moves your
market; the golden set prices what the model is worth on our workload. two
markets, same discipline: measure, don't trust the headline.

curious whether you've seen anyone track cost per verified answer — cost per
golden-set pass — across generations. feels like the metric that connects your
market lens to the eval lens, and i haven't seen it written up anywhere.