Claude Opus 5 is out, and Artificial Analysis — who supported Anthropic's pre-release evaluation — just dropped their full benchmark breakdown. The headline: new top model for agentic knowledge work, and cheaper per task than Fable 5.
That combination doesn't come along often at the frontier.
"Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59)"
What actually changed
- New agentic leader: 1861 Elo on GDPval-AA v2 — more than 100 points ahead of both Fable 5 and GPT-5.6 Sol. On AA-Briefcase (agentic knowledge work), it's +146 Elo over Fable 5.
- Joint first on coding: Opus 5 (xhigh) with Claude Code tops the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA.
- 89% on Terminal-Bench v2.1: Roughly in line with the current terminal leader, GPT-5.6 Sol.
- Cost per task: $2.03 at max effort — vs Fable 5's $2.75. That's 26% less for equivalent or better intelligence on agentic benchmarks.
- 1M token context window (same as Opus 4.8), 5 effort settings (low → max), and server-side fallback support.
- Pricing: $5/$25 per million input/output tokens — same rate as previous Opus launches.
The cost-intelligence shift
For agentic workloads — the things most teams are actually building on right now — Opus 5 doesn't just match Fable 5. It beats it, and charges less to do it.
Fable 5 was the "throw more at it" option. Opus 5 reframes the trade-off: better agentic outcomes and a lower bill. At mid-tier effort settings (high, xhigh), it can outperform both Opus 4.8 and Sonnet 5 on a cost-per-task basis. That's a lot of headroom to play with before you're even at max effort.
The caveat worth flagging: factual knowledge still lags. Opus 5 improved +7 points on AA-Omniscience over Opus 4.8, but its hallucination rate climbed 14 points to 50% — it guesses more confidently when uncertain. For retrieval-heavy or factual precision tasks, Fable 5 still holds the edge.
What to do
-
Running agentic pipelines? Opus 5 is the new default to benchmark. Start at
highorxhigheffort before committing tomax. - On Claude Code? You're already getting the benefit — joint first on the Coding Agent Index.
- Cost-sensitive on frontier models? Max-effort Opus 5 undercuts Fable 5 by 26%. Re-run your cost model — this changes the calculus.
- Factual knowledge tasks? Hold off. A 50% hallucination rate is a hard limit for anything knowledge-intensive. Fable 5 still wins there.
Full benchmark breakdown: Artificial Analysis — Claude Opus 5
✏️ Drafted with KewBot (AI), edited and approved by Drew.
Top comments (0)