Key Takeaways
Claude Opus 5 pricing is $5 per million input tokens and $25 per million output tokens. Identical to Opus 4.8, and exactly half of Claude Fable 5.
Anthropic's own announcement argues cost per completed task, not price per token. Those are different numbers and only one of them is on your invoice.
Run a 2,000 conversation a month support agent and the Fable to Opus 5 move saves about $55. Dropping that same agent to Haiku 4.5 saves another $44.
Caching moved my numbers harder than the launch did. A cache hit on Opus 5 bills at $0.50 per million, a tenth of the base input rate.
The launch item nobody led with is automatic fallbacks. For a production agent, a safety refusal used to be an outage. Now it reroutes.
If you run a general business workload, this is a reason to audit your model tier. It is not a reason to upgrade it.
Anthropic shipped Claude Opus 5 on July 24, 2026, and Claude Opus 5 pricing came in at $5 per million input tokens and $25 per million output tokens. Half of Fable 5. Four outlets covered it within hours and only two of them put the cost story in the headline at all. I spent Saturday morning recalculating what my clients actually pay.
The bill barely moved.
Anthropic's launch post leads with "half the price" of Fable 5, not with a benchmark score. That framing is the actual news.
What is Claude Opus 5 pricing, exactly?
Claude Opus 5 pricing sits at $5 per million input tokens and $25 per million output tokens on every platform, which is the same rate Opus 4.8 has carried since May, and it is precisely half of what Claude Fable 5 charges at $10 and $50 per million. Nothing got cheaper. A cheaper model got better.
That distinction matters more than the headlines suggested. Anthropic did not cut the price of Opus. It moved Fable-class capability down into the Opus price band and left the band alone.
| Model | Input per million | Output per million | Cache hit per million |
|---|---|---|---|
| Claude Fable 5 | $10 | $50 | $1 |
| Claude Opus 5 | $5 | $25 | $0.50 |
| Claude Sonnet 5 (through Aug 31, 2026) | $2 | $10 | $0.20 |
| Claude Sonnet 5 (from Sep 1, 2026) | $3 | $15 | $0.30 |
| Claude Haiku 4.5 | $1 | $5 | $0.10 |
The column most people skip is "Cache Hits and Refreshes". On Opus 5 that column reads $0.50, one tenth of the base input rate.
There's a Fast mode too, running around 2.5 times the default speed at twice the base rate, per Anthropic's launch post. Most business workloads don't need it. A support agent that answers in 4 seconds instead of 10 doesn't convert better.
Why did Anthropic cut its frontier price now?
Anthropic didn't frame this as a price cut at all, and the framing it did choose says a lot about where the market pressure is coming from, because every headline number in the announcement is a ratio of performance to cost rather than a raw benchmark score. Cost per task. Not capability per task.
The numbers back that up. On Frontier-Bench v0.1 Opus 5 more than doubles Opus 4.8's score at a lower cost per task. On CursorBench 3.2 at max effort it lands within 0.5% of Fable 5's peak while costing half as much per task. On OSWorld 2.0 it beats Fable 5's best result at just over a third of the cost.
Ars Technica read it the same way and said so in the headline: this is about token efficiency, not a capability leap. Samuel Axon also pointed at the competitive squeeze, noting that Kimi K3 comes in at $15 per million output tokens at similar performance, and that Cursor and Meta are building routers that pick a small model when a small model will do.
Ars framed the cost story as token efficiency rather than a price cut. Of the two readings, that one survives contact with an invoice.
I wrote about this same pressure when Kimi K3 topped the coding arenas and almost nothing changed for production stacks. The pattern repeats. Frontier labs are now competing on the denominator.
Does a cheaper frontier model actually lower your bill?
Only if you were already paying frontier rates. Most businesses running a support or booking agent never were. Which is why a headline about halved Claude Opus 5 pricing produces a lot of excitement and very little invoice movement. Price per token is not what you pay. Tokens times price is what you pay.
Take a workload I see constantly. A support agent handling 2,000 conversations a month, roughly 3,500 input tokens per conversation once you count the system prompt, retrieved documents and the running history, and about 400 output tokens in the reply.
| Model | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Claude Fable 5 | $70 | $40 | $110 |
| Claude Opus 5 | $35 | $20 | $55 |
| Claude Sonnet 5 (intro rate) | $14 | $8 | $22 |
| Claude Haiku 4.5 | $7 | $4 | $11 |
So the launch saves that business $55 a month. Real money, and I'll take it. But the same agent on Haiku 4.5 runs $11, and the tier decision you make in an afternoon is a bigger lever than the price change you waited three months for.
Then there's caching. ZDNET was the only outlet to touch it, and only sideways, as a note about tool changes not nuking the cache. Nobody priced it out. On Opus 5 a cache hit bills at $0.50 per million against a $5 base rate. Push 3,000 of those 3,500 input tokens behind a stable cached prefix, the system prompt and the policy docs that never change between turns, and the input line falls from $35 to roughly $8 before write costs. That's a config change. It saves nearly as much as the launch did, it works on Opus 4.8 today, and it worked last month.
Check your cache hit rate before you check the release notes.
Where do the four reports disagree?
Reading all four pieces side by side turned up three genuine contradictions. Worth knowing if you're briefing a board off a single article. A model launch gets covered in about four hours, and the errors land in the boring details rather than the analysis. Small stuff. It still matters.
The Verge said Anthropic released Opus 5 on Thursday. TechCrunch said on Friday. Anthropic's own post is dated July 24, 2026, which was a Friday, so TechCrunch has it right.
ZDNET's David Gewirtz put Sonnet 5's arrival on July 1, while TechCrunch grouped Mythos 5, Fable 5 and Sonnet 5 as June releases. Trivial on its own. Less trivial if you're tracking how fast a vendor deprecates the model your agent is pinned to.
The third disagreement is the interesting one. The Verge framed Opus 5 as carrying more cyber safeguards than the previous Opus. TechCrunch called it less restrictive than Fable and therefore preferable in most cases. Both are true, and Anthropic's own numbers reconcile them: the cyber classifiers are expected to intervene around 85% less often than Fable 5's, while still blocking binary-based vulnerability scanning and exploit generation that Opus 4.8 allowed through.
How much does a small business AI agent cost per month?
Across the 126 systems I've shipped, the model line has almost never been the number that decides whether a project pays for itself. Call it 15% model tokens against 85% everything else: retrieval infrastructure, telephony minutes if there's voice, monitoring, and the human hours spent fixing what the agent got wrong in week one. Model cost is a rounding error on most builds.
A client came to me in a similar spot last quarter, convinced their AI bill was the problem. It was $60 a month. Their actual problem was that the agent escalated 40% of conversations to a human because nobody had written the escalation rules properly, so they were paying a person to redo the agent's work.
We fixed the routing. The token bill went up. The total cost went down.
The calculator defaults a support agent to 2,000 calls a day, which is why the model line looks small next to build and infrastructure.
If you want your own number rather than mine, the AI agent cost calculator models all of it, and I keep the model rates in it verified against the vendor pricing pages. I've also written the longer version of this arithmetic in the AI chatbot pricing breakdown and in what the Army's unlimited token contract taught everyone about budgeting.
Which launch feature actually matters for production agents?
Automatic fallbacks, and it's buried at the bottom of the announcement under "Getting started" while the coverage argued about price, which tells you how far the gap has grown between what makes a good headline and what makes a good on-call night. Requests flagged by the safety classifiers can now route to another model instead of returning an error.
Think about what a hard refusal does to a live agent. A customer asks something the classifier dislikes, the API returns a refusal, and your booking flow dies mid conversation at 11pm. I've been paged for exactly that. The fallback turns an outage into a slightly worse answer, and slightly worse answers are survivable.
The second beta is mid conversation tool changes, letting you swap which tools the model can reach without invalidating the prompt cache. If you've ever watched a long agent session blow its cache because one tool definition changed, you know what that's worth. If you haven't, it reads like nothing.
Neither feature got more than a passing line in the coverage. Both change more about running an agent than the price did.
Should you switch your agent to Opus 5?
Here's my honest decision gate, and it has four questions, which is three more than most people ask before pushing a model change to production on a Monday morning. If you answer no to the first one, stop reading and go do something with better returns.
Are you currently on Fable 5 or Mythos 5 for a general business workload? If yes, move. You're paying double for capability the benchmarks say you get back at Opus 5 within half a percent on coding work.
Are you on Opus 4.8? The price is identical, so the only question is quality, and Box reported Opus 5 outperforming 4.8 by 8% overall with 17% better due diligence results. Test it against your own evals, not theirs.
Are you on Sonnet 5 or Haiku 4.5 and happy? Stay. Moving up a tier multiplies your token line by 5 to 25 times for gains most support and booking workloads never surface.
Is your agent actually failing on reasoning, or is it failing on retrieval? Nine times out of ten I trace a "the model is too dumb" complaint back to bad chunking. A better model on bad context is a more expensive wrong answer.
That last one is where most of the money leaks. If you're not sure which side of it you're on, my free AI readiness assessment walks the same questions I'd ask on a call, and it takes about six minutes.
What I got wrong about model upgrades
For roughly the first two years of doing this work I treated every frontier release as a maintenance task and moved client agents onto the newest model within a week of launch, because it felt like the responsible thing to do and because the benchmark charts were genuinely better every time.
It cost me. Twice I shipped an upgrade that regressed a prompt tuned tightly to the older model's quirks, and both times the client noticed before I did, which is the worst possible order for that to happen in.
What I do now is duller. New model goes behind a flag, runs against a saved set of 40 to 60 real conversations from that client's own history, and gets compared on the two things they care about: resolution rate and escalation rate. If it wins, it ships. If it ties, it waits, because the cheapest model that clears your bar is the right model and a tie is not a reason to change anything.
Opus 5 is currently in that flagged state on two builds. I'll know in a week.
Frequently asked questions about Claude Opus 5 pricing
Is Claude Opus 5 cheaper than Opus 4.8?
No. Both bill at $5 per million input tokens and $25 per million output tokens. Opus 5 is positioned as better work for the same rate, and Anthropic's claim is a lower cost per completed task because the model uses fewer tokens to finish the job.
How much cheaper is Opus 5 than Claude Fable 5?
Exactly half on both sides of the meter. Fable 5 charges $10 per million input and $50 per million output. Opus 5 charges $5 and $25. On a 2,000 conversation a month support agent that difference works out to about $55 a month.
What is the effort setting and does it change my bill?
Effort is a dial that trades intelligence against token consumption, with settings running from low up through high, xhigh and max. Anthropic publishes separate performance curves per effort level. It moves your bill more than any model swap, and only ZDNET got near it, secondhand, through a customer quote about holding quality at lower reasoning levels.
Should a small business use Opus 5 for a customer support chatbot?
Usually not. Most support and booking workloads run fine on Haiku 4.5 at $1 and $5 per million, roughly a fifth of the Opus rate. Reach for Opus when the task involves multi step reasoning over messy documents, not when it involves answering the same 30 questions.
Does prompt caching still work with Opus 5?
Yes, and it's the biggest lever in the whole pricing table. Cache hits bill at $0.50 per million on Opus 5, a tenth of the base input rate. Anthropic also shipped mid conversation tool changes in beta so swapping tools no longer invalidates the cache.
Is Opus 5 safe to use for cybersecurity work?
Partly. It can find vulnerabilities in source code, but blocks binary based vulnerability scanning, penetration testing and exploit generation. Anthropic expects its classifiers to fire around 85% less often than Fable 5's, and runs a Cyber Verification Program for teams that need the restrictions lifted.
What to do this week
Pull last month's token bill and split it into input, output and cache hits. If cache hits are under half your input tokens on a workload with a fixed system prompt, you have a bigger saving available than this launch offered and you can capture it today.
Then check what tier you're actually on. If it's Fable 5 or Mythos 5 for general business work, move to Opus 5 and pocket the difference. If it's Sonnet 5 or Haiku 4.5 and your resolution rate is fine, do nothing at all, which is the correct answer far more often than the release notes imply.
Still unsure whether your problem is the model or the plumbing? That's the question I answer on most first calls, and the free readiness assessment gets you most of the way there without one. If you'd rather see what a built system looks like first, the agents I ship and the AI glossary cover the vocabulary and the shapes.
More on picking the right size of system: when to use AI agents versus plain automation, agents versus chatbots, what RAG actually does for a business, and the automations I'd put in a small business first.
Citation Capsule: Claude Opus 5 launched July 24, 2026 at $5 per million input tokens and $25 per million output tokens, half of Claude Fable 5's $10 and $50, with cache hits at $0.50 per million and cyber classifiers expected to intervene around 85% less often than Fable 5's. Anthropic, Introducing Claude Opus 5 (July 24, 2026) · Claude Platform Pricing Docs (July 2026) · Ars Technica (July 25, 2026) · TechCrunch (July 24, 2026) · The Verge (July 24, 2026) · ZDNET (July 24, 2026) · Frontier-Bench · GDPval-AA, Artificial Analysis.
Last updated July 26, 2026. Written by Jahanzaib Ahmed, AI Systems Engineer. I've shipped 126 production AI systems for businesses in the US, Canada and Australia. If your AI bill stopped making sense, start with the readiness assessment.
Top comments (0)