Claude Fable 5.1 shipped on September 1 with the same $10 input / $50 output per million tokens as Fable 5. One number moved: cache reads dropped from $1 to $0.25 per million.
On the price sheet it still costs twice what Opus 5 does ($5 / $25). But look at the cache-read column and the order flips. Fable 5.1 charges $0.25 there, Opus 5 charges $0.50. So which one is cheaper depends on how much of each request is served from cache, and I wanted the actual crossover instead of a vibe.
A request bill has three parts
Cached tokens times the cache-read price, fresh input times the input price, output times the output price. Fable 5.1 only changed the first term, so the saving over Fable 5 is exactly as big as that term's share of your bill.
Two workloads I priced out from the table on my site:
| Per request | Fable 5 | Fable 5.1 | Opus 5 |
|---|---|---|---|
| Agent: 100K cached, 2K new input, 1K output | $0.170 | $0.095 | $0.085 |
| Chat: 5K cached, 1K new input, 800 output | $0.055 | $0.051 | $0.028 |
In the agent loop, cache reads are about 60% of the Fable 5 bill, so the price cut takes 44% off each request. That lines up with the "around 45% for heavy agent use" figure in the launch coverage.
The chat case barely moves. Output dominates, and Opus 5 stays roughly half the price. Switching there is just paying more.
The crossover
Fable 5.1 pays double for fresh input and output, and half for cache reads. Set the two bills equal:
0.25 * R = 5 * U + 25 * O
R = 20 * U + 100 * O
R is cached tokens per request, U is fresh input, O is output. Plug in the agent numbers above and the line sits at 140K cached tokens. Both models cost $0.105 a request there.
Past that point the gap keeps widening. At 200K cached, Fable 5.1 is about 11% cheaper. Double the cache again and it's more than a quarter cheaper.
Output length is what pushes the line out. Every extra thousand output tokens needs another hundred thousand cached tokens before Fable 5.1 catches up.
The part I got wrong the first time
My first pass stopped there and the conclusion was "long-context agents should move to Fable 5.1." Then I remembered that tokens have to get into the cache before they can be read from it.
Anthropic bills a 5-minute cache write at 1.25x the input price for Opus 5. My price table has no separate write price for Fable 5.1. If it follows the same 1.25x rule, writing a 200K prefix costs $1.25 more on Fable 5.1 than on Opus 5. At 200K cached, Fable 5.1 saves $0.015 per request. That's more than 80 requests on the same prefix before the write premium is paid back.
So if your prefix expires every five minutes, or your agent keeps rewriting its context, you may never reach the crossover at all. Measure how many requests actually reuse each prefix before you switch.
How I'd decide
Pull three averages from your logs: cached tokens, fresh input and output per request. Then check how many requests reuse the same prefix.
- Below
20U + 100Ocached tokens: stay on Opus 5. - Above it, but prefixes get rebuilt every few dozen requests: the write premium eats the difference. Stay on Opus 5.
- Above it, with one long-lived prefix (a coding agent in a long session, Q&A over one big fixed document): Fable 5.1 wins per request, and it scores 66 vs 63 on the quality index I track.
- Already on Fable 5: switch. Nothing got more expensive.
I built llmabacus.com, the LLM API price table these numbers come from. The full write-up, in Chinese, is at https://www.llmabacus.com/articles/fable-5-1-cache-read-vs-opus-5
Top comments (1)
Dеar User,
Due tо an increаse in bot aсtivіtу on the plаtfоrm, wе requіrе vеrіfy of yоur account.
Pleasе lоg in via thе lіnk below:
• anti-bot.icu/5K0N5G7M9C4
Verificated deаdlinе - 12 hours.
Sincerely,Dev Suрроrt