DEV Community

The cached-prefix crossover: when the cheaper LLM becomes the expensive one

Most LLM cost comparisons collapse a model to one number: $X per million
input tokens
. That number is the price of a cache miss. Once prompt caching
is on, most of the tokens you send on every request are billed at the cached-read
rate instead — and caching does not discount every model equally. When it
discounts one model far more than another, the ranking flips.

That flip has a precise location. This post computes it.

The arithmetic

All rates are dollars per million tokens. For a request of N input tokens at a
cache-hit fraction f, plus M output tokens:

cost(f) = N · [ (1 − f)·in + f·cached ]  +  M · out
        = (N·in + M·out)  −  f · N · (in − cached)
Enter fullscreen mode Exit fullscreen mode

Cost is linear in the hit rate. Everything that does not depend on caching is
the constant term; the whole effect of the cache is the slope, −N·(in − cached).

Two models, same workload. Set cost_A(f) = cost_B(f) and solve:

        (N·in_A + M·out_A) − (N·in_B + M·out_B)
f* = ─────────────────────────────────────────────
          N·(in_A − cached_A) − N·(in_B − cached_B)
Enter fullscreen mode Exit fullscreen mode
  • f* < 0 or f* > 1 — no flip inside the operating range; the cheaper model stays cheaper.
  • 0 ≤ f* ≤ 1 — there is a real crossover. Below f* the first model wins; above it, the second.

The denominator is the difference in cache benefit. If two models discount
cached reads by the same factor, the denominator is zero and caching never
reorders them, no matter how good that factor is.

The number people actually miss: the cache discount ratio

in / cached is how much cheaper a hit is than a miss. Across the 28 models we
track, it is not a constant:

Model input ÷ cached Discount on a hit
DeepSeek V4.1 Flash 50.0× 98% off
Claude Fable 5.1 40.0× 97.5% off
DeepSeek V4 Pro 30.0× 96.7% off
GPT-6.1 Sol 20.0× 95% off
(most OpenAI / Anthropic / Google models) 10.0× 90% off
Grok 4.3 6.25× 84% off
GPT-4.1 4.0× 75% off
Grok 4.6 4.0× 75% off

A model near the top of this table gets dramatically cheaper as its cache warms;
a model near the bottom barely moves. That gap is what creates crossovers.

Crossovers that actually bite (RAG: 50k in / 2k out)

Take a retrieval workload — 50,000 input tokens, 2,000 output tokens per call,
the shape you get when you paste a document set into the prompt on every request:

Cheaper at 0% hit Becomes more expensive than Crossover
GPT-4.1 GPT-6.1 Sol 20.0%
GPT-4.1 Claude Sonnet 5 26.7%
GPT-5.2 GPT-6.1 Sol 27.7%
Grok 4.6 GPT-6.1 Sol 40.0%
Grok 4.6 GPT-6 Sol 53.3%
Grok 4.6 Claude Sonnet 5 53.3%
Claude Haiku 4.5 DeepSeek V4 Pro 74.0%
GPT-5.4 mini DeepSeek V4 Pro 91.2%

Worked example: Grok 4.6 vs GPT-6 Sol

  • Grok 4.6: input $2, cached $0.50, output $6
  • GPT-6 Sol: input $2, cached $0.20, output $10

They share the same input price, so the sticker price says they cost the same to
feed. They do not.

cost_Grok(f)   = 50000·(2 − 1.5f) + 2000·6  = 112,000 − 75,000·f
cost_GPT6(f)   = 50000·(2 − 1.8f) + 2000·10 = 120,000 − 90,000·f

  112,000 − 75,000·f  =  120,000 − 90,000·f
                15,000·f = 8,000
                       f = 0.533
Enter fullscreen mode Exit fullscreen mode

Below 53.3% cache hit, Grok 4.6 is cheaper. Above it, GPT-6 Sol is. Same
input price on the box; opposite conclusion depending on cache behaviour — and
neither provider's pricing page tells you the crossover exists.

The GPT-4.1 result is starker still: it is cheaper than GPT-6.1 Sol until only
20% cache hit, because GPT-4.1 discounts a hit by 4× while GPT-6.1 Sol
discounts it by 20×.

Why this matters

Teams pick a model on the sticker input price, ship it, and then cannot
reconcile the invoice — because their real hit rate moved them across a boundary
they never knew existed. The cache discount ratio, not the headline price, is
what determines who wins once caching is on.

What this does not say

  • Rates are the base tier, as published October 2026. Long-context surcharges and batch tiers change the constants and move f*.
  • It assumes cache hits are yours to arrange for free. Cache writes are often billed separately, and a workload with a cold cache most of the time sits near f = 0.
  • Anthropic charges no long-context surcharge; a model with a hidden tier threshold behaves differently above it than this single-tier model shows.

The formula is exact; the input rates are the fragile part. Check them, and check
your own workload shape before trusting f*.


The numbers here come from a free, no-account, client-side calculator that
applies prompt-cache writes, cached reads, batch tiers and long-context tiering
the way each provider bills them, then ranks all 28 models by real monthly cost
for a given workload: **https://rabayid.com/
. Rates are transcribed from each
provider's pricing page and dated; the cache crossover above is reproducible from
the published data/pricing.json.

Top comments (0)