DEV Community

chenyu
chenyu

Posted on AI-assisted

Your system prompt is silently killing your prompt cache

A benchmark on DeepSeek. Moving roughly 30 tokens from the top of a system message to the bottom cut steady-state inference cost by 96%.


If you run a chat or roleplay app, your system message is probably the largest thing you send to the model. A character card, a lorebook, a memory summary — tens of thousands of tokens, resent on every single turn.

DeepSeek's context caching is supposed to make that cheap. It's on by default, it requires no code changes, and cached input costs 50x less than uncached input ($0.003 vs $0.15 per million tokens, off-peak).

In practice, most apps get almost none of it. Not because caching is broken, but because of one line of code near the top of the system prompt.

I built a benchmark to measure exactly how much that line costs. Here's what it found.

The setup

I took a realistic roleplay prompt — character card, world book with 20 lorebook entries, long-term memory, relationship state — and built a stable block of roughly 20,000 tokens.

Then I ran three variants. The stable content is byte-identical in all three. The only thing that changes is where a small volatile header goes:

Session context: local time <ISO timestamp>, turn <n>, session <uuid>, mood index <0.00>.
Enter fullscreen mode Exit fullscreen mode

That's about 30 tokens. Almost nothing. Here's where each variant puts it:

Variant Where the volatile header goes Typical of
A — naive Front of the system message Most first implementations
B — minimal fix End of the system message A one-line change from A
C — optimized End of the newest user turn; the system message never changes Append-only architecture

Each variant ran 8 turns of a real conversation against deepseek-flash, with thinking mode disabled. Every number below comes from the API's own usage fields — prompt_cache_hit_tokens and prompt_cache_miss_tokens. Nothing is modelled.

Results

Variant Cache hit rate Total input tokens Cost for 8 turns Cost per turn vs A
A 0.0% 138,237 $0.021006 $0.002626 —
B 85.8% 137,788 $0.003465 $0.000433 −83.5%
C 86.8% 138,987 $0.003326 $0.000416 −84.2%

Turn 1 is always a cache miss in every variant — the cache is cold, there's nothing to match against yet. If you exclude that cold start and look at steady state:

Variant Cost per turn vs A
A $0.002633 —
B $0.000127 −95.2%
C $0.000107 −95.9%

Variant A never had a single cache hit. Not one turn out of eight. The entire 17,000-token prefix was reprocessed at full price, every time, because the first ~30 tokens of the request were different.

Two details worth looking at

1. The one-line fix gets you almost everything.

B and C land within 0.7 percentage points of each other. You do not need to redesign your architecture to capture most of this. You need to move one string.

2. But C pulls ahead as the conversation grows.

Look at the absolute cache hit tokens per turn:

Turn B hit tokens C hit tokens
2 16,896 16,896
4 16,896 17,152
6 16,896 17,280
8 16,896 17,536

B flatlines at exactly 16,896 tokens — which is 264 × 64, and 64 tokens appears to be the cache granularity. Because B's system message changes every turn, only the stable block ahead of the header can ever be reused. The conversation history is dead weight that gets recomputed forever.

C keeps climbing, because its system message is frozen and its history is append-only, so each turn's cached prefix includes everything before it.

Over an 8-turn conversation that difference is small. Over a 50-turn session it isn't.

Why this happens

DeepSeek documents the rule clearly, and it's easy to miss:

A cache hit requires that the corresponding prefix has already been persisted... A subsequent request can only hit the cache if it fully matches a cache prefix unit.

Fully is the operative word. Prefix caching is all-or-nothing at the point of divergence. Put a timestamp at position zero and everything after it is new content as far as the cache is concerned — even if 99.99% of the bytes are identical to the last request.

I also verified this without spending anything, by diffing the raw request bodies of turn 1 and turn 2:

Variant Request bytes Identical prefix Share
A 84,665 75 0.1%
B 84,665 84,493 99.8%
C 84,665 84,664 100.0%

75 bytes out of 84,665. That's the whole story.

What it costs at scale

Extrapolating the measured steady-state per-turn cost to an app with 1,000 daily active users averaging 50 turns each:

Variant Monthly cost Monthly saving vs A
A $3,948.78 —
B $190.78 $3,758.00
C $161.12 $3,787.67

That's a 96% reduction from moving a header. The ratio is the finding; the absolute dollars depend on your prompt size and volume.

The other free win

If you're on this model family, check one more thing: thinking mode is enabled by default, with effort set to high, and reasoning tokens are billed as output.

Roleplay and chat apps don't need a chain of thought. Users want a reply in character, not an internal monologue. If your integration never explicitly disables it, you may be paying for reasoning on every message.

{ "thinking": { "type": "disabled" } }
Enter fullscreen mode Exit fullscreen mode

Caveats

I'd rather you trust the parts that hold up than be surprised later:

  1. Caching is best-effort. DeepSeek makes no guarantee of a 100% hit rate, and repeated runs will vary.
  2. Cache entries expire automatically, typically within hours to days. A quiet app caches less than a busy one.
  3. The first request of any conversation is always a miss.
  4. A real app compresses history rather than letting it grow, so absolute figures will differ from the projection above.
  5. Peak pricing is twice off-peak. This run was entirely off-peak.
  6. This measured one provider. The specific numbers are DeepSeek's; the principle — volatile content at the front destroys prefix caching — applies to any provider with prefix-based caching.

What to check in your own app

You don't need a benchmark harness for the first pass. Just look at your request construction and ask:

  • Does anything in my system message change between turns? Timestamps, turn counters, user state, random IDs, "current mood" — anything.
  • Do I rebuild the message array each turn, or append to it?
  • Is my history being re-summarised in a way that changes earlier tokens?

If the answer to the first one is yes, you have a one-line fix and it's probably worth more than any model swap you could make this quarter.


Measured on deepseek-flash with thinking disabled, 8 turns per variant, ~20k-token stable block. Raw per-turn usage data available on request.

If you run an app like this, I'd genuinely like to know whether these numbers match your real bill — especially if your prompts are larger than 20k tokens, where I'd expect the gap to widen.

Benchmark and raw data: https://github.com/chenyu520-ai/rp-cache-lab

Top comments (2)

Collapse
 
max_quimby profile image
Max Quimby •

This matches what we see running scheduled multi-agent jobs where a large fixed preamble gets re-read every turn — the single biggest cost lever is keeping that prefix byte-identical, and the fastest way to kill it is exactly what you showed: one volatile field (timestamp, turn counter, uuid) sitting near the top. Variant C is the right instinct — treat the system block as append-only and let only the newest user turn carry the volatile bits.

Two things I'd add from other providers. The prefix-cache behavior generalizes — Anthropic's is prefix-based with explicit breakpoints plus a TTL, so a stray dynamic field high in the system block invalidates everything downstream the same way. And the amortization depends heavily on session length: if most of your sessions are only 2–3 turns, the cold turn-1 miss dominates and the win shrinks fast. Did you look at the hit-rate curve as a function of conversation length? That distribution is what decides whether this is a 95% win or a rounding error for a given app.

Collapse
 
tomveber profile image
Tom Veber •

Was the 85.8% vs 86.8% gap between B and C stable across reruns, or is one point inside the noise of 8 turns?