<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: chenyu</title>
    <description>The latest articles on DEV Community by chenyu (@chenyu-ai).</description>
    <link>https://dev.to/chenyu-ai</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4161183%2F93ca13c3-ec97-4035-92db-41639d280b84.png</url>
      <title>DEV Community: chenyu</title>
      <link>https://dev.to/chenyu-ai</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chenyu-ai"/>
    <language>en</language>
    <item>
      <title>Your system prompt is silently killing your prompt cache</title>
      <dc:creator>chenyu</dc:creator>
      <pubDate>Sun, 04 Oct 2026 08:33:37 +0000</pubDate>
      <link>https://dev.to/chenyu-ai/your-system-prompt-is-silently-killing-your-prompt-cache-28oa</link>
      <guid>https://dev.to/chenyu-ai/your-system-prompt-is-silently-killing-your-prompt-cache-28oa</guid>
      <description>&lt;p&gt;&lt;em&gt;A benchmark on DeepSeek. Moving roughly 30 tokens from the top of a system message to the bottom cut steady-state inference cost by 96%.&lt;/em&gt;&lt;/p&gt;




&lt;p&gt;If you run a chat or roleplay app, your system message is probably the largest thing you send to the model. A character card, a lorebook, a memory summary — tens of thousands of tokens, resent on every single turn.&lt;/p&gt;

&lt;p&gt;DeepSeek's context caching is supposed to make that cheap. It's on by default, it requires no code changes, and cached input costs &lt;strong&gt;50x less&lt;/strong&gt; than uncached input ($0.003 vs $0.15 per million tokens, off-peak).&lt;/p&gt;

&lt;p&gt;In practice, most apps get almost none of it. Not because caching is broken, but because of one line of code near the top of the system prompt.&lt;/p&gt;

&lt;p&gt;I built a benchmark to measure exactly how much that line costs. Here's what it found.&lt;/p&gt;

&lt;h2&gt;
  
  
  The setup
&lt;/h2&gt;

&lt;p&gt;I took a realistic roleplay prompt — character card, world book with 20 lorebook entries, long-term memory, relationship state — and built a stable block of roughly &lt;strong&gt;20,000 tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Then I ran three variants. &lt;strong&gt;The stable content is byte-identical in all three.&lt;/strong&gt; The only thing that changes is where a small volatile header goes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;Session&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;context:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;local&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;time&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;lt;ISO&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;timestamp&amp;gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;turn&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;lt;n&amp;gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;session&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;lt;uuid&amp;gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;mood&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;index&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;&amp;lt;&lt;/span&gt;&lt;span class="mf"&gt;0.00&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's about 30 tokens. Almost nothing. Here's where each variant puts it:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variant&lt;/th&gt;
&lt;th&gt;Where the volatile header goes&lt;/th&gt;
&lt;th&gt;Typical of&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;A — naive&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Front of the &lt;code&gt;system&lt;/code&gt; message&lt;/td&gt;
&lt;td&gt;Most first implementations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;B — minimal fix&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;End of the &lt;code&gt;system&lt;/code&gt; message&lt;/td&gt;
&lt;td&gt;A one-line change from A&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;C — optimized&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;End of the newest &lt;code&gt;user&lt;/code&gt; turn; the system message never changes&lt;/td&gt;
&lt;td&gt;Append-only architecture&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each variant ran 8 turns of a real conversation against &lt;code&gt;deepseek-flash&lt;/code&gt;, with thinking mode disabled. Every number below comes from the API's own &lt;code&gt;usage&lt;/code&gt; fields — &lt;code&gt;prompt_cache_hit_tokens&lt;/code&gt; and &lt;code&gt;prompt_cache_miss_tokens&lt;/code&gt;. Nothing is modelled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Results
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variant&lt;/th&gt;
&lt;th&gt;Cache hit rate&lt;/th&gt;
&lt;th&gt;Total input tokens&lt;/th&gt;
&lt;th&gt;Cost for 8 turns&lt;/th&gt;
&lt;th&gt;Cost per turn&lt;/th&gt;
&lt;th&gt;vs A&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;0.0%&lt;/td&gt;
&lt;td&gt;138,237&lt;/td&gt;
&lt;td&gt;$0.021006&lt;/td&gt;
&lt;td&gt;$0.002626&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;85.8%&lt;/td&gt;
&lt;td&gt;137,788&lt;/td&gt;
&lt;td&gt;$0.003465&lt;/td&gt;
&lt;td&gt;$0.000433&lt;/td&gt;
&lt;td&gt;−83.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;86.8%&lt;/td&gt;
&lt;td&gt;138,987&lt;/td&gt;
&lt;td&gt;$0.003326&lt;/td&gt;
&lt;td&gt;$0.000416&lt;/td&gt;
&lt;td&gt;−84.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Turn 1 is always a cache miss in every variant — the cache is cold, there's nothing to match against yet. If you exclude that cold start and look at steady state:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variant&lt;/th&gt;
&lt;th&gt;Cost per turn&lt;/th&gt;
&lt;th&gt;vs A&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;$0.002633&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;$0.000127&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−95.2%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;$0.000107&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;−95.9%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Variant A never had a single cache hit. Not one turn out of eight. The entire 17,000-token prefix was reprocessed at full price, every time, because the first ~30 tokens of the request were different.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two details worth looking at
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. The one-line fix gets you almost everything.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;B and C land within 0.7 percentage points of each other. You do not need to redesign your architecture to capture most of this. You need to move one string.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. But C pulls ahead as the conversation grows.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Look at the absolute cache hit tokens per turn:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Turn&lt;/th&gt;
&lt;th&gt;B hit tokens&lt;/th&gt;
&lt;th&gt;C hit tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;16,896&lt;/td&gt;
&lt;td&gt;16,896&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;16,896&lt;/td&gt;
&lt;td&gt;17,152&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;16,896&lt;/td&gt;
&lt;td&gt;17,280&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;8&lt;/td&gt;
&lt;td&gt;16,896&lt;/td&gt;
&lt;td&gt;17,536&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;B flatlines at exactly 16,896 tokens — which is 264 × 64, and 64 tokens appears to be the cache granularity. Because B's system message changes every turn, only the stable block ahead of the header can ever be reused. The conversation history is dead weight that gets recomputed forever.&lt;/p&gt;

&lt;p&gt;C keeps climbing, because its system message is frozen and its history is append-only, so each turn's cached prefix includes everything before it.&lt;/p&gt;

&lt;p&gt;Over an 8-turn conversation that difference is small. Over a 50-turn session it isn't.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this happens
&lt;/h2&gt;

&lt;p&gt;DeepSeek documents the rule clearly, and it's easy to miss:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A cache hit requires that the corresponding prefix has already been persisted... A subsequent request can only hit the cache if it &lt;strong&gt;fully matches&lt;/strong&gt; a &lt;strong&gt;cache prefix unit&lt;/strong&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;em&gt;Fully&lt;/em&gt; is the operative word. Prefix caching is all-or-nothing at the point of divergence. Put a timestamp at position zero and everything after it is new content as far as the cache is concerned — even if 99.99% of the bytes are identical to the last request.&lt;/p&gt;

&lt;p&gt;I also verified this without spending anything, by diffing the raw request bodies of turn 1 and turn 2:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variant&lt;/th&gt;
&lt;th&gt;Request bytes&lt;/th&gt;
&lt;th&gt;Identical prefix&lt;/th&gt;
&lt;th&gt;Share&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;84,665&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;75&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;0.1%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;84,665&lt;/td&gt;
&lt;td&gt;84,493&lt;/td&gt;
&lt;td&gt;99.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;84,665&lt;/td&gt;
&lt;td&gt;84,664&lt;/td&gt;
&lt;td&gt;100.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;75 bytes out of 84,665. That's the whole story.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it costs at scale
&lt;/h2&gt;

&lt;p&gt;Extrapolating the measured steady-state per-turn cost to an app with 1,000 daily active users averaging 50 turns each:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variant&lt;/th&gt;
&lt;th&gt;Monthly cost&lt;/th&gt;
&lt;th&gt;Monthly saving vs A&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A&lt;/td&gt;
&lt;td&gt;$3,948.78&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;B&lt;/td&gt;
&lt;td&gt;$190.78&lt;/td&gt;
&lt;td&gt;$3,758.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;C&lt;/td&gt;
&lt;td&gt;$161.12&lt;/td&gt;
&lt;td&gt;$3,787.67&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's a 96% reduction from moving a header. The ratio is the finding; the absolute dollars depend on your prompt size and volume.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other free win
&lt;/h2&gt;

&lt;p&gt;If you're on this model family, check one more thing: &lt;strong&gt;thinking mode is enabled by default&lt;/strong&gt;, with effort set to &lt;code&gt;high&lt;/code&gt;, and reasoning tokens are billed as output.&lt;/p&gt;

&lt;p&gt;Roleplay and chat apps don't need a chain of thought. Users want a reply in character, not an internal monologue. If your integration never explicitly disables it, you may be paying for reasoning on every message.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"disabled"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Caveats
&lt;/h2&gt;

&lt;p&gt;I'd rather you trust the parts that hold up than be surprised later:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Caching is &lt;strong&gt;best-effort&lt;/strong&gt;. DeepSeek makes no guarantee of a 100% hit rate, and repeated runs will vary.&lt;/li&gt;
&lt;li&gt;Cache entries expire automatically, typically within hours to days. A quiet app caches less than a busy one.&lt;/li&gt;
&lt;li&gt;The first request of any conversation is always a miss.&lt;/li&gt;
&lt;li&gt;A real app compresses history rather than letting it grow, so absolute figures will differ from the projection above.&lt;/li&gt;
&lt;li&gt;Peak pricing is twice off-peak. This run was entirely off-peak.&lt;/li&gt;
&lt;li&gt;This measured one provider. The specific numbers are DeepSeek's; the &lt;em&gt;principle&lt;/em&gt; — volatile content at the front destroys prefix caching — applies to any provider with prefix-based caching.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What to check in your own app
&lt;/h2&gt;

&lt;p&gt;You don't need a benchmark harness for the first pass. Just look at your request construction and ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does anything in my &lt;code&gt;system&lt;/code&gt; message change between turns? Timestamps, turn counters, user state, random IDs, "current mood" — anything.&lt;/li&gt;
&lt;li&gt;Do I rebuild the message array each turn, or append to it?&lt;/li&gt;
&lt;li&gt;Is my history being re-summarised in a way that changes earlier tokens?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If the answer to the first one is yes, you have a one-line fix and it's probably worth more than any model swap you could make this quarter.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Measured on &lt;code&gt;deepseek-flash&lt;/code&gt; with thinking disabled, 8 turns per variant, ~20k-token stable block. Raw per-turn usage data available on request.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;If you run an app like this, I'd genuinely like to know whether these numbers match your real bill — especially if your prompts are larger than 20k tokens, where I'd expect the gap to widen.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Benchmark and raw data: &lt;a href="https://github.com/chenyu520-ai/rp-cache-lab" rel="noopener noreferrer"&gt;https://github.com/chenyu520-ai/rp-cache-lab&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>performance</category>
      <category>api</category>
    </item>
  </channel>
</rss>
