Headline: Prompt caching saves money only when the cached prefix is byte-identical across requests, and a single interpolated timestamp in a system prompt is enough to make every call a cache write instead of a cache read. The three things that fixed my bill were ordering the request static-first (tools, then system, then messages), putting the
cache_controlbreakpoint at the end of the stable region, and readingcache_read_input_tokensfrom every response instead of assuming caching was working.
Key takeaways
- Prompt caching is a provider-side store of the processed prefix of a request, so a later request with the identical prefix skips re-processing it. With the Anthropic Messages API it is opt-in per content block: add
cache_control: { type: "ephemeral" }to a block and everything up to and including that block becomes the cached prefix. A request may carry at most four breakpoints. - A cache read is billed at 0.1x the base input token price. A five-minute cache write is billed at 1.25x and a one-hour write at 2x. The extra 0.25x paid on a five-minute write is repaid by the first hit inside the window, because each hit saves 0.9x.
- The cache key is the exact serialized prefix in the fixed order tools, then system, then messages. A timestamp in the system prompt, a tool array built by iterating a
Map, or a swapped model id all invalidate from the point of change onward. - Every Messages API response reports
input_tokens,cache_creation_input_tokens, andcache_read_input_tokens. Cache hit rate isread / (read + creation + input), and it is the only honest way to know whether caching is on. - Prompt caching does not reduce output token cost and does not fix a bloated prompt. Deleting four thousand tokens of context nobody reads beats caching them at 0.1x forever.
The feature was a support agent with a large tool schema and a long API reference pinned into the system prompt, which is exactly the shape prompt caching is built for. I added a breakpoint, deployed, and the cost per conversation did not move. The usage object was blunt about it: cache_creation_input_tokens was populated on every single request and cache_read_input_tokens was zero on every single request. I was paying 1.25x to write a cache entry that nothing ever read.
What is prompt caching, and what am I actually paying for?
Prompt caching lets the provider reuse the work it already did on a prefix of your prompt, so identical leading tokens are billed at a steep discount instead of being processed again. It splits your input bill into three separate line items rather than one: fresh uncached input, cache writes, and cache reads. Reading those three numbers is the whole discipline.
Two constraints decide whether a breakpoint does anything at all. The first is the minimum cacheable prefix, which is 1024 tokens on the larger Claude models and 2048 on the smallest ones; below the minimum the cache_control marker is ignored and no error is raised, which looks exactly like caching being broken. The second is the time to live: the default entry lives five minutes and every read refreshes that window, so a steady stream of requests keeps a prefix warm indefinitely without paying to write it again. Passing ttl: "1h" extends the window at a higher write price.
Where do I put the cache breakpoint?
Put the breakpoint on the boundary between what never changes per request and what always changes. The provider serializes tools first, then the system prompt, then messages, so anything you want cached must live before anything volatile. A large tool schema is usually the biggest static block in an agent request and is the best first thing to get behind a breakpoint.
const response = await anthropic.messages.create({
model: "claude-sonnet-5",
max_tokens: 1024,
tools: TOOLS, // serialized first, so it sits inside the cached prefix
system: [
{ type: "text", text: AGENT_INSTRUCTIONS },
{
type: "text",
text: API_REFERENCE_DOC,
cache_control: { type: "ephemeral" }, // breakpoint: cache everything above
},
],
messages: [{ role: "user", content: question }],
});
console.log(response.usage.cache_creation_input_tokens); // non-zero on the cold call
console.log(response.usage.cache_read_input_tokens); // non-zero on every warm call
For a multi-turn conversation I use two breakpoints rather than one. The first sits at the end of the tools and system region, which never changes for the life of the deployment. The second sits at the end of the last completed turn, so the growing transcript is cached incrementally and each new turn only pays full price for the tokens the user just added.
Why did my cache hit rate drop to zero?
A cache hit requires the prefix to be identical byte for byte, so anything that varies per request destroys every hit from that point onward. My own failure was one interpolated clock reading in the system prompt, added months earlier so the agent could reason about business hours.
- Interpolated volatile values. A timestamp, a request id, a user name, or a session id rendered into the system prompt changes the prefix on every call.
-
Non-deterministic ordering. A tool array built by iterating an object or a
Set, or a JSON body serialized with unstable key order, produces a different prefix from identical data. - Model changes. Cache entries are keyed per model, so moving from one model id to another starts cold, and so does an alias that silently points somewhere new.
- Middleware that rewrites the request. A proxy or gateway that injects a header-derived block or reorders messages breaks the prefix even though your application code is unchanged.
- Trimming a conversation from the front. Dropping the oldest messages to fit a context budget rewrites the start of the prefix and invalidates everything.
// Breaks every hit: the prefix is different on every request
const system = `You are a support agent. Current time: ${new Date().toISOString()}`;
// Fix: keep the system prompt frozen, move volatile facts after the breakpoint
const system = [
{ type: "text", text: STATIC_INSTRUCTIONS, cache_control: { type: "ephemeral" } },
];
const messages = [
{ role: "user", content: `Current time: ${now}\n\n${question}` },
];
When is a one-hour cache TTL worth 2x the write price?
The break-even point is pure arithmetic on the published multipliers, not a benchmark. A five-minute write costs 0.25x more than an uncached request and each hit saves 0.9x, so the first hit inside the window already pays for the write. A one-hour write costs 1x more, so it needs roughly two hits to come out ahead.
| Mode | Write price | Read price | Hits to break even | Where I use it |
|---|---|---|---|---|
| No caching | 1x input | n/a | n/a | Prefixes below the model minimum, or genuinely one-shot calls |
| Five-minute cache | 1.25x input | 0.1x input | 1 | Chat turns and agent loops, where a live user keeps traffic flowing |
| One-hour cache | 2x input | 0.1x input | 2 | A long document reused across a work session, or a batch spread over an hour |
Concurrency changes the math in one specific way. The entry is created by the request that writes it, so firing ten cold requests in parallel means paying the write price ten times. When I fan out, I send one warm-up call first and start the fan-out after it returns.
How do I measure cache hits instead of guessing?
Log the three input counters on every response and compute a hit rate, because caching fails silently and looks identical to caching that was never enabled. I emit the ratio as a metric and treat a drop after a deploy as expected rather than alarming, since editing a prompt is by definition a cache invalidation.
const {
input_tokens: fresh,
cache_creation_input_tokens: written = 0,
cache_read_input_tokens: read = 0,
} = response.usage;
const hitRate = read / (read + written + fresh);
logger.info({ fresh, written, read, hitRate }, "llm.usage");
Through the Vercel AI SDK the same numbers arrive under provider metadata instead of a top-level usage object. Marking a message with providerOptions.anthropic.cacheControl sets the breakpoint, and providerMetadata.anthropic on the result carries cacheCreationInputTokens and cacheReadInputTokens. If those fields are absent, the request never reached the provider in a cacheable shape.
What does prompt caching not fix, and how do other providers do it?
Caching is an input-side optimization only. It does not change the model output, it does not reduce output token price, and it does not make a badly scoped prompt cheap. The single biggest cost win I found was not the cache at all; it was deleting a stale section of the reference document that no answer had ever cited.
- OpenAI applies prompt caching automatically to prompts above roughly 1024 tokens, with no explicit breakpoint to place, and reports cached tokens in the usage payload.
- Google Gemini exposes both implicit caching and an explicit context caching API where you create a cached content handle with its own TTL and reference it by name.
- Anthropic is the explicit one: you choose the breakpoints, which costs a design decision and buys precise control over what is cached.
The portable lesson survives all three: structure the request static-first, keep the volatile parts last, and verify with the provider's own token counters.
FAQ
Q: Does prompt caching change the model's output?
A: No. Prompt caching reuses the processed prefix and does not alter sampling, so the same prompt and parameters behave the same whether the prefix was read from cache or processed fresh.
Q: Can another organization read my cached prompt?
A: No. Cache entries are scoped to your own organization and keyed on the exact prefix, so there is no cross-organization sharing of cached content.
Q: Does the five-minute TTL reset on every hit?
A: Yes. Each cache read refreshes the five-minute window, so continuous traffic keeps a prefix alive without ever paying the write price again.
Q: Why do I still see cache creation tokens on every request when nothing changed?
A: Something in the prefix is changing, most often an interpolated value or a non-deterministic key order. Hash the serialized prefix of two consecutive requests and diff the two strings; the difference is always visible once you look at the bytes.
Q: Should I cache tool definitions or the system prompt first?
A: Tools, because they are serialized before the system prompt and are usually the largest static block in an agent request. Placing the breakpoint after the system prompt covers the tools as well, since the cached prefix is everything up to the breakpoint.
Originally published on devya.dev. Also on eng-ahmed.com. Built by Devya Solutions.
Top comments (1)
The most useful takeaway here is that prompt caching is really an architecture property of the request, not just a switch you turn on. If the prefix isn't stable, the cache can be technically configured and still provide almost no benefit.
This is something we pay close attention to at IT Path Solutions when optimizing production AI workflows. I’d treat the serialized prompt prefix almost like an API contract: static tool definitions, stable system instructions, and then volatile request/session data. That makes cache behavior much easier to reason about and prevents an innocent timestamp or middleware transformation from silently destroying the hit rate.
The recommendation to log
cache_read_input_tokensandcache_creation_input_tokensis probably the most important operational detail. Without those counters, a cost regression can look like “the model got expensive” when the actual problem is simply that every request became a cache write.I’d also add cache-hit rate to deployment-level observability alongside latency and token usage. A prompt change that improves model quality but drops cache reuse should be visible as an intentional cost/quality tradeoff, not discovered later on the bill.
And the stale-reference example is a good reminder that caching is secondary to context hygiene. Paying 0.1x for tokens nobody needs is still more expensive than removing those tokens in the first place.