Most agents throw a chunk of them away every single turn, and never see it happen.
There are two levers on an AI coding bill. The big one is routing: send the easy work to a cheaper model, keep the frontier model for the hard slice, and measure whether the cheap one was actually good enough. That's where the large savings are, and you have to earn them.
This is about the small lever. It's duller and it needs no judgement at all, which is rather the point. If you run an agent, you're almost certainly paying full price for tokens you'd already bought at a discount, over and over, and you can't see it happening.
Here's the mechanism. When you send a request to Claude, the big unchanging part at the front (your system prompt, your tool definitions, the file context an agent reloads every turn) can be cached. The next request that starts with the same prefix reads it back at about a tenth of the normal input price. A cache hit is cheap. That's the whole deal, and it's a good one.
A cache miss is the opposite. You pay the full fresh-read price for a prefix you'd already paid to store. Same tokens in, same answer out, higher price. You get nothing for the difference. It's the purest waste in the whole bill.
And the cache is fragile in a way nobody warns you about. Change one byte near the front and the hit becomes a miss. The usual culprits are dull: a line ending that flipped from LF to CRLF, a trailing space, your tool definitions serialised in a different order this run, a timestamp injected into the system prompt, a volatile header, the order two chunks got concatenated. None of these change what the model sees in any way that matters. All of them bust the cache.
A miss isn't a small, contained thing either. The cache matches your prefix from the front, token by token, and the instant it meets a byte that's changed, everything after it has to be recomputed. Claude Code builds that prefix in a fixed order: system prompt and tool definitions, then your static project memory (your CLAUDE.md and docs), then the conversation, then the new input. Put a volatile line near the top, a git branch, a coverage figure, a "last updated" stamp in the first few lines of a CLAUDE.md, and you don't pay for those few tokens. You pay to recompute the entire cached prefix beneath it, on every request that carries it. One changed byte before a breakpoint takes every cached segment after it down with it.
On a single request the cost is pennies, sometimes less. Easy to wave away. But an agent doesn't send one request. It sends the same fat prefix thousands of times a day, and if a stray character busts the cache every turn, you re-pay every turn. It might not be much each time. It accumulates.
Worth doing the sum once. Say an agent carries a 20k-token cached prefix, turns over 2,000 times a day, and a single stray CRLF busts the cache each turn. On a mid-tier Claude model at roughly $3 per million input tokens, with cached reads about a tenth of that, the gap between hit and miss on that prefix runs to something near $100 a day. From one character, on one agent. That's a deliberately unkind worst case, and your real figure will be smaller and messier, but you can see the shape of it. Nobody notices, because it never shows up as a line item. It's the bill being quietly larger than it needed to be.
So we built the part that makes it visible, and where we safely can, the part that fixes it.
OmnisRouter now labels why each cache miss happened. It watches the traffic going through it, works out the cause of every miss, and tags it. This is content-free: it classifies the reason, it doesn't read or keep your prompts. It hands that up to OmnisVigil as a receipt.
OmnisVigil turns those receipts into a number you can act on, and it's honest about which misses count. A genuine edit to your prompt is a cache miss too, but it isn't waste, you meant to change it. Same for switching model or rewriting the system prompt. OmnisVigil separates those out and puts only the avoidable misses in the headline, attributed to the team, repo and commit that caused them. So the number you're looking at is the part you could actually recover, not a scary total you can't do anything about.
Three of the avoidable causes fix themselves. Line endings, trailing whitespace and tool ordering are safe to normalise without changing a thing the model reads, so you flip them on in OmnisVigil and OmnisRouter cleans up future requests on the way through. The misses turn back into hits. No code change, no tradeoff. The other avoidable causes (an injected timestamp, a volatile header, a concatenation order that drifts) need a one-off change at your end, so OmnisVigil flags those with the evidence rather than touching them silently.
I'm not going to quote you a savings figure. It's your workload, not ours, and measured beats claimed. What I'll say is that the dashboard shows you your own avoidable-waste number, and for most of it the fix is a toggle.
Routing is the big lever. Cache hygiene is the small one, and it's the bit still leaking after you've done the routing and told yourself you're optimised. Both are open. Both are measured. If your agent bill is bigger than you can explain, this is a cheap place to start.
- Route: OmnisRouter
- Govern: OmnisVigil
- Measure: OmnisBench

Top comments (1)
A cache hit only happens if the prefix stays identical, so a timestamp or a request id near the top of the system prompt quietly turns every call into a miss
Curious whether you track cache hit ratio per agent