DEV Community

Ivan
Ivan

Posted on Originally published at ivanped.ai

How I verified prompt-cache reads in a generation-repair workflow

The cache was receiving writes without delivering reuse

I found the caching problem in Launcherry's production usage records: repeated generation calls wrote prompt tokens to cache but read nothing back within the run. The Google Search copy bank and analysis records showed the same pattern. I needed to verify what later calls reused through the application, rather than rely on the fact that caching was configured.

I treated that as an operating question that needed evidence. The relevant outcome was whether later calls actually reused suitable context. A cache setting, a successful response or a lower estimate in a spreadsheet could not establish that behaviour.

The September 17 verification records the investigation and correction. In each of two later generation-repair calls, 3,879 of 3,932 input tokens were read from cache. That is 98.65%, rounded to 99%. The result describes prompt-token reuse in those measured calls, not a product-wide reuse rate or a 99% cost saving.

Ask what is being reused and what is changing

Generation and repair have a useful relationship for this investigation. They concern the same underlying work, while a repair asks for a specific correction. Some context can remain relevant across those calls; other parts need to change. The request design has to preserve the distinction.

I examined the behaviour through the application’s own adapter and the production provider path. The recorded probe compared requests that held relevant conditions steady with requests that changed them. That helped establish which changes interfered with reuse, rather than treating every cache miss as the same problem.

My starting point is the request the product actually sends. Two prompts that seem similar to a human can differ in ways that matter to the provider’s cache. The model, request configuration and surrounding adapter behaviour belong in the investigation alongside the written instructions.

Provider conditions also differ. OpenRouter’s prompt-caching documentation distinguishes cached-token reads from cache writes and describes provider-specific behaviour. I use that documentation to frame a check, then rely on the application’s measured response to establish its result.

Keep the correction bounded by product requirements

I directed a correction to the generation request design so repeated repair work could reuse its shared context. The work retained the generation requirements and validation responsibilities. Reuse needed to support the product’s existing job, rather than become a reason to weaken its checks.

The implementation was accompanied by regression cases covering the relevant request behaviour and the relationship between draft and repair. Those checks provide repeatable evidence about how the application constructs its work. A live rerun then provided the evidence about cache reads.

I keep those checks separate. A unit test can establish the application’s request structure under its test conditions. It cannot prove that a remote provider returned a cache hit. Conversely, one provider hit cannot establish that all the surrounding application paths remain correct. The investigation needed both kinds of evidence.

Read the measured result without inflating it

The recorded rerun includes preparation calls before the two successful repair reads. Those earlier calls wrote cache entries. Repairs 2 and 3 then each read 3,879 of 3,932 input tokens from cache and wrote no new cache tokens. The numerator and denominator belong to each of those calls.

I calculate the ratio by dividing 3,879 by 3,932 and multiplying by 100. The resulting 98.65% can be rounded to 99% for a compact outcome label, provided the sample stays visible. Reporting the denominator makes the label inspectable.

It would be misleading to apply that ratio to the whole generation run. The earlier cache writes are part of the run, and the recorded ratio concerns input tokens rather than every billed token. A broader claim would require broader measurement.

I also avoid equating reuse with latency improvement. The cache evidence establishes reads. A latency claim needs timing measurements under relevant conditions. The same discipline applies to claims about throughput, margins or customer savings.

Connect caching to economics without confusing the measures

AI usage is an operating cost in Launcherry. Caching can affect that cost, but total economics also depend on the provider’s rates, output generation, the number of calls and the amount of repair required. A highly reused input can still accompany an unhelpful result or an expensive sequence.

My acceptance decision therefore keeps quality and economics connected but separately measured. Output evaluation assesses the deliverable. Production usage establishes the cost and cache behaviour observed in that environment. The workflow architecture determines which work happens and how often.

For a product team considering optimisation, I would begin with a complete user outcome and its actual usage records. Locate repeated work, identify reusable context and measure the effect of a bounded change. Preserve the original quality requirements so a cheaper run cannot pass simply by doing less useful work.

Require observed reuse before claiming an optimisation

This investigation changed how I describe the result. I can point to a specific diagnosis, a request-design correction, regression coverage and live cache-read evidence. I cannot use those two calls to establish a universal saving, and I do not need that larger claim to explain the value of the work.

I want to be able to trace an operating assumption into application behaviour, test it and correct it. It connects AI workflow construction to the cost of delivering the product, while keeping the evidence understandable.

The Launcherry case study carries the compact outcome. This article preserves the scope behind it. For another system, I would ask the same concrete question before calling caching successful: which later calls read cached context, how much did they read and what did the complete run cost to produce an acceptable result?

Originally published at ivanped.ai: Prompt caching in practice: Diagnosing writes without reuse. More of my work: ivanped.ai

Top comments (0)