Anthropic cut the Claude Sonnet 5.5 prompt-cache read rate on October 7, 2026, from $0.20 to $0.10 per million tokens. The change halves the price of a cache read, but it does not halve an application’s model bill. Your impact depends on how many cache-read tokens your traffic actually generates, how much new caching writes cost, and how often requests miss.
Start with the usage counters your application already receives. Compare the same workload before and after the rate change, and keep the calculation separate from output tokens, uncached input, retries, and any other price changes. That gives you a defensible read-cost delta without turning a vendor price cut into an unmeasured savings claim.
What changed on October 7
Anthropic’s release notes say Sonnet 5.5 cache reads moved from $0.20 to $0.10 per million tokens. The note also says cache writes and all other prices are unchanged. The new read rate is $0.10 per million tokens; the old rate was $0.20.
The direct difference is therefore $0.10 for each million cache-read tokens, assuming the same token volume and billing terms. This is a change to one line of usage. It is not a claim that your total Anthropic spend falls by 50 percent: output generation, uncached input, cache creation, request volume, retries, and other models remain separate parts of the bill.
For an existing workload, first calculate the counterfactual on the cache reads you actually recorded:
read_cost_delta = cache_read_tokens / 1,000,000 × ($0.20 - $0.10)
For example, if a hypothetical month contains 25 million Sonnet 5.5 cache-read tokens, that same read volume cost $5 at the former rate and $2.50 at the new rate. The $2.50 difference is arithmetic from the published rates, not a forecast for your application. Use your own counters and invoice to establish your result.
Find the calls that can benefit
The change matters most to a route that already gets cache hits on a substantial, repeated prefix. Examples include a stable system instruction, shared reference material, or tool definitions reused across requests. The important evidence is not how large a prompt looks in a source file. It is how many of its tokens are reported as cache reads after the request runs.
Do not add caching solely because reads are cheaper now. A first request can create cached content, and cache writes are unchanged by this price change. A prompt that changes on every call may create writes without earning enough later reads to offset them. Put stable content before request-specific data, and measure whether the exact route is reusing that prefix.
Start by grouping traffic by model ID and operation: for example, support classification, code review, or a repository assistant’s repeated project context. Keep the grouping stable across the comparison. If you combine Sonnet 5.5 with other models, or combine unrelated request types, a workload shift can look like a price effect.
Log reads, writes, and ordinary tokens separately
The response usage object distinguishes cache reads from cache creation and ordinary input. Persist those counters with the model ID, route, timestamp, request outcome, and an internal request class. Avoid storing customer prompts or secrets just to calculate billing. Aggregated counters are enough for the price comparison.
A small helper can calculate the read line and the rate-change delta from one response:
type Usage = {
cache_read_input_tokens?: number;
cache_creation_input_tokens?: number;
input_tokens: number;
output_tokens: number;
};
function sonnet55CacheReadCosts(usage: Usage) {
const readTokens = usage.cache_read_input_tokens ?? 0;
return {
readTokens,
writeTokens: usage.cache_creation_input_tokens ?? 0,
currentReadUsd: (readTokens / 1_000_000) * 0.10,
previousRateReadUsd: (readTokens / 1_000_000) * 0.20,
rateChangeDeltaUsd: (readTokens / 1_000_000) * 0.10,
};
}
Use this as an estimate of the cache-read component at the published rates, not as a substitute for the billed total. Retain the raw counters so you can recompute the comparison if the rate or a negotiated billing term changes. Also keep cache creation and regular input in the report: a larger read saving can coexist with higher write or uncached-input spend.
Compare like with like
Choose a representative window for one stable operation, then compare it with a window containing the same operation and a similar request mix. Record the number of completed calls, cache-read tokens, cache-creation tokens, ordinary input tokens, output tokens, errors, and retries. If traffic volume differs, report both totals and per-call averages; do not attribute a change in request count to the new read rate.
For the cleanest check, replay an owned, non-sensitive sample through the same prompt and model configuration, or use production aggregates without retaining prompt contents. Keep the model, prompt prefix, cache settings, output limit, and application behavior fixed. If you change a breakpoint, TTL, model, or prompt at the same time, run that as a separate experiment. Otherwise you cannot tell whether the read-price change or your implementation caused the result.
Then reconcile the token calculation with the provider’s billed usage for the same period. Differences can come from discounts, platform billing, adjustments, or traffic that your application did not tag. Keep the measured result narrow: “cache-read spend changed by this amount on this route” is supportable; “the application is now cheaper by this percentage” requires the full bill and a stable denominator.
Decide whether to change the cache
If your route already has frequent cache reads, update your cost model with the new rate and leave a working cache configuration alone until the data suggests a change. If reads are rare, investigate misses and cache creation before expecting this price cut to move the total. If the request is mostly one-off content, the price announcement does not make it a caching candidate.
For a proposed cache change, test one route at a time. Confirm the response reports reads on repeated requests, check that the cached prefix is actually shared, and inspect the change in writes and ordinary input. Preserve correctness and latency checks alongside spend; a cache configuration that raises stale-context risk or breaks the request is not a saving.
Keep a dated before-and-after note with the raw counters, rate used, calculation, and billing source. When your change is a prompt or harness adjustment rather than cache pricing, use the separate fixed-task measurement guide to keep that comparison honest.
Keep the rate effect separate from the total bill
The price change applies to eligible Sonnet 5.5 cache-read usage; you do not need to redesign a working prompt to receive that lower rate. Separate two questions in your report. First, what would the same observed cache-read tokens cost at the old and new prices? That isolates the rate effect. Second, how did the route’s complete billed cost change over the comparison window? That includes traffic, writes, uncached input, output, retries, and any platform-specific terms.
If the total fell by more than the cache-read calculation, check whether request volume or output tokens also fell. If it fell by less, inspect write tokens, cache misses, or a changed request mix before concluding that the published rate did not apply. Keep these as separate lines in a dashboard or review note so a lower cache-read rate does not become a claim about overall cost without evidence.
For a rollout decision, keep the test reversible: preserve the old route configuration, compare one operation at a time, and set a threshold using your own correctness and latency requirements. The price cut changes the arithmetic; your measured workload decides whether a cache adjustment is worthwhile.
Sources
- Anthropic Claude Platform release notes — October 7, 2026 entry for the Sonnet 5.5 prompt-cache read price change and unchanged cache-write and other prices.


Top comments (2)
I like how this separates the cache-read price reduction from actual application-wide savings. Tracking reads, writes, retries, and per-call costs separately makes the results much more trustworthy than simply claiming a 50% bill reduction.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.