You step away for lunch with a big Claude Code session open. You come back, type one message, and continue. Nothing looks different. But that one message may have cost 20 times more than the one before it.
I wanted a real number, so I scanned a month of my own Claude Code logs. The short answer: a typical break past the one-hour mark cost me about 125k–135k tokens of cold cache writes. Over the month, these restarts added about 15% to my input usage.
This post explains why it happens, shows the script I used, and lists what I changed.
How prompt caching works in Claude Code
Every turn in a Claude Code session sends the whole conversation to the model again: system prompt, tool definitions, project files you read, tool output, everything. A long session can carry 100k–400k tokens of context.
Prompt caching makes this affordable. After the first request, the stable prefix of the prompt is stored in a cache. Later requests read it from the cache instead of processing it again. In API pricing terms, relative to the base input price:
| Token type | Relative price |
|---|---|
| Uncached input | 1.0x |
| Cache write, 1-hour TTL | 2.0x |
| Cache read | 0.1x |
A warm turn reads the context at 0.1x. A cold turn writes the whole context again at 2.0x. That is where the roughly 20x comes from.
The cache has a time to live (TTL). In my Claude Code sessions it is one hour, and each request resets the timer. Tool calls count as requests, so an agent working on its own keeps the cache warm. The risk is when you are the slow part: a meeting, lunch, the end of the day.
Note: the 1-hour TTL is what my sessions used. The default API TTL is 5 minutes, and Claude Code can drop to 5 minutes in some cases, for example usage overage. The script below reports which TTL each write used, so you can check your own.
Where the data lives
Claude Code stores each session as a JSONL file under ~/.claude/projects/<project>/<session-id>.jsonl. Every assistant turn includes a usage block. Trimmed:
{
"type": "assistant",
"timestamp": "2026-10-01T13:28:05.993Z",
"message": {
"id": "msg_...",
"usage": {
"input_tokens": 2,
"cache_creation_input_tokens": 32912,
"cache_read_input_tokens": 0,
"output_tokens": 211,
"cache_creation": {
"ephemeral_1h_input_tokens": 32912,
"ephemeral_5m_input_tokens": 0
}
}
}
}
Three fields matter:
-
cache_read_input_tokens: context served from the cache (cheap). -
cache_creation_input_tokens: context written to the cache (expensive). -
cache_creation.ephemeral_1h_input_tokens/ephemeral_5m_input_tokens: which TTL the write used.
On a normal warm turn, reads are large and writes are small (just the new tool output). A cache miss looks the other way round: a large write, little or no read. If the previous turn in the same file was more than an hour earlier, the idle gap is the cause.
The script
The script walks every session file, removes duplicate messages by ID, and flags turns where:
-
cache_creation_input_tokensis at least a threshold (default 50,000), and - the write is larger than the read.
It records the gap since the previous turn, groups the misses by gap length, and estimates a weighted cost share using the prices above.
Click to see the full script (100 lines)
import json, glob, os, sys
from datetime import datetime, timedelta
from collections import defaultdict
ROOT = os.path.expanduser("~/.claude/projects")
MIN_MISS = int(sys.argv[1]) if len(sys.argv) > 1 else 50_000
def parse_ts(s):
return datetime.fromisoformat(s.replace("Z", "+00:00"))
misses = []
totals = defaultdict(int)
per_project = defaultdict(lambda: defaultdict(int))
for path in glob.glob(os.path.join(ROOT, "*", "*.jsonl")):
project = os.path.basename(os.path.dirname(path))
session = os.path.basename(path)[:8]
prev_ts = None
seen_ids = set()
with open(path, encoding="utf-8", errors="replace") as f:
for line in f:
try:
d = json.loads(line)
except json.JSONDecodeError:
continue
if d.get("type") != "assistant":
continue
msg = d.get("message") or {}
u = msg.get("usage")
ts = d.get("timestamp")
if not u or not ts:
continue
mid = msg.get("id")
if mid in seen_ids:
continue
seen_ids.add(mid)
t = parse_ts(ts)
read = u.get("cache_read_input_tokens", 0) or 0
created = u.get("cache_creation_input_tokens", 0) or 0
raw = u.get("input_tokens", 0) or 0
cc = u.get("cache_creation") or {}
w1h = cc.get("ephemeral_1h_input_tokens", 0) or 0
w5m = cc.get("ephemeral_5m_input_tokens", 0) or 0
totals["turns"] += 1
totals["read"] += read
totals["created"] += created
totals["raw"] += raw
per_project[project]["read"] += read
per_project[project]["created"] += created
gap = (t - prev_ts) if prev_ts else None
if created >= MIN_MISS and created > read:
misses.append(dict(project=project, session=session, ts=t,
gap=gap, created=created, read=read,
w1h=w1h, w5m=w5m))
prev_ts = t
misses.sort(key=lambda m: m["ts"])
def fmt_gap(g):
if g is None: return "start"
h = g.total_seconds() / 3600
return f"{h:6.1f}h" if h >= 1 else f"{g.total_seconds()/60:5.0f}m "
print(f"Cache misses >= {MIN_MISS:,} created tokens, read < created\n")
print(f"{'date':16} {'gap':8} {'created':>9} {'read':>8} {'ttl':4} {'project':28} sess")
buckets = defaultdict(lambda: [0, 0])
for m in misses:
ttl = "1h" if m["w1h"] >= m["w5m"] else "5m"
print(f"{m['ts'].strftime('%Y-%m-%d %H:%M'):16} {fmt_gap(m['gap']):8} {m['created']:>9,} {m['read']:>8,} {ttl:4} {m['project'][:28]:28} {m['session']}")
g = m["gap"]
key = ("session start" if g is None else
"<5m" if g < timedelta(minutes=5) else
"5m-1h" if g < timedelta(hours=1) else
"1h-24h" if g < timedelta(hours=24) else ">24h")
buckets[key][0] += 1
buckets[key][1] += m["created"]
print("\nMisses by idle gap before the request:")
for k in ["session start", "<5m", "5m-1h", "1h-24h", ">24h"]:
n, tok = buckets[k]
print(f" {k:14} {n:4} misses {tok:>12,} tokens written")
miss_tok = sum(m["created"] for m in misses)
print(f"\nAll assistant turns: {totals['turns']:,}")
print(f" cache read tokens : {totals['read']:>14,}")
print(f" cache created tokens : {totals['created']:>14,} ({miss_tok:,} in the misses listed above)")
print(f" uncached input tokens: {totals['raw']:>14,}")
# cost weights: read 0.1, 1h write 2.0, raw 1.0
w_read = totals['read'] * 0.1
w_created = totals['created'] * 2.0
w_raw = totals['raw'] * 1.0
w_miss = miss_tok * 2.0
tot = w_read + w_created + w_raw
print(f"\nWeighted input cost share (read=0.1, write=2.0, raw=1.0):")
print(f" cache reads : {w_read/tot*100:5.1f}%")
print(f" cache writes: {w_created/tot*100:5.1f}% of which listed misses: {w_miss/tot*100:5.1f}%")
print(f" raw input : {w_raw/tot*100:5.1f}%")
print("\nTop projects by cache-write tokens:")
for p, v in sorted(per_project.items(), key=lambda kv: -kv[1]["created"])[:8]:
print(f" {p[:40]:40} written {v['created']:>12,} read {v['read']:>14,}")
Run it with an optional threshold:
python cache_misses.py # default 50,000
python cache_misses.py 30000 # lower threshold
My results
The data covers 2026-08-31 to 2026-10-01: 2,410 assistant turns across five projects.
The misses
Restarts after an idle gap of over one hour, threshold 50k:
| Idle gap | Tokens re-written |
|---|---|
| 1.1h | 170k |
| 1.1h | 80k |
| 1.2h | 137k |
| 1.2h | 77k |
| 1.2h | 193k |
| 1.3h | 80k |
| 1.4h | 142k |
| 1.5h | 58k |
| 1.6h | 110k |
| 1.6h | 164k |
| 2.2h | 148k |
| 2.5h | 53k |
| 2.6h | 171k |
| 2.6h | 103k |
| 2.8h | 131k |
| 15.1h | 124k |
| 18.7h | 233k |
| 61.8h | 476k |
| Metric | Value |
|---|---|
| Idle restarts over 1h | 18 |
| Tokens re-written cold | 2.65M |
| Mean per restart | 147k |
| Median per restart | 134k |
| Median for gaps of 1.1h–1.6h (10 restarts) | 123k |
The mean is pulled up by one outlier: a session resumed after a 2.5-day weekend that re-wrote 476k tokens in one request. The median is the more honest "cost of a break".
The pattern
- Most gaps barely passed the limit. 10 of the 18 were between 1.1 and 1.6 hours. These are lunch breaks and meetings, not abandoned sessions.
- One stale session cost as much as three or four normal breaks. Resuming a huge session after days away is the worst case.
-
Session starts are a separate, fixed cost. At a 30k threshold the script also shows 14 session starts, each 31k–36k tokens. That is the startup context: system prompt,
CLAUDE.md, memory, skills, and MCP tool definitions. Every new session pays it once, and no timing habit removes it.
The cost share
Weighted by relative price across all 2,410 turns:
| Bucket | Tokens | Weighted share |
|---|---|---|
| Cache reads | 234M | 55% |
| Cache writes, normal incremental | 6.8M | 32% |
| Cache writes, idle misses | 2.7M | 13% |
| Uncached input | 82k | 0.2% |
Idle misses were 13% of the total weighted input. Put another way: without them, input cost would have been 87% of what it was, so the misses added about 15% on top.
The 32% for normal writes is the ordinary cost of new tool output going into the cache each turn. Timing does not change it. Only the 13% slice can be avoided.
What I changed
1. Come back within the hour when I can. Most of my misses were 5 to 35 minutes past the limit. If I know I will be back soon, I make the next message the first thing I do.
2. Run /compact before a long break. If I will be away for more than an hour and the session is large, I compact while the cache is still warm. The summary is cheap to produce on a warm cache. The restart after the break then re-writes a small summary, not 150k tokens of history, and every turn after that is cheaper too.
3. Do not revive stale sessions. For anything older than a day, I start a fresh session with a short brief: goal, current state, key files. That costs about 35k tokens of startup context, not 400k+ of old history.
4. Do not bother with keep-alive pings. Sending a "heartbeat" message every hour does keep the cache warm, and each one costs only about 0.1x of the context. On paper it wins for gaps up to about 20 hours (20 pings at 0.1x equal one cold write at 2.0x). In practice it adds noise to the conversation, it stops when the machine sleeps, and it keeps a large context alive that you may not need. /compact is simpler.
Caveats
- Plan differences. On a subscription plan, tokens count against usage limits, not a bill. The relative weights still show what eats your limits faster, but your plan may not count tokens exactly like API prices.
- The weights are an estimate. Read = 0.1, 1h write = 2.0, uncached = 1.0 follows published API pricing ratios. Output tokens are left out on purpose, because they do not depend on caching.
- The heuristic is simple. "Write > read and write ≥ 50k" catches idle restarts well. It also catches a few non-idle events, such as one write under 5 minutes in my data, so always look at the gap column.
- One user, one month. These are numbers from my own work: several infrastructure and app repos, long agentic sessions, normal office hours.
Try it on your own logs
Save the script, run it, and look at two things: the 1h-24h bucket and the >24h bucket. If they hold a large share of your cache writes, a habit change will save you more than any prompt tweak.
Over to you: is saving ~15% of tokens worth changing how you work? Or is the extra attention not worth it? I'd like to know how you handle long sessions and breaks.
These numbers are averages from one month of my own work. Your numbers could vary.
Top comments (0)