DEV Community

Cover image for How Much Does Your Lunch Break Cost in Claude Code Tokens?
vadim albarov
vadim albarov

Posted on

How Much Does Your Lunch Break Cost in Claude Code Tokens?

You step away for lunch with a big Claude Code session open. You come back, type one message, and continue. Nothing looks different. But that one message may have cost 20 times more than the one before it.

I wanted a real number, so I scanned a month of my own Claude Code logs. The short answer: a typical break past the one-hour mark cost me about 125k–135k tokens of cold cache writes. Over the month, these restarts added about 15% to my input usage.

This post explains why it happens, shows the script I used, and lists what I changed.

How prompt caching works in Claude Code

Every turn in a Claude Code session sends the whole conversation to the model again: system prompt, tool definitions, project files you read, tool output, everything. A long session can carry 100k–400k tokens of context.

Prompt caching makes this affordable. After the first request, the stable prefix of the prompt is stored in a cache. Later requests read it from the cache instead of processing it again. In API pricing terms, relative to the base input price:

Token type Relative price
Uncached input 1.0x
Cache write, 1-hour TTL 2.0x
Cache read 0.1x

A warm turn reads the context at 0.1x. A cold turn writes the whole context again at 2.0x. That is where the roughly 20x comes from.

The cache has a time to live (TTL). In my Claude Code sessions it is one hour, and each request resets the timer. Tool calls count as requests, so an agent working on its own keeps the cache warm. The risk is when you are the slow part: a meeting, lunch, the end of the day.

Note: the 1-hour TTL is what my sessions used. The default API TTL is 5 minutes, and Claude Code can drop to 5 minutes in some cases, for example usage overage. The script below reports which TTL each write used, so you can check your own.

Where the data lives

Claude Code stores each session as a JSONL file under ~/.claude/projects/<project>/<session-id>.jsonl. Every assistant turn includes a usage block. Trimmed:

{
  "type": "assistant",
  "timestamp": "2026-10-01T13:28:05.993Z",
  "message": {
    "id": "msg_...",
    "usage": {
      "input_tokens": 2,
      "cache_creation_input_tokens": 32912,
      "cache_read_input_tokens": 0,
      "output_tokens": 211,
      "cache_creation": {
        "ephemeral_1h_input_tokens": 32912,
        "ephemeral_5m_input_tokens": 0
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

Three fields matter:

  • cache_read_input_tokens: context served from the cache (cheap).
  • cache_creation_input_tokens: context written to the cache (expensive).
  • cache_creation.ephemeral_1h_input_tokens / ephemeral_5m_input_tokens: which TTL the write used.

On a normal warm turn, reads are large and writes are small (just the new tool output). A cache miss looks the other way round: a large write, little or no read. If the previous turn in the same file was more than an hour earlier, the idle gap is the cause.

The script

The script walks every session file, removes duplicate messages by ID, and flags turns where:

  • cache_creation_input_tokens is at least a threshold (default 50,000), and
  • the write is larger than the read.

It records the gap since the previous turn, groups the misses by gap length, and estimates a weighted cost share using the prices above.

Click to see the full script (100 lines)
import json, glob, os, sys
from datetime import datetime, timedelta
from collections import defaultdict

ROOT = os.path.expanduser("~/.claude/projects")
MIN_MISS = int(sys.argv[1]) if len(sys.argv) > 1 else 50_000

def parse_ts(s):
    return datetime.fromisoformat(s.replace("Z", "+00:00"))

misses = []
totals = defaultdict(int)
per_project = defaultdict(lambda: defaultdict(int))

for path in glob.glob(os.path.join(ROOT, "*", "*.jsonl")):
    project = os.path.basename(os.path.dirname(path))
    session = os.path.basename(path)[:8]
    prev_ts = None
    seen_ids = set()
    with open(path, encoding="utf-8", errors="replace") as f:
        for line in f:
            try:
                d = json.loads(line)
            except json.JSONDecodeError:
                continue
            if d.get("type") != "assistant":
                continue
            msg = d.get("message") or {}
            u = msg.get("usage")
            ts = d.get("timestamp")
            if not u or not ts:
                continue
            mid = msg.get("id")
            if mid in seen_ids:
                continue
            seen_ids.add(mid)
            t = parse_ts(ts)
            read = u.get("cache_read_input_tokens", 0) or 0
            created = u.get("cache_creation_input_tokens", 0) or 0
            raw = u.get("input_tokens", 0) or 0
            cc = u.get("cache_creation") or {}
            w1h = cc.get("ephemeral_1h_input_tokens", 0) or 0
            w5m = cc.get("ephemeral_5m_input_tokens", 0) or 0
            totals["turns"] += 1
            totals["read"] += read
            totals["created"] += created
            totals["raw"] += raw
            per_project[project]["read"] += read
            per_project[project]["created"] += created
            gap = (t - prev_ts) if prev_ts else None
            if created >= MIN_MISS and created > read:
                misses.append(dict(project=project, session=session, ts=t,
                                   gap=gap, created=created, read=read,
                                   w1h=w1h, w5m=w5m))
            prev_ts = t

misses.sort(key=lambda m: m["ts"])

def fmt_gap(g):
    if g is None: return "start"
    h = g.total_seconds() / 3600
    return f"{h:6.1f}h" if h >= 1 else f"{g.total_seconds()/60:5.0f}m "

print(f"Cache misses >= {MIN_MISS:,} created tokens, read < created\n")
print(f"{'date':16} {'gap':8} {'created':>9} {'read':>8} {'ttl':4} {'project':28} sess")
buckets = defaultdict(lambda: [0, 0])
for m in misses:
    ttl = "1h" if m["w1h"] >= m["w5m"] else "5m"
    print(f"{m['ts'].strftime('%Y-%m-%d %H:%M'):16} {fmt_gap(m['gap']):8} {m['created']:>9,} {m['read']:>8,} {ttl:4} {m['project'][:28]:28} {m['session']}")
    g = m["gap"]
    key = ("session start" if g is None else
           "<5m" if g < timedelta(minutes=5) else
           "5m-1h" if g < timedelta(hours=1) else
           "1h-24h" if g < timedelta(hours=24) else ">24h")
    buckets[key][0] += 1
    buckets[key][1] += m["created"]

print("\nMisses by idle gap before the request:")
for k in ["session start", "<5m", "5m-1h", "1h-24h", ">24h"]:
    n, tok = buckets[k]
    print(f"  {k:14} {n:4} misses  {tok:>12,} tokens written")

miss_tok = sum(m["created"] for m in misses)
print(f"\nAll assistant turns: {totals['turns']:,}")
print(f"  cache read tokens    : {totals['read']:>14,}")
print(f"  cache created tokens : {totals['created']:>14,}  ({miss_tok:,} in the misses listed above)")
print(f"  uncached input tokens: {totals['raw']:>14,}")
# cost weights: read 0.1, 1h write 2.0, raw 1.0
w_read = totals['read'] * 0.1
w_created = totals['created'] * 2.0
w_raw = totals['raw'] * 1.0
w_miss = miss_tok * 2.0
tot = w_read + w_created + w_raw
print(f"\nWeighted input cost share (read=0.1, write=2.0, raw=1.0):")
print(f"  cache reads : {w_read/tot*100:5.1f}%")
print(f"  cache writes: {w_created/tot*100:5.1f}%   of which listed misses: {w_miss/tot*100:5.1f}%")
print(f"  raw input   : {w_raw/tot*100:5.1f}%")

print("\nTop projects by cache-write tokens:")
for p, v in sorted(per_project.items(), key=lambda kv: -kv[1]["created"])[:8]:
    print(f"  {p[:40]:40} written {v['created']:>12,}  read {v['read']:>14,}")
Enter fullscreen mode Exit fullscreen mode

Run it with an optional threshold:

python cache_misses.py          # default 50,000
python cache_misses.py 30000    # lower threshold
Enter fullscreen mode Exit fullscreen mode

My results

The data covers 2026-08-31 to 2026-10-01: 2,410 assistant turns across five projects.

The misses

Restarts after an idle gap of over one hour, threshold 50k:

Idle gap Tokens re-written
1.1h 170k
1.1h 80k
1.2h 137k
1.2h 77k
1.2h 193k
1.3h 80k
1.4h 142k
1.5h 58k
1.6h 110k
1.6h 164k
2.2h 148k
2.5h 53k
2.6h 171k
2.6h 103k
2.8h 131k
15.1h 124k
18.7h 233k
61.8h 476k
Metric Value
Idle restarts over 1h 18
Tokens re-written cold 2.65M
Mean per restart 147k
Median per restart 134k
Median for gaps of 1.1h–1.6h (10 restarts) 123k

The mean is pulled up by one outlier: a session resumed after a 2.5-day weekend that re-wrote 476k tokens in one request. The median is the more honest "cost of a break".

The pattern

  • Most gaps barely passed the limit. 10 of the 18 were between 1.1 and 1.6 hours. These are lunch breaks and meetings, not abandoned sessions.
  • One stale session cost as much as three or four normal breaks. Resuming a huge session after days away is the worst case.
  • Session starts are a separate, fixed cost. At a 30k threshold the script also shows 14 session starts, each 31k–36k tokens. That is the startup context: system prompt, CLAUDE.md, memory, skills, and MCP tool definitions. Every new session pays it once, and no timing habit removes it.

The cost share

Weighted by relative price across all 2,410 turns:

Bucket Tokens Weighted share
Cache reads 234M 55%
Cache writes, normal incremental 6.8M 32%
Cache writes, idle misses 2.7M 13%
Uncached input 82k 0.2%

Idle misses were 13% of the total weighted input. Put another way: without them, input cost would have been 87% of what it was, so the misses added about 15% on top.

The 32% for normal writes is the ordinary cost of new tool output going into the cache each turn. Timing does not change it. Only the 13% slice can be avoided.

What I changed

1. Come back within the hour when I can. Most of my misses were 5 to 35 minutes past the limit. If I know I will be back soon, I make the next message the first thing I do.

2. Run /compact before a long break. If I will be away for more than an hour and the session is large, I compact while the cache is still warm. The summary is cheap to produce on a warm cache. The restart after the break then re-writes a small summary, not 150k tokens of history, and every turn after that is cheaper too.

3. Do not revive stale sessions. For anything older than a day, I start a fresh session with a short brief: goal, current state, key files. That costs about 35k tokens of startup context, not 400k+ of old history.

4. Do not bother with keep-alive pings. Sending a "heartbeat" message every hour does keep the cache warm, and each one costs only about 0.1x of the context. On paper it wins for gaps up to about 20 hours (20 pings at 0.1x equal one cold write at 2.0x). In practice it adds noise to the conversation, it stops when the machine sleeps, and it keeps a large context alive that you may not need. /compact is simpler.

Caveats

  • Plan differences. On a subscription plan, tokens count against usage limits, not a bill. The relative weights still show what eats your limits faster, but your plan may not count tokens exactly like API prices.
  • The weights are an estimate. Read = 0.1, 1h write = 2.0, uncached = 1.0 follows published API pricing ratios. Output tokens are left out on purpose, because they do not depend on caching.
  • The heuristic is simple. "Write > read and write ≥ 50k" catches idle restarts well. It also catches a few non-idle events, such as one write under 5 minutes in my data, so always look at the gap column.
  • One user, one month. These are numbers from my own work: several infrastructure and app repos, long agentic sessions, normal office hours.

Try it on your own logs

Save the script, run it, and look at two things: the 1h-24h bucket and the >24h bucket. If they hold a large share of your cache writes, a habit change will save you more than any prompt tweak.

Over to you: is saving ~15% of tokens worth changing how you work? Or is the extra attention not worth it? I'd like to know how you handle long sessions and breaks.

These numbers are averages from one month of my own work. Your numbers could vary.

Top comments (0)