I kept hitting Claude Code's weekly limit, so before keeping any token-saving tool I now measure it on a fixed task set.
How I measured
- 5 fixed tasks: fix a bug, scan a log, check git changes, summarize a CSV, summarize test results
- 3 runs per task, same day, same test repo,
claude -p --output-format json - Cost-weighted tokens = input 1, cache write 1.25, cache read 0.1, output 5
- Answer quality counted as part of correctness
Results
| Tool | What it does | Cost-weighted tokens vs baseline | Correct |
|---|---|---|---|
| rtk v0.51.0 | compresses shell output via a hook | +7.1% | 15/15 |
| caveman v3.1.0 | makes replies ultra-terse | +10.9% | 15/15 |
| pxpipe v0.14.0 | renders bulky context as images via a local proxy | +196.7% | 15/15 |
Why all three went up
- Fewer tokens per step can mean more steps. rtk trimmed output by ~63%, but Claude ran extra commands to see what was cut. Each turn re-reads the context, so cache reads rose 19%.
- Anything added to the system prompt is paid on every call. caveman's rules (~2,900 tokens) cost more than the 12% output it saved.
- Images are not free. pxpipe's rendered context cost more than the original text; cache writes on the first call went from ~27K to ~102K.
Where my tokens actually go
About 42% of my cost is the always-loaded context (system prompt, instructions, tool list). Command output is about 6%.
Raw numbers and method: https://github.com/awesomefred0827-hash/claude-token-log
Next: vendor and patch
I'm thinking of vendoring all three (pulling the source and patching it myself):
- rtk: keep error lines uncompressed
- caveman: one "be brief" rule, no skill list
- pxpipe: compress only past a context-size threshold
What do you think? Would these patches work, or would you fix them differently? I'll re-test whatever you suggest and post the results.
Top comments (0)