You've done it too: pasted a 40k-token debug log into a conversation, watched your context budget evaporate, and spent the rest of the session fighting the model for attention. The paste felt free. It wasn't.
tokw weighs your files in tokens before you paste them in — so you know exactly what each file costs your budget, ahead of time.
git clone https://github.com/hahahahahahahahah6/token-weigh
cd token-weigh
python3 tokw.py -r ./src --ext .py,.md
path chars est. tokens
------------------------------------------------
src/legacy_parser.py 48210 12052 << over 10% of budget
src/handlers.py 31508 7877
docs/design.md 12440 3110
src/utils.py 8902 2226
------------------------------------------------
TOTAL: 25265 est. tokens = 12.6% of budget (200000)
The estimate is a documented heuristic, not a tokenizer: each CJK character ≈ 1 token, every other character ≈ 1/4 token — which mirrors how modern tokenizers actually behave on mixed content. It won't match any specific model's exact count; it tells you the order of magnitude, which is what you need to decide what goes in. Set a budget with --budget 200000 --json for machine-readable output. It walks directories recursively (skipping hidden dirs, __pycache__, node_modules, .git, and binary files), and the exit code is always 0 — it's informational.
Same household as mcp-tax, different axis: mcp-tax weighs MCP server schemas — the hidden context tax your agent pays every turn for tools it never calls. tokw weighs your files — the explicit cost of what you paste or attach. One audits what the agent brings to the table; the other audits what you bring. Run both and you know where your whole context budget actually goes.
Stdlib only, Python 3.9+. 9/9 tests pass. MIT licensed.
Repo: https://github.com/hahahahahahahahah6/token-weigh
What's the biggest thing you've ever pasted into a chat and immediately regretted?
Top comments (0)