DEV Community

Rosehsu
Rosehsu

Posted on

I wanted to know which project was eating my Claude Code quota

When I hit my Claude Code 5-hour or weekly limit, I want to know which project used it up. /usage doesn't answer that. It gives account-wide percentages plus a few tags like ">150k context" or "4+ parallel sessions". Nothing is broken down per project, and I can't check those percentages against anything.

Claude Code already writes per-message token usage into every local session transcript under ~/.claude/projects/*/*.jsonl. So I wrote a plugin, quota-audit, that reads those files and breaks the usage down per project and per skill. It also shows the exact times a rate limit was hit and which repo was blocked. Everything stays on your machine. It only reads those files and sends nothing anywhere.

What the report looks like

It opens with a few computed findings. After that comes one overview row per window (5h, 24h, 7d), then detail for the longest window. Here's a trimmed example:

KEY FINDINGS (7d window)
  • Estimated cost $147.31, 65% of it in my-app ($95.20)
  • Top skill is ai-work at 54%
  • 11 sessions crossed 150k context, $88.40 combined (60% of window cost)
  • 3 rate-limit hits in 7d; 14 all-time, most often blocking my-app (9 times)

window     est. cost  sessions   >150k   hits   top project
5h            $12.40         5       2      0   my-app 81%
24h           $38.75        21       4      1   my-app 70%
7d           $147.31        99      11      3   my-app 65%
Enter fullscreen mode Exit fullscreen mode

Skills are attributed per message, not by whatever command a session started with, so the split between named skills and general interactive use holds up. Sessions that crossed 150k context are listed by name with the most expensive first, because that's usually the biggest lever. Rate-limit hits are grouped by repo and time of day.

There's also a check for polling loops, like a browser tab or script left open and calling something every minute or two. Those inflate session counts without spending real quota. To see every window's detail and every individual hit, pass --full to render.py.

Trying it

The current version is v1.3.0. Inside an interactive Claude Code session:

/plugin marketplace add rosehsu47/claude-code-quota-audit
/plugin install quota-audit@rosehsu47
Enter fullscreen mode Exit fullscreen mode

Or non-interactively:

claude plugin marketplace add rosehsu47/claude-code-quota-audit
claude plugin install quota-audit@rosehsu47
Enter fullscreen mode Exit fullscreen mode

If the install summary says Run /reload-plugins to activate, run that. After that, ask something like "who's eating my quota" or "which repo is burning through my 5-hour limit". You can also call /quota-audit directly. The report comes out in English or Traditional Chinese depending on your locale. Set QUOTA_AUDIT_LANG=en or QUOTA_AUDIT_LANG=zh to force one.

What it can't do

The dollar figures are estimates from a pricing snapshot bundled with the plugin. They are not your bill. I'd trust the ratios between projects more than the absolute total, and I'd check the total against a live claude -p "/usage". Per-project subagent cost isn't available. Like /usage, it only sees the machine it runs on, not other devices or claude.ai.

The code is at github.com/rosehsu47/claude-code-quota-audit, and I'd like to hear how other people figure out where their quota goes.

Top comments (2)

Collapse
 
deanlee profile image
Dean Lee •

Breaking down usage by local transcripts exposes the compounding cost of long context. In extended coding sessions, the rate limit hit is rarely driven by output tokens or message count alone. Once a conversation crosses 150k tokens, every minor bash check or file grep re-transmits the cumulative context window. Even with prompt caching discounts, the rolling five-hour token bucket meters total processed volume. A secondary factor worth parsing in those JSONL files is the split between cache read and cache creation tokens. When tool definitions change mid-run or terminal outputs fluctuate, cache invalidations force cold writes that consume allowance substantially faster than steady-state reads.

Collapse
 
rosehsu47 profile image
Rosehsu •

Thanks for reading! Small correction: the tool already reads cache read and cache write tokens separately from the transcripts, and prices each write at the TTL recorded for it (5m at 1.25x, 1h at 2x), so the cost estimates already include the cold-write penalty. Run it with --full and the totals show up.

What it doesn't do yet is break that down per session, and that's a good idea! I tried a quick version on my own data: most of the expensive rewrites came from coming back to a long session after more than an hour idle, once the cache had expired. I'm adding this to the report next.