How I Built a Unified Cost Tracker for AI Coding Agents
AI coding agents are amazing — until the bill arrives. And it will arrive from three places at once: Claude Code in one directory, Cursor in another, Codex in a third. Each logs usage in its own JSONL format, with different field names and different pricing models.
So I built a small local tool that reads all three, normalizes them, and produces one cost report — by model, by day, by agent — with budget guardrails that warn at 80% and flag at 100%. Here's what it took.
The Hard Part: Three Different Log Formats
Each agent logs usage differently:
// Claude Code — usage nested under message, with cache fields
{"type":"assistant","message":{"model":"claude-sonnet-4-20250514",
"usage":{"input_tokens":523,"output_tokens":187,"cache_read_input_tokens":1200}},
"timestamp":"2026-08-01T10:00:00.000Z"}
// Codex CLI — usage nested under payload
{"type":"response_item","payload":{"type":"message","model":"gpt-5-codex",
"usage":{"input_tokens":2000,"output_tokens":800}}}
// Cursor — different field names entirely
{"type":"assistant","message":{"model":"gpt-4o",
"usage":{"prompt_tokens":800,"completion_tokens":300}}}
The parser can't assume a shape. Instead of three hand-written parsers, I wrote one recursive extractor that walks the object tree, finds any usage object, and pulls the tokens with fallbacks:
function extractUsage(obj, out, parentTs) {
if (!obj || typeof obj !== 'object') return;
if (obj.usage) {
out.push({
model: obj.model || obj.message?.model || '',
inputTokens: obj.usage.input_tokens ?? obj.usage.prompt_tokens ?? 0,
outputTokens: obj.usage.output_tokens ?? obj.usage.completion_tokens ?? 0,
cachedReadTokens: obj.usage.cache_read_input_tokens ?? 0,
ts: parentTs || obj.timestamp || ''
});
return;
}
for (const k of Object.keys(obj)) {
if (['usage','message','payload','request'].includes(k)) {
extractUsage(obj[k], out, obj.timestamp);
}
}
}
That handles all three formats (and future ones) with one code path.
The Pricing Engine: 23 Providers, Cache Discounts
The real insight: you don't need per-model accuracy to be useful — you need good enough to spot runaway spend. I reuse a pricing table of 23 providers (Anthropic, OpenAI, DeepSeek, Google, etc.), match by exact name or prefix, and fall back to a provider-name guess for unknown models:
function estimateCost({ model, inputTokens, outputTokens, cachedReadTokens = 0 }) {
const price = lookupModel(model); // $/1K tokens
const cached = Number(cachedReadTokens) || 0;
// cache reads are ~10x cheaper than fresh input
const cost = (inputTokens * price.input
+ cached * price.input * 0.1
+ outputTokens * price.output) / 1000;
return { cost, estimated: !!price.estimated };
}
Unknown models get an estimated rate (flagged in the report), so nothing silently disappears from the total.
Budget Guardrails
The killer feature isn't the report — it's the limit. Set a monthly budget, and the tool tells you where you stand:
- < 50% — OK
- ≥ 80% — WARN (you'll overspend at this rate)
- ≥ 100% — OVERRUN
Guardrail: $0.0698 / $50 (0%) → OK
Two Forms: CLI + MCP
Developers live in terminals, but agents live in the MCP ecosystem. So the tool ships as both:
-
CLI:
agentcost scan ~/.claude/projects 50→ terminal report -
MCP server: exposes
scan_cost,get_budget,set_budgettools — so your coding agent can answer "how much have I spent this month?" itself
What I Learned
-
Field-name drift is the real cost —
input_tokensvsprompt_tokensvsinputTokens. One recursive extractor beats three parsers. - Cache accounting matters — Claude's cached reads are ~10x cheaper; ignoring them overstates costs badly on long sessions.
- Estimates beat silence — an unknown model that shows up as "$0" teaches users to ignore the tool. Flag it as estimated instead.
- Local-first is a feature — "nothing leaves your machine" is a privacy story users actually care about, especially for cost data.
Status
The project is in early release (v0.1, 22/22 tests passing) and available on npm: agentcost-cli. This is my first open-source tool from the SellerTools family, and I'm looking for feedback from heavy Claude Code / Cursor / Codex users: what's missing for your workflow? Budgets per project? Team aggregation? Anomaly alerts? Tell me in the comments.
Top comments (2)
A unified tracker is useful when it normalizes the unit of work, not only the token bill. Comparing tools by raw cost can be misleading if one burns tokens exploring and another burns human time recovering from weaker context.
A unified cost view gets valuable when it explains behavior, not just spend. The useful split is usually planning, context reloads, tool failures, retries, and actual implementation work.