Your AI agent ran 47 times yesterday. Cost stayed flat, right on budget—until execution #34, which cost 8 times the daily average. The logs say it succeeded. The output looks reasonable. But something quietly broke, and most monitoring misses it until the bill arrives.
The gap is token accounting.
What tokens actually tell you
Every LLM API call returns three distinct numbers:
- Input tokens — the prompt + context fed into the model
- Output tokens — the response the model generated
- Total tokens — input + output, the unit that gets billed
Most teams watch total tokens. Some teams set a cost budget and call it done. But token cardinality—the shape of the input/output ratio—is where silent failures announce themselves.
Here's the pattern:
Your agent normally runs with 800 input tokens and 200 output tokens per execution. One run jumps to 800 input and 1,600 output. Cost spikes proportionally. The logs show success—no error, no timeout, no warning. But the output token explosion is a diagnostic signal: something made the model generate 8x more text than it should have.
Why?
- A tool-use loop where the model keeps calling a tool without hitting the termination condition
- A prompt that accidentally triggered verbose output (a jailbreak, or a prompt injection in user input)
- A degraded model response that requires follow-up calls, cascading the token count
- Retrieval-augmented generation (RAG) pulling in massive context by mistake
None of these are visible in execution logs. All of them show up first in token cardinality.
The three failure patterns token accounting catches
Pattern 1: Output inflation without visible errors
An agent that normally outputs 150 tokens suddenly outputs 1,200 tokens. The API call succeeded. The response field is populated. But the model generated verbosity instead of signal—maybe a confused response that tried to explain itself at length, maybe a jailbreak that forced the model to keep writing.
Token cost spiked. The failure is silent because there was no exception, no HTTP error, no timeout. A monitoring system that only watches success/failure rates misses it entirely.
Pattern 2: Repetitive tool calls (the loop without exit)
A tool-use agent calls a tool, gets a response, calls the same tool again—five times. Each call adds input tokens (the previous responses now part of the context), and each adds output tokens. The total token spend becomes 3x normal, but the execution still reports success because the final loop iteration returned a valid response.
This is a silent failure compounded by token accounting visibility. The agent didn't error. It just worked inefficiently and burned budget.
Pattern 3: Cumulative token bleed across chained calls
Some agents split work across multiple API calls: first call to gather context, second call to reason, third call to act. If the first call returns something unexpectedly large (a context window that got polluted, an embedding search that returned 500 results instead of 5), all downstream calls inherit that bloat.
On execution 1: 400 + 400 + 400 = 1,200 total.
On execution 2: 400 + 400 + 400 = 1,200 total.
On execution 3: 4,000 + 400 + 400 = 4,800 total.
The third execution cost 4x more, but the logs show three successful API calls in sequence. Without token cardinality breakdown, you never see that the first call was the culprit.
Why baseline token spend is your fastest diagnostic
Most monitoring systems alert on cost thresholds: "alert if execution costs >$5." That's late—by then the damage is done. Baseline-aware token monitoring is earlier.
Here's what it looks like:
Establish the baseline — over the first 30 days, collect every execution's input tokens, output tokens, and total cost. Compute the mean, standard deviation, and a realistic ceiling (mean + 2 sigma).
Watch for divergence — on each new execution, compare its token counts to the baseline. If output tokens spike 3+ standard deviations above the mean, or if the input/output ratio inverts wildly from the historical pattern, flag it immediately.
-
Attribute the failure class — token patterns tell you what went wrong faster than reading logs:
- Output inflation → model degradation or jailbreak
- Input growth → context bloat or retrieval overfetch
- Cumulative bleed across calls → earlier call exceeded its budget
Kill the agent before it loops — if an agent starts exhibiting token patterns consistent with repetitive tool calls (same input token count, output grows, pattern repeats), you can pause it before the fourth iteration burns 10x budget.
On https://agents.opsveritas.com, this is built into the platform: every agent execution shows input, output, and total tokens, and a cost anomaly detection watches for the 3-sigma spike on that agent's own 30-day baseline. You can set a cost anomaly trigger so the agent pauses automatically on the third consecutive execution that exceeds the baseline.
The point isn't the specific numbers—it's that token cardinality is faster feedback than cost alone, and cost baselines are faster feedback than static thresholds.
The mechanics: why it matters for your agent
Let's walk through a real pattern:
Your agent scores leads. It normally takes a 4-line lead record, generates 3-5 bullets of reasoning, and returns a score 0–100. Tokens: 120 input, 60 output, 180 total. Cost: $0.0007 per execution.
One day, execution #42 comes through. Same lead record. The agent runs. Output arrives. Success. Cost: $0.005.
Token breakdown: 800 input, 4,200 output, 5,000 total.
What happened?
The lead record that day included a very long company description (a PDF copy-pasted as text). The agent's retrieval step pulled in all of it as context. Then, when the agent tried to reason about the lead, it generated an explanation for every fact, creating massive output. No error. No timeout. Just silent inefficiency.
Without token cardinality visibility, you'd see the cost spike and start debugging the lead record, the model, the prompt. With token accounting, you see immediately that input tokens jumped from 120 to 800—the context got bloated—and output tokens followed. The fix is clear: trim the retrieval step, not the prompt.
This is diagnostic signal in its purest form: the token pattern tells you the failure class before you read a line of code.
Where token accounting breaks down
Token accounting is powerful for cost signal, but it has limits:
It doesn't catch semantic failures — an agent that generates well-formed, token-efficient output that's wrong won't show up in cardinality data. A hallucination that uses exactly the expected number of tokens is silent to this signal.
It doesn't distinguish between inefficiency and correctness — an agent that generates long output may be verbose, or it may be providing necessary detail. Token growth doesn't tell you whether the extra tokens added value.
Baselines drift — if your agent's behavior naturally changes (more complex leads, richer output required, evolved prompts), the old baseline becomes noise. Baseline-aware systems need retraining windows or admin overrides.
The signal is cost-as-failure-early-warning, not cost-as-correctness-guarantee.
The pattern in summary
Token accounting works because it makes invisible inefficiency visible. An agent that loops silently, a retrieval step that over-fetches, a prompt that was accidentally jailbroken—none of these create error logs. They create token patterns.
If you're building or operating an LLM agent:
- Instrument every API call to capture input, output, and total tokens.
- Baseline that agent's token cardinality over its first month of real execution.
- Alert on divergence, not on static thresholds.
- Use the token pattern to diagnose failure class, not just severity.
Cost visibility becomes cost governance the moment you stop asking "how much?" and start asking "why did this execution's token cardinality diverge from baseline?"
That's when you catch failures before they compound.
Top comments (0)