Drafted with AI help, human-reviewed by The Agent Loop.
Short version: My own analytics dashboard said 52 views this morning while every per-post row printed null. The total was right, the attribution was broken, and the footnote even claimed the data didn't exist. That is what most agent cost tracking looks like: a correct number you can't decompose. A provider invoice tells you that spending rose and never which agent run, retry, or tool call raised it. OpenTelemetry's GenAI conventions define no cost attribute in the published spec (the proposal has been open since August 2026) and mark their token counters as "a proxy for cost approximation", and the span that wraps your tool call carries no usage fields. Price each span yourself, roll it up from the child inference spans, and alarm at runtime instead of waiting for the invoice.
For skimmers
- AWS anomaly detection reads Cost Explorer data with a delay of up to 24 hours and re-evaluates about 3 times a day: your dollar alarm always fires after the money is gone
- The same input token costs 0.1× as a cache read and 1.25× as a cache write: a 12.5× swing on identical bytes
- OpenTelemetry GenAI: token counters exist,
execute_toolspans exist, cost attributes don't — registry grep = 0, and the proposal (PR #443) has been open since 09 Aug 2026 - OpenAI's Agents API says its own usage counts are "not a final bill" and that missing usage "does not mean zero usage"
- Spend limits sit at organization/project scope, not per span. Nothing in the default stack attributes one tool call
- Trajectory trimming (AgentDiet) cut input tokens 39.9–59.7% and total cost 21.1–35.9% at equal performance, which you can't find without per-span numbers
The invoice is a lagging indicator
Two days ago I wrote up why an agent bill grows even when the model price doesn't move. A reader, @hannune, replied with the follow-up I deserved: “Tag cost per call is the one that took me too long to add” — and the question underneath it: how do you actually attribute spend to a span, not to the account?
Start with what attribution gives you today. Amazon's Cost Anomaly Detection uses Cost Explorer data, and the documentation states the delay plainly: up to 24 hours, with the detector running roughly three times a day. Cloud cost monitors fire when "today's total cost exceeds $10,000". Useful. Also, by construction, a day late.
Braintrust's cost guide puts the limitation better than I can: the invoice "can show that spending increased, but it cannot explain which customer, feature, prompt change, retry pattern, or agent run caused the increase."
The scale question is no longer academic. One widely reported month-long run of roughly a hundred agent instances billed $1,305,088.81 across 603 billion tokens and 7.6 million requests, with a single day at $19,985.84 (reported figures, treat as un-audited). At that size, "the total went up" is not an explanation.
The scary part isn't the size of the bill. It's that you can't name the line that produced it.
Tokens are not money
Here is the step everyone skips. You count tokens, then you treat the count as cost. It isn't, because the same token bills at different rates depending on state.
- OpenAI prompt caching: cache read at 0.1× the uncached input rate, cache write at 1.25×, minimum cacheable prefix 1,024 tokens. Their own multi-turn agent example reports more than 90% cache-hit.
-
Anthropic: five-minute write 1.25×, one-hour write 2×, read 0.1×. And
input_tokenscounts only tokens after the last cache breakpoint: the real identity istotal_input_tokens = cache_read + cache_creation + input_tokens. - The timestamped-prefix trap, straight from Anthropic's docs: put a date at the front of the prompt and "you pay for a fresh cache write on every request and never get a read."
So "input tokens" in your log isn't a cost. It's a count that has to be split into three buckets and multiplied by three different rates, with a price table that changes per model and per date. Miss the split and you are off by an order of magnitude, in either direction.
The volume makes the split matter. In September 2025 Claude 4 Sonnet reportedly hit 100 billion tokens a day on OpenRouter, of which 99% were input tokens accumulated in a trajectory and 1% was what the model generated. Nearly all of your spend lives in the tokens you keep re-sending, and those are exactly the tokens cache state re-prices.
The span that runs your tool has no price
This is about tool- and run-level granularity. Vendor dashboards attribute what they are able to price; the question is what happens at the tool call, where the span you need has no fields to carry it.
This is the part I did not expect when I went looking for standards support.
OpenTelemetry's GenAI semantic conventions define gen_ai.usage.input_tokens and gen_ai.usage.output_tokens, and the registry page marks them Deprecated, moved to the GenAI repository. In that repository every GenAI attribute still carries a Development badge, not Stable. Fine, that's how specs mature.
What's not fine: the token-metrics document says the counters "serve as a proxy for cost approximation", and there's no cost, USD, or price attribute in the published registry. I grepped the whole conventions repo: nine mentions of "cost", every one of them prose, zero attributes to set.
There is a proposal. PR #443, "Add per-operation cost conventions (gen_ai.usage.cost.*)", has been open since 9 August 2026, with issue #287 asking for the same thing since May 2025. The standard knows about the gap and has not closed it. Meanwhile two vendors already emit gen_ai.usage.cost on their own — and are filing issues against themselves because the value ships as a bare number with no currency field (OpenLIT, Laminer). A cost with no unit is half an attribute.
Worse for anyone hoping to attribute per tool: the execute_tool {gen_ai.tool.name} span exists, is typed INTERNAL, and its attribute table carries gen_ai.tool.* fields plus error.type. No usage, no cost. The span that says this tool ran cannot say this tool cost.
The spec's own discussion explains why nobody has bolted one on: tool and retrieval charges "are not model cost to begin with", so they were left out of the cost work entirely. Fair, as far as model billing goes. It also means that for the line item your CFO actually asks about — which tool cost money — there is no standard home for the number, only child inference spans you can roll up yourself.
The commercial tools are honest about the same gap:
-
Langfuse: only
generationandembeddingobservations track cost; other observation types carry none. Tool and agent spans are unpriced by design. - Langfuse, on overlapping buckets: usage and inferred cost "will be counted double", and the displayed cost "overstates what your provider actually charged."
- Datadog LLM Observability: prices each span independently and reports PARTIAL COST when a span lacks price data. Its own example alert threshold is $10,000 a day.
- OpenAI Agents API: usage counts are "not a final bill", "missing usage does not mean zero usage", and the response does not expose cache-write counts, so a caller "cannot determine the exact model charge."
Add one more trap from the spec itself: a retried request's span SHOULD cover the logical operation "with all retries", which means SDK-level retries (the Python SDK retries certain errors twice by default) are folded into one span and invisible in the trace. Three billed attempts, one visible operation.
Nobody's standard says what a call cost. Everyone's dashboard prints a number anyway.
Three beliefs that keep teams shipping totals
Belief 1: the invoice is attribution. It is a ledger, not a trace. AWS gives you up to 24 hours of lag and no run identity. You will always be able to say that week was expensive and never which agent.
Belief 2: our observability tool tracks cost, so we're attributed. It tracks cost on the spans it can price. Generation spans, maybe embeddings. The tool-call span that actually explains the spend is unpriced, so your per-feature rollup silently under-counts and your total silently over-counts. Datadog calls the result PARTIAL COST for a reason.
Belief 3: the spend limit protects us. OpenAI's spend limits and alerts are enforced at organization or project scope. Nothing there stops one runaway loop inside one project from being the whole bill. A budget at the account level is a smoke alarm in the building's lobby.
What I would wire up
tool call
|
v
inference spans --> split billed tokens: cache read | cache write | plain input
| |
v v
price table <----- keyed by model + date (not a constant)
|
v
price attributes on the PARENT span --> rollup: tool -> run -> feature
|
v
runtime budget alarm (seconds, not 24h)
invoice --> reconciliation only, never the source of truth
-
Billable, not used. The spec's own rule: when a system reports both, instrumentation MUST report billable tokens. Log
cache_read,cache_creation, andinput_tokensseparately, never a single merged input number. -
Put the price on the parent.
execute_toolcarries no usage fields, so compute cost from the child inference spans and writecost.usd(your own attribute, until the spec grows one) onto the tool span. That is the only place the tool's spend becomes queryable. -
Key the price table by model and date. A constant
$3/Mis how you quietly drift 20% off reality. Store it as data next to the trace, not in code. -
Tag the dimensions you'll be asked about. Datadog's
cost_tagspattern: team, feature, customer, run id. The invoice will never carry these. - Alarm at runtime, at the unit that can be stopped. Account-level spend limits stay on as a backstop, but your real threshold belongs on the run: kill the loop when this run passes a ceiling, not when the project passes $10,000 a day.
- Keep idempotency in the money path. The spec folds retries into one span, and the SDK retries twice by default. Forward an idempotency key with the write so a retried run can't bill you twice for one action (the Stripe-shaped fix from last week's approval post).
- Reconcile monthly, don't investigate monthly. Compare your attributed total to the invoice and treat the delta as a bug in attribution, not as a surprise bill.
The payoff is measurable. Inference-time trajectory trimming (AgentDiet) strips the useless, redundant and expired context agents keep re-sending: across two LLMs and two benchmarks it cut input tokens 39.9–59.7% and total cost 21.1–35.9% while agent performance stayed the same. You cannot see that win on an invoice. You can see it the moment cost lives on the span.
Admitted limit: I have not run this on a production bill. What I did run is smaller and embarrassing: this morning's dashboard printed a correct 52 with a broken per-post column, because one field name was wrong and a footnote asserted the data was unavailable. Totals that nobody can decompose will always lie to you eventually.
FAQ
How do you attribute LLM cost to a specific tool or span?
Count billable tokens on each inference span, split into cache read, cache write, and plain input, multiply each by a price looked up for that model and date, then write the sum onto the parent execute_tool span and roll up by tool, run, and feature. No standard does this for you.
Does OpenTelemetry have a cost attribute for GenAI?
Not yet. The published conventions define token counters and describe them as "a proxy for cost approximation", with no cost, USD, or price attribute, and execute_tool spans carry no usage fields either. A proposal (gen_ai.usage.cost.*, PR #443) has been open since August 2026 and is still open, so treat any gen_ai.usage.cost you see in a vendor's output as a private extension with no agreed unit.
Why does my observability tool show a different total than the invoice?
Three usual causes: overlapping usage buckets double-counting, spans that carry no price data being silently skipped (Datadog reports this as PARTIAL COST), and usage counts that never included cache writes: OpenAI states its Agents API usage "cannot determine the exact model charge".
What's the difference between tokens used and tokens billed?
Used is what the model saw. Billed is what you pay for: cached input re-reads at 0.1×, fresh cache writes at 1.25×, and plain input at 1×. Anthropic's input_tokens counts only tokens after the last cache breakpoint, so the two numbers are not comparable.
How do you alert on agent spend before the invoice arrives?
Put a budget on the unit that can be stopped (per run or per loop), evaluated against your attributed per-span cost. Account-level spend limits and cloud anomaly detectors work as backstops, but AWS's detector itself has up to a 24-hour delay.
Sources
-
OpenTelemetry GenAI PR #443 —
gen_ai.usage.cost.*conventions (open since 2026-08-09) · issue #484 · issue #287 (open since 2025-05-30) -
OpenLIT:
gen_ai.usage.costhas no currency attribute · Laminer: same gap - OpenTelemetry GenAI attribute registry (token counters marked deprecated)
- OpenTelemetry GenAI semantic conventions, new repository (Development badges)
- OTel GenAI token metrics (proxy for cost approximation; billable-token rule)
- OTel GenAI spans,
execute_tool(no usage/cost attributes; retries folded in) - OpenAI: prompt caching (0.1× read, 1.25× write, 90% discount)
- OpenAI Agents API: observability ("not a final bill")
- OpenAI Python SDK (
max_retriesdefault 2) - OpenAI production best practices (spend limits at org/project scope)
- Anthropic: prompt caching (1.25× / 2× writes, 0.1× reads, cache breakpoints)
- Langfuse: token and cost tracking (double-counting, priced observation types)
- Datadog LLM Observability: cost (PARTIAL COST, cost tags, per-span pricing)
- Datadog Cloud Cost Management: cost monitors
- Braintrust: how to track LLM costs in 2026 (invoice cannot explain the increase)
- AWS Cost Anomaly Detection (up to 24 hours of delay)
- Reducing Cost of LLM Agents with Trajectory Reduction (arXiv 2509.23586) · HTML
- Reported $1.3M / 603B-token agent month (secondary report, un-audited)
Related on The Agent Loop
- Your agent's cost problem isn't the model. It's the loop.
- Your agent's approval needs an expiry date
- Why your MCP approval gate never fires (and what to do instead)
Over to you: if one tool call in your agent tripled its spend last week, what would you check first, and would that check actually name the tool? Reply below, I read every one.
And if this saved you from shipping another unattributed total, tap the unicorn (or like) and follow The Agent Loop — I read every reply, and the next one is about the thing nobody wants to audit.
Prefer an inbox to a feed? Every post also goes out by email: subscribe at buttondown.com/theagentloop.

Top comments (0)