DEV Community

Cover image for The Most Useful Line on Your AI Cost Report Is the One You Can't Explain
Ken W Alger
Ken W Alger

Posted on Originally published at kenwalger.com

The Most Useful Line on Your AI Cost Report Is the One You Can't Explain

Attribution, allocation, and why "unknown" belongs in the schema.

This piece grew out of a comment thread on Sarvar Nadaf's Per-Agent Cost Tracking for Multi-Agent AI on AWS. The schema below was worked out in that conversation, in public, and it is better for it. Where a specific idea came from the exchange, I have tried to say so.


Most AI cost dashboards answer one question well: how much did this run cost. Total tokens, model spend, per-agent spend, latency, tool usage. Those are real and useful numbers, and for a while they are enough.

They stop being enough the moment your system becomes a composition. Once a request flows through a retriever, a knowledge graph, three specialist agents, and a supervisor that synthesizes their output, the total tells you almost nothing about what to change. A run can be correct, return HTTP 200, look healthy in every latency-and-errors panel, and still cost forty percent more than an identical run that produced the same answer. The overspend is real. It is just not anywhere you are looking.

To find it, you have to stop asking where the money was spent and start asking what caused it to be spent. Those are different questions, and the gap between them is the whole subject of this piece.

Where a Cost Is Incurred Is Not What Caused It

Consider a retrieval operation that costs $0.004: searching, ranking, fetching. That is the direct cost, and it is easy to attribute. It happened on that span, you can measure it, done.

Now suppose that retrieval returned 20,000 tokens, and all of them were hydrated into a supervisor's context on the next step. The supervisor then costs $0.009. How much did the retrieval really cost?

The direct answer is still $0.004. But that is no longer the interesting answer, because the retrieval also caused cost somewhere else. It inflated the supervisor's context, and some portion of that $0.009 exists only because the retrieval handed it too much material. The cost was incurred at the supervisor. It was caused, in part, at the retriever.

This gives you two distinct dimensions, and a useful cost model has to carry both:

  • Where the cost was incurred. This is just the span. Directly observed, low ambiguity.
  • What caused or contributed to it. This is the interesting axis, and it is the one no aggregate dashboard shows.

A record that captures both might look like this:

span_id: retrieval-104
direct_cost: $0.0040
primary: retrieval
returned_tokens: 20000

downstream:
  span_id: supervisor-105
  attributed_cost: $0.0021
  caused_by: retrieval-104
  attribution_method: proportional
Enter fullscreen mode Exit fullscreen mode

The dollars are single-counted. We do not charge the $0.0021 twice. The supervisor genuinely incurred it; the retrieval genuinely contributed to causing it; and the record says both without inventing money. What we have added is lineage: a link from a downstream cost back to the decision that helped produce it.

The Hard Part Is Honesty About How You Know

Here is the question that breaks naive versions of this: how do we know retrieval-104 actually caused $0.0021 of the supervisor's cost, and not some other amount?

Sometimes you can measure it. If you have a controlled comparison where the only meaningful change is that retrieval result, the delta is real evidence. Supervisor costs $0.006 without the retrieved material and $0.009 with it, so roughly $0.003 of downstream cost is attributable to that retrieval. That is measured causation, and it is the strongest claim you can make.

Most production traces do not give you that. In a real run the supervisor is carrying system instructions, conversation state, the outputs of other agents, tool results, and the retrieved material, all at once. There is no clean counterfactual. So you fall back on allocation: split the supervisor's context cost proportionally by the tokens each source contributed. That is a reasonable method. It is not measurement, and the receipt must not pretend it is.

This is why the single most important field in the whole schema is not a dollar amount. It is this:

attribution_method:
  - measured_delta
  - proportional
  - estimated
  - unknown
Enter fullscreen mode Exit fullscreen mode

That field is what keeps the entire model honest. It stops a proportional guess from masquerading as measured causation. With it, a line can say "retrieval span 104 contributed an estimated $0.0021 of downstream context cost, allocated proportionally by hydrated token share," and every word in that sentence is defensible, because the method is stated. Without it, the same $0.0021 acquires a precision the evidence never earned.

Resist collapsing this into a confidence score. A number like confidence: 0.82 feels rigorous and gives you nothing, because now you have a second number whose provenance you have to go investigate. measured_delta, proportional, estimated, and unknown each tell you why you are entitled to believe the figure. The method is the provenance. A score would hide it.

Show the Method Where the Decision Is Made

A natural instinct is to keep the attribution method as drill-down metadata, out of the main view, so the report stays clean. That instinct is wrong, and it is wrong for the same reason aggregate dashboards are wrong: it makes two different claims look equivalent.

The method belongs inline, next to any attributed cost, with one sensible exception. A directly observed cost carries no ambiguity and needs no method tag:

retrieval-104   RETRIEVAL   $0.0040
Enter fullscreen mode Exit fullscreen mode

There is nothing to disclose there; it was measured on the span. But the moment a number is attributed rather than observed, the method has to ride along:

retrieval-104 -> downstream CONTEXT   $0.0021   proportional
Enter fullscreen mode Exit fullscreen mode

Drop the word proportional and that $0.0021 visually becomes as solid as the $0.0040 above it, which is a lie of formatting. A report that hides the distinction between what it measured and what it allocated has committed the same sin as the dashboard that only shows a total. If two numbers make materially different claims, the interface must not make them look the same.

So the main report shows amount, category, direct versus downstream, and method. The drill-down holds the evidence behind the method: hydrated token counts, comparison runs, parent-child span references, the assumptions the allocation rests on. The decision surface stays readable; the receipt is one click away, not dumped into the table.

"Unknown" Is Not a Gap in the Accounting

Every honest version of this schema has to allow unknown as a real value, not a placeholder you feel bad about. And once you sit with it, the unknown rows turn out to be the most useful rows in the report.

A high downstream cost with unknown attribution is not incomplete bookkeeping. It is the system telling you exactly where your observability boundary stops letting you explain its own behavior. It is pointing at the place where you cannot yet answer "what caused this," which is the place most worth instrumenting next. A tidy report with no unknown rows has usually not achieved understanding. It has hidden its ignorance behind confident allocation.

There is also a real reason unknown is sometimes the only honest answer: context is not additive. An extra 5,000 tokens of context does not simply add a proportional slice of cost. It can change caching behavior, alter the reasoning path the model takes, or change how much output the model generates downstream. When that happens, token share and cost share stop mapping to each other cleanly, and any proportional number you report is a polite fiction. In those cases the schema should say unknown and mean it, rather than allocate a figure it cannot defend.

Read that way, the report stops being a statement of where you paid and becomes a map of two things at once: what you can explain about your spending, and where your ability to explain it runs out. The second map is the one that tells you what to build.

What This Actually Costs to Build

The reason this is not a research project is that most of the structure already exists. If your traces are parent-child spans, the causal lineage is physically present already; a downstream cost sits under the decision that produced it. You are not inventing a new tracing mechanism. You are making the attribution semantics explicit on top of a trace you already record.

Concretely, that is a small number of additions. Stamp a primary category on each span. Add a contributes_to link and a cause on spans that produce downstream effects. Add the attribution_method on any attributed cost. Then roll the report up along two axes, category and direct-versus-downstream, and let unknown be a first-class row rather than a swept-under one. The intelligence is not in the plumbing. It is in having the honest method field and being willing to publish the unknown rows.

One warning from the same conversation that produced all this: how you record and how you attribute are coupled. Change the way spans are emitted and you can silently break the logic that reads them. The defense is the same one that makes the whole model trustworthy, a known-good baseline you compare against, so that when your instrumentation shifts under you, the numbers move and you notice.

The Point

We have spent a lot of effort making AI spend visible. Better token counts, per-agent breakdowns, nested traces. All of it answers "how much." Almost none of it answers "why," and "why" is the only version of the question you can act on.

The move from one to the other is not a bigger dashboard. It is a small, honest schema: separate where a cost was incurred from what caused it, state the method behind every attributed number, and treat the costs you cannot explain as signal rather than embarrassment. Do that, and the report stops telling you what you spent and starts telling you what to fix, including, in the unknown rows, where to look first.

Even the cost report, it turns out, needs provenance. Not just the amount and the category, but how sure you are about who to blame. That last column may be the most useful one on the page.


With thanks to Sarvar Nadaf, whose post and the conversation under it produced this schema, and to the commenters in that thread who pushed on the baseline and the propagation. The receipt is better for the argument.

Top comments (3)

Collapse
 
max_quimby profile image
Max Quimby •

The incurred-vs-caused split is the right frame, and the part people underestimate is how fast the causal chain gets laundered. Your retrieval→supervisor example is clean because the 20k tokens flow straight through. But the moment you put a summarization or compaction step in between — retriever returns 20k, a cheap model condenses it to 2k, supervisor reads the 2k — the supervisor's cost now looks small and well-behaved, and the real culprit (a retriever with no token budget) is invisible on every span. The cost got attributed to the step that cleaned up the mess.

The thing that's worked for us is carrying a contributed_by list down the chain rather than trying to reconstruct it after the fact, so the condense step inherits the retriever's id even though it spent almost nothing itself. Curious how you handle fan-out — one retrieval that feeds three agents in parallel. Do you split the downstream cost across them, or attribute the full contribution to each? We never found a split that wasn't arbitrary.

Collapse
 
reidmarlow profile image
Reid Marlow •

The observation about context not being additive hits the exact wall where proportional allocation falls apart in production. Prompt caching makes this especially sharp. When an upstream retrieval span injects unpinned dictionary keys, dynamic timestamps, or fluctuating headers near the front of a supervisor prompt, it invalidates the prefix cache for the supervisor and every subsequent turn in that trajectory. That converts cheap cache reads into full cache writes across the entire conversation history. On a standard trace, that invalidation shows up as a massive cost spike directly on the supervisor span, even though the supervisor ran the exact same logic as before. If the schema forces a proportional token split, the upstream retriever looks harmless because it only passed four hundred bytes, while the supervisor gets blamed for burning the budget. Having unknown as an explicit attribution method keeps telemetry honest when cache boundary changes make token volume decouple from actual spend.

Collapse
 
prpatel05 profile image
Pratik Patel •

We've started treating unknown attribution as a weekly triage queue, not a dashboard embarrassment. The failure mode I keep hitting is proportional allocation quietly becoming "truth" in finance reviews once someone pastes the number without the method, you can't walk it back. Making attribution_method a required column on any attributed line stopped that faster than adding more span tags.