DEV Community

Cover image for The Most Useful Line on Your AI Cost Report Is the One You Can't Explain

The Most Useful Line on Your AI Cost Report Is the One You Can't Explain

Ken W Alger on October 01, 2026

Attribution, allocation, and why "unknown" belongs in the schema. This piece grew out of a comment thread on Sarvar Nadaf's Per-Agent Cost Trackin...
Collapse
 
max_quimby profile image
Max Quimby •

The incurred-vs-caused split is the right frame, and the part people underestimate is how fast the causal chain gets laundered. Your retrieval→supervisor example is clean because the 20k tokens flow straight through. But the moment you put a summarization or compaction step in between — retriever returns 20k, a cheap model condenses it to 2k, supervisor reads the 2k — the supervisor's cost now looks small and well-behaved, and the real culprit (a retriever with no token budget) is invisible on every span. The cost got attributed to the step that cleaned up the mess.

The thing that's worked for us is carrying a contributed_by list down the chain rather than trying to reconstruct it after the fact, so the condense step inherits the retriever's id even though it spent almost nothing itself. Curious how you handle fan-out — one retrieval that feeds three agents in parallel. Do you split the downstream cost across them, or attribute the full contribution to each? We never found a split that wasn't arbitrary.

Collapse
 
kenwalger profile image
Ken W Alger •

“Causal laundering” is exactly the problem. The compaction step can make the downstream span look healthier while simultaneously hiding the event that made compaction necessary in the first place.

I like contributed_by because it preserves lineage without pretending lineage is allocation. For the fan-out case, I don't think I'd split the downstream cost unless there were an independent reason to believe the split represented causation. Three consumers receiving the same retrieval result means that retrieval contributed to three causal chains. Saying it caused 33.3% of each feels like we've manufactured precision.

So I'd be inclined to preserve the full contribution relationship on each branch, then keep monetary allocation as a separate question with its own attribution_method. That means contribution edges don't have to sum to 100%, which I increasingly think is a feature rather than a defect.

You've also given me another failure mode for the model: transformation can reduce an input's visible size without reducing its causal significance.

Collapse
 
reidmarlow profile image
Reid Marlow •

The observation about context not being additive hits the exact wall where proportional allocation falls apart in production. Prompt caching makes this especially sharp. When an upstream retrieval span injects unpinned dictionary keys, dynamic timestamps, or fluctuating headers near the front of a supervisor prompt, it invalidates the prefix cache for the supervisor and every subsequent turn in that trajectory. That converts cheap cache reads into full cache writes across the entire conversation history. On a standard trace, that invalidation shows up as a massive cost spike directly on the supervisor span, even though the supervisor ran the exact same logic as before. If the schema forces a proportional token split, the upstream retriever looks harmless because it only passed four hundred bytes, while the supervisor gets blamed for burning the budget. Having unknown as an explicit attribution method keeps telemetry honest when cache boundary changes make token volume decouple from actual spend.

Collapse
 
kenwalger profile image
Ken W Alger •

That's a fantastic counterexample to proportional allocation because the upstream contribution can be tiny in bytes and enormous in consequence.

It also suggests the causal object isn't always the tokens themselves. In your example it's the cache-boundary change caused by those tokens. The supervisor incurred the additional spend, but nothing about the supervisor's own behavior explains the delta.

I think that strengthens the case for keeping unknown rather than forcing allocation, but it may also require richer causal records eventually: not just retrieval → supervisor, but something closer to retrieval changed prefix → cache invalidated → supervisor incurred full-price processing.

That's much closer to an explanation than “retriever 3%, supervisor 97%,” even if the explanation refuses to produce percentages.

Collapse
 
sizzlebop profile image
Jessica Doering •

I like the idea of treating unknown as an actual useful result instead of something that needs to be cleaned up or estimated away. I think there’s a tendency with AI systems to make dashboards look more precise than the underlying data really is, especially once you have agents, retrieval, tools, and context all affecting each other.

Knowing where money was spent is useful, but knowing what actually caused that cost is a much harder and more interesting problem. And if you can’t explain part of it yet, that’s probably exactly where you should be looking next. I’d much rather see an honest unknown than a very confident-looking number that’s basically an educated guess.

Collapse
 
aifrontierpost profile image
AI Frontier Post •

The confidence: 0.82 line is the one worth stealing. A proportional guess wearing a score feels rigorous while telling you nothing, which is exactly why the method belongs inline next to the number instead of buried in drill-down.

Collapse
 
kenwalger profile image
Ken W Alger •

Exactly. confidence: 0.82 looks quantitative enough that it's very easy to stop asking what the 0.82 is confidence in.

A highly confident proportional allocation is still a proportional allocation. That's why I'm increasingly convinced the method is part of the value, not metadata about the value. Remove the method, and you've changed what the number is entitled to mean.

Collapse
 
prpatel05 profile image
Pratik Patel •

We've started treating unknown attribution as a weekly triage queue, not a dashboard embarrassment. The failure mode I keep hitting is proportional allocation quietly becoming "truth" in finance reviews once someone pastes the number without the method, you can't walk it back. Making attribution_method a required column on any attributed line stopped that faster than adding more span tags.

Collapse
 
kenwalger profile image
Ken W Alger •

I really like treating unknown as a queue rather than an embarrassment. That changes it from “telemetry we haven't cleaned up yet” into a visible backlog of causal questions worth investigating.

And your finance-review example is exactly why I wanted attribution_method beside the number. Once 42% escapes the system without “proportional allocation” attached, it becomes organizational fact remarkably quickly.

That suggests another useful metric: not merely unknown cost, but unknown attribution resolved per period and what method replaced it. Then shrinking unknown means the organization actually learned something rather than somebody finding a convenient bucket to dump it into.

Collapse
 
izgorodin profile image
Edward Izgorodin •

The lineage is in the trace for causes inside one run. Two causes sit outside it, and both land in unknown unless a record can point into another trace. One is the prompt cache. If the provider caches prompt prefixes, whether the supervisor's first call in a run reads its prefix from cache or pays for a fresh write depends on whether another run used the same prefix within the cache lifetime, which runs from minutes to hours depending on provider and settings. So two identical runs can differ in cost with nothing in either trace that names the cause. Reid's case is a cause inside the trajectory. This one is the timing of a different run.

The other is persistent memory. A note or summary hydrated into context today was written by an earlier run, and if it was written long, every later run that recalls it pays for that length, while none of that cost appears in the trace of the run that wrote it. The decision that caused the cost is a write in a past trace.

Both fit the schema without new attribution fields. Record cache read and cache write tokens on each span where the provider reports them, so a cold start shows as its own line instead of hiding inside the supervisor's total, and stamp each stored memory with the id of the span that wrote it, so caused_by can name a span in a past trace. The cache case even admits a measured_delta: replay the same request while the cache is warm, confirm from the read field that it hit, and compare the input side of the two bills.

Collapse
 
blobdole profile image
Doug •

This is an interesting problem we have been thinking about with our crash debugging software.

ForensicDbg does a pile of things with the crash data to fix up missing and incorrect information, add additional known labeling and context, then links it all together and presents it clearly. When it comes to our MCP server though, it is that last part that becomes interesting.

We can send you the bare minimum amount of information needed to solve most basic crashes. Efficient! Unless it is not enough to get the job done, then the multiple back-and-forth trips to the tool are WAY more expensive than if we just sent more the first time.

We can send you a complex blob of nicely formatted and analyzed data that should be able to solve 98% of crashes in a single return. If it takes additional data you are still six steps ahead in terms of efficiency. But you are also 6x more expensive on any crash that would have been solved with the simple return above.

Figuring out the middle ground has been a very interesting, and ongoing challenge.

Collapse
 
presango profile image
Priya Nair (Presango) •

The retrieval example is the clearest statement of this I have read: the span that incurs the cost is not the span that caused it. We see a small version of it with an AI presenter that takes live questions during a session (I work on Presango, so bias declared). The expensive step is always the answer, but the cause is usually upstream: how much of the deck and the supporting documents got hydrated into context for a question that needed one slide. Attributing that to the answering step makes the dashboard say the wrong thing to fix. Keeping an explicit unknown bucket rather than forcing every cost onto a cause is the honest version, and I suspect its size over time is a better health metric than the total. Do you track how the unknown share trends run to run, or just report it?

Collapse
 
kenwalger profile image
Ken W Alger •

I haven't taken it as far as a run-to-run operational metric yet, but I think your instinct is right with one important caveat: I'd want to track why the unknown share changed, not just whether it went down.

A falling unknown rate could mean the system got better at causal attribution. It could also mean somebody replaced unknown with increasingly aggressive proportional guesses. Those are opposite outcomes hiding behind the same improving chart.

So I'd probably track unknown share alongside transitions out of unknown: unknown → measured_delta, unknown → proportional, unknown → estimated, etc. Then the trend tells us whether we're accumulating evidence or merely accumulating confidence.

Your presenter example is also a great illustration of incurred versus caused. The answer generation span can be perfectly healthy while repeatedly paying for an upstream decision to hydrate far more context than the question required.

Collapse
 
_firelinks profile image
Mike Dabydeen •

The incurred-versus-caused split is useful because it keeps the receipt honest. The distinction I would add to this proposed schema is the decision context that allowed the expensive branch to run. A retrieval span can be correctly attributed and still exceed a declared budget if no policy sees its returned-token count before hydration. Keeping the task, budget, observed count, and resulting decision beside the causal edge would let a reviewer ask two separate questions: what caused the spend, and what allowed it to continue? How are you thinking about the point where attribution becomes a control rather than a report?

Collapse
 
hannune profile image
Tae Kim •

This showed up for me in a specific shape: the same retrieval step ran identically for five different requests and charged to five separate causal chains because nothing connected them back to a shared trigger. Five lines in the report, one real cause. What actually fixed it was adding a source_request_id tag to the retrieval call and grouping by that before any allocation ran, which dropped the unexplained bucket by a noticeable amount. I'm curious whether your schema has anything that surfaces when two attribution chains are pointing at the same underlying event, because that's the version of the problem that's hardest to see from the dashboard alone.

Collapse
 
kenwalger profile image
Ken W Alger •

That's a really useful failure mode because it happens before allocation. If five attribution chains actually descend from one source event, no allocation method can repair the report until you've established that identity.

source_request_id sounds like exactly the kind of provenance I'd want carried forward rather than reconstructed later. I'd probably keep both identities: the immediate causal parent and the originating request/event. Then you can ask “what directly caused this?” and “what common event are these chains ultimately descended from?” without collapsing the two.

I don't have anything in the schema yet that explicitly detects two chains converging on the same originating event, and I think that's a gap. Before asking how to divide cost, we may have to establish how many causes we're actually looking at.