I don't know exactly how much that weekend of retry loops cost us. I want to lead with that, because "we didn't know" is the entire point of this post. It was some hundreds of dollars — visible later as an unusually bad daily line on the provider invoice — but which agent, which feature, which user cohort triggered it: unrecoverable. The invoice aggregates. Incidents don't live in aggregates.
Agents fail in financially weird ways
Traditional services fail by erroring. Agents often fail by succeeding expensively: retry loops that burn tokens per attempt, fallback chains that escalate to a bigger model because the cheap one hiccuped, a scheduled job nobody remembered that summarizes things no one reads, per-user cost that scales with your least-considerate power user. Every one of these looks like normal traffic on a provider dashboard. None of them look normal when attributed properly.
The monitoring gap is attribution, not alerts
We had alerting. We had alerting on error rates, latency, uptime — all the things that go wrong in a normal service. We had nothing that answered "what is this agent spending right now, per agent, per feature, per user?" When the invoice surprise arrived, the honest postmortem conclusion was: the signal existed in our logs, buried, and no one was watching for it because cost wasn't a monitored dimension.
AgentCost exists to make cost a monitored dimension: attribution of LLM and agent inference spend to the agent, feature, and user that triggered it, a running ledger, per-agent budget guardrails, and early-warning alerts — the class of alert that says "this loop has burned 40% of its budget in an hour" while it's happening, not on the 1st of next month.
The cheap version, if you build it yourself
Honest alternative, since not everyone wants a tool: tag every model call with agent/feature/user identifiers, log tokens and model per call, and build one daily job that aggregates by tag and diffs against a threshold. It's a day of work and covers the worst surprises. What you don't get cheaply is the guardrail part — killing or degrading a runaway agent automatically at a budget line — and the audit-ready ledger, which finance will eventually ask for.
What it is not
Attribution doesn't make anything cheaper by itself; it tells you where to look. Budget guardrails need your pipeline to tag calls correctly — garbage tags, garbage attribution. And cost is one axis: a cheap agent that's also wrong in every output is a different tool's problem.
Do one thing this week
Pick your single most expensive agent and answer: what did it cost last month, per user? If that takes more than a minute to answer, wire up attribution for that one agent — AgentCost is free to try, and the first week of its ledger against your gut estimate is usually instructive. Mine was off by more than I'd like to admit, in a direction I didn't predict.
FAQ
What is AgentCost?
AgentCost exists to make cost a monitored dimension: attribution of LLM and agent inference spend to the agent, feature, and user that triggered it, a running ledger, per-agent budget guardrails, and early-warning alerts — the…
Why does "Agents fail in financially weird ways" matter?
Traditional services fail by erroring. Agents often fail by succeeding expensively: retry loops that burn tokens per attempt, fallback chains that escalate to a bigger model because the cheap one hiccuped, a scheduled job nobody…
Is AgentCost free to try?
If that takes more than a minute to answer, wire up attribution for that one agent — AgentCost is free to try, and the first week of its ledger against your gut estimate is usually instructive.
References
- FinOps Foundation — The industry framework for allocation and governance of variable spend, applied here to per-agent cost attribution.
Top comments (0)