DEV Community

LXSAIHUB
LXSAIHUB

Posted on

An agent kept retrying all weekend. Nobody knew until the invoice.

I don't know exactly how much that weekend of retry loops cost us. I want to lead with that, because "we didn't know" is the entire point of this post. It was some hundreds of dollars — visible later as an unusually bad daily line on the provider invoice — but which agent, which feature, which user cohort triggered it: unrecoverable. The invoice aggregates. Incidents don't live in aggregates.

Agents fail in financially weird ways

Traditional services fail by erroring. Agents often fail by succeeding expensively: retry loops that burn tokens per attempt, fallback chains that escalate to a bigger model because the cheap one hiccuped, a scheduled job nobody remembered that summarizes things no one reads, per-user cost that scales with your least-considerate power user. Every one of these looks like normal traffic on a provider dashboard. None of them look normal when attributed properly.

The monitoring gap is attribution, not alerts

We had alerting. We had alerting on error rates, latency, uptime — all the things that go wrong in a normal service. We had nothing that answered "what is this agent spending right now, per agent, per feature, per user?" When the invoice surprise arrived, the honest postmortem conclusion was: the signal existed in our logs, buried, and no one was watching for it because cost wasn't a monitored dimension.

AgentCost exists to make cost a monitored dimension: attribution of LLM and agent inference spend to the agent, feature, and user that triggered it, a running ledger, per-agent budget guardrails, and early-warning alerts — the class of alert that says "this loop has burned 40% of its budget in an hour" while it's happening, not on the 1st of next month.

The cheap version, if you build it yourself

Honest alternative, since not everyone wants a tool: tag every model call with agent/feature/user identifiers, log tokens and model per call, and build one daily job that aggregates by tag and diffs against a threshold. It's a day of work and covers the worst surprises. What you don't get cheaply is the guardrail part — killing or degrading a runaway agent automatically at a budget line — and the audit-ready ledger, which finance will eventually ask for.

What it is not

Attribution doesn't make anything cheaper by itself; it tells you where to look. Budget guardrails need your pipeline to tag calls correctly — garbage tags, garbage attribution. And cost is one axis: a cheap agent that's also wrong in every output is a different tool's problem.

Do one thing this week

Pick your single most expensive agent and answer: what did it cost last month, per user? If that takes more than a minute to answer, wire up attribution for that one agent — AgentCost is free to try, and the first week of its ledger against your gut estimate is usually instructive. Mine was off by more than I'd like to admit, in a direction I didn't predict.

FAQ

What is AgentCost?

AgentCost exists to make cost a monitored dimension: attribution of LLM and agent inference spend to the agent, feature, and user that triggered it, a running ledger, per-agent budget guardrails, and early-warning alerts — the…

Why does "Agents fail in financially weird ways" matter?

Traditional services fail by erroring. Agents often fail by succeeding expensively: retry loops that burn tokens per attempt, fallback chains that escalate to a bigger model because the cheap one hiccuped, a scheduled job nobody…

Is AgentCost free to try?

If that takes more than a minute to answer, wire up attribution for that one agent — AgentCost is free to try, and the first week of its ledger against your gut estimate is usually instructive.

References

  • FinOps Foundation — The industry framework for allocation and governance of variable spend, applied here to per-agent cost attribution.

Top comments (0)