DEV Community

lixingliangsy
lixingliangsy

Posted on

We had no idea what our AI agents were costing us. On purpose, apparently.

An uncomfortable exercise for any team running LLM agents: try to answer "what did agent X cost us last month?" If your answer involves opening the model provider's billing dashboard and squinting at a total, you're in the same place we were. We knew the total. We had no idea what it was made of.

The total is the least useful number

The provider invoice tells you what you spent. It doesn't tell you:

  • which agent spent it,
  • whether that spend was triggered by a paying customer or a free-tier user looping on retries,
  • whether one feature's background summarization job is quietly the biggest line item you run.

Without attribution, every cost conversation is a guess. When I proposed cutting a feature, the honest answer to "but what does it cost us?" was "some fraction of the invoice, unclear which." That's not a decision-making position.

Attribution is the fix, and it's boring on purpose

AgentCost does the unglamorous part: it attributes LLM and agent inference cost to the agent, the feature, and the user that triggered it, and keeps a running ledger. On top of that sit per-agent budget guardrails and early-warning alerts — the point is finding out about a runaway loop from an alert at 2pm, not from the invoice at the end of the month.

The ledger being audit-ready matters more than I expected. The first time finance asks engineering to justify AI spend — and they will — "here's the per-agent breakdown with timestamps" ends that conversation in one email instead of a week.

What changed for us

Nothing dramatic, which is the point. No agent turned out to be secretly catastrophic. But two things surfaced: one internal tool was burning meaningful tokens on a schedule nobody remembered setting, and our retry logic on one path was multiplying cost in a way the provider total had smoothed into invisibility. Attribution turned both from vibes into line items with numbers attached.

What it is not

It doesn't optimize your prompts or cache your calls — cost reduction is still your job; it just tells you where to aim. It tracks inference spend, so a SaaS subscription your team forgot about is someone else's problem (and honestly, a different tool of ours). And attribution requires your pipeline to actually tag calls with agent/feature/user context; if nothing labels the calls, there's nothing to attribute.

The two-question test

Ask your team today: what does our most expensive agent cost per month, and per active user? If both answers take longer than a minute to produce, you're flying on a total. AgentCost is free to try — the fastest way to see its value is wiring up one agent and comparing its ledger against your gut feeling about what that agent spends.

FAQ

What is AgentCost?

AgentCost does the unglamorous part: it attributes LLM and agent inference cost to the agent, the feature, and the user that triggered it, and keeps a running ledger. On top of that sit per-agent budget guardrails and…

Why does "The total is the least useful number" matter?

The provider invoice tells you what you spent. It doesn't tell you:

Is AgentCost free to try?

AgentCost is free to try — the fastest way to see its value is wiring up one agent and comparing its ledger against your gut feeling about what that agent spends.

References

  • FinOps Foundation — The industry framework for allocating and governing variable spend — the practice this article applies to agent costs.

Top comments (0)