DEV Community

ssamuels8
ssamuels8

Posted on

Your AI agents are bleeding money and your provider's dashboard won't tell you where

Your AI agents are bleeding money and your provider's dashboard won't tell you where

You shipped the agents. They're working. Users are happy. Then the invoice arrives and it's three times what you modelled.

You open the OpenAI dashboard. Total spend by model. Total tokens by day. A number that tells you everything went up and nothing about why.

This is the production AI cost problem nobody talks about enough. Not "which model is cheapest" — there are a hundred comparison tables for that. The real problem is that once your agents are running, you have almost no visibility into what's actually driving your bill. And without visibility, every optimisation is a guess.

This is what I've learned running AI agents in production, what the actual cost drivers are, and the specific things that made a meaningful difference — starting with one that requires zero code changes.

The invoice tells you nothing useful
Every major LLM provider — OpenAI, Anthropic, Google, Mistral, DeepSeek — gives you roughly the same dashboard. Total tokens consumed. Total cost. Broken down by model, by day.

That's useful for a monthly finance review. It's useless for operations.

Here's what it doesn't tell you:

Which of your agents is responsible for which share of the bill. Whether a cost spike happened because call volume increased, because prompts got longer, or because a deployment changed something. Which feature is the cost driver. Whether your prompt caching is actually working or silently failing. Which agents are generating more output tokens than they need to.

When you're running one agent calling one model, aggregate spend is fine. When you're running five agents — a support agent, a summarisation agent, a code review agent, a classification agent, a data extraction agent — all potentially calling different models at different rates, aggregate spend hides everything.

The first thing any serious production AI setup needs is per-agent cost attribution. Not total spend. Cost by agent, by call, in real time.

Where the money actually goes
Before you can cut costs, you need to understand the cost structure. It's more nuanced than most teams realise when they're first getting started.

Every LLM API call has three token categories: input tokens, output tokens, and cached input tokens. Input tokens are your prompt — the system message, any context you inject, the conversation history, the user's message. Output tokens are what the model generates in response. Cached tokens are input tokens that were served from the provider's prompt cache rather than processed fresh.

Output tokens cost more than input tokens on most models. Cached input tokens cost significantly less than uncached ones — on Anthropic, cached tokens are charged at approximately 10% of the standard input rate. On OpenAI, similar savings apply for eligible prompts.

This matters enormously in practice. An agent with a 3,000-token system prompt making 5,000 calls per day is sending 15 million input tokens daily — just in the system prompt, before counting any user messages or context. If those tokens are cached, you're paying for roughly 1.5 million effective token-equivalents instead of 15 million. If they're not cached, you're paying full price on every single call.

Most teams aren't caching. Most teams don't even know whether their caching is working.

The three biggest cost drivers in production
After tracking a lot of agent calls, the same patterns come up repeatedly.

Repeated stable context. The system prompt, background instructions, a knowledge base, a product catalogue — content that's the same on every call. If this isn't cached, you're paying to process it fresh thousands of times per day. This is almost always the biggest single cost driver, and it's also the easiest to fix.

Prompt growth over time. Prompts tend to grow. Teams add instructions, add examples, add edge-case handling, add clarifications, add new features. Nobody removes anything. A prompt that was 800 tokens at launch might be 2,400 tokens a year later. Every addition compounds across every call, every day. Most teams have never audited their prompts against their original intent.

Unbounded output tokens. Some agents are generating significantly more output than their task requires. A classification agent that should return a label and a confidence score but instead explains its reasoning at length. A summarisation agent with no length constraint producing 800-word summaries when 200 words would serve the user equally well. Output tokens are expensive. Unconstrained output is a quiet, consistent cost drain.

The fix that requires no code changes
Prompt caching is the highest-impact LLM cost optimisation available right now. On agents with a stable system prompt, it can cut input token costs by 80 to 90 percent, permanently, without changing a single output.

The mechanism is simple: if the first portion of your prompt is identical across calls, the provider stores it. On subsequent calls that share the same prefix, the provider reads from cache rather than reprocessing the full input — and charges you a fraction of the normal rate.

The requirements are specific. The cacheable prefix must be long enough — Anthropic requires at least 1,024 tokens. The prefix must be stable — any change invalidates the cache, including a single character difference. The stable content must come first in the prompt; dynamic content like user messages and variable context goes at the end. On Anthropic you need to add explicit cache control markers at the right breakpoint. On OpenAI, caching happens automatically for eligible prompts but requires the right prompt structure to trigger.

This sounds straightforward. In practice there are a dozen ways to accidentally break it.

A timestamp injected at the top of your system prompt. A user's name or role personalised into the beginning of the instructions. A deployment that changes the wording of the system prompt by one sentence. Tool definitions that get reordered between calls. Any of these bust the cache silently — you pay full price and your monitoring dashboard doesn't tell you why.

The cache hit rate is the metric that tells you whether your caching is actually working. A stable system prompt with a high call volume should be achieving 90 percent or higher cache hits at steady state. If it's lower, something is invalidating the cache and you need to find it.

Per-agent visibility changes everything
Here's what changes when you can see cost per agent rather than cost in aggregate.

You find your expensive agents immediately. In most production setups, 20 percent of agents are responsible for 80 percent of spend. Without per-agent attribution, you don't know which 20 percent. With it, you know in the first five minutes.

You can set realistic budgets. Once you know your p50 and p95 cost per call for each agent, you know what normal looks like. You can detect drift — an agent whose average call cost is creeping up — before it becomes an invoice surprise.

You make better model decisions. A summarisation agent and a complex reasoning agent have fundamentally different requirements. The summarisation agent might perform equally well on a model that costs ten times less. But you only find this out by testing — and you only know it's working by measuring cost per agent after the switch.

You catch bugs faster. A cost spike on one agent is often a bug signal — a runaway loop, a context that's growing unbounded, a change to the prompt that broke caching. Per-agent monitoring surfaces these as anomalies rather than hiding them in aggregate numbers.

The right order to do this
Most teams try to optimise before they can see. They switch to a cheaper model, shorten a prompt, try to implement caching — and then can't tell whether it worked, because they have no baseline and no per-call data.

The order that actually works:

First, get visibility. Per-agent cost, token breakdown — input, output, cached — latency, and error rates. This is your baseline. Everything else depends on it.

Second, identify your targets. Which agents are most expensive? Which have the worst cache hit rates? Which are generating unexpectedly long outputs? The data tells you where to focus.

Third, enable prompt caching. For any agent with a stable system prompt over 1,024 tokens, prompt caching should be the first optimisation. It's the highest ROI change you can make, and it doesn't touch your outputs at all.

Fourth, audit prompts for waste. Sample real production prompts. Look for duplicate instructions, legacy copy from features that no longer exist, context that's injected on every call but only needed sometimes. A 20 percent reduction in prompt size is a permanent 20 percent reduction in input costs.

Fifth, constrain outputs where appropriate. Add explicit length guidance to agents that are generating more than they need to. Use structured output formats for tasks that don't need prose. Set reasonable max token limits.

Sixth, evaluate model switches. Once you have per-call quality signals and cost data, test cheaper models against your actual production prompts. GPT-4o mini is 10 to 15 times cheaper than GPT-4o and performs comparably on many production tasks. DeepSeek V3 has among the lowest input pricing available for a frontier-class model. The question is always cost per acceptable output, not cost per token.

What "monitoring is free" actually means for your stack
The barrier to getting this visibility is lower than most teams think.

A proxy layer sits between your application and your AI providers. You change one URL — your provider's base URL to the proxy's endpoint — and every call is logged. Cost, tokens, latency, errors, per agent. Your application code doesn't change. Your prompts don't change. Your outputs don't change. It adds a few milliseconds of network overhead and gives you complete operational visibility into your AI layer.

This is how Vigil works. You change one URL per provider — OpenAI, Anthropic, Gemini, Mistral, DeepSeek, all supported — label your agents, and have per-agent cost monitoring in real time. Automatic prompt caching activates immediately across all your agents without you managing cache breakpoints manually. The savings appear in your cost dashboard from day one.

Monitoring is free. You don't need to instrument your code. You don't need to build dashboards. You don't need to maintain infrastructure. The cost visibility and the first round of optimisation — prompt caching — are live in five minutes.

Model routing and history trimming are on the roadmap but not live yet. What's available today — per-agent observability, prompt caching, cost and token breakdown per call — is already the foundation that makes every other optimisation decision possible.

The compounding problem nobody budgets for
One more thing worth saying plainly.

LLM costs compound in ways that traditional infrastructure costs don't. A server you provision costs what it costs. An LLM API call costs what it costs today — but next month your prompts are longer, your call volume is higher, and you've added three agents. The same architecture that cost $400 a month at launch costs $4,000 a month eighteen months later, and the team is genuinely confused about why.

The teams that manage this well aren't necessarily running cheaper models or shorter prompts. They're running with visibility from the beginning. They see costs grow in real time. They know which agents are scaling well and which are scaling badly. They optimise continuously rather than in a panic when the invoice arrives.

Per-agent monitoring is the foundation. Everything else — caching, model selection, prompt engineering, context management — is built on top of it.

If you're running AI agents in production without per-agent cost visibility, you're flying blind. That's fixable in five minutes.

→ vigil.wtf

Top comments (0)