DEV Community

Babar Hayat for OpsVeritas

Posted on

The Three Cost-Baseline Patterns Every AI Agent Builder Should Monitor

Your agent's daily cost hits $150. Last week it was $12. You have no idea why, or when it'll happen again.

Token budgets sound like the answer. Set a hard ceiling—say, $100/day—and stop the agent if it breaches. Reasonable, right?

Except your agent costs $8 most days, $15 on heavy-load days, and occasionally spikes to $45 during legitimate peak traffic. A $100 budget would never catch the moment it suddenly starts running $80 consistently—by then you've already burned $560 that month beyond your normal baseline.

The real problem isn't the ceiling; it's that you can't see the floor. You need to know what normal looks like for this specific agent, so you notice when it drifts.

Here's what that looks like in practice.

Pattern 1: The baseline mean

Before anything else, you need to know: what does this agent typically cost?

Pull your last 30 days of executions. Calculate the mean cost-per-run.

For a customer-support agent that runs ~100 times a day at $0.08 per run, that's roughly $8/day. Write that down. That's your baseline mean.

Why 30 days? Because a week is noisy. An agent that runs on business days only, or spikes during monthly reporting, needs more history to show the real center of gravity.

What you're doing: establishing what "healthy" looks like. When you look at tomorrow's costs, you'll compare against this.

Pattern 2: The variability (standard deviation)

Your baseline mean is $8/day. But days aren't identical.

Some days the agent runs 95 times. Some run 105. Some run 80 (someone was on vacation). Pull your 30 days of daily costs, calculate the standard deviation, and you get, say, $1.20.

Now you have two numbers:

  • Mean: $8
  • Std dev: $1.20

What this tells you: on 68% of days, expect $6.80–$9.20 (mean ± 1σ). On 95% of days, expect $5.60–$10.40 (mean ± 2σ). The remaining 5% of days will be weirder than that—and that's okay if it's rare.

What you're doing: learning the shape of normal variation so you don't alarm on it.

Pattern 3: The 3-sigma spike (the actual alert threshold)

Here's where you actually catch the break.

A spike is a day that costs more than mean + 3σ.

With your baseline ($8 ± $1.20), that's:

  • Mean + 3σ = $8 + ($1.20 × 3) = $11.60

If a day costs $15 or more, that's now a genuine outlier. Something changed. The agent is looping, a prompt is longer, a model call is repeating.

One spike is information. It could be a one-off surge in traffic, a larger-than-usual customer request, something legitimate. Let it pass.

Three spikes in a row? Your agent's baseline has shifted. It's now sustainably more expensive than it used to be. That's when you need to investigate and decide: is this okay, or does it point to a bug?

Why not a hard token budget?

Token budgets are simple—$50/day, stop at $50.01. But they're context-blind.

An agent that normally costs $5 gets a $50 budget and drifts to $48 without ever triggering. An agent that legitimately costs $30 on load days gets a $50 budget and gets killed on the first real traffic spike.

Baselines adapt to each agent's actual behavior. They let you notice the direction of change, not just absolute size.

How to implement this yourself

  1. Pull 30 days of costs — via your provider's API, logs, or CSV export. One row per day (or per execution, then aggregate).
  2. Calculate mean and standard deviation — any spreadsheet or quick Python script will do:
   import statistics
   costs = [...]  # your 30 days
   mean = statistics.mean(costs)
   stdev = statistics.stdev(costs)
   threshold = mean + (3 * stdev)
   print(f"Alert if daily cost exceeds ${threshold:.2f}")
Enter fullscreen mode Exit fullscreen mode
  1. Set a recurring check — weekly or daily, depending on your execution volume. Compare today's cost to your threshold.
  2. Investigate spikes — when a day crosses 3σ, ask: Did traffic surge? Did a prompt get longer? Did a model call start repeating? Was there a bug in the tool-use loop? Write it down so you see patterns.
  3. Recalibrate quarterly — your agent's normal behavior changes. Every 90 days, recalculate mean and stdev from the most recent 30 days.

The deeper point

Token budgets assume all cost growth is bad. But cost growth that comes from legitimately serving more requests—more customers, larger queries—isn't a failure. It's your business working.

Baselines separate signal from noise. They let you build agents with confidence: you'll see the moment something breaks (the spike), and you'll ignore the moment something is just busier (normal variation).

This is exactly what platforms like https://agents.opsveritas.com automate: they calculate these baselines for you, flag genuine spikes across your whole agent portfolio, and let you investigate without writing a single script. But the thinking is the same whether you're doing this by hand or with a tool—know your baseline, know your variation, and treat outliers as the signal they are.

Start with one agent. Pull its history. Do the math. You'll see it immediately.

Top comments (0)