You set a cost limit: "Stop my agent if it exceeds $50 per run." Reasonable. But then your agent starts calling a larger model, or hits an edge case that loops the prompt five times. The cost goes to $60, hits your threshold, and stops. Except it could have been $200 by the time you noticed—because the threshold is static, and your agent's baseline wasn't.
This is the fundamental flaw with hard thresholds: they don't adapt to what "normal" actually is for your agent.
The better approach isn't a bigger number—it's a baseline.
Why Baselines Matter: Context Is Everything
An agent that costs $5 per run is fine. An agent that never cost more than $2 is in trouble at $5. Same absolute number, two completely different signals.
Hard thresholds assume all agents are identical. But they're not. A customer-support agent calling GPT-4 on long documents might healthily hit $20 per run. A code-generation agent might run for $0.50 and spike at $5. A financial-analysis agent that handles large datasets might normalcy sit at $15—and a real anomaly might be $30, not $50.
Baselines capture this variability. Instead of a fixed number, you measure: "What does this specific agent usually cost?" Then you detect when it deviates significantly from its own pattern.
Enter 3-Sigma (The Math That Catches Anomalies)
In statistics, a sigma is a standard deviation—a measure of how spread out your data is. The "3-sigma rule" says: if a value is more than 3 standard deviations away from the mean, it's statistically unusual (roughly 99.7% of the time, normal values fall within 3 sigma).
Here's a concrete example.
Imagine an agent that, over its last 30 runs, has costs like:
$1.20, $1.18, $1.25, $1.19, $1.22, $1.20, $1.23, $1.21, …
(average: $1.21, spread very tight)
The standard deviation here is tiny—maybe $0.02. So 3-sigma would be:
Mean + (3 × StdDev) = $1.21 + (3 × $0.02) = $1.27
Now a run costs $1.80. Is that concerning? Absolutely. It's $0.59 away from the baseline, or about 29 sigma—wildly anomalous. A hard threshold of $2 would have missed this entirely.
Contrast with our second agent, the customer-support chatbot with a mean cost of $18 and high variability:
Mean: $18, StdDev: $4 (costs swing widely based on query complexity)
3-sigma threshold: $18 + (3 × $4) = $30
A $25 run here is totally normal—the agent is well within its own historical envelope. The hard threshold of $50 would never catch its real anomaly, which might be $40 or $45.
Why This Catches What Hard Thresholds Miss
The power of 3-sigma is that it adapts to the agent's behavior and reduces false alarms.
- For tight, consistent agents: a tiny spike is caught immediately, because even small deviations from a tight pattern are statistically significant.
- For variable agents: it doesn't page you for every fluctuation. The threshold floats with the agent's normal variance.
- For newly deployed agents: you need a warmup period to build a baseline (the KB-grounded implementation uses ~5 historical runs before activation), but once you have it, anomalies are caught automatically—no tuning required.
Hard thresholds force you to choose: set it low, and variable agents trigger false alarms every week. Set it high, and you miss real problems in consistent agents. 3-sigma gives you both: sensitivity where it matters, and tolerance for natural variation.
Implementation Notes
In practice, a production system usually adds a few refinements:
- Minimum sample size. You need at least a handful of historical runs before a baseline is statistically sound. (The typical threshold is 5+ data points.)
- Outlier removal. If an agent had one catastrophic run, that run can skew the standard deviation. Smart systems exclude extreme outliers when building the baseline, so one bad run doesn't permanently raise the alarm threshold.
- Grace period / rolling window. Don't trigger on a single run. Many implementations require 2–3 consecutive anomalous runs before alerting, reducing noise. Some use a rolling 30-day window so the baseline heals as soon as the expensive period ends.
- Hybrid approach. You can layer 3-sigma detection with a safety hard cap—"catch anomalies via 3-sigma and always stop if cost exceeds $500 in a single run," no matter what the baseline says. This catches both drift and catastrophic failure.
In Code
The math itself is straightforward. Here's a skeleton:
import statistics
def is_cost_anomaly(current_cost, historical_costs, threshold_sigmas=3):
if len(historical_costs) < 5:
return False # Not enough data yet
mean = statistics.mean(historical_costs)
stdev = statistics.stdev(historical_costs)
upper_bound = mean + (threshold_sigmas * stdev)
return current_cost > upper_bound
In a real system, you'd:
- Store historical costs in a database alongside agent metadata.
- Compute the mean/stdev on a schedule (or lazily, on each new run).
- Flag the run if it exceeds the bound, and decide whether to kill the agent, alert the user, or log it for review.
The threshold (3 sigma, or 2.5, or 2 depending on your risk tolerance) is tunable—more sensitive thresholds catch earlier, but increase false alarms. Most production systems start at 3-sigma and adjust if the signal-to-noise ratio doesn't feel right.
Why Builders Should Care
If you're running LLM agents in production—whether internal tools, customer-facing chatbots, or data-processing pipelines—cost anomalies are your early warning for:
- Prompt loops. An agent re-querying the model because it didn't understand the response the first time.
- Model fallback. Graceful degradation that tries GPT-3.5, fails, then silently retries with GPT-4.
- Runaway tool use. An agent calling an external API in a tight loop because it didn't parse the response.
- Data explosion. An otherwise-normal agent processing an unexpectedly large document or dataset.
Hard thresholds won't catch these until the damage is done. Statistical baselines let you catch the pattern of abnormality—the signature of something going wrong—before the bill becomes catastrophic.
You can implement this yourself in a few hours. Or, if you're using a platform like the AI Agents Control Tower, cost-anomaly detection via 3-sigma is built in: the system learns your agent's baseline automatically and alerts you (or pauses the agent, if you enable the kill switch) when it detects a spike. But the principle—baseline-aware detection beats fixed thresholds—is universal.
Start by asking: "What does normal cost look like for my agent?" The answer to that question is the beginning of real cost governance.
Top comments (1)
Five runs isn't enough. I know because I burned a week thinking our baseline was solid, then the agent started routing different query types and the whole variance picture shifted. The technical term side of this is that LLM costs are right-skewed, so 3-sigma on raw dollar amounts will fire too often on the high end. Log-transforming the cost values first and then computing the threshold dropped our false positive rate by about half.