In 1978, researchers asked 60 people at Harvard Medical School a single question. A disease affects 1 in 1,000 people. The test for it has a 5% false positive rate. A patient tests positive. What's the chance they actually have it?
The most common answer was 95%. The right answer is about 2%. Only 11 of the 60 got it.
Now swap "disease" for "runaway agent loop" and "test" for "the spend-anomaly alert in your monitoring stack." The math is the same, and so is the bad intuition. Odds are it's paging someone on your team this week.
The arithmetic nobody runs before shipping an alert
Base rate neglect means judging a signal by how accurate it sounds while ignoring how rare the thing it detects is. The detector's quality matters less than you'd expect. The prior matters more.
So do the sum:
def precision(base_rate, sensitivity, false_positive_rate):
"""P(real incident | alert fired), via Bayes."""
true_alerts = base_rate * sensitivity
false_alerts = (1 - base_rate) * false_positive_rate
return true_alerts / (true_alerts + false_alerts)
# The 1978 Harvard question
print(precision(0.001, 1.0, 0.05)) # 0.0196
# A "95% accurate" cost-anomaly check, evaluated hourly.
# Real runaway loops: about 1 per 2,000 agent-hours.
print(precision(1 / 2000, 0.95, 0.05)) # 0.0094
That's under 1%. When your "95% accurate" detector pages someone, it's wrong 99 times out of 100. Nobody stays sharp through that. By the third false page at 2am the alert gets muted. Then the channel gets muted. Then the one real loop shows up and nobody's listening.
Your team isn't being lazy. The math makes that outcome close to certain.
Google Flu Trends was the same bug at planet scale
The 2009 Nature paper behind Google Flu Trends describes an impressive pipeline. Ginsberg and colleagues screened 50 million candidate search queries against CDC regional flu data, kept the 45 that fit best, and hit a mean correlation of 0.90 on validation data. A number like that usually ends the argument.
Then, between August 2011 and September 2013, it over-predicted flu in 100 of 108 weeks. At the 2012–13 peak it estimated flu-like illness at roughly 11% of US doctor visits. The CDC's eventual figure was around 6%. Lazer, Kennedy, King, and Vespignani's 2014 Science post-mortem blamed "big data hubris" and "algorithm dynamics."
Here's the base-rate angle. Search 50 million candidates for the best fit and the winners will mostly be things that rise and fall with winter, or with flu news, rather than with flu. Among the top scorers, true signal is rare and coincidental correlation is common. A 0.90 found by searching through 50 million options tells you much less than a 0.90 from a hypothesis you wrote down in advance. And search volume tracks attention, so a scary news cycle could inflate the estimate without anyone getting sicker.
The structural failure was that nothing pulled the model back to the slow, boring ground truth. Nobody kept asking how often the model's high readings had actually matched a high CDC count.
If you've ever picked "leading indicators" for incidents by correlating a few hundred metrics against past outages, you've built a smaller version of the same thing.
Where it hides in your dashboards
Anomaly alerts on rare events. Covered above. If the event is rare and you check often, almost every alert is a false positive, however good the model is.
A/B test "wins." Say 1 in 10 of your experiments has a real effect, and you run at α = 0.05 with 80% power. Plug it in: precision(0.10, 0.80, 0.05) comes out to 0.64. More than a third of your statistically significant wins are noise, and some of them are now in production.
Per-endpoint error panels. Put 200 endpoints on a grid, each with a 5% chance of crossing its threshold by luck, and you'll see about 10 red tiles on any given check. Someone will investigate each one.
The fix is the same every time: write down the base rate before trusting the signal. If you can't estimate it, you can't interpret the alert.
Stop predicting the failure. Make it impossible.
For LLM spend there's a better option than a smarter detector: skip detection entirely. A pre-flight budget check has no base rate problem because it isn't guessing whether a loop is runaway. It compares the worst-case cost of the next call with the money left, and refuses if the call doesn't fit. There's nothing probabilistic to get wrong.
| Mechanism | When it acts | What it fails to stop |
|---|---|---|
| Billing / budget alert | After spend is recorded, often hours later | Everything spent between the threshold and someone reading the email |
| Anomaly detector | On statistical deviation | Real incidents buried under false pages, plus slow ramps that never look anomalous |
| Rate limit | Requests per minute exceeded | A slow loop under the limit running all weekend, or a few very expensive calls |
| Pre-flight enforcement | Before the request is sent | Only cost you underestimated, and max_tokens bounds that |
Here's a minimal single-process version wrapping the Anthropic client:
import threading
from anthropic import Anthropic
class BudgetExceeded(Exception):
pass
class BudgetGuard:
def __init__(self, cap_usd, usd_per_mtok_in, usd_per_mtok_out):
self.cap = cap_usd
self.spent = 0.0
self.reserved = 0.0
self.p_in = usd_per_mtok_in / 1_000_000
self.p_out = usd_per_mtok_out / 1_000_000
self._lock = threading.Lock()
def reserve(self, input_tokens, max_output_tokens):
worst_case = input_tokens * self.p_in + max_output_tokens * self.p_out
with self._lock:
left = self.cap - self.spent - self.reserved
if worst_case > left:
raise BudgetExceeded(f"need ${worst_case:.4f}, ${left:.4f} left")
self.reserved += worst_case
return worst_case
def release(self, hold):
with self._lock:
self.reserved -= hold
def settle(self, hold, usage):
actual = usage.input_tokens * self.p_in + usage.output_tokens * self.p_out
with self._lock:
self.reserved -= hold
self.spent += actual
client = Anthropic()
# Use your model's current list prices here.
guard = BudgetGuard(cap_usd=20.0, usd_per_mtok_in=3.0, usd_per_mtok_out=15.0)
def guarded_create(**kwargs):
countable = {k: kwargs[k] for k in ("model", "messages", "system", "tools") if k in kwargs}
tokens_in = client.messages.count_tokens(**countable).input_tokens
hold = guard.reserve(tokens_in, kwargs["max_tokens"])
try:
resp = client.messages.create(**kwargs)
except Exception:
guard.release(hold)
raise
guard.settle(hold, resp.usage)
return resp
The lock does more work than it seems to. Without reservation, 20 parallel agents each read spent, each see the same stale number, each decide they're under budget, and together they blow through the cap. Folding the check and the claim into one atomic step closes that gap. When the call returns, the hold is swapped for the actual cost, so unused headroom goes back into the pool.
This guard lives in one process's memory. Once you have multiple workers, the reservation has to live somewhere shared and atomic, like a Redis Lua script or a single UPDATE ... WHERE spent + reserved + $1 <= cap in Postgres. baar-core packages this pattern as an open-source library: atomic reservation across concurrent calls, and a 402 before the provider is ever contacted (pip install baar-core). For teams that need per-user caps in a SaaS product, noburn.dev is what we built on top of it. It enforces budgets pre-flight and blocks a user's API calls before they fire once that user is over budget.
Keep the detector, demote it
You don't need to delete your anomaly alerts. Change what they're for. A detector with 1% precision is a fine diagnostic: it tells you where to look when you're already looking. As protection it's useless, because the people it protects you from have learned to ignore it.
Put hard constraints where being wrong is expensive, and keep probabilistic signals where being wrong costs a glance. Before you ship any new alert, run it through precision() with an honest base rate. If the result is under 10%, it belongs on a dashboard, not on a pager.
Google had 50 million queries and a 0.90 correlation, and still forgot to ask how rare the real thing was. Most of us have fewer excuses and far less data.
Has anyone on your team ever worked out the precision of your noisiest on-call alert? What was the number?
Originally published at https://robatdasorvi.com/stories/why-humans-are-terrible-at-evaluating-probability-and-what-that-means-for-product
Top comments (0)