DEV Community

Hive80-lab
Hive80-lab

Posted on Originally published at hive80-lab.github.io

Burn-rate alerts: know you're burning before the budget is gone

An error budget checked at the end of its window is not a budget - it's an autopsy. The 28 days run out, the number is red, and the only thing anyone learned is that the month is unsavable. Burn-rate alerts fix the TIMING: they watch the pace of spending, not the total, and they speak while there is still budget left to save.

The whole scheme is one formula and one pair.

The formula. Burn rate = fraction of budget spent / fraction of window elapsed. 1.0x is on pace (121 bad minutes a month arriving at 11 seconds an hour). 6x means the budget is gone in about 4.7 days - ticket territory. 14.4x means gone in under 2 days - pager territory. Do the arithmetic once for your own budget: 121 minutes over 28 days = allowed pace of 11s/hour, so 14.4x = 2.6 bad minutes per hour. Two thresholds. That's the entire alert scheme.

Why not alert on the raw error rate? Because a 0.4% error rate against a 99.7% target means nothing by itself: for ten minutes it's a rounding error, for two days it's the whole budget. Raw-rate alerts flap with traffic. Burn rate normalizes spending against time, which is the quantity you actually decide on.

Two windows, one truth. A long window ignores disasters; a short window amplifies noise. So pair them: fast window 1h at 14.4x (catches the spike), slow window 6h at 6x (confirms it's real). The alert fires only when BOTH agree. Fast-only = Slack note. Slow-only = ticket. The pair is what turns a noisy metric into a trustworthy one - a false alert now has to defeat two independent windows at once. Add a 24h/2x tier for the chronic drip that empties a budget without any dramatic hour.

The page should carry the math. "Checkout SLO burning at 16x - 7 of 121 budget minutes spent, projected breach in ~26h if unchanged." An on-call who reads the stakes in five seconds makes a better decision than one who opens three dashboards.

Worked example: a twelve-person SaaS, checkout SLO 99.7% -> 121 budget minutes. Saturday 14:00, a config regression starts burning 3 budget-minutes/hour. 1h window crosses 14.4x at 14:45; the 6h window crosses 6x at 16:15; the pair fires; on-call rolls back at 16:25. Damage: 7 of 121 minutes. Month survives, no policy pause, ten-minute post-mortem. The counter-example - same regression, no burn alerts: raw error rate sits under the alert line all weekend (0.4% < 1%), the 28-day graph looks flat, and Monday's policy pause lands mid-sprint with no culprit and no graph showing the moment it started. The difference isn't discipline. It's an alert on the pace instead of the total.

The full template - the formula, both window pairs, the ticket tiers, and the wiring steps - lives on my ops-notes site: Burn-Rate Alerts. Free companion: The First 30 Minutes.

Top comments (0)