We promised a customer "99.9% uptime." Nobody on the team had done the math.
99.9% is 43 minutes and 50 seconds of downtime per month. That's the entire budget. Every deploy that 502s for two minutes, every migration that runs long — it all spends from the same 43 minutes. The month we finally wrote the number on the whiteboard was the month the pager got quieter.
The downtime budget, in one table
| Promise | Downtime per month |
|---|---|
| 99% | 7h 12m |
| 99.5% | 3h 36m |
| 99.9% | 43m 50s |
| 99.95% | 21m 55s |
Pick the number your customers actually experience today — not the aspirational one.
The 7-step checklist
- Publish one number. An SLO nobody believes is worse than an honest one.
- Measure from the customer's side. An internal replica failing while users see nothing spends nothing. No external check = finding #0.
- Keep a one-line ledger per incident. Date, minutes, cause, customer-visible y/n. Sum one column monthly. Ten lines a quarter is the whole SLO governance a small team needs.
- Plan maintenance inside the budget. A 20-minute migration window is half of 99.9% in one shot. Either schedule it deliberately or drop the promise to 99.5% and sleep better.
- Spend deploys from the same wallet. Ten deploys a week at 30s of blip each = 25 minutes/month — over half the budget. Zero-downtime deploys aren't a luxury; they're how you keep budget for real incidents.
- Set the burn alarm, not the perfection alarm. Alert at pace: 50% of budget burned in the first week means you'll breach.
- Breach → 3-question review, not shame. What spent it? What would have prevented it? Is the SLO still the right promise?
The rules that keep it honest
- The budget is shared. Deploys, maintenance, incidents — one wallet. The moment one team believes their downtime "doesn't count", the number means nothing.
- Round up in the customer's favor. 3 minutes of errors counted as 5. Cheap insurance against lying to yourself.
- Renew the promise yearly, out loud. An SLO reviewed yearly is a contract you can keep; one set in stone is a breach you scheduled.
We keep the full ledger template, the burn-rate alert thresholds, and the maintenance-window calendar in the ops kits we build for small teams:
- The First 30 Minutes — free one-page incident quick-start
- Ops Starter Kit — $14 — incident response for small teams
- Ops Starter Kit Vol. 2 — $27 — advanced incident command, monitoring pack included
- Automation Starter Pack — $19 — pick-first workflows so uptime checks update themselves
Launch week: 30% off any paid kit with code HIVE-LAUNCH30 at checkout.
The full write-up with the ledger example lives on our ops notes: The Uptime Budget — HIVE80lab Ops Notes.
What's your number — and do you actually know how much of it you spent last month?
Top comments (0)