Why Your On-Call Schedule Is Costing You Money
Bad on-call schedules don't just burn out engineers — they directly cost your business money through slower incident response, higher turnover, and missed SLAs.
After auditing on-call setups for a dozen small teams, I found the same three mistakes everywhere.
Mistake #1: The Hero Schedule
One person is on-call 24/7. They're the "expert." When they're unavailable, incidents pile up.
The cost: When the hero sleeps, takes a day off, or quits, your MTTR (mean time to resolve) spikes. I've seen 4-hour incidents that should have been 20-minute fixes because the only person who knew the system was on vacation.
The fix: Minimum 2-person rotation. Even if the second person is less experienced, they can acknowledge, triage, and escalate. The hero shouldn't be the only path.
Mistake #2: No Runbook = No Delegation
When the on-call engineer gets paged, they need to know what to do. Without runbooks, only the expert can respond. This creates a single point of failure.
The cost: Every incident requires the senior engineer. Junior engineers can't contribute. The senior engineer burns out. Turnover costs you $50K-$100K per replacement.
The fix: Document your top 5 incident types with step-by-step runbooks. A junior engineer should be able to handle a SEV3 incident by following the runbook.
Mistake #3: Alert Fatigue
Your monitoring sends 200 alerts/day. 190 are noise. The on-call engineer stops reading alerts. The 10 real incidents get missed.
The cost: Missed incidents = downtime = lost revenue. I've seen a $10K/hour revenue service go down for 3 hours because the on-call engineer muted their alerts after the 50th false positive that day.
The fix:
- Group related alerts (don't page 5 times for the same underlying issue)
- Set alert thresholds based on business impact, not technical metrics
- Review alert volume weekly — if it's over 10 pages/day, you're over-alerting
- Use warning-level alerts for non-paging notifications
The ROI of Fixing Your On-Call
| Metric | Bad On-Call | Good On-Call |
|---|---|---|
| MTTR | 2-4 hours | 15-30 minutes |
| Alert volume | 200+/day | 5-10/day |
| Engineer turnover | High | Low |
| SEV1 incidents/quarter | 5-10 | 1-2 |
| On-call satisfaction | "Hate it" | "Manageable" |
Building a Better On-Call Schedule
Week 1: Audit
- List all alerts and their frequency
- Identify your top 5 incident types
- Map who currently responds to each
Week 2: Document
- Write runbooks for your top 5 incidents
- Create communication templates
- Define escalation paths
Week 3: Rotate
- Set up a 2+ person rotation
- Give each person the runbook library
- Do a practice incident (game day)
Week 4: Tune
- Review alert volume
- Adjust thresholds
- Gather feedback from the team
Fix your on-call today: The Ops Mega Bundle ($49) includes everything you need — 5 complete kits with runbooks, escalation templates, on-call schedule templates, and field cards.
Start small: The Ops Starter Kit ($14) gives you the essential runbooks and triage system.
For your team: Ops Field Cards ($4) — 12 printable incident checkcards. Tape them to monitors, hand them to new on-call engineers.
Free starter: The First 30 Minutes Checklist — download free, use today.
How does your on-call rotation work? What would you change?
Top comments (0)