90 Days. 47 Incidents. Zero Guesswork.
I tracked every single outage, degradation, and near-miss across our infrastructure for 90 days. Not because I'm obsessive — because I was tired of the same post-mortem meetings where nobody could agree on what actually happened.
The results surprised me. And they'll probably surprise you too.
The Setup
I created a simple tracking system (just a spreadsheet, honestly) with these columns:
- Date and time
- Severity (SEV1-SEV4)
- Root cause category
- Time to detect
- Time to resolve
- Customer impact (users affected, revenue at risk)
- Prevention (what we could have done)
That's it. No fancy tools. No APM integration. Just discipline.
The Results: What Actually Causes Outages
Here's the breakdown of 47 incidents over 90 days:
| Root Cause | Count | % of Total |
|---|---|---|
| Config changes (bad deploy) | 14 | 30% |
| Third-party API failures | 9 | 19% |
| Database issues (locks, slow queries) | 7 | 15% |
| DNS/networking | 5 | 11% |
| Resource exhaustion (disk, memory) | 4 | 8% |
| Certificate expiry | 3 | 6% |
| Human error (manual ops) | 3 | 6% |
| Security events | 2 | 4% |
The big surprise: Config changes caused nearly a third of all incidents. Not infrastructure failure. Not DDoS attacks. Just someone pushing a bad config.
The 5 Patterns That Repeat
1. The Friday Afternoon Deploy (14 incidents)
Someone pushes a change at 4:30 PM on Friday. It works in staging. It fails in production at 6 PM when traffic patterns differ.
Prevention: Deploy freeze after 2 PM on Fridays. No exceptions. I wrote this into our CI/CD pipeline as a hard gate.
2. The Silent Certificate Expiry (3 incidents)
SSL/TLS certificates expire. Nobody notices until the dashboard goes red and customers complain.
Prevention: Certificate monitoring with 30/14/7-day alerts. I set this up in 20 minutes using a free checklist from the Ops Starter Kit.
3. The Third-Party Cascade (9 incidents)
Your payment provider has an outage. Your auth service times out. Your app hangs because you didn't implement circuit breakers.
Prevention: Circuit breakers on every external call. Timeout at 5 seconds. Fallback to cached data. This alone cut our third-party incident impact by 60%.
4. The Database Lock (7 incidents)
A long-running query locks a table. Everything queues behind it. The app appears down.
Prevention: Query timeout at 30 seconds. Lock monitoring with automated kill of queries over 60 seconds. Read replicas for reporting queries.
5. The Disk Full Surprise (4 incidents)
Logs fill the disk. The database can't write. Everything stops.
Prevention: Log rotation. Disk space alerts at 80%. Automated cleanup of temp files. 15-minute setup.
The 30-Minute Incident Response Checklist
When an incident hits, the first 30 minutes determine everything. Here's what I do:
Minutes 0-5: Identify
- [ ] Confirm the incident is real (check dashboards, not just alerts)
- [ ] Determine severity (SEV1 = revenue down, SEV2 = degraded, SEV3 = minor)
- [ ] Page the right people (not everyone)
Minutes 5-15: Scope
- [ ] What's affected? (which services, which users)
- [ ] What changed recently? (check deploy log, config changes)
- [ ] Is it us or a dependency? (check status pages)
Minutes 15-25: Stabilize
- [ ] Roll back the last change if applicable
- [ ] Scale resources if it's capacity-related
- [ ] Switch to fallback/circuit breaker mode
- [ ] Communicate to stakeholders (internal + external)
Minutes 25-30: Prepare Recovery
- [ ] Document what you've tried
- [ ] Identify the root cause hypothesis
- [ ] Plan the fix (not the permanent fix — the get-us-back-up fix)
The ROI of Tracking
After 90 days of tracking, we:
- Reduced incident frequency by 40% (from 47 to ~28 per quarter)
- Cut mean time to resolve by 55% (from 45 min to 20 min)
- Eliminated Friday deploys (saved 14 incidents/quarter)
- Set up certificate monitoring (saved 3 incidents/quarter)
Total time invested: about 2 hours per week reviewing the tracker.
Free Resources
I've put together the actual templates I use for incident tracking, the 30-minute response checklist, and the certificate monitoring setup. They're all in the Ops Starter Kit — $14 for the full bundle, or you can grab individual checklists from our free template library.
The kit includes:
- The 90-day incident tracker template
- The 30-minute response checklist (printable)
- Certificate monitoring setup guide
- Circuit breaker implementation patterns
- Post-mortem template that doesn't waste time
TL;DR
- Track every incident. You can't fix what you don't measure.
- Config changes are your #1 enemy. Gate them, review them, freeze on Fridays.
- The first 30 minutes matter. Have a checklist, not a panic.
- Third-party failures are your problem. Circuit breakers, timeouts, fallbacks.
- Certificates expire silently. Monitor them or suffer.
What's your most common incident cause? I'm curious if your breakdown matches mine.
This article is part of a series on practical ops for small teams. Follow for more real-world playbooks, not theory.
Top comments (0)