Nobody trains for the page
Your phone lights up at 2:07am. Service is down. What you do in the next 15 minutes decides whether this is a 20-minute blip or a lost customer. Here is the exact sequence — steal it, tape it to your desk.
Minutes 0-2: Acknowledge and stabilize your own head
Acknowledge the page in your alerting tool so nobody else piles in. Say out loud what you know: what's broken, since when, who's affected. Panic is just unstructured information — structure it.
Minutes 2-5: Stop the bleeding, don't find the cause
The question is not "why did this break." The question is "can I make it stop being broken." Rollback beats debugging. Restart beats root-cause. Scale-up beats optimization. Every minute customers spend down costs more than every minute of elegance you skip.
Minutes 5-10: Communicate before you're asked
One message to stakeholders: "We know, we're on it, next update at 2:30." A status page or a pinned channel message. Silence is how a technical problem becomes a trust problem.
Minutes 10-15: Capture the state
Screenshot the dashboards. Copy the relevant logs to a file. Note what you changed and when. Tomorrow-you, writing the post-mortem, will have nothing unless 2am-you spends 60 seconds on this.
After: the part that makes you better
The post-mortem, the timeline, the action items — this is where the incident pays you back. Teams that skip it have the same incident twice. Teams that do it well turn incidents into runbooks, and runbooks into calm.
If you want the operational side pre-built — the first-15-minutes checklist, the post-mortem template, the on-call drill schedule — that's exactly what's in the Ops Starter Kit:
- 🛠️ Ops Starter Kit (incident checklists, drills, post-mortem templates): https://hive80lab.gumroad.com/l/ops-starter-kit
- ⚙️ Automation Starter Pack (scripts that prevent the 2am pages): https://hive80lab.gumroad.com/l/automation-starter-pack
Use code LAUNCH50 for 50% off at checkout.
What's the worst page you ever got? Comments — let's compare scars.
Top comments (0)