Stop diagnosing, start stabilizing: a field guide for live incidents
Your error rate just spiked. Someone is already asking "what changed?" in Slack. Here's the uncomfortable truth: the fastest way to make an incident worse is to start debugging the root cause before you've stopped the bleeding. This is the sequence we use to stabilize production systems first and investigate second, whether that's a single Node server or a full multi-node setup behind a load balancer.
The two mistakes that make incidents worse
Under pressure, teams almost always do one of these:
- Change too many things at once, so nobody can tell what fixed (or broke) what
- Start root-causing before the system is even stable
Everything below exists to prevent those two mistakes.
What you need before an incident happens
You can't stabilize what you can't see or touch. Minimum bar:
- Shell access to affected hosts (SSH or a bastion)
- Basic monitoring: CPU, memory, disk I/O, error rates (Prometheus, Datadog, or even
htopplus logs) - A load balancer or reverse proxy in front of the app (Nginx, HAProxy, cloud LB)
- A known, tested rollback procedure
- A second person to sanity-check decisions, even just in a Slack thread
If you're missing any of this, that's your next project, not something to build mid-incident.
Step 1: freeze all changes
This isn't technical, it's procedural. Post this immediately:
INCIDENT: elevated error rate on checkout-api since 14:32 UTC.
No deploys, no config changes, no manual DB edits until stabilized.
Incident channel: #inc-2024-checkout
This stops the classic "quick fix" from a well-meaning teammate that stacks a second problem on top of the first one you haven't diagnosed yet.
Step 2: check what changed in the last 24 hours
Most live failures trace back to something recent: a deploy, a config edit, a traffic spike, or an upstream update.
# Recent deploys
git log --since="24 hours ago" --oneline
# Recent config changes (if infra is version controlled)
git -C /etc/nginx log --since="24 hours ago"
# System-level changes
last -x | head -20
cat /var/log/dpkg.log | grep "$(date +%Y-%m-%d)"
If timing lines up, roll back before you dig further:
# Tagged release rollback
git checkout tags/v2.14.1
./deploy.sh production
# Kubernetes rollback
kubectl rollout undo deployment/checkout-api
Step 3: shed load before you preserve correctness
No obvious recent change, or the rollback didn't help? Reduce load on the failing component. You're trading functionality for stability, temporarily.
Ordered roughly by disruption:
- Rate limit at the edge (per IP or per token)
- Kill non-critical endpoints: recommendations, analytics beacons, background reports
- Maintenance page for non-essential routes, keep checkout/login alive
- Scale horizontally if the bottleneck is compute, not a shared resource
Emergency Nginx rate limit:
limit_req_zone $binary_remote_addr zone=emergency:10m rate=5r/s;
server {
location /api/ {
limit_req zone=emergency burst=10 nodelay;
limit_req_status 503;
proxy_pass http://backend;
}
}
nginx -t && nginx -s reload
Step 4: the database goes first
Databases usually fail last and recover slowest. A saturated DB will keep everything else unstable no matter what you fix upstream.
-- Connections vs max
psql -c "SELECT count(*) FROM pg_stat_activity;"
psql -c "SHOW max_connections;"
-- Long-running queries
psql -c "SELECT pid, now() - query_start AS duration, query
FROM pg_stat_activity
WHERE state = 'active'
ORDER BY duration DESC LIMIT 10;"
Kill runaway queries holding locks:
SELECT pg_terminate_backend(pid)
FROM pg_stat_activity
WHERE pid = 14832;
If connection exhaustion is the pattern, PgBouncer fixes this without touching app code. No pooler running? That's your action item once the fire's out.
Step 5: fail over instead of fighting a bad node
Disk pressure, memory leak, bad kernel state on one box: pull it from rotation instead of debugging it live.
# HAProxy
echo "disable server backend/web03" | socat stdio /var/run/haproxy.sock
# Kubernetes
kubectl cordon node-3
kubectl drain node-3 --ignore-daemonsets --delete-emptydir-data
Once it's out of rotation, it's a forensic artifact you can inspect at your own pace, not a liability sitting in the traffic path.
Step 6: communicate facts, not guesses
STATUS 15:10 UTC: error rate back to baseline (0.2%) after rolling back
deploy v2.14.2 and draining node-3. Root cause under investigation.
Next update in 30 min or on change.
Don't close the incident the second the dashboard looks calm.
How to actually verify stability
One green graph isn't proof. Hold these steady for 30 to 60 minutes at real traffic:
- Error rate: back to baseline, checked in 5-minute buckets, not instant values
- Latency: p50, p95, and p99 all recovered, a good average can hide a bad p99
- Resource headroom: CPU, memory, connection pools have actual margin
- Queue depth: backlogs draining, not just growing slower
# Nginx status code breakdown
tail -n 5000 /var/log/nginx/access.log | awk '{print $9}' | sort | uniq -c
# Live connection stats
watch -n 2 "ss -s"
# DB connection headroom
psql -c "SELECT count(*), max_conn FROM pg_stat_activity, (SELECT setting::int AS max_conn FROM pg_settings WHERE name='max_connections') s GROUP BY max_conn;"
Compare current traffic-normalized error rate to the same time last week, not to five minutes ago. Post-load-shedding quiet periods lie.
Pitfalls to avoid
- Root-causing mid-incident: adds risk, rarely speeds recovery
- Rolling back and rolling forward in the same window: give one change time to prove itself
- Declaring victory too early: five calm minutes isn't a stable system
- Skipping the change freeze: uncoordinated fixes are how you get two incidents instead of one
Stabilize first. Root-cause after. Every time.
Full original write-up with more context: binadit.com
Originally published on binadit.com
Top comments (0)