DEV Community

Hive80-lab
Hive80-lab

Posted on Originally published at hive80-lab.github.io

We cut our pager volume 31 to 14 in one month. Here's the exact checklist.

We cut our pager volume 31→14 in one month. Here's the exact checklist.

Last spring our on-call started muting the alert channel. Not because things were fine — because 2 of every 3 pages asked a human to do nothing. That's the moment alert fatigue stops being an annoyance and becomes an outage with no witness. The fix cost us one audit meeting a month, not a tooling budget.

The ratio that tells the truth

Pull last month's pager log. Count total pages, and count pages that led to an action. Under 1-in-3 action rate? The system is spamming, and it gets worse on its own. Ours was 31 pages, 9 actions. That number is your baseline; you'll only know you're winning if you track it.

The 3-question test

For every alert that fired more than 3 times last month:

  1. Is this actionable right now?
  2. Would a human act differently at 3AM than at 3PM?
  3. If the alert vanished, would anything get worse?

Two "no" answers = demote or delete. We deleted 6 alerts and nobody has missed them. Deleting an alert is allowed — that sentence took our team a while to believe.

Tame the flaky three

Every team has 2–3 alerts everyone ignores on sight. Write down the tell ("disk alert double-fires after patching; it's real if it persists 15 min"), then move that condition INTO the alert (a for: 5m clause beats a mental note). The tell goes in the handoff doc until the fix ships.

Thresholds with data, not vibes

Our 80% CPU alert fired every afternoon at 2PM. It was describing Tuesday, not danger. Set lines above the routine peak and let routine stay silent. A threshold you've never tuned is a guess wearing a pager.

Route by severity, not curiosity

SEV1 pages at any hour. SEV2 business hours or chat queue. SEV3 never pages — morning digest. If your SEV3s wake people, your severity matrix is decorative.

The fixed noise budget

Every noisy alert you delete earns the right to add exactly one sharper one. 15 pages that all matter beat 40 that don't — and the trend line ("pages 31→14, action rate 28%→61%") is the report that keeps the discipline alive.

The full monthly checklist (including the flaky-alert taming table) lives here, free:

https://hive80-lab.github.io/ops-notes/alert-fatigue-checklist.html

If you're building the alert-routing layer that makes severity routing automatic, we packaged our pick-first workflows as the Automation Starter Pack ($19) — and launch week everything is 30% off with code HIVE-LAUNCH30 at https://hive80lab.gumroad.com

Free starting point: The First 30 Minutes — the one-page incident quick-start we run first on every page: https://hive80lab.gumroad.com/l/first-30-minutes

What's the noisiest alert in your shop right now — and does it pass the 3-question test?

Top comments (0)