DEV Community

晖莫
晖莫

Posted on

The Runbook That Sent Me to a Dashboard That No Longer Exists

The alert fired at 02:14. Payments queue depth climbing, no consumers draining it. The runbook told me to check the "Payments Worker Health" dashboard in Grafana. It did not exist. Someone had renamed it to svc-payments-consumer; the runbook still pointed at the old name.

Fine. Annoying, but survivable. Two steps later it told me to restart the consumer with payments-worker restart --queue=main. The binary is now ledger-consumer. That command has not existed for months.

I typed it anyway. command not found. I sat at 02:19 with a paging incident and a document that had lied to me twice. Then I closed the runbook and stopped reading it.

The cost is not the two wrong steps

Here is what happened next. Everything else in the runbook — checking broker backlog before restarting anything, the DLQ inspection, the note that a stuck consumer is usually a broker problem — I skipped. I went to memory and guesswork.

I restarted the consumer. That was the wrong move. The broker had partitioned and the consumer was holding a lock it could not release. Restarting made it worse. The runbook had a section warning about exactly that. I never got to it.

That is the real failure mode. It is not that step four is stale. It is that a responder who has been lied to twice treats the remaining eight steps as suspect, including the correct ones. A bad runbook destroys the value of the good parts.

Why they rot

Runbooks get written once, right after an incident, by the person with the most context. That person already understood the system, so they wrote an index into their own head. "Check the Payments Worker Health dashboard" made sense to them, not to a stranger.

Then the system moved. Dashboards get renamed during cleanup. Commands get consolidated when a repo splits. Nobody deletes the old line, because deleting it is not part of any ticket, and whoever would notice never needs the runbook.

No bug, no alerting on documentation drift. It drifts until an incident finds out.

What makes a runbook survive

A step survives if a tired person can run it and see, without judgment, whether it worked. That means executable and verifiable:

# Is the consumer actually draining? Expect lag to fall over ~30s.
ledger-consumer status --queue payments --format json | jq '.lag_seconds'

# Broker-side backlog. If this is climbing while consumer lag is flat,
# the consumer is not the problem. Do not restart it.
rabbitmqctl list_queues name messages_ready messages_unacknowledged | grep payments
Enter fullscreen mode Exit fullscreen mode

The commands can be copied and run, and each has an expected result next to it. "Expect lag to fall" tells the responder what done looks like. No interpretation required.

A link beats a description. Do not name a dashboard — paste the URL. Do not describe where the log lives — paste the query. A link that 404s is a signal you can act on. A name that quietly points at a renamed thing is a trap.

Every runbook needs a named owner and a review date at the top. Not a team. A person:

owner: @dana
reviewed: 2026-08-14
review_by: 2027-02-14
Enter fullscreen mode Exit fullscreen mode

The date converts an invisible problem into a calendar item, and tells a responder how much to trust the page. A runbook reviewed four months ago is different from one reviewed three years ago.

"Contact X" is not a step

"Page the payments team" is not a step. It is an admission that the runbook has nothing to say, and it fails at the moment it matters: the named person is asleep, on a plane, or already in the incident channel asking what you tried.

Replace it with what you would ask them to do. If the answer is "ask Dana to run the replay tool," then the step is how to run the replay tool, what a good run looks like, and what to do if it fails. If Dana is genuinely the only one who can do it safely, write the threshold that makes it Dana's problem and what to do meanwhile. Escalation is a fallback, not a step.

Fix it during the incident, not after

Runbooks rot because updating them is a separate chore with its own ticket, and separate chores lose. Do not make it separate. Make the runbook part of the fix.

My rule now: the incident is not closed until the runbook the responder actually used has been corrected. Not in a follow-up. Before the postmortem, before the channel goes quiet. If you hit a stale step, fix it in the same session while the context is still in your head.

Concretely, the closing checklist becomes:

  • Every step you ran: did it work as written?
  • Every step that failed: corrected now, with a command that exists.
  • Any new step you improvised: added.
  • Owner and review date bumped.

Anyone who fought an incident has already debugged the documentation. That knowledge lasts about an hour. Spend it in the same session, on the same page. Otherwise the next responder at 02:14 gets the same broken dashboard, the same dead command, and one less reason to trust the rest of it.


I write about production failures in Postgres, queues, and distributed systems.

Subscribe by email · RSS · Bluesky

Top comments (0)