DEV Community

Abhishek Singh
Abhishek Singh

Posted on AI-assisted

Your Disaster Recovery Runbook Is Wrong, and the Auditor Knows It

Every enterprise I have worked with has a disaster recovery runbook. Almost none of them have one that would work at 3 a.m.

I know because I have been on the bridge when a real outage hit a site whose runbook was reviewed and approved eleven months earlier. The IP addresses were wrong. The database had been migrated to a new cluster. The vendor contact had left. The restore order listed a system that had been decommissioned and omitted the identity provider that everything else now depended on. We recovered, but we recovered from memory and from Slack, and the runbook was open on someone's second monitor as a formality.

Then, two months later, the auditor asked for evidence that the DR plan had been tested, and the evidence pack was a folder of screenshots with no timestamps.

This article is about the gap between a runbook that satisfies a review and a runbook that recovers a hospital, and the small set of practices that close it.

Why runbooks rot

A DR runbook is a snapshot of an environment on the day it was written. The environment changes every week: new VMs, a migrated database, a changed replication target, a new dependency introduced by an application upgrade. Nobody updates the runbook for each change because the runbook lives in a Word document on SharePoint and the change lives in a ticket.

The result is that a runbook is most accurate on the day it is approved and steadily less accurate every day after. By the yearly test it is wrong in ways that only show up during the test, and by the real disaster it is wrong in ways that show up at the worst possible moment.

What a runbook actually has to contain

Strip away the boilerplate and a working DR runbook has six parts. If any is missing, the document is a policy, not a runbook.

  1. Declaration criteria and authority. Who decides this is a disaster, on what evidence, and who they call first. Name the role and the person, with a phone number that was verified in the last ninety days.
  2. The dependency order. Identity and network first, then data, then application, then edge. Written as a list that a tired engineer can follow, not a diagram they have to interpret. Every system has a tier, and no tier starts until the previous one is verified.
  3. Per-system restore steps with the exact commands, console paths or recovery-plan names, the expected duration, the owner role, and the verification check that proves it worked. "Fail over the database" is not a step. "Run ASR recovery plan HIS-DB-DR; confirm SQL listener answers on 10.x.x.x; run the smoke query" is.
  4. Contacts — internal, vendor, client — with an escalation path. Verified, dated.
  5. Communication templates. What the client's leadership is told at declaration, hourly, and at resolution. Written in advance because nobody writes well at 3 a.m.
  6. Failback. How you get back to the primary site, which is the half of DR that nobody tests and everyone needs.

The practices that keep it alive

Generate the dependency order from the inventory, not from memory

The single biggest source of runbook error is the restore order being wrong because dependencies changed. The fix is to stop writing the order by hand. Pull the inventory from the source of truth — the cloud provider's resource graph, the CMDB, a maintained CSV for on-prem — and derive the order from the dependency edges. When a new database appears, it appears in the order; when a system is decommissioned, it drops out.

This is not exotic. A spreadsheet with two columns, system and depends on, and a sort is enough for a small site. For a large one it is a query against the CMDB. Either way the order is regenerated, not remembered.

Diff the runbook against reality every week

Once the inventory drives the runbook, checking for drift is cheap: compare this week's inventory to the one the approved runbook was built from. New systems, removed systems, changed replication status. Anything that differs is a flag on the runbook, and the runbook is not "current" until someone has looked at the flag.

In practice this turns the yearly panic into a ten-minute weekly task, and it means the runbook the team opens during a real incident was last reconciled days ago, not months.

Make the test produce evidence as a by-product

The auditor's question is not "did you test DR?" It is "show me". Show the steps that were run, by whom, when, with what result, mapped to the control that requires it — HIPAA §164.308(a)(7)(ii)(D) for testing and revision, ISO 27001 A.17.1.3 for verifying continuity.

The way to have that evidence is to not collect it afterwards. Run the drill from a checklist where each step records the operator, the start and end time, the pass or fail, and a screenshot or log excerpt at the moment of execution. When the drill ends, the evidence pack exists. Nobody spends the following week hunting screenshots.

Two details make the pack audit-proof. Timestamps come from the system, not typed. And the finished pack is hashed — a SHA-256 printed on the cover — so that the auditor, or anyone, can confirm it was not edited after the fact.

Run tabletops between the real tests

A full failover test once a year is expensive and disruptive, which is why it happens once a year. A tabletop — walking the runbook step by step against the current inventory, with the on-call team, without touching anything — takes ninety minutes and can happen quarterly. It finds the wrong IP and the departed vendor contact in a conference room instead of on a bridge.

Close the findings before the next test

Every test produces findings: a step that took twice as long as expected, a system that would not fail over cleanly, an RTO missed. The runbook is only as good as the fixes. Put each finding into the ticketing system with an owner and a due date, and make "open DR findings" a number that appears in the client's quarterly review. Findings with no owner are still open a year later; findings in a review deck get closed.

What the client sees

When this works, the client-side experience changes in three ways. The IT director sees a readiness status that is green, amber or red today, not a PDF dated last spring. The auditor receives a pack that answers the question without a follow-up. And the on-call engineer, at 3 a.m., opens a runbook that was reconciled last Tuesday and follows it.

That last one is the point. Everything else is how you get there.

A checklist

  • [ ] Every contact in the runbook verified in the last 90 days
  • [ ] Restore order generated from the inventory, not written by hand
  • [ ] Weekly drift check between inventory and approved runbook
  • [ ] Every step has command, owner role, expected duration, verification check
  • [ ] Drill run from a checklist that captures operator, timestamps, result, evidence
  • [ ] Evidence pack mapped to controls and hashed
  • [ ] Quarterly tabletop, annual full test
  • [ ] Findings tracked with owners and reviewed with the client

If four or more are unticked, the runbook will fail when it is needed. Fix the order and the drift check first; they remove most of the risk for the least effort.


Abhishek Singh leads incident management and monitoring for enterprise healthcare clients on Azure, with seven years in enterprise IT operations including five at Acquia as a Senior Support Engineer.

Top comments (0)