I have written a great many root cause analyses for enterprise healthcare clients. For the first year, almost none of them changed anything. The document went to the client, the client filed it, and three months later the same class of incident happened again.
The problem was not the analysis. The problem was that I was writing the RCA for the wrong reader, in the wrong order, with the wrong ending. This is what I changed, and the structure I now use for every major incident.
Who actually reads an RCA
Four people read a root cause analysis, and they want different things.
The client's IT director wants to know two things in the first thirty seconds: is it fixed, and will it happen again. They will not read past the first paragraph if the first paragraph does not answer both.
The on-call engineer who will face the next similar incident wants the timeline and the detection gap: how long between the fault and the first alert, and what would have caught it sooner.
The change manager wants to know whether a change caused it and whether the change process failed.
The auditor, months later, wants evidence that an incident was investigated, actions were assigned, and actions were closed.
Most RCAs are written for none of them. They are written for the engineer who did the investigation, in the order the investigation happened, ending with a list of everything that could conceivably be improved. That document serves the writer and no one else.
The structure that works
1. Two-sentence summary, up top
Before the title if you can get away with it. Sentence one: what broke, for whom, for how long. Sentence two: the root cause and whether the permanent fix is in place.
Between 02:14 and 04:51 IST on a weeknight, the interface engine at a hospital site stopped delivering lab results to the EHR, affecting several departments. Root cause was a transaction log that filled the disk after a backup job failed silently; the log has been moved to a monitored volume and the backup job now alerts on failure.
The IT director stops here, satisfied. Everyone else keeps reading.
2. Impact in the client's terms, not yours
"Service degraded for 157 minutes" means nothing to a hospital. "Lab results for the emergency department were delayed by up to two and a half hours; no results were lost; a small number were re-sent manually by the lab" is what they need. Say what happened to the people who use the system. If you do not know, ask the client before you write the RCA, not after.
3. Timeline: fault, detection, response, resolution
A table, not prose. Four columns: time, what happened, who or what noticed, what was done. Every row sourced from a log, a ticket or a chat timestamp, never from memory.
| Time | Event | Detected by | Action |
|---|---|---|---|
| 01:40 | Nightly backup job fails; exits without error code | Nobody | None |
| 02:14 | Transaction log fills disk; interface engine stops writing | Nobody | None |
| 02:51 | Lab calls the help desk: results not appearing | Human, 37 min later | P1 raised |
| 03:05 | On-call engineer joins bridge | Ticket | Investigation |
| 03:40 | Disk identified as full; log truncated | Engineer | Service restored 03:52 |
| 04:51 | Queued messages fully replayed | Engineer | Incident closed |
The timeline is where the real findings hide. In this one, the fault happened at 01:40 and the first human knew at 02:51. That seventy-one-minute detection gap is a bigger finding than the disk being full.
4. Root cause: one sentence, then the chain
State the root cause in one sentence a non-engineer can repeat. Then show the chain of contributing causes, because there is always a chain.
Root cause: the backup job failed without raising an alert, so the transaction log grew unchecked until the disk was full.
Contributing causes: (a) the backup job's exit code was not monitored; (b) the disk hosting the transaction log had no growth-rate alert, only a static 90% threshold that fired at 02:11 into an unmonitored mailbox; (c) the interface engine does not raise its own alert when it stops writing.
Use "five whys" if it helps you find the chain, but do not put "Why 1, Why 2, Why 3" in the document. The client does not care about your method.
5. What we did, and what we will do
Two lists, clearly separated. Immediate actions are done and dated. Preventive actions each have an owner, a due date and a way to verify closure. An action without an owner and a date is a wish.
| Action | Type | Owner | Due | Verified by |
|---|---|---|---|---|
| Transaction log moved to dedicated 200 GB volume | Immediate | DBA lead | Done | Change record |
| Backup job wrapped to alert on non-zero exit | Preventive | Backup owner | Two weeks | Test failure raises P2 in ServiceNow |
| Disk growth-rate alert added to all DB volumes at the site | Preventive | Monitoring lead | Two weeks | Alert definition attached |
| Interface engine "no writes for 5 min" alert added | Preventive | Monitoring lead | Two weeks | Alert definition attached |
Three to five actions. If you have twelve, you have not decided what matters.
6. Detection gap, called out on its own
I now give this its own heading because it is the finding clients act on most. "The fault was detectable at 01:40 and was detected by a human at 02:51. With the two alerts above in place, detection would have been at 01:41." A number like that gets budget approved.
Three things to stop doing
Stop writing it chronologically. The investigation went: symptom, hypothesis, dead end, second hypothesis, cause. The RCA goes: cause, impact, timeline, actions. Nobody needs your dead ends.
Stop blaming, and stop hiding. "Engineer X truncated the log without approval" belongs in the timeline as a fact, if it happened, with no adjective attached. But "human error" is never a root cause; it is a symptom of a missing control. Ask what would have made the error impossible or harmless, and write that as the action.
Stop sending it a week later. A draft RCA within 24 hours, with "under investigation" where needed, beats a perfect one after seven days. The client's leadership is being asked what happened today, not next Thursday.
The test
Before you send an RCA, read only the first paragraph and the actions table. If a reader who saw nothing else would know what happened, why, and what will stop it recurring, the document works. If they would need to read the middle to find that out, rewrite the top.
A root cause analysis is not a record of your investigation. It is the instrument that makes the next incident not happen. Write it that way and people will read it.
Abhishek Singh leads incident management and monitoring for enterprise healthcare clients on Azure, with seven years in enterprise IT operations including five at Acquia as a Senior Support Engineer.
Top comments (0)