An outage does two kinds of damage at once. The technical damage is what the incident is about. The trust damage is what the silence is about. Customers forgive a broken checkout for an hour; they forgive a vendor who said nothing for an hour much more slowly — because silence reads as "they don't know."
The fix is not eloquence. It is a schedule, written in daylight: who gets told what, at which minute.
The rule that beats every template: tell them before they ask
The moment a customer has to ask what is wrong, you have already lost the update race. So the timeline starts earlier than most teams think — the first update goes out before diagnosis is complete.
"We know checkout is failing, we are on it, next update by 14:30."
That is a complete first update. No diagnosis needed. What it contains is proof of existence: someone is awake, the thing is owned, and there is a time when you will hear more.
Three audiences, three clocks
- Users / customers — first word <=15 min, then every 30 min while degraded. Status page or in-app banner.
- Your own staff — briefed immediately at declaration, same cadence as customers. A support rep who learns about the outage from a customer tweet will improvise a promise engineering has to honor.
- Leadership — at declaration and every severity change. Severity, business impact, decision asks.
The internal audience is the one small teams forget — and it's the one that burns you.
The five-part update skeleton
Every update — first, middle, or final — has the same five slots:
- Status line — one sentence: what is degraded, for whom.
- What we know — facts, no speculation.
- What we're doing — present tense, no heroics.
- Workaround — if one exists, formatted as an instruction.
- Next update time — a clock time, never "shortly".
Slot 5 is the single highest-leverage sentence in incident communications. It converts your audience from refreshers into waiters. And if you can't make the next update time, the update saying you'll be late is itself an update — it preserves more trust than a silent twenty minutes past the promise.
The single-writer rule
All customer-facing words come from one comms lead per incident. Three writers means three versions of "what we know," and the audience will notice the differences and read them as concealment.
What never to write mid-incident
- ETAs you invented — "fixed in 20 minutes" becomes a promise with your name on it. Give a next-update time (which you control), not a next-fix time (which you don't).
- "Everything is fine" — if customers can see the failure, this sentence costs you the next three updates' credibility.
- Names and blame — causes with blame stripped out belong in the blameless review, published later, on purpose.
- Jargon — "p99 latency on the ingest path" means nothing to a shop owner whose card reader is down. Write the customer's sentence.
Worked example: 47-minute gateway outage, zero "what is happening?" tickets
A twelve-person e-commerce SaaS lost card checkout at 14:02. Declared at 14:05, comms lead named in the same message. First customer update at 14:08 — before anyone knew the cause. Updates at 14:30, 14:49, 14:55. Total customer-visible silence during the entire outage: zero. Support received no "what is happening?" tickets. A churn-risk customer replied to the final update: "first vendor that told me before I noticed."
Metrics that keep the timeline honest
- Time to first update — target under 15 minutes, every incident.
- Cadence kept % — promised next-update times met. Below 90%, update times have become aspirations.
- Post-mortem published <=72 h — the loop-closer that makes your next first update read as information, not PR.
The full template — the timeline table, the copy-paste update skeleton, channel decisions, and the worked example — is on the site: Incident Communication Timeline Template for Small Teams.
Top comments (0)