We closed just over eleven thousand incidents last year and I can quote the average resolution time for every category of them. What I could not tell anybody, until a finance colleague asked an awkward question about our support costs, was how many of those tickets were the same incident arriving again.
We sampled a single month properly. A little over a third of everything raised traced back to nine underlying faults, each of which had been repaired, correctly and within target, dozens of separate times.
A warehouse job that fails whenever a bank holiday lands on a Monday, resolved by rerunning it, eleven occurrences. A print queue on the second floor that stops accepting work after a few days, resolved by restarting the spooler, fourteen. A reporting tool that locks people out when they open a second session, resolved by clearing the session, twenty-two. A supplier portal that times out for anybody sitting in one particular office, resolved by suggesting they try again later, thirty-one.
Every one of those tickets was closed inside its target. Our figures looked healthy all year, because a measure built on resolutions rewards a department for becoming extremely good at repeating a repair.
The reason none of them were ever corrected is that the correction was in nobody's job. First line closes the ticket, which is precisely what we pay them to do. Second line meets the same fault a few times and writes a knowledge article, which makes the repair faster and the defect permanent. The engineers who could actually fix the warehouse job are committed to project work a year ahead, and there was no route to claim an hour of that on the strength of a pattern that appeared in no report anybody read.
So we started producing the report. Incidents are grouped monthly by cause rather than by category, and the worst ten leave the operational pack and enter the delivery plan with an owner and a date against each. Five percent of the team's capacity is reserved for that work, a figure small enough that nobody argued about it.
Eight of the nine are gone. Ticket volume has fallen by roughly eighteen percent, and all of the fixes together cost less than a fortnight of effort.
Being excellent at recovery is how an organisation learns to tolerate a fault indefinitely.
– Serguey Shinder
Top comments (0)