When a production line stops, a customer reports a defect, or a service goes down, the first explanation is usually the closest one: a bad part, a tired operator, a misconfigured server. Fixing that explanation feels productive, and it often is, for a day. Then the same problem returns, because the thing that produced the symptom was never addressed.
The 5 Whys is a simple method for looking past the first explanation. You ask "why" repeatedly, usually five times, until you reach a cause that you can act on. It was developed within the Toyota Production System and is now used in manufacturing, software operations, incident management, and quality programs. Its simplicity is its strength, and also its main risk: a shallow chain of whys produces a confident, wrong answer.
This article explains how to run a 5 Whys analysis well, where it fails, and how software engineers and engineering managers can use it after incidents and defects. It includes a small Java model that records a why chain and checks whether it reaches an actionable root cause.
The method in one example
Consider a web service that returns errors for twelve minutes every night.
- Why did the service return errors? The database connection pool was exhausted.
- Why was the pool exhausted? A nightly batch job held connections for longer than usual.
- Why did the batch job hold connections longer? It processed a larger file than it was designed for.
- Why was the file larger? An upstream system started sending a full export instead of incremental changes.
- Why was the change not caught? The upstream team does not notify downstream teams about export format changes, and the batch job has no check on input size.
The first answer ("the pool was exhausted") is a symptom. The fifth answer points to a process gap and a missing control, which are things a team can change. Adding more connections to the pool would hide the problem for a while. Adding an input size check and a change notification agreement addresses the cause.
Why the chain matters more than the number five
Five is a convention, not a rule. Some chains stop after three whys because the cause is clear and verified. Others need seven. The goal is to continue until you reach a cause that is:
- Specific, so someone can describe what would have prevented it.
- Controllable, so the team can change it without waiting for someone else's decision.
- Verifiable, so you can confirm that it was actually the cause.
Stopping at "human error" is almost never the right answer. People make errors because systems allow them to, and the useful question is what made the error easy to make and hard to notice.
Where the method fails
The 5 Whys is useful and fragile at the same time. Several failure modes appear again and again.
Single-cause thinking. Real incidents usually have several contributing causes. A linear chain forces one of them to the top and hides the others. When an incident has multiple branches, draw them as a tree rather than forcing a line.
Blame disguised as analysis. If the answers name people, the analysis becomes a search for someone to blame, and people stop reporting honestly. Ask why the system allowed the action, not why the person took it.
Confirmation bias. Teams often reach the answer they expected. Ask for evidence at each step: logs, measurements, or a reproduction. A "why" that is not supported by evidence is a hypothesis.
Stopping at the cause that is easiest to fix. A chain can end at "the team did not have time to write tests," which is true but hard to act on, or at "the test suite does not cover the batch path," which is specific and fixable. Push toward the fixable cause, even if it is less comfortable.
No follow-through. The most common failure is that the analysis produces a good document and no change. Every root cause should lead to an action with an owner and a date, and the action should be checked later.
Using 5 Whys well in software and operations
For software teams, the method fits naturally into incident reviews. A few practices make it more reliable:
Separate the timeline from the analysis. First establish what happened and when, using logs and alerts. Only then ask why. Chains built from memory tend to be wrong.
Write each why as a claim with evidence. "The pool was exhausted, as shown by the connection metric reaching the configured maximum at 02:14." This makes the chain checkable.
Involve the people closest to the work. The engineer who was on call often knows the step that was not written down. Their account is more valuable than a reconstruction from tickets.
Turn causes into controls. Each root cause should map to one of a small set of control types: prevent the condition, detect it earlier, or reduce the impact when it happens. An action that only adds a reminder is the weakest control.
A small model for why chains
The following Java model stores a why chain, checks that each step has evidence, and flags a chain whose final cause is not actionable. It is deliberately simple, because the discipline matters more than the tooling.
import java.util.ArrayList;
import java.util.List;
public class FiveWhysChain {
public enum Kind { SYMPTOM, CONTRIBUTING, PROCESS_GAP, ROOT_CAUSE }
public record Step(String question, String answer, String evidence, Kind kind, boolean actionable) {
public Step {
if (answer == null || answer.isBlank()) {
throw new IllegalArgumentException("every why needs an answer");
}
}
}
public record Verdict(boolean complete, boolean hasEvidence, boolean endsActionable, List<String> issues) {}
public static Verdict review(List<Step> chain) {
List<String> issues = new ArrayList<>();
if (chain.size() < 3) {
issues.add("chain has fewer than 3 whys; the analysis is probably too shallow");
}
boolean hasEvidence = true;
for (int i = 0; i < chain.size(); i++) {
Step s = chain.get(i);
if (s.evidence() == null || s.evidence().isBlank()) {
hasEvidence = false;
issues.add("why " + (i + 1) + " has no evidence: "" + s.answer() + """);
}
}
boolean endsActionable = false;
if (!chain.isEmpty()) {
Step last = chain.get(chain.size() - 1);
endsActionable = last.actionable() && last.kind() != Kind.SYMPTOM;
if (!endsActionable) {
issues.add("final cause is not actionable; keep asking why or reframe the cause");
}
if (last.answer().toLowerCase().contains("human error")) {
issues.add("final cause names human error; ask what allowed the error to happen");
}
}
boolean complete = issues.isEmpty();
return new Verdict(complete, hasEvidence, endsActionable, List.copyOf(issues));
}
public static void main(String[] args) {
List<Step> chain = List.of(
new Step("Why did the service return errors?",
"The database connection pool was exhausted.",
"Connection metric reached max at 02:14", Kind.SYMPTOM, false),
new Step("Why was the pool exhausted?",
"The nightly batch job held connections longer than usual.",
"Batch duration rose from 4 to 11 minutes", Kind.CONTRIBUTING, false),
new Step("Why did the job run longer?",
"It processed a full export instead of incremental changes.",
"Input file size 14x the normal size", Kind.CONTRIBUTING, false),
new Step("Why was the export full?",
"The upstream team changed the export format without notice.",
"Upstream change log, no notification sent", Kind.PROCESS_GAP, false),
new Step("Why was the change not caught?",
"The batch job has no input size check and no change agreement exists.",
"Batch code has no size validation; no contract in the integration docs",
Kind.ROOT_CAUSE, true));
Verdict v = review(chain);
System.out.println("complete=" + v.complete() + " evidence=" + v.hasEvidence()
+ " actionable=" + v.endsActionable());
v.issues().forEach(i -> System.out.println(" - " + i));
}
}
Running this with the example chain returns a complete analysis: each step has evidence, and the final cause is actionable. Change the last step to "human error" with no evidence, and the review flags both problems. The model cannot judge whether the causes are true, but it can refuse to accept a chain that is unsupported or that ends where no one can act.
Practical guidance for engineering leaders
After your next incident or defect review, ask:
- Is the timeline established from data, separate from the analysis?
- Does every why have evidence, or is some of it memory?
- Does the final cause name a control the team can build, or a person to blame?
- Does each root cause have an action, an owner, and a date?
- Did we check, after the change, that the cause is gone?
If the answers to questions 3 and 4 are no, the review was probably a story, not an analysis.
Key takeaways
The 5 Whys is a simple method with a sharp edge. It works when each answer is supported by evidence, when the chain continues until it reaches something a team can change, and when the result leads to a real control rather than a reminder. The number five matters less than the discipline behind it: follow the evidence past the first explanation, avoid blaming people for what systems allowed, and make sure the analysis ends in action that someone verifies.
Top comments (0)