DEV Community

Cover image for The Outcome Is Not Evidence: How to Grade a Decision You Will Only Make Once
TuringCorp
TuringCorp

Posted on

The Outcome Is Not Evidence: How to Grade a Decision You Will Only Make Once

The Outcome Is Not Evidence: How to Grade a Decision You Will Only Make Once

Twenty months ago a company passed on acquiring a small competitor. The competitor was bought by somebody else, grew faster than the model predicted, and is now the reason this quarter's numbers look soft. In the review, the sentence used is: we got that one wrong.

The decision may well have been wrong. The outcome does not establish it.

A decision made once produces exactly one result, and that result is a single draw from a distribution the team never got to sample. Where a call is repeated a thousand times, the rate is the evidence and the individual result is noise. Where a call happens once, the individual result is all you have, and it is the weakest evidence in the room. Reviewing it as though it were a verdict on the thinking is how an organization learns the wrong lesson from the few observations it will ever get.

Two questions wearing the same sentence

"Was the decision good?" and "Did it work out?" are different questions. In a high-volume process they collapse into each other: make the same call ten thousand times at a stated confidence, and the count of good outcomes becomes a test of the decision rule. That is what a calibration table is. It converts a decision into a rate, and a rate into something that can be graded.

For a decision made once, nothing collapses. The call was a bet with stated odds, or it was not. The outcome is one card turned over. A 30 percent chance that arrives is not a mistake, and a 40 percent call that happens to win is not a triumph. If your review cannot separate those two sentences, it is not reviewing the decision. It is narrating the result backwards and calling the narration wisdom.

This is not an argument that outcomes do not matter. It is an argument about what they can be asked to prove. For a one-off decision, the outcome tells you what happened. It cannot tell you whether what happened was knowable, whether the option set was complete, whether the reasoning held, or whether the person who signed had done the work and got unlucky.

The only gradeable object was written before the result

To review a decision rather than an outcome, you need the claim as it was made, not as it is remembered. Four things have to exist before anyone knows how it ends: the options that were actually on the table, the pick, the confidence attached to the pick, and the reasons, in prose, with the assumptions named.

With those four, a stranger can do something real a year later. They can ask whether the confidence was honest, which is a question with an answer. They can ask whether the conclusion followed from the reasons given. They can ask which of the written assumptions turned out false, which is the most instructive question in any post-mortem, because it produces something to carry forward instead of a verdict on a person.

Without them, the review has only memory, and memory is reconstructed from the outcome. Everyone will sincerely recall having had doubts. A doubt that went unrecorded is not evidence, and the honest version of that sentence hurts: if the process left no record, the outcome is the only thing left to grade, so the outcome is what gets graded, every time.

What a confidence number buys a review

A stated confidence is the part of the record that makes a single decision reviewable, and it is why we publish ours as bands rather than as one headline figure.

The numbers are self-run on named benchmarks, with the failures disclosed rather than quietly retried. On JudgeBench, 620 judgments, with the 6 first-verdict failures disclosed, raw accuracy came in at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we report it as a tie; the point here is not that the pick is sharper. The bands are the useful part: calls we reported at 90% confidence or above were right 99.6% of the time, the 80–90% band 94.0%, the 70–80% band 84.1%, and calls below 70% were right 67.7% of the time.

Read the bottom band as a review instruction. A decision reported at 67.7% that went the other way has told you almost nothing. Three of them, read together, have told you something. A decision reported at 94% that went the other way is worth an afternoon, because the interesting question is no longer whether the outcome was bad but what the instrument missed. A single bad outcome and a broken process are different findings, and only a stated probability lets you tell them apart after the fact.

Outcome bias is worse when the news is good

Bad news gets reviewed. The failure is loud, somebody asks for a meeting, and the decision gets examined, often unfairly, but examined.

Good news gets no review at all, and that is the more expensive failure. A win confirms the process without testing it, so an organization can spend years repeating a coin flip that happened to land well and call it a method. The person who got lucky learns the wrong lesson with more confidence than the person who got unlucky, because nothing contradicts them. If your post-mortems only fire on failure, your process is corrected in one direction only.

The inverse is quieter and just as damaging. When a defensible decision is punished for a bad outcome, the lesson the room takes is not "run a better process." It is "do not be the one holding a defensible decision when the number is bad." That lesson produces fewer hard decisions, more decisions deferred to a calendar, and a portfolio of choices that were only ever safe because nobody with a name on them took a risk.

What a review should actually ask

Six questions, and none of them is "did it work."

Was the option set complete, and who was allowed to add to it? Would this decision have looked different if the result had been known in advance, and if so, what was actually written down beforehand? What confidence was stated, and is the outcome inside that band? Which written assumption turned out false? What would have changed the decision at the time, and was that thing knowable then? And finally: what does the next person facing this question get from this record?

That last question is the standard. A review is not a court. It is the mechanism by which one decision becomes usable by somebody who was not in the room, which is also the only way a decision made once pays for itself more than once.

The practical habit is small. Before the call, write half a page: the options, the pick, the confidence as a number, the reasons, and the trigger that would reopen it. Date it. After the call, read the page before you read the outcome. That ordering is the whole discipline. Do it in the other order and the outcome quietly rewrites what you think you wrote.

Why a verdict alone cannot be reviewed

This is where our own product is on the line.

A judge that returns a bare conclusion hands the review one bit, and the review will always interpret that bit through the result. A judge that returns a pick, a calibrated confidence, and a written argument hands the review something it can grade independently of how it ended, which is the only kind of grading available when the answer arrives once.

This series has spent weeks arguing about which decision belongs to which regime. This is the piece about what happens to a decision afterwards, which is where most of the learning is supposed to happen and where almost none of the writing exists. Decider exists for the decision you will make only once: one question, two candidate answers, a pick, a confidence you can hold it to later, and an argument long enough to disagree with, which is also long enough to be reviewed by somebody who was not there.

You will get one observation. Keep the claim.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to