When Both Answers Are Defensible, the Right Answer Is That It Is a Coin Flip
The scene below is a composite, written for this piece. It is not a recording of one particular engagement, and the people in it are invented. The shape of the decision is not: it is the kind of question people actually bring to a judge.
Two people run a small studio. A former colleague, now running operations at a company that grew faster than its systems, offers them a year-long retainer. One client, paid monthly, roughly what their three current clients pay together for the same number of hours. The work is unglamorous and stable. The alternative is what they already have: three clients, none of whom could end their year on their own, a pipeline that has to be refilled every few months, and enough variety that neither of them has to become a specialist in something they stopped enjoying.
Write the two columns out and nothing breaks. The retainer buys a floor and costs optionality: sign it, and for a year their capacity is spoken for, and the next interesting inbound email gets answered with "not now". Staying as they are buys variety and costs the floor: a bad quarter is survivable, a bad half-year is not, and the two of them carry that risk personally, in a business with no investors and no cushion. Nobody in the room is short of information. They can list every client, every rate, every hour, and what each branch would do to next year's cash. What is missing is a criterion that decides between the columns, and there is a real possibility that no such criterion exists yet.
That is not a failure of preparation. It is the most common shape of a decision that changes a year: both branches survive an adversarial reading, the costs are asymmetric, and the thing that would break the tie is not a fact about the world but a preference that has not been written down yet.
What we ask a judge to do, and why it breaks here
A judge, human or machine, is usually asked some version of: which one is better? The question quietly assumes there is a fact of the matter. Where the gap is wide, that assumption is safe and the answer is worth having. In the scene above it is not safe. One option is better on stability, focus and sleep; the other is better on variety, upside and the freedom to say yes to something better in eight months. Those are not the same units, and nothing converts them.
This is where a decision tool is most often asked to do something it cannot do. Pressed for an answer, it will give one, because giving one is what it was built to do. It will pick the retainer, or pick staying independent, and it will say so in a confident voice, and the person reading it will feel that a question has been resolved. Nothing has been resolved. A tie has been relabelled as a verdict.
The most useful sentence a judge can produce
Here is what we think the useful output looks like in that situation: Option A, 60%. The two are close.
That is not timidity, and it is not the judge refusing to work. It is the judge reporting the shape of the problem: the candidates are close enough that the pick is being decided by something outside them, which means the deciding criterion is one the person holds and has not yet articulated. Knowing that before you sign is worth as much as knowing which way a lopsided comparison went. It changes what you do next. You stop hunting for the missing fact, because there is no missing fact, and you start asking what you are actually optimising for.
The opposite behaviour is easy to miss because it sounds like strength. A judge that reports high confidence on everything is not a better judge. Its number has been flattened into decoration: it separates no cases, it cannot be routed on, and a threshold rule built on it will fire the same way every time. A confidence value earns its place by moving. If it never moves, it is not information about the decision. It is a tone of voice.
What the number looks like when someone checks it
Confidence is only worth reading if someone has measured what it meant. Ours is calibrated on JudgeBench, self-run under the official judging protocol over 620 pairs, with the six pairs whose first verdict failed disclosed rather than imputed. On raw accuracy that run gives 92.5% for Decider against 92.2% for a plain direct model baseline. That is a tie. We publish it as a tie, and we claim no accuracy advantage over a direct model call, or over anyone else.
What we do publish is the table the number came from:
| Reported confidence | Judgments | Share of run | Actually right |
|---|---|---|---|
| 90% or above | 283 | 45.6% | 99.6% |
| 80–90% | 184 | 29.7% | 94.0% |
| 70–80% | 82 | 13.2% | 84.1% |
| Below 70% | 65 | 10.5% | 67.7% |
Read the bottom row first, because it is the row this article is about. When the judge reported below 70% confidence, it was right 67.7% of the time. On a two-option comparison, that is a system telling you it is staring at something close to a coin flip, and being roughly honest about it. The low number is a measurement of how little separates the two candidates.
Then look at the distance between the top row and the bottom row: it is a little under 32 points. That distance is the information. A judge whose confidence is always high produces a flat version of this table, where every bin reads the same and the number tells you nothing you did not already assume.
The benchmark built to make this hard
Our second published set is ContextualJudgeBench, self-run over the full official set of 2,000 pairs under the official pairwise protocol, with 1,991 pairs completed. In our scoring, a pair counts as correct only when the judge picks the right answer in both presentation orders, which is why the random floor is 25% rather than 50%. Twelve orders, 0.3% of the total, were excluded after repeated platform failures, and that exclusion is disclosed next to the result rather than buried. On that set our consistent accuracy is 67.1% against the benchmark's official reference value of 65.4%, measured on the same judged set.
The interesting part is not the top line. It is that the benchmark deliberately contains near-tie splits, and on those the consistent accuracy sits in the 46–60% range. That is the difficulty its own authors designed in, and the benchmark file says so. It also does not help, because nothing helps on a pair where both answers are defensible.
The same confidence bands mean something different there. Judgments reported at 90% or above came back right 83.3% of the time across 789 judgments. Judgments reported below 70% came back right 55.4% of the time across 529 judgments. Same interface, same scoring rule, and the high band is worth more than 16 points less than it was on JudgeBench.
Sit with that for a moment, because it is the reason this article does not end with "the number is reliable". Calibration is a property of a benchmark and a population, not a property that travels for free between them. On a set full of genuinely close pairs, below-70% confidence really does mean close to a coin flip, and the number is still doing its job: it is telling you that this benchmark, and possibly your problem, is full of decisions the answers do not contain.
Use the number as a valve, not as a verdict
This is the practical consequence, and it is the reason a low number is worth paying for.
Treat confidence as a routing signal, not as a badge. In the high band, act on the pick: the comparison was decisive and the reasoning is there if you want to check it. In the middle band, read the argument and move: the call holds up but a quick look is cheap insurance. In the low band, stop treating the judge as the decision-maker. The pick is nearly arbitrary, and the useful work has moved to you: write the criterion down, take the branch you can reverse, ask the person who pays if it goes wrong, or accept that this is a coin flip and flip it on purpose and move on.
A judge built to run at volume is designed to act above a threshold and hand the rest upward; that hand-off is only as good as the number that triggers it. Ours is the branch that gets handed the low numbers, which is exactly why the low bins are the ones we publish in full.
The failure this prevents is not being wrong. It is acting as if a bare conclusion were a decision. "Option A" with no number and no sense of the spread has no downstream branch: either you obey it or you ignore it. A calibrated pick has two branches built in, and the branch you take tells you what kind of problem you are actually holding. Sometimes that is a comparison a machine can settle. Sometimes it is a preference only you can supply. A judge that always sounds certain cannot tell you which one you are in.
So when both answers hold up and the judge says the two are close, that is the job being done, not the job being dodged. The evidence stops; your criterion starts; and the next move is to name that criterion out loud. If, after naming it, the two options still balance, then it really is a coin flip, and the honest version of that sentence is worth more than a fabricated verdict.
What this does not fix
Calibration is not correctness. On the JudgeBench run, judgments reported at 90% or above were still wrong roughly four times in a thousand. On the ContextualJudgeBench run, the same band was wrong far more often than that. A single judgment can be wrong, a well-calibrated judge can be wrong on the one case you cared about, and a 60% call is by construction the kind of call that goes the other way four times out of ten.
For anything expensive, one-way, or reputationally exposed, the position written into our own benchmark file applies: set your own threshold from the accuracy column, and apply your own review policy on top of it. Our published bands describe a comparison on a named benchmark under a named protocol with a stated number of exclusions. They are not a statement about the outcome of your particular decision, and we do not promise that an output is ready to use as it lands. The signature is still yours.
If you want a hard question and two candidate answers put in front of a judge that returns a pick, a calibrated confidence, and a 900–1,700 character argument for it, start here: Decider.
Top comments (0)