Who Wrote the Ballot: The Judgment That Happens Before the Judge Is Called
Two sites are on the shortlist for a second location. The rent is modeled, the foot traffic has been counted for a month, the fit-out has three quotes, and the spreadsheet has been open for six weeks. Both sites work. The argument in the room is genuinely about which one, and it is a real argument, decided on real numbers.
What nobody says out loud is that the option do not open a second location this year, and fix the throughput problem at the first one was never written on the list.
That is not a failure of analysis. The analysis was careful. It is a failure at a step that happens before the analysis, that is almost never recorded as a step, and that has already determined more of the outcome than every row of the spreadsheet beneath it. The step is deciding what gets to be on the ballot.
Two acts, and only one of them can be bought
Split any decision into two acts.
The first is nomination: which options are admissible. The second is ranking: of the options admitted, which is better. They are performed by different processes, at different times, with different evidence. Only the second one is what every judgment product on the market, ours included, actually does.
A judge receives the option set as an input, the way a calculator receives numbers. It cannot widen the set, because the set is not in its input space. It cannot report that the set is wrong, because nothing in the interface has a slot for that finding. What it can do is answer the question it was handed, precisely and quickly. That is not a limitation anybody is concealing. It is the shape of the category.
Earlier in this series I argued that it is a virtue of typed judgment that the set of possible answers is defined by the caller. This piece is the invoice for that virtue.
Every confidence number is conditional on the list
Read a probability the way a statistician reads it and the missing part becomes obvious. A confidence of 0.9 is not a property of the world. It is a claim that among cases resembling this one, and among the options you supplied, the answer named is right about nine times in ten. The conditioning event is doing half the work, and it is invisible in the output. The number prints to two decimals either way.
So a figure can be honestly measured, published, and still describe a comparison that should never have been run. The instrument is not lying. The question was malformed before the instrument was switched on, and no amount of care inside the comparison recovers an option that was never compared.
Our published measurements are a clean illustration of how tightly the conditioning binds. They are self-run on named benchmarks, with the failures disclosed rather than quietly retried. On JudgeBench: 620 judgments, with the 6 first-verdict failures disclosed. Raw accuracy came out at 92.5% against 92.2% for a plain direct baseline. That is a tie, and we publish it as a tie, with no accuracy claim resting on it. The bands are the part worth keeping: calls we reported at 90% confidence or above came back right 99.6% of the time, and the 80–90% band 94.0%. On ContextualJudgeBench, self-run under the official pairwise protocol in which a pair counts as correct only if both presentation orders are judged correctly, which is why the random floor is 25% rather than 50%, consistent accuracy is 67.1% against the benchmark's official reference of 65.4%, with 12 orders excluded after repeated platform failures and that exclusion stated beside the result. The deliberately constructed near-tie splits sit at 46–60%.
Now look through that paragraph for the row that would tell you whether the right answer was among the two candidates. It is not there, and it cannot be. Every figure measures the ranking act, on questions whose admissible answers were fixed in advance by somebody who does not appear in the measurement at all. That is a genuine measurement, and it covers roughly half of what a decision is.
Better ranking makes a bad ballot harder to challenge
There is a comfortable story in which framing is simply the next unmeasured frontier, and better instruments will eventually reach it. The uncomfortable version is worse than that.
Improving the ranking act does not merely fail to catch a bad ballot. It launders it. A pick reported at 0.91, supported by a published band table, an exclusion count and an owner's name, is a much harder thing to challenge in a review than a shrug. The measured apparatus confers authority on the pair it was pointed at, and the pair inherited that authority by being the only pair on the page. The better the ranking, the more legitimate the framing looks.
That is a risk of our own product, not only of somebody else's, and it is the reason this piece exists. A careful comparison is the best disguise an unexamined list has ever had. A series that has spent two weeks asking judges to publish their curves should say plainly what a curve measures: the comparison, not the choice of what to compare.
Where the options actually go missing
Omissions are not random, and naming the patterns is most of the defense.
The status quo is not printed. Doing nothing, keeping the current system, waiting a quarter, running a two-month pilot before committing — the most common real outcomes are rarely written on a ballot, because they look like the absence of a decision rather than a decision. Once they are off the page, the deliberation proceeds as though one of the listed items must win. And it usually does.
The option that costs somebody in the room is the one missing. A missing item is disproportionately the one that would require a person present to conclude that an earlier call of theirs was wrong, or that a capability they own has quietly become the problem. Omission is a political act performed in analytical vocabulary, which is exactly why it survives review: no scoring rubric can flag a row that was never created.
The third way gets read as a compromise and dismissed. Real options often sit in a different frame rather than between the two on the list: not vendor A against vendor B, but neither, plus one internal quarter spent learning the domain. A two-item ballot converts a framing question into a preference question, and a false dilemma is durable because it looks like focus.
An artifact for the part nobody measures
A ballot is a document, so it can be written, dated and signed like other documents.
Write the list before the criteria. Criteria written first quietly determine which options are capable of scoring well.
Force the null onto every ballot: do nothing, keep the status quo, decide later. Print it. If it is genuinely not viable, the reason belongs on the page beside it.
Give the omission an owner. Before the comparison starts, ask who pays if something is missing from the list, and give that person a documented way to add a row. A line on the page, not a veto.
Commission the list from somebody who disagrees with the likely conclusion, and ask what they would add. Two people independently producing the same list is the cheapest framing check available.
Then ask what the options have in common. If both candidates assume the same thing — that the capability is needed at all, that the market is where it was last year, that this is a hiring problem rather than a scoping problem — the shared assumption is the decision, and the comparison is downstream of it.
Where the judge's responsibility ends
A pairwise judge takes one question and two candidate answers. It can tell you which of the two is stronger and how much confidence it places in that. It cannot tell you the pair was incomplete, and it should not pretend otherwise. The ballot is not an optional preamble to the work. It is the input, and it is the half of the decision the instrument is silent about.
So the honest division is this. Counting carefully is a solved problem, and it keeps getting cheaper; that is the whole achievement of the fast end of this field. What stays expensive, and stays unmeasured, is who was allowed onto the ballot. Write the list down first, put the null option on it, let the person who would pay for an omission sign the page, and only then run the comparison.
If you are the one who has to defend a list as well as a pick — one question, two candidate answers you cannot separate by inspection — the judge that spends minutes on the ranking is here: https://api.turingcorp.net/platform/go/decider?src=jev-13
Top comments (2)
The invoice has a metric-side twin, and it is worth separating from the philosophical one, because "the conditioning is invisible in the output" is only true of a measurement that has no reference whose value you knew in advance.
Put one in and the row you say is missing prints. This week I ran the same procedure on the same pairs and changed exactly one thing: which side's null the reference was scored against. Across the constructions where the reference stays a draw from its own null, the control read 0.469–0.500 — it has to read 0.5 by construction, so that residual is benign. Under the asymmetric conditioning — the null carrying one member's length profile, the reference the other — the same control read 0.778. Same corpus, same pairs, same code: a 0.28 absolute move on a quantity that a priori cannot move at all. That is the framing step showing up inside the measurement, and the reason it is legible is not extra care in the comparison; it is that 0.5 was known before the run.
So the ballot artifact has a companion on the instrument side, and it is mechanical rather than a matter of diligence: force a reference whose value you knew before the run onto every measurement, the way you force the null onto every ballot. Then "was the question malformed" stops being a residue of the ranking act and becomes a printed row — it prints as a displacement, and a displacement of a fixed point is the one error the comparison underneath cannot absorb.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.