Originally published on hexisteme notes.
I run a review step that sends the same question to two models from different vendors and reads back structured verdicts. On one batch of four rulings the two legs picked different answers on two of them.
A 1–1 split between two voters is not a tie you can break by counting. It is a coin flip.
Both splits resolved cleanly anyway, and neither resolution involved the picks. It involved what both legs had thrown away.
A panel of two is not a smaller panel of three
The panel was supposed to have three legs. The third — a CLI worker from a third vendor — had hit its usage cap that morning, with a reset date three days out. I recorded the gap instead of quietly shipping a two-leg result as if it were the designed one, and then had to actually work out what a two-leg split means, because the majority I would normally have reached for did not exist.
It is worth saying plainly, because the arithmetic is easy to skip: majority voting needs at least three independent voters. With two, "2–0" is agreement and "1–1" carries no information if the only thing you recorded is the pick. The fix is not a third leg. The fix is to record more than the pick.
The brief format that makes a split readable
Every option in the brief is numbered, and every answer has to come back in one shape:
N. <choice> — <2–3 lines of reasoning> — falsified if: <one line>
That last field looks like paperwork and does half the work in this post. A model asked to name the condition under which its own answer would be wrong will, often enough to matter, name a condition that is already true. You cannot read that if you never asked for it, and no amount of re-reading the choice will recover it.
Split one: the picks overlapped even though the letters didn't
The question was whether an abrupt change in a character's behaviour needed setup on the page. Four options went out:
- A — strengthen an existing one-clause link back to an earlier chapter
- B — put a new fact on the page so the acquisition becomes the protagonist's own choice
- C — A and B together
- D — leave it, and make the change a question for a later chapter
Leg one picked A. Leg two picked C.
Tally: 1–1, deadlock. Overlap: both picks contain A. C is A+B. The only thing actually in dispute was B.
B then lost to a domain invariant rather than a vote. It required inventing a fact the project's canon did not have, and the project has an explicit rule against speculative canon — not a style preference, but the constraint that keeps a long series from contradicting itself twenty chapters later. An option that can only be executed by breaking a standing invariant isn't a candidate; it's a bug report about the brief.
Then the falsification field paid for itself. The leg that chose C had written its own falsifier as, in substance, "wrong if the coincidence is later revealed as a deliberate setup by the adults." That is the central motif of the work. The leg had handed over the condition that invalidated its own pick, in the same answer, unprompted.
Three independent reasons converged on A. None of them was the tally.
The part neither leg saw
Executing A meant opening the manuscript at the clause to be strengthened. It pointed at the wrong place: the clause named one location, and the scene it referenced happens somewhere else entirely.
The flow reviewer had read eight chapters end to end and had not caught it. That axis reads who, when, and why; it does not check where. A panel that agrees is still only as wide as the axes you gave it, and one leg can raise an objection without being able to settle one.
Split two: this time it was the rejections that overlapped
Second ruling: whether a transgression at the climax leaves a visible price on the page. Options:
- A — leave it; the price is already there as a transfer (a chill moving from the wrist to the chest) rather than a mark
- B — add a mark on the body
- C — make the outcome worse
- D — add one more beat of the character noticing
Leg one picked A. Leg two picked D. 1–1 again.
Both rejected B and C — and rejected them for the same structural reason. B and C each reframe the event as an incomplete repayment, and an earlier ruling in the same project had fixed the opposite invariant: handing the object back is a registration, not a repayment. That distinction is the engine the whole series runs on. Two models arriving independently at "these two options switch off the engine" is a far stronger signal than either one's preference between the survivors.
That left A and D, and ground truth cut it. D asked for one more beat of the character noticing. Grep the chapter: the beat is already there, twice.
Why rejection is the low-noise channel
Across those four rulings one leg picked option A every single time — 4 for 4. The other picked A twice, C once, D once.
That is disposition. How interventionist a model is by default is a real, stable property, and it rides directly on the choice. If you tally choices from a two-model panel, a good fraction of what you are measuring is which vendor's model is more eager to change things.
It does not ride the same way on the rejection. In those four rulings neither model ever picked an option the other had explicitly ruled out. The two unanimous rulings were unanimous on the reject side too: in one, both legs named the same bad option — make a required phrase "appear" by relocating it into an unrelated scene — and gave nearly the same reason. The string crosses; the institution behind the string does not.
Four rulings is an anecdote, not a study, and I'll take the correction if the next twenty go the other way. What makes me willing to act on it now is the mechanism rather than the count. A pick is one sample from a distribution that training shaped. A rejection with a cited reason is a claim that a specific constraint was checked and failed. Those are different kinds of statement, and counting them in the same unit is the actual error.
When the overlap doesn't decide
Same week, a different gate, two legs split on a diagnosis rather than a menu. The question was whether a batch of worldbuilding passages sat inert next to the main conflict. One leg read eight of eight sampled passages as unrelated to conflict. The other read five of eight as conflict-generating and named which conflict each one fed — this rule powers that theft, that rule powers that concealment.
There was no shared rejection to read, because there was no menu to reject from. The tiebreak was evidence: one leg cited passages, the other asserted a summary judgement.
Citations beat verbosity, and this is the tiebreak most worth naming out loud, because length is not evidence and the longer answer always feels more considered. The cited answer also changed the diagnosis rather than settling it: the disease was not inert exposition. It was re-introducing a rule that had already done work in an earlier chapter — a duplicate, not a filler. No vote-counting rule would have produced that, because the winning answer wasn't one of the options.
The procedure I use now
- Require a falsification line per option in the brief. Without it, step 5 has nothing to read.
- Transcribe the answers into rejections, not picks. One row per ruling:
| Ruling | Leg A picked | Leg B picked | Both discarded | Still standing |
|---|---|---|---|---|
| Climax price | A | D | B · C | A · D |
- One option standing → done. Two or more → step 4.
- Cut the survivors against a domain invariant — a contract, a schema, an architectural constraint, anything the project already committed to. An option that can only be executed by breaking one is out regardless of who picked it.
- Read each leg's own falsifier. If the condition it names is already true, that leg has invalidated its own pick for you.
- Record what decided the rejection, not what was chosen. The rejection reason is what gets reused next time; the choice is a fact about one ruling.
What would prove this wrong
- If, over the next several splits, the common rejection routinely leaves two or more options standing, then this is not a decision rule — it is a menu-narrowing step, and the real work lives in steps 4 and 5. Still useful, but I would stop describing rejection as the thing that decides.
- If I hit a split where both legs rejected the same option and the rejection was wrong, both for the same reason, then rejections are correlated in a way this post does not model and the "the noise cancels" claim goes with it. Correlated error is exactly what a same-vendor panel produces, which is why the legs come from different vendors — and why routing a claim to the verifier that can actually check it matters more than adding legs.
- If a leg's falsification lines turn out to be decorative — never naming a condition that is already true, or naming one so vague it can't be evaluated — then step 5 is theatre and the brief needs a different field.
This isn't really about models
Two human reviewers on a pull request. Two static analysers with overlapping rule sets. Two vendors answering the same architecture question. The move is the same: ask each one what they would rule out and why, not only what they would do.
Vetoes intersect more cleanly than recommendations, because a recommendation is partly a statement about the recommender and a veto with a reason attached is a claim about the artifact. One of those is auditable. And when you have exactly two reviewers, auditable is all you have — there is no majority to hide behind.
More notes at hexisteme.github.io/notes.
Top comments (1)
The falsified-if field is the part I'm going to pull out of this. We run a similar two-leg pattern for entity merge decisions and kept getting stuck on 1-1 splits because we were only recording the merge/no-merge vote. After reading this I want to add a reject-if field to the structured output: the condition under which the model's own recommendation fails. The part about the leg handing over its own invalidation unprompted is exactly what we've seen when we ask models to explain their confidence, they often volunteer the counterargument before we ask for it.