DEV Community

Cover image for Order invariance is not a spot check. It is the denominator.
TuringCorp
TuringCorp

Posted on

Order invariance is not a spot check. It is the denominator.

Order Invariance Is Not a Spot Check. It Is the Denominator.

For the past two weeks, the most useful question about Jev has not been "how fast is it" or "how cheap is it." It has been: what happens when you swap the two candidates?

The reports coming back from independent community audits say the answer is "something": changing the order of the options visibly moves the output probabilities. That is not gossip — several of these audits are public, reproducible, and some were pre-registered (see Sources). I want to be careful about how I frame it, because the interesting part is not that a model behaves this way. Pairwise judges are sensitive to presentation order in general. It is one of the oldest known failure modes in this corner of evaluation, it shows up across model families, and it is a property you either design around or you do not. It is not a character flaw.

The interesting part is what a team does once it knows.

Two kinds of parts

There are two things you can put in a pipeline, and they look identical from the outside.

The first is a judge you permute. It runs at high volume, code acts on its output, and order sensitivity is a tuning problem: you randomize the presentation, you average over a few permutations, you move on. For that job, a cheap fast judge with a wobble is still an excellent part. This is the job Jev was built for, and it is genuinely well built for it.

The second is a judge you threshold on. Here the confidence field is not a diagnostic — it is load-bearing. Somebody downstream writes a rule against it: above this band act, in the middle band ask, below that band escalate to a human or to a slower system. The moment that rule exists, the number stops being a metric and becomes an input to policy.

Those are different kinds of parts, and they demand different kinds of evidence. A benchmark accuracy score is evidence for the first. For the second, a score is close to irrelevant — what you need to know is how the confidence number itself was validated, in the exact form you are about to consume it. A number that was measured under order A and applied under order B has not been validated for your use.

That is the whole argument. It is not a claim about anyone's model. It is a claim about what "confidence" has to survive before you are allowed to build a switch statement on it.

What our protocol actually says

We publish a benchmark file at api.turingcorp.net/benchmarks/latest.json, and the line that matters here is the protocol field for ContextualJudgeBench. Verbatim:

Self-run with the official vanilla pairwise protocol; consistent accuracy (both response orders judged correctly); random floor 25%; failures disclosed (12 orders rerun-excluded). Reference model measured on the same judged set.

Read the parenthetical twice, because it is the entire point. In our benchmark, a pair counts as correct only when the judge picks the right answer in both presentation orders. Order invariance is not an audit we run afterward, and it is not a caveat in a footnote. It is inside the definition of "correct." A single flipped order converts a scored success into a scored failure.

The consequence is visible in an unusual place: the random floor. Guess at chance on a two-option pairwise task and you get 50%. Our published floor is 25%, precisely because a coin-flip judge has to get two independent orders right, and 25% is what that costs. We think that is the honest denominator for a system whose output feeds a threshold rule, and we would rather publish a lower floor against a stricter definition than a higher number against a looser one.

The same file's footnote states the rule a second time and then discloses the damage:

Self-run with the official protocol over the full 2,000 official pairs. Consistent accuracy requires the same correct pick in both presentation orders (random floor 25%). 12 orders (0.3%) were rerun-excluded after repeated platform failures; the reference model was measured on the same judged set. Previously published results measured on a subset are archived in the repository changelog.

Twelve orders were excluded rather than counted, and rather than quietly imputed. On the full 2,000-pair official set, 1,991 pairs completed, with a self-run consistent accuracy of 67.1% against the benchmark's official reference value of 65.4%.

I want to flag the least flattering number in that file before anyone else does. ContextualJudgeBench contains splits that were deliberately built as near-ties, and on those the consistent accuracy sits in the 46–60% range. That is not a bug we are hiding; it is the benchmark's own difficulty design, and order invariance does not rescue a pair where both answers are defensible. Invariance buys you the right to compare numbers across orders. It does not buy you the answer.

The part we are not allowed to skip

The other benchmark we publish is JudgeBench: 620 pairs. Our protocol line reads:

Self-run with the official judging protocol. Both columns use first successful verdict per pair; the six pairs whose first verdict failed are disclosed rather than imputed. Reference model measured on the same judged set.

Six pairs failed on first verdict, and we disclose them instead of scoring a retry as if it were the first attempt. On that set, Decider measured 92.5%. The plain direct baseline measured 92.2%.

That is a tie. We report it as a tie, and we do not make an accuracy claim over it — not against a direct model call, and not against anyone else. If you came here for a number that lets us look sharper, this is the wrong article.

What we do put weight on is a different column, also from that file and also a self-run with the failures disclosed: judgments reported at ≥90% confidence were right 99.6% of the time on that benchmark, and judgments reported in the 80–90% band were right 94.0% of the time. Those are the bins a threshold rule would actually read.

That is not a statement that we are sharper than a direct model call — we measured a tie, and the tie is the finding. It is a statement about what the number attached to a judgment has meant historically, on a named benchmark under a named protocol with a stated exclusion count.

What is public on the other side

TypeSafe has not published a reliability diagram or an expected calibration error for Jev. That is a statement about what is publicly available, not a statement about what is true internally — the training method is described as reinforcement learning aimed at calibrated decisions, and it may well be producing exactly what it claims. Independent audits have started filling that gap from the outside, which is a healthy sign for the ecosystem (Sources); but if you are the person writing the threshold, "not published by the vendor" is the answer you need to have.

The reason this matters is not that Jev is dubious. It is that Jev's own recommended usage is threshold routing: act on high confidence, and send low confidence to a human or to a stronger system. That guidance is sound, and it puts the reliability question directly on the critical path of the intended deployment. Nobody is being ambushed here; the question is simply upstream of the use case.

So the two systems are not rivals. One is a decision primitive optimized for volume, and one of them — ours — spends its budget on the two things you cannot get from a cheaper judgment: a written argument, and a confidence number whose bins are published with the exclusions that produced them.

Three tests you can run on any confidence field

If a decision interface hands you a confidence value and invites you to build policy on it, you can check it yourself. Three tests, in order of cost:

1. Permute the input. Send the same question with the candidates in the opposite order. Then send it again with a different permutation. A confidence number that moves when the order moves is reporting formatting, not comparative strength. If your pipeline consumes that number, the movement is now inside your policy.

2. Ask for the reliability diagram. Not the headline accuracy — the calibration curve. Which bins exist, how many cases fall in each, and what fraction of each bin was actually right. If the diagram does not exist, you have learned something load-bearing about the interface, and you have learned it before you shipped.

3. Read the denominator. Find out what happened to the cases that failed. Were they rerun, excluded, imputed, or dropped from the count? A protocol that names its exclusions is telling you how much to trust the number above them. A protocol that has no exclusions to name is usually not a protocol with no failures; it is a protocol that has not looked.

Apply these to any judge, including ours. Everything the three tests ask for is spelled out above: the permutation requirement is inside our scoring rule, the bins and their case counts are in the published file, and the failure counts — 6 for JudgeBench, 12 rerun-excluded orders for ContextualJudgeBench — are printed next to the results rather than in a footnote nobody reads. That is deliberate. A reliability claim you cannot audit is a feeling with a decimal point.

Try to break the bins

Here is the honest limit of everything above. Order invariance raises the bar for being counted correct; it does not certify that the judge is right, and it certainly does not remove the need for your own review policy on decisions that are expensive or irreversible. Our published near-tie band still runs 46–60%, and that is our own benchmark telling on us.

So the useful thing you can do with this article is adversarial, not appreciative. Our confidence bins are published, our exclusions are counted, and the protocol that produces them is quoted verbatim above. If you run decisions through a judge that reports confidence and you find our bins do not hold up, that is a more valuable result than another leaderboard position, and we would like to hear it.

If you want to put a hard question and two candidate answers in front of the judge whose numbers are on the record, start here: https://api.turingcorp.net/platform/go/decider?src=jev-2

Then run the three tests on it. That is the invitation, and it is a real one.

Sources

Everything of ours quoted above is in one file, published and machine-readable: https://api.turingcorp.net/benchmarks/latest.json

The independent Jev audits referenced in the opening, all public:

TypeSafe's own description of the model and its recommended threshold-routing usage: https://typesafe.ai/blog/introducing-system-one-models-and-jev

Top comments (0)