DEV Community

Cover image for Show Me the Diagram
TuringCorp
TuringCorp

Posted on

Show Me the Diagram

Show Me the Diagram

A vendor sends over an integration guide. It has three columns: a confidence value the system emits, the action you are meant to take at each level, and what each call costs. There is a cut-off at 0.8 and a note that below 0.6 you should escalate. The person who wrote it is competent and the system works.

Ask one question: show me the reliability diagram.

Not the average accuracy. Not a chart of how many calls the system makes. A plot with reported confidence on one axis, how often the system was actually right on the other, and the number of cases behind each bin. The artifact that says: when this thing said 0.9, it was right this often; when it said 0.6, this often; and here is how many times it said each.

If that plot does not exist, you have not received a measurement. You have received three columns of design opinion with numbers in them. The guidance may still be good. But the one thing you were going to build a policy on — what a 0.9 means — has not been established anywhere, by anyone, at any point. And you will still be switching on it in production, because the number is right there and it looks like the other numbers your systems use.

That is the whole complaint, and it applies to every judge, including ours. A confidence value with no curve behind it is not a judgment. It is a feeling with a decimal point.

What the number is claiming

Read a confidence value slowly. When a system reports 0.9, it is claiming that among cases that look like this one, about 90 in 100 turn out the way it says. That is a factual claim about a frequency, and frequencies can be counted. Nothing about the interface tells you whether anyone counted.

The claim is also not the same as accuracy, which is why swapping one for the other goes wrong so quietly. An average score answers "how often is this thing right across everything we tested." A calibration curve answers "what does this particular number buy me in error rate." Only the second one is usable at a cut-off, and it is almost never the one that ships with the product.

Two systems can post identical accuracy and be worth very different amounts to you. One can be right 88% of the time and honest about which cases it does not know. The other can be right 88% of the time and report 0.95 on everything. The first one can be routed. The second one cannot, because a constant cannot separate cases — a gate built on it fires the same way every time, which means the gate is decoration. The score did not distinguish them. Only the curve does.

Calibration belongs to the population, not to the model

Here is the part that turns "show me the diagram" from a procurement ritual into an actual question.

We publish ours, and the least convenient number in it is a comparison. On JudgeBench — self-run under the official judging protocol, 620 judgments, with the 6 pairs whose first verdict failed disclosed rather than quietly retried — calls we reported at 90% confidence or above came back right 99.6% of the time, and calls in the 80–90% band came back right 94.0%.

Then the same interface, the same scoring rule, a different corpus. On ContextualJudgeBench, self-run over the full official set under the pairwise protocol where a pair counts as correct only if both presentation orders are judged correctly — which is why the random floor is 25% rather than 50% — with 12 orders excluded after repeated platform failures and that exclusion stated next to the result, our consistent accuracy is 67.1% against the benchmark's official reference value of 65.4%. The deliberately constructed near-tie splits in that benchmark sit at 46–60%.

Now put the two runs side by side. Same judge. Same confidence field. On the second corpus, the top reported band is worth dramatically less than it is on the first. That benchmark was built to contain genuinely close pairs, and a population full of close pairs is a population where "90% sure" does not buy what it bought last week.

Nobody behaved badly to produce that gap. It falls out of the arithmetic. Calibration is a property of a model and a population: it depends on the mix of cases you feed in, how many of them are genuinely decidable, and how the base rate sits. The number is not a fixed property of the software, the way latency or price are. It is a measurement of a distribution, and the distribution is partly yours.

Which gives you the three things you actually wanted from that vendor in the first place. The curve is a claim about a corpus that is not your traffic. The curve is a claim about a corpus you can partly describe but cannot fully inspect. And the curve is only as current as the model version it was measured on — a retrained model with an unchanged interface invalidates it silently.

What a diagram has to contain

Not every chart is evidence. Three details carry most of the weight, and their absence is more informative than the headline number.

The bin edges and the count in each bin. A bin with four cases in it is an anecdote with axis labels. A bin with three hundred cases is a rate. Most published calibration summaries show neither, which leaves you unable to tell which of the two you are reading.

What happened to the failures. Which cases were excluded, rerun, dropped or imputed, and how many. Ours, printed next to the results rather than in a footnote: 6 pairs on JudgeBench whose first verdict failed, 12 orders rerun-excluded on ContextualJudgeBench, out of the full 2,000-pair official set. A protocol that names its exclusions is telling you how far to trust the rows above them. A protocol with no exclusions to name has usually not looked, rather than having nothing to report.

The label set, the population and the date. A calibration curve is meaningless without the corpus it was measured on and when. This is the detail that kills almost every diagram in circulation, ours included, because it is the one that makes the curve somebody's measurement of something specific instead of a general property of the product.

One caution that applies to any single-number summary: a scalar can hide the only bin you care about. An aggregate error figure that averages across a corpus tells you about average performance and nothing about the band your threshold actually sits in. The bin table is the artifact; the single number is a convenience that should never be the thing a decision rests on.

We are asking to be measured with the same ruler

It would be easy to read all of this as an argument against somebody else's product, so let us remove that reading.

TypeSafe has not published a reliability diagram or an expected calibration error for Jev. That is a statement about what is publicly available, not about what is true internally. Their recommended usage is threshold routing — act above a line, send the rest to a person or to a slower judge — and that recommendation is sound. It simply puts the reliability question directly on the critical path of the intended deployment, and leaves the curve unpublished while it does.

The reason I am comfortable writing that sentence is that the same ruler is pointed at us, and it reads worse in one specific place. The near-tie splits on our own second benchmark come back at 46–60%. That is the least flattering figure we have, it is on our own page, and it is the one that makes the other rows mean anything, because a vendor that publishes only its best row has told you which row it wants read.

So the invitation is literal, and it is not a rhetorical flourish. Our bins are published with their counts. Our exclusions are numbered. The near-ties are published alongside the good rows. If you take our curve and redraw it on your own traffic, your own bin edges and your own consequences, and the bands do not hold, that is a more useful result than any of our rows — and we would rather have it than not. We have been asked for the diagram, we published it, and anyone holding the same ruler should apply it here first.

The practical version, if you are the one signing: for any judge you are considering, ask for the plot, the counts, the exclusions and the date. Price the bands against your own error budget rather than someone else's average. Then write the cut-off down with the name of the person who chose it, because a threshold is a commitment and not a setting.

And if the diagram turns out to exist and looks ordinary — an honest curve with a soft middle and a low band that says this one is close — do not treat that as a disappointment. It is the most valuable thing a decision tool can hand you. It tells you where the machine stops and the minutes begin.

If you want to point that ruler at a judge that takes one question and two candidate answers and returns a pick with a calibrated confidence and a written argument you can disagree with, the entry point is here: https://api.turingcorp.net/platform/go/decider?src=jev-10

Top comments (0)