Swap a cheap model in behind an expensive one and the obvious way to check it is shadow traffic: send the same request to both, compare the answers, count how often they agree. High agreement, ship it.
I measured that on 60 live calls while building SuperRouter. The routed model agreed with the reference 100% of the time. It was correct 75% of the time.
Both numbers are real. The gap between them is the whole problem.
Agreement is not correctness
Agreement asks whether two models produced the same answer. It never asks whether the answer was right. And two models from the same era, trained on overlapping data, fail in the same direction far more often than they fail independently, so agreement is highest exactly where it protects you least.
A shadow run reporting 100% is consistent with two different worlds: a cheaper model that is genuinely as good, and a cheaper model that is wrong in precisely the same places as the expensive one. Nothing in the agreement number separates them.
What I do instead
Score against ground truth, per failure mode, and never against the other model.
- Ground truth is generated, not hand-labelled. Faults are planted deliberately, so the right answer is known before any model sees the case.
- A planted defect has to move pixels. One of 18 fault classes returned success and changed nothing on screen. Every fixture is now gated against a healthy frame of the same screen, because a fault no model could possibly have caught was quietly inflating every score.
- Both halves of the exam have to be hard. The faithful cases were verbatim copies of the source, so false-alarm rates sat at 0-3% across seven models and that axis measured nothing at all.
- Runs from different versions of a set are never compared. Every run fingerprints the exact cases it sat. Without that, the table ranked a model measured on 90 easy cases above one measured on 592.
The part that surprised me
I assumed a published leaderboard could stand in for measuring your own product. Across two products, rank order mostly transfers when a model is judging (0.83), but every model dropped a median 22 points in absolute terms. When the task was pointing at the right control rather than judging, the order barely transferred at all (0.49).
So a leaderboard tells you roughly who is good in general. It does not tell you who is good at your thing, and the second question is the only one that decides your bill.
The takeaway
If you are routing to save money, agreement rate is the metric you will reach for first and the one that will mislead you hardest. Measure correctness against something you constructed, split by the ways your product can actually break. Otherwise you are measuring how similar two models are and calling it quality.
SuperRouter is open source, Apache-2.0, with no runtime dependencies: https://github.com/M19K/superrouter
Top comments (3)
Correlated blind spots are the quiet killer of shadow evals. If an agent receives a malformed tool payload or an ambiguous date range, two models from the same training vintage almost always produce the exact same plausible hallucination. An output diff scores that as full agreement. Gating test fixtures on deterministic state mutations instead of cross-model diffing is the only way to catch when both models trip over the same edge case.
Saw this same gap in a document classification pipeline where we ran a smaller model in shadow against the reference for two weeks and got 96% agreement, so we swapped it in. The error rate on production edge cases doubled. What the shadow numbers hid was that both models had the same blind spot on a specific document layout we only see in one customer's uploads, so they always agreed and were always wrong together. Ground truth labels on a stratified sample of those edge cases would have caught it in day two.
The agreement gap is the signal here. I would add a small adversarial slice and track abstentions separately so a high match rate cannot mask confident wrong answers.