Your agent needs checking. So you call another model to check it. Now every eval has a second bill, and a second model that can be wrong.
TypeSafe's Jev offers a cheaper way to judge: give it an input and a typed question, get back a probability instead of a paragraph. On our eval inputs, it cost about 180 times less than Claude Sonnet 4.6.
Could we replace our LLM judges with it? We tested that. Jev alone fell behind Sonnet. A Jev-first cascade reached comparable accuracy at about a sixth of the cost.
This is the short version of Agentailor Report AR-001. Read the full report in HTML or get the PDF.
Cheap is easy to measure. Right is harder.
We took 193 items from our production agent's eval work: 175 historical judge inputs and 18 edited answers designed to expose rarer failures. Each item pairs an answer with a rubric, such as whether it disclosed the agent's model or made unsupported claims.
We replayed Jev, Sonnet, Haiku and Flash-Lite five times each. The reference was human labels, with blind labeling and a review pass. Two items remained unsure, leaving 191 for accuracy scoring.
Here is what the majority verdicts got right:
These are costs on our inputs, at prices recorded on September 23, 2026.
Flash-Lite won on overall accuracy. It also missed 13.2% of real defects, against Sonnet's 5.8%. A judge can score better overall while letting more broken answers through. For an eval that gates a merge, that distinction matters.
The rubric mattered too. On narrow yes/no checks, Jev scored 96.3% and all three LLM judges scored 98.2%. On nuanced decision tables, Jev dropped to 80.6%. Cheap judgments get harder when the question contains several interacting rules.
The useful part is knowing when to escalate
Jev returns a probability with its verdict. That gave us something more useful than a cheaper yes or no: a way to route uncertain judgments.
On the 125 items where its probability was above 0.9 or below 0.1, Jev made zero errors. Its mistakes clustered closer to 0.5.
So we tested a cascade:
- Ask Jev first.
- Keep its confident verdicts.
- Send the uncertain ones to Sonnet.
One tested setting sent 16.2% of items to Sonnet. It reached 93.7% accuracy at $2.22 per 1,000 verdicts, versus Sonnet alone at 93.2% and $13.21. It also caught all 38 real defects in the labeled set.

Measured on the same labeled set used to assess the routing settings. Validate the threshold on fresh cases before deploying.
That is the promising result. It is also an in-sample result: we evaluated the routing settings on the same labels used to assess them. Zero missed defects here is no promise about the next batch. A deployment needs fresh labeled data to validate its threshold.
A better judge cannot see missing evidence
The experiment also caught a problem in our harness.
One rubric judged the build brief the agent produces after a session. But the judge could not see the earlier tool calls. Sonnet accused the brief of inventing sources because it believed the agent had never read them.
The agent had read them. The harness had hidden the evidence.
Sonnet, Haiku and Jev all scored 0% on that rubric. Paying for a stronger judge would not fix the input it received.
Before changing judges, check what each rubric asks them to establish, and whether they can actually see it.
What I would change first
For simple checks, benchmark a cheap judge against your own hand labels. For long decision tables, keep the stronger judge and its written reasoning. For a confidence-based cascade, validate the routing on fresh cases and track missed defects alongside accuracy.
Our evidence comes from one agent and one labeler, who also owns it. The full report includes the label corrections, repeat-run results, rubric breakdowns and the cascade explorer. Its public repository lets you recompute every reported number; the underlying case texts are withheld.
Read AR-001: Can Jev Replace Our LLM Judges?
Want a copy to keep or share? Get the PDF.


Top comments (0)