Drafted with AI help, human-reviewed by The Agent Loop.
Short version: Your model bill has a shadow bill: every rollout you run before you trust a number, every judge pass over its output, every trace you keep around. It never arrives as its own line, so nobody budgets it. Arize writes the arithmetic out: production eval cost = traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention. Four multipliers and two human costs. This post reads those terms from primary sources, then shows the ladder that keeps them small.
For skimmers
- A vendor reports one data leader told them LLM-as-judge evaluation cost 10× the baseline agent workload (anecdote, not a benchmark)
- τ-bench: the best gpt-4o agent has >60% average task success, but pass^8 drops below 25% — proving reliability means rollouts, and the paper prices its own loop at "around 200 dollars" for one trial per task ($0.38 agent + $0.23 simulated user per task)
- Judge tuning has been done for ~$2K instead of ~$2M: 4,480 configurations, multi-fidelity search vs full Alpaca-Eval-style evaluation; one Alpaca-Eval annotation ≈ $24
- Ladder order: deterministic checks first, sampled LLM judges second, humans for escalation on uncertainty × consequence
- We found no primary source for a universal "evals cost 5–30× a run" constant; decompose your own bill instead
The judge is a second workload
Your agent's run has a bill. The evaluation of that run has a second one, computed the same way: tokens in, tokens out, repeated per surface you score. Arize's model makes it explicit. Offline: dataset size × system variants × evaluator runs × cost per evaluation. Production: traffic volume × sampling rate × evaluation surfaces × evaluator cost + human review + retention (Arize, 2026). Nothing exotic. The judge is another inference workload, and retention is storage you pay for monthly.
How big can it get? Monte Carlo reports that "one data + Ai leader confessed to us their evaluation cost was 10 times as expensive as the baseline agent workload" (Monte Carlo, 2026). Read that sentence carefully: one leader, one vendor, one confession. It is not a benchmark and this post will not use it as one. It is evidence that the shadow bill can exceed the run bill when scoring is naive.
Practitioners say the same thing in the open. In an Ask HN thread on evals, one engineer called their AI-as-judge setup "wildly better than one-off testing" and added the catch: "it can get costly if the iteration count is high." Another: "it's quite expensive to do well. I find that for every hypothesis I might have to run a thousand prompts to collect enough data for a conclusion." Self-reports, not measurements, but they point the same direction as the vendor anecdote: verification scales with how often you iterate, and teams iterate constantly.
The multipliers you can actually defend
Nobody has published "evals cost 5–30× a run." We looked for it in the primary literature and it is not there. What is there is a set of study-specific multipliers you can quote with their method attached.
Reliability multiplies rollouts. τ-bench defines pass^k as the probability that all k independent trials of a task succeed. Its headline: "Even for the best-performing gpt-4o function calling agent which has a >60% average task success, pass^8 drops to <25%" (τ-bench, 2024). To compute that number the paper runs "at least 3 trials per task," caps episodes at 30 agent actions, and simulates the user with a language model. It also prints its own cost line: agent and user simulation cost "$0.38 / $0.23 per task respectively, so running one trial per task costs around 200 dollars." That is a real dollar figure for a real benchmark loop. It is also where the discipline matters: pass^8 means an 8× rollout opportunity, not an 8× dollar bill. Trajectory lengths, caching, sampling, judge passes, and storage all sit between those two statements, and the paper measures none of them for you.
Candidate generation multiplies coverage. The Codex paper solved 28.8% of HumanEval with one sample, then "generating 100 samples per problem and selecting the sample that passes the unit tests (77.5% solved)" (OpenAI, 2021). A 100× generation step bought ~49 points of coverage, with a unit test doing the selecting. That is 2021 code benchmarks, not agent bills; the transferable idea is that the sampling multiplier is measurable when someone bothers to publish it.
Judge search multiplies cheaply when you stop evaluating everything. "Tuning LLM Judge Design Decisions for 1/1000 of the Cost" estimates "approximately 2K$ to search through 4,480 judge configurations" versus "around 2M$" if each were evaluated the Alpaca-Eval or Arena-Hard way, with one annotation costing about $24 (arXiv 2501.17178, 2025). A thousand-fold reduction comes from early stopping and cheaper fidelity levels, not from cheaper tokens.
Three numbers, three methods, three papers. The honest summary: multipliers exist per surface, they are measurable per surface, and anyone selling you one universal ratio is skipping the measurement.
The cheap-verification ladder
The way to keep the shadow bill small is to refuse to pay the expensive layer for work the cheap layer can do.
1. Deterministic checks first. Exit codes, schema validation, exact tool arguments, database state. τ-bench itself judges by "an efficient and faithful evaluation process that compares the database state at the end of a conversation with the annotated goal state" (τ-bench, 2024). Monte Carlo's example of the right tool for the job: a "code based monitor to ensure an output is a valid US zip code." No model required.
2. Sample instead of scoring everything. Arize: "Sampling rate controls how many traces are scored. Filtering controls which production traffic receives that coverage," and "randomly sample some cases from the cheap path and re-evaluate them with stronger judges or humans." Broad monitoring at a low rate, dense evaluation on the slices that are risky or changing.
3. Escalate on uncertainty and consequence. Arize's rule: "The expensive layers of the stack should receive only the cases that remain unresolved after cheaper checks," with escalation judged on two variables, uncertainty and consequence. High consequence gets strong verification even when the case looks simple.
4. Only then a small judge. If you need an LLM judge, use a small one for the easy stratum. GPT-4o mini's launch price, posted July 18, 2024: 15 cents per million input tokens and 60 cents per million output tokens, "more than 60% cheaper than GPT-3.5 Turbo" (OpenAI, 2024). Dated launch price: pricing pages mutate, so re-check before it enters your budget. Keep humans for calibration and the high-consequence tail; one commenter in the HN thread noted human review "accumulates with each new model that gets released," which is exactly what the sampling tier is for.
What skipping actually costs
Anthropic's description of the no-evals loop: "Absent evals, debugging is reactive: wait for complaints, reproduce manually, fix the bug, and hope nothing else regressed." The fix does not start with a judge fleet. "20-50 simple tasks drawn from real failures is a great start," and regression suites "should have a nearly 100% pass rate" (Anthropic, 2026).
The uncomfortable coda: verification itself can be wrong. The same article documents Opus 4.5 scoring 42% on CORE-Bench until a researcher found rigid grading that "penalized '96.12' when expecting '96.124991…'", ambiguous task specs, and unreproducible stochastic tasks. After fixing the harness, the score jumped to 95%. A cheap evaluation bill does not make the number true; it just makes it cheap.
Put your own number on it
You already know what a run cost (What one agent run actually costs), and you know where the budget check belongs (Where does the budget check go?). The third slot is proving the run worked: rollouts × judge passes × judged tokens × retention. Multiply it out for your last release and write the number down next to the run cost.
Where does your evaluation cost live today: a number you can defend, a guess, or nothing because nobody counted? Reply below, I read every one.
Related: What one agent run actually costs · Where does the budget check go? · Your agent's tests pass. That's the problem. · Your agent's cost problem isn't the model. It's the loop.
If this gave you the third slot on the bill, tap the unicorn below; it takes one click and it is the only metric Dev.to actually shows me. And follow The Agent Loop if you want tomorrow's post in your feed: I am working through where money actually moves in an agent, one post a day.
Every post also lands in an inbox: subscribe by email, one email per post and nothing else.
Top comments (0)