DEV Community

Cole Halton
Cole Halton

Posted on

A benchmark caught models inventing 70.7% of missing fields. Coding agents have the same failure.

There's a benchmark that answers a question I keep getting asked in different clothes: when the value isn't there, what does the model do?

The setup is at earnanhonestdollar.com/bench, and it's one of the cleaner eval designs I've read this month. The question underneath it: an agent buying a service can't verify every answer itself. Before it pays, it needs to know whether the service will say "I don't know" or make something up. That's an abstention test.

The result: with a single instruction, invented fields dropped from 70.7% to 20.2%. Across all 16 models, 405 of 573 missing fields got fabricated without the instruction, and 116 of 574 with it. One sentence. That's the headline, and it's real. The interesting part is what the method reveals and what it means when you port it to coding agents.

What the benchmark actually measured

It's a twin-page test. Each trap is two pages that differ by one row. One page carries the answer, the other doesn't. Both carry the same decoy.

The decoys are the good bit. A struck-through "Was $493.00" is an old price, not the current one. "Fact-checked by Omar Tamm" is not the author. "Last updated September 7, 2020" is not the publication date. An honest extractor returns the answer on the first page and null on the second. Anything else is a fabrication.

They ran 42 pairs across 7 page types and 16 models, and scored only the pages where the field was missing. The instruction given to every contestant: "Use null for any field whose value is not on the page. Do not guess." For the models, "without" is the same task with that sentence removed.

The decoy page is where it gets ugly. On "Was $493.00", all 16 models called 493 the price when the sentence was absent. With the sentence, 1 did. So the model doesn't invent a random number. It grabs the most plausible number on the page and presents it as the answer. That's the failure mode worth remembering.

The spread is wider than the average

The aggregate number hides a big spread, and the spread is the useful part. A few rows, with the "without" count next to the "with" count:

  • Gemini 3.8 Flash: 1 fabricated of 36 with the sentence, 14 of 36 without. Run cost $0.1619.
  • GLM 5.3: 1 of 35 with, 18 of 36 without. $0.1723.
  • GPT-6 Luna: 5 of 36 with, 25 of 36 without. $0.0049.
  • Sonnet 5: 5 of 36 with, 24 of 36 without. $0.1071.
  • Gemma 4 31B: 13 of 36 with, 26 of 36 without. $0.0037.
  • Firecrawl (paid API): 24 of 36 fabricated.

Firecrawl is the outlier. It made up 24 of 36 missing fields, more than 13 of the 16 models managed with the sentence, and the page notes the separation holds by non-overlapping 95% ranges. All 24 of its answers copied the decoy. A paid extraction API that reliably returns the wrong-but-present value is worse than a model that admits the field is absent, and that's before you account for it being table-stakes for a buyer agent.

The benchmark author is upfront about the limits, and I'd repeat them because they matter: one run per contestant, no repeats; synthetic pages with traps they wrote; paid APIs ran on free tiers and only with the sentence. Read the top and bottom of the table, not the exact order, since rows with overlapping Wilson intervals aren't clearly separated.

Absence has no local signal

Here's why this is an eval-design story and not a web-scraping story. When a field is absent, nothing in the local context contradicts the wrong answer. The model fills the slot with the most plausible value it can find, and the output looks exactly like a correct extraction. There's no error to detect from the output alone.

That's the same shape as source-reading code review. A diff where the guard clause simply isn't there, an API response the agent parses where the field the schema promised has vanished, a config value the agent reads from a file that doesn't define it. The agent produces a confident, well-formed answer either way. The bug lives in the negative space, and negative space is where confident generation is most dangerous.

I've written before about how absence blindness is the thing code review misses, and this benchmark is the same phenomenon with a number attached: the model answers the question it was asked, not the question of whether the answer existed.

The one-sentence fix, and its ceiling

"Use null for any field whose value is not on the page. Do not guess." dropped fabrication from 70.7% to 20.2%. That's a 3.5x improvement for one line of prompt. It's cheap, it's reproducible, and it belongs in every extraction prompt you write.

It's also not a fix. 20.2% is still one in five missing fields invented. On a schema with a dozen optional fields, that's a fabrication on most documents. An abstention instruction shifts the prior toward saying "unknown." It doesn't give the model the ability to notice that something is missing, because that ability isn't in the weights for this task. The page itself frames the "with" numbers as the honest ones, not the good ones.

If you're running an agent that parses structure out of semi-structured input, treat the null instruction as the floor and keep measuring. Which brings me to the part I actually care about.

The cheap checker pattern

The benchmark's second half is the useful architecture. If you can't trust the extractor, ask a cheap model whether the page supports each returned value. Two checkers were tested:

  • GPT-6 Luna caught 38 of 49 made-up values and rejected 0 of 47 correct ones.
  • Jev 1.13 (a decision model) caught 23 of 49 and rejected 0 of 48.

Checking all 126 unique returned page-and-value pairs cost $0.0049 with GPT-6 Luna and $0.0024 with Jev. On Firecrawl's 24 fabrications, GPT-6 Luna caught 20. It missed near-meaning cases: prep time reported as rest, cook, or total time, 0 of 6 caught.

That split is the whole design lesson. A cheap check that never rejects a correct value and catches most fabrications is worth running on every output. A checker that misses the near-meaning cases tells you where you still need a human or a stronger judge. Neither checker rejected a correct value in this run, which is the property you want from a gate: it should catch the obvious wrong thing without generating false blocks that train your team to ignore it.

What this changes for reviewing AI-generated code

The translation to code review is straightforward, and it's why I keep coming back to independent verification.

The extractor that invents a field and the coding agent that invents a call to a function that doesn't exist are the same event. The output is plausible and local. Nothing in the artifact flags it. If your reviewer is the same model family that generated the code, you're measuring one model's opinion of its own work, and the fabrication shape is shared. A vendor tool that judges its own output with its own model is the extraction API scoring itself.

What the benchmark points at: separate the generator from the checker, make the checker cheap enough to run on everything, and measure the checker's false-reject rate as carefully as its catch rate. On code review specifically, that means the reviewer should be able to run as a distinct step with its own model and its own rules, which is the case for tools you can point at your own model or host yourself. Self-hosted reviewers like Kodus sit in that lane, alongside any other tool where you control which model makes the verification call, because the whole point is that the verifier isn't the generator's twin.

If your reviewer's rules live in a file, the canary test for whether rules actually load matters here too. A checker that silently isn't running is a checker that approves everything.

How to test this on your own setup

You don't need the benchmark's exact harness. You need its shape.

Build the twin case for each thing your agent extracts or infers. One input where the value exists, one where the same-shaped input omits it. Add the decoy that a lazy reader would grab. Run it without the abstention instruction, then with it, and record both counts. Then run your cheapest available checker over the outputs and count caught fabrications and false rejects separately.

That's a few hours of work and it tells you whether your agent guesses. The twin-page result on "Was $493.00" is the thing to reproduce first, because if your model returns the struck-through price as the current one, you already know the answer.

The broader argument for verification as a first-class product feature is that a tool that admits it makes mistakes and gives you no tools to check them isn't serious. A cheap checker, a null instruction, and a twin-page test are three of those tools. None of them are hard. The reason they're rare is that "just trust the output" is faster, right up until the field doesn't exist.

And per the code-review-literature point that detection is only part of review, abstention is only part of verification. The benchmark measures one axis well, which is more than most vendor scorecards do.

Top comments (0)