DEV Community

Matteo Poli
Matteo Poli

Posted on Originally published at reado.watermelon-studio.it

AI review benchmarks score the bot. Nobody scores the reader

This week GitHub published ReviewBench, an open benchmark for AI code reviewers. It is careful work: 219 pull requests from 187 open-source repositories across 19 languages, picked so their size matches what actually gets reviewed on GitHub, based on an analysis of 103.9 million pull requests. The ground truth comes from human reviewers, from bugs that authors fixed in later commits, from static analysis and from several frontier models, and senior engineers who didn't build the dataset re-labelled every finding, agreeing with it 96.6% of the time.

You get precision, recall and F1, so you can say "this reviewer finds 40% of the real issues and half of what it says is noise." That is a real improvement on vibes.

But read what the benchmark is measuring: the bot. How many true issues it catches, how many false ones it raises. It cannot measure the other half of every review, which is a person reading the finding and deciding whether it is right. That half is where review time goes now.

The code got cheap. Checking it didn't

Anthropic's own research puts the gap plainly: developers use AI in roughly 60% of their work, but say they can "fully delegate" only 0–20% of tasks (summary of Anthropic's 2026 Agentic Coding Trends Report). Everything between those two numbers is work an agent did and a human still has to look at.

And there is a lot more of it. At the AI Engineer World's Fair, Sourcegraph's CEO described agents producing a "tidal wave of code" that wears down old codebases with duplicated helpers, inconsistent standards and fragile dependencies (recap). GitHub's own guide to reviewing agent pull requests reads like a list of things a person has to go and look at: tests that quietly disappeared, utilities duplicated because the agent didn't find the existing one, a permission check on the path that matters, traced from input to output.

None of that is a syntax error a bot can flag. It is reading.

A finding is a bill, not a gift

Every comment an AI reviewer leaves comes with a cost attached: someone has to open the file, find the place, understand enough of the code around it to judge the claim, and then decide. A precise comment on three lines can be settled in seconds. A vague one ("this function may not handle all edge cases") costs a re-read of the whole file, and is often right in a way that doesn't help.

Two reviewers with the same F1 can cost a team very different amounts of attention. One leaves eight sharp comments, each pinned to the line that matters. The other leaves the same eight insights spread over twenty comments, half of them pointing at the wrong line after the next commit moved things around. The benchmark scores them the same. The person reading them doesn't.

GitHub seems to know this. The ReviewBench post says the useful distinction "is not simply how many comments a tool produces", and they let you tilt the score towards precision. That's the right instinct. It still stops at the bot's output, because the human side is hard to put in a dataset.

What would scoring the reader look like?

I don't think it needs a benchmark. It needs us to treat reading as the main job instead of the cleanup after the real job:

  • Time to verify, per finding. Not how many issues were found, but how long it took a person to confirm or dismiss each one. That number punishes vague comments the way they deserve.
  • Findings that survive the next commit. A comment that pointed at line 212 and now points at an unrelated line is worse than no comment. Anchors should move with the code, or say they lost their place.
  • What the human actually read. "Approved" says nothing about whether the risky file was opened. Coverage of the reading, not just of the tests.
  • Human notes as first-class input. The reviewer who spots the wrong thing on line 212 should be able to point at it as precisely as the bot does, and hand it back to the agent without describing a location in prose.

Where this is going for me

I build Reado, an IDE for that half of the job, so take this with the obvious bias. Reado is built around the reader: you read the code your agent wrote, leave comments on the exact lines, and the agent resolves them. Comments re-anchor when the code moves, and orphan themselves when it is gone. Reado tracks what you have actually read and what changed since. In a guided review the agent can propose findings, but they land on the line as proposals until you accept them.

I'm not arguing against AI reviewers. Let them do the mechanical pass first, as GitHub recommends. I'm arguing that the bottleneck has moved, and our tools and our metrics haven't followed it yet. We're getting very good at scoring the bot. The scarce resource is still the person who has to read.

If you review agent code every day, I'd like to hear how you keep up. Tell me in the comments below, or on Discord.

Originally published on reado.watermelon-studio.it.

Top comments (2)

Collapse
 
madhabcoder profile image
Nilamadhab Senapati •

"Time-to-verify" is the metric I've been missing in every flashy review benchmark.

A finding that takes five minutes to confirm is basically a tax on the reviewer. I hit a small version of this building a PR roaster: when the same model wrote and graded the roast, it approved jokes about files that weren't in the diff, and I burned time checking fiction. A separate scorer (Jev) against the real diff — plus a plain file-name pre-check — didn't make the bot smarter. It made me cheaper. Curious how you'd instrument verify-time without turning it into another vanity number.

Collapse
 
urion profile image
Matteo Poli •

Thanks, "it made me cheaper" is exactly it. I wouldn't time people. I'd log two timestamps per finding: when someone first opens its location, and when they accept or dismiss it. Then only look at the distribution per source (which bot, which rule), never per person. A source whose findings mostly get dismissed after a long look is the one costing you. And the file-name pre-check is great: the cheapest filter is the one that runs before anyone reads anything.