I run a small team of AI agents that answers research questions. One agent searches the web and takes notes, one writes the answer, and a third, the reviewer, checks the write-up before I see it. The reviewer is the expensive part, so I tested three versions of it on the same draft. That draft had a real mistake in it: it credited a hardware claim to a source that didn't say it.
The expensive reviewer, Claude Sonnet at high effort, caught the mistake for about three cents. The cheap one, Claude Haiku, missed it for about one cent. I kept Haiku and gave it a new ability: before deciding, open two or three of the pages the draft cites and check they really say what the draft claims. Then I wrote that this closed the gap.
I never re-ran the test.
The comment
A reader summed up the idea better than I had: the cheap reviewer didn't need to be smarter, it needed something to be wrong against. A reviewer that only reads the draft can judge whether it sounds right. One that holds the source can actually fail it.
That's true, and it made me notice I had described a mechanism, not a result. So I pulled the original flawed draft out of the run log and gave it to the new reviewer. Same draft, same notes, same mistake. The only change was the ability to open pages.
What happened
I ran it twice, plus once with the page-opening tool switched off to recreate the original setup.
| Reviewer | Verdict | Caught the mistake? |
|---|---|---|
| Haiku, no tools (the original setup) | revise, for other reasons | no |
| Haiku that can open pages, run 1 | approve, no notes | no |
| Haiku that can open pages, run 2 | revise, for other reasons | no |
The version with the new ability did worse. On one run it approved the flawed draft outright.
The reason was in the draft itself. It cited sources as "[5]", with no link. The citations only started carrying URLs the next day, so the reviewer had nothing it could actually open. Instead of saying so, it went looking. It opened an Apple specification page, a hardware review, a code repository: six pages that seemed relevant to the topic and weren't the source of anything in the draft. It read them, found nothing that contradicted the draft, and signed off.
That is worse than the reviewer with no tools. The no-tools reviewer never claimed to have checked a source. This one did the motions of a source check against pages that couldn't fail the claim, and came back more confident, not less.
The fix
Two rules, and the second one had to live in code rather than in the instructions:
- Only the cited sources. Before reviewing, the system collects every URL that actually appears in the draft and the research notes, and the reviewer is told to open only those. A request to open any other page is refused before anything is fetched.
- Say "not checkable" when you can't check. If the draft cites nothing with a link, the reviewer is told not to search or guess. Its notes always begin with "Citations not checkable: no source URLs", and that line is added by the code if the model leaves it out. A sign-off can no longer look like a source check that didn't happen.
On the same link-less draft, the fixed reviewer opened nothing, both times, and both reviews started with that line.
What the source check actually adds
Fixing the failure still left the original question: does opening the sources help the cheap reviewer catch mistakes? So I took a recent draft that does cite real links and planted a false claim in it: "Node 22 also dropped support for Windows 10 [1]." The cited source is the official release announcement, which says nothing of the kind, and the research notes don't mention Windows at all.
Every version caught it, including the one with no tools. The no-tools reviewer noticed that the research notes never mention Windows. The version that opened the sources went further: it named the official announcement and said the claim wasn't in it, which is stronger evidence and cost about two and a half times as much.
So in my tests, opening the sources didn't catch anything that comparing the draft against the research notes missed. It produced better proof, not more catches. The mistake that started all this, a hardware claim sitting right next to figures the notes really did contain, was missed by the cheap reviewer in every setup I tried. Only the expensive one has caught it.
What I'd tell myself three days ago
A tool is not a check. A reviewer that can open pages will open some, and if the right ones aren't available it will find pages that look close enough, and treat not finding a problem as finding none. The ability to look only helps when there is something specific to look at, and when "I couldn't look" is an answer the system is allowed to give.
The smaller lesson is the one the reader's comment forced on me: the only way to know what a fix does is to re-run the case that made you want it. I had the flawed draft saved the whole time. Writing "this closed the gap" took less effort than checking whether it had.
Top comments (0)