DEV Community

Cole Halton
Cole Halton

Posted on

The code was stamped by Claude driven by a prompt. So what is code review actually checking?

A comment on the front page of HN this week described something I keep running into when I test review tools, and it's stuck with me. Someone tried to trace where a change at their company actually came from. The code was generated by Claude from a prompt. The prompt was for a ticket that Atlassian's AI integration wrote. The ticket digested docs that were themselves AI-written, sourced from strategy memos the commenter was "90% sure were written entirely by Claude." The full comment is in the Hacker News thread on Simon Späti's post about intent going missing, and Späti's original note is at ssp.sh, "The Problem is not the AI Code, but Nobody Knows Anything Anymore".

The line that matters for anyone selling or buying code review tooling: nobody in that chain could name who decided the change should exist.

Every AI reviewer I've tested checks the patch against something. Usually it's the wrong something.

Open any AI code review product and it does roughly the same four things. It reads the diff, compares it to the surrounding code and to whatever conventions it can infer, checks that tests exist and pass, and flags patterns that look like security problems or duplicate logic. I've run this loop on CodeRabbit, Greptile, Qodo Merge, Pullfrog and a few self-hosted options enough times to know the shape of the output.

All of those checks assume a ground truth outside the diff. The convention exists because someone decided the codebase should work that way. The test exists because someone decided the behavior matters. The architecture exists because someone chose it over the alternative.

When the ticket, the doc, the prompt and the code are all generated, the reviewer is checking the output against a spec that another model invented on the way in. It can confirm the patch matches its own upstream prompt. It cannot confirm the upstream prompt was a good idea, because nothing in the toolchain ever recorded whether a human had an opinion about it.

This is a different failure than the ones I usually write about. It isn't that the reviewer reads the patch and not the execution, which I've covered before. It's that even a reviewer that reads execution, runs the tests and traces the call graph is checking self-consistent fiction when the whole decision trail is synthetic.

The product tells you to check its output and hands you nothing to check it with

Glyph's post What Would A Serious AI Product Look Like? makes the point that the chat UIs already admit fallibility in grey fine print ("AI can make mistakes, double-check responses") and then offer zero affordances to do the checking. The essay's proposed fix is a two-column worksheet: model output on one side, human verification notes and a checkbox on the other.

Coding is worse, because the checking gets silently pushed downstream into code review. The author of the patch never looks hard. The reviewer inherits the job of verifying claims they didn't write, about a decision they weren't in the room for. Glyph's phrase for this is a dark pattern that encourages offloading, and I think that's fair. The reviewer becomes the last human in the loop for decisions that stopped being human several steps earlier.

So when a team asks how to review the growing volume of AI-generated code, the honest answer is that volume is the second problem. The first one is that there's often no human-authored claim left to review against. A bigger queue of self-consistent patches is easy to describe and hard to actually fix, because you can add reviewers and you can't add intent after the fact.

What still works: anchor the review to a human decision, not to the patch

The one lever I've seen hold up across tools is requiring a human-authored artifact at the ticket level before a patch is mergeable. Not a generated summary of the diff. An actual statement, written by a person, of what the change is for and what would make it wrong.

Once that exists, the review has something real to check. The reviewer can flag that the patch does more than the stated intent, or that the tests don't cover the failure mode the intent named, or that the change quietly contradicts an earlier decision. None of that is possible when the intent line reads "implement feature per AI-generated ticket."

This is the part of setup I'd verify before trusting any tool. Kodus takes team standards as a config file, the same class of artifact as CodeRabbit's path instructions or a Greptile rules file, and the whole category is only as good as whether the rules file is human-authored and actually loaded at review time. My earlier protocol for checking whether your reviewer even reads your rules applies here directly: if the standards are generated too, you're back to a model checking a model.

For the specific mechanics, I've written up how to tell whether an AI reviewer actually follows your team's rules, and the broader problem is the same one I flagged in who reviews the AI-generated code before the human does.

What to actually measure instead of catch rate

Catch rate on a patch tells you how good the model is at spotting code that looks wrong in isolation. It says nothing about whether the change should have been made. If you want a number that maps to this problem, the proxy is intent coverage, and you can measure it on your own merged PRs without a vendor's help.

Pull a fixed two-week slice of merged PRs. For each one, check whether a human-authored intent artifact exists at the ticket or PR description level, or whether the description is a restatement of the diff. Count the share that are the latter. That number is your exposure to this failure, and it will probably be higher than anyone on the team expects, because generated descriptions read fine.

Then track review comment density split by kind: comments about whether the change should exist, versus comments about how it's written. If nearly all the review traffic is style and almost none is intent, your process is optimizing the part that's already cheap to generate.

The reason I trust this over a vendor benchmark is the same reason I don't trust any single time-saved percentage. The category's own published numbers swing wildly depending on who runs the test. In the sourced stats I maintain, Greptile's self-run benchmark reports an 82% catch rate while Augment Code's re-run on the same repositories scored the same tool at 45%. A tool that can't reproduce its own headline number on someone else's hardware is not a tool whose "94% fewer bugs" claim should set your review policy.

And the productivity premise underneath all of this deserves the same skepticism. The stats page collects METR's 2025 randomized trial, where experienced open-source maintainers were 19% slower with AI tools on mature codebases while believing they were 20% faster. Feeling faster and being faster diverge, and the same gap likely applies to feeling like the review is covered.

The concrete ask

Before the next sprint, find one merged PR that traces back through a generated ticket and ask two people who should know: who decided this should exist, and what would make it wrong. If nobody can answer, you don't have a review problem you can fix with a better diff-reader. You have a record-keeping problem, and the fix is upstream of the code. The artifact that makes AI review useful is a human sentence about intent. Everything downstream is just checking that sentence for consistency, which a model is genuinely good at, one you can make reproducible, and one that stops being fiction the moment a person wrote the line.

Top comments (2)

Collapse
 
muhannad_salkini_1036ea1b profile image
Muhannad Salkini •

This matches what I keep seeing: the reviewer's ground truth is inferred conventions, and an inferred convention has no author. It might be something the team chose on purpose, or something three agents happened to repeat.

The piece I'd add is provenance on the decisions themselves. A useful reviewer needs to know, for each rule it enforces, who ratified it and where: "agreed in PR #412 review, by the person who owns billing" is a different kind of fact from "the last 20 files do it this way". Once decisions carry that, two checks become possible that today's tools can't do:

  • Flag implicit new decisions. A diff that introduces a new pattern (new dependency direction, new data path, new retry policy) with no recorded decision behind it gets routed to a human, even if it's small and the tests pass. That's the "nobody decided this should exist" case from the HN thread, caught at the PR instead of discovered months later.
  • Weight violations by who decided. Breaking a rule a human explicitly ratified should block. Breaking a pattern the model merely inferred should be a question, not a finding.

You can't recover intent from a fully synthetic chain after the fact. But you can stop new changes from adding to it, by making "which human decision does this rest on?" a question the review step has to answer.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.