DEV Community

Cover image for AI Approved the PR. Nobody Knows Why the Code Works.
Dvir Segal
Dvir Segal

Posted on Originally published at dvirsegal.medium.com

AI Approved the PR. Nobody Knows Why the Code Works.

Picture this, you open a PR and the AI report is already waiting. Three findings, all minor. You skim the diff, approve, and move on. A month later someone asks why that service retries three times instead of failing fast. The author doesn’t know, because an agent wrote it. You don’t know either, and you approved it.

LEGO-style scene, one month later: a developer asks why the service retries three times instead of failing fast, while the author shrugs and the reviewer holds an APPROVED stamp

I review PRs with an AI code review workflow (gist) that maps the blast radius of each change with CodeGraph, a local tool that turns the codebase into a queryable graph of callers and callees, and reports findings by severity. It catches real bugs. It can’t tell me why the code is the way it is.

The gap

Review was never mostly about bugs. Microsoft found that most review comments are about maintainability, alternatives, and knowledge sharing, what Brian Houck calls the invisible output of code review. I wrote about it back in 2020, in Every code has its audience: review is how knowledge moves between engineers, juniors reviewing seniors included.

An AI reviewer covers the visible output. The invisible one is getting harder to keep. The agent throws its reasoning away before anyone sees the PR, and PRs grew 64% while review time didn’t. Houck, citing Margaret-Anne Storey, calls the result intent debt: the gap between what a system does and what anyone understands about why. There’s a personal cost too. In my experience, using AI dulls the brain, especially for juniors, and review was one of the few places they picked up judgment.

So I changed where I spend my attention.

The Intent Header

A reply to one of my threads on X put a finger on the problem: design-time intent usually stays in a meeting or a chat thread and never reaches the diff. Without it, an AI review can only confirm that the code is consistent with itself. The fix it suggested, which I adopted, is to attach a short intent statement to the PR. I call it the Intent Header:

  • Intent: the input, the expected result, and what must not happen.
  • Decisions: what was chosen, what was rejected, and why.
  • Acceptance criteria: derived from the intent.
  • Proof: a log line that shows the behavior actually happened.

For a password reset, it might look like this:

Intent: A user who requests a reset gets one email with a link valid for 30 minutes.
Must not happen: a second email for a repeated request within a minute; a link that works twice.
Decision: rate limit in the API, not the mail queue, because queue retries would bypass it.
Acceptance: two requests within a minute -> one email; a used link -> 410.
Proof: password_reset.sent user=<id> deduped=false
Enter fullscreen mode Exit fullscreen mode

LEGO-style scene: a developer and a robot agent point at an open booklet titled Intent Header listing Intent, Decisions, Acceptance and Proof

The author, the reviewer, and the agent on either side all check against the same statement, and the reasoning stays in the history after the merge. It’s close to what Addy Osmani calls a “statement of purpose” for agentic review, and it’s the written “why” I argued for in Code comments are your code autobiography. It also answers the question an engineer told me they ask on every PR: what would break if this didn’t work, and how did you check that it didn’t?

Read the tests first

The sharpest pushback on that thread said review is a transition-era practice: AI acts as the gatekeeper, end-to-end tests validate, and the merge goes straight to production. The first half isn’t far from reality. Meta’s RADAR already auto-merges low-risk changes, though it still sends the risky ones to people.

Tests have a limit of their own. They check only what someone thought to test, and an agent that can’t make a test pass can change the test until it does. Another reply pointed out that coverage won’t catch this, since it says which lines ran, not whether any assertion would fail if the logic were wrong. Mutation testing does: it breaks the code on purpose and checks whether the tests notice. I’m a fan of short feedback loops, but a loop is only as good as the tests in it. That’s why I think the test changes in a PR should be read before the code changes.

LEGO-style scene: a robot agent swaps a red brick in a wall labeled Tests while a developer with a magnifying glass asks: Did you fix the code, or the test?

And even if one day nobody reads code, someone still has to define what the system should do and why. That’s exactly what the Intent Header holds.

Where the rest of my attention goes

  • I still read the code. As I wrote in The Joy of Negative Code Lines, AI is for acceleration, not abdication. I still click Merge, so I need to be able to explain the change.
  • Design before code. Understand the requirements, walk through them, and raise risks and gaps before anyone writes a line.
  • Data over hunches. Real numbers from the database or traffic analytics. Same rule as in effort estimation: don’t assume, measure.
  • Deep design reviews. Hard questions, so the reasoning gets written down and challenged by someone other than its author.
  • Challenge the task. Some tasks shouldn’t be built at all, as I wrote about complex features and cargo culting. The cheapest code to review is code nobody needed to write.

The Bottom Line

Houck closes his piece with this: “AI should absolutely reduce the time we spend reviewing code. It just shouldn’t reduce the amount we learn from it.”

If you automate your reviews, decide where the learning goes. For me, it goes into design reviews and the Intent Header at the top of every PR.

Top comments (0)