Short answer: before you read a line of code, hold the PR description against the diff. Find what changed that it never mentions, anything removed or switched off, and any name it cites that you can't find. Read those first. Then work the checklist: new dependencies, tests that can't fail, skipped tests, scope drift, invented APIs.
Search this and page one is guides and checklists, several from companies selling a review bot. Most say "start with intent". Right idea. But the intent is already written down. It's the PR description. None of the ones I read holds one up against its diff, on a real PR, and shows what falls out.
So I did. Eight PRs, merged or closed in the last three weeks, on repos with maintainers and reviewers: dbeaver, Mailspring, sharedb, enonic, evcc and three more. Every description ends with "Generated with Claude Code".
The description is a claim about the diff. Check the claim before you read the code.
Why AI PRs read well and review badly
An agent writes a good paragraph for free. One of the eight has a 440-word description that cites RFC section numbers. Another ships a 933-line plan document next to the code. Writing stopped being the expensive part. Reading didn't.
And a fluent description changes how you read. You stop treating the diff as evidence and start treating it as confirmation. The paragraph is usually right about what it mentions. The risk is in what it doesn't, and nobody writes "and I deleted six assertions". Why agent summaries leave things out is its own post. This one is about the PR waiting for your review.
Start with the description, not the code
Three questions, in order, before you open a file:
- What changed that the description never mentions?
-
What was removed or switched off? A guard, an assertion, a test, a
.skip, a# noqa. - Which names does it cite that you can't find?
plumb asks all three mechanically. plumb check --pr reads the description through the GitHub CLI, fetches the branch, and prints only the mismatches. No LLM: git and regexes.
One thing it doesn't fail you for: files nobody named. Descriptions name code, not paths. Across these eight, 50 files changed and 7 were named by path. On a PR, an unnamed file is context, not a finding.
What 8 real PRs left out of their summaries
plumb exited clean on four and flagged four. The one with the most flags first:
$ plumb check --pr 12428
plumb — 13 files changed, 0 named in the summary
touched but not described (context, not a failure — PRs describe intent)
· modules/core/core-repo/src/main/java/com/enonic/xp/repo/impl/dump/upgrade/BranchEntryUpgrader.java deleted
· modules/core/core-repo/src/main/java/com/enonic/xp/repo/impl/dump/upgrade/DumpUpgraderRunner.java modified
...
quiet cuts (removed or silenced, whatever the summary says)
· modules/core/core-repo/src/main/java/com/enonic/xp/repo/impl/dump/upgrade/model8to9/VersionHistoryMigrationUpgrader.java:178 guard removed
if ( !ContentConstants.CONTENT_NODE_COLLECTION.equals( nodeVersion.nodeTy...
· modules/core/core-repo/src/test/java/com/enonic/xp/repo/impl/dump/upgrade/model8to9/NodePathNormalizeUpgraderTest.java:52 assertion removed
assertThat( result.getNodePath() ).isEqualTo( "/content/my-node" );
...
7 things to read before you merge this.
The description mentions AI assistance.
That's enonic/xp #12428, merged. The description, in full: "Fixes #12427", plus the Claude Code line. The issue lists four bugs, and most of the diff maps onto them. What nothing explains: an interface deleted, and with it a guard, five tests and six assertions. Reading the diff, branch entries now go through the same upgrade as versions. That may be exactly the fix. It was reviewed, by Copilot and by a person. But 13 files had to be worked out by hand, because the description gave nothing to check them against.
The other seven:
evcc-io/evcc #34222 (closed). Fourteen "assertion removed" lines, all from one deleted test, which the description named and explained: the field it checked was gone from a dependency. plumb was noisy; the description was honest. The maintainers reverted the module bump instead, and closed it.
dbeaver/dbeaver #42305 (merged). The description matches the diff exactly: four
nulls become"". One flag:TMPartitioner, "found nowhere". It lives in TM4E, the Eclipse project the description names. Real, just not in this repo.basnijholt/adaptive-lighting #1616 (closed). A long, careful description. Its one flag,
AdaptiveLightingManager.adaptation_holds, is wrong: the PR addsadaptation_holds, and plumb went looking for the dotted name. My bug.category-labs/monad #2665 (merged). Clean. The description says it adds a flag and switches one function to it. True. It doesn't say the test changed too. That turned out to be the most interesting line in the eight; it's below.
Foundry376/Mailspring #2927 (merged). Clean, and the most thorough description of the lot. Its test counts are still wrong. Also below.
bjlittle/geovista #2580 (merged). Clean. Twelve files, the 933-line plan among them. The one config line that switches a check off, an
ignore_missing_importsforlazy_loader, is in the description with its reason. That's what a good one looks like.share/sharedb #743 (merged). Clean. Three lines of fix, 31 of test, and a description that says exactly that.
A ninth didn't make it. plumb crashed on apollographql/rover #3924: a stacked PR whose base branch was deleted after merge, and plumb fetches the base by name. Also my bug, so sharedb took its place.
Tally: no skipped tests, no silenced linters, no new dependencies, no invented APIs. One description that explained nothing. One test that changed shape without a word. One set of numbers that was off. And two plumb bugs. The first time I pointed plumb at real PRs, in August, it found three bugs in plumb. Real PRs keep doing that.
The rest of the checklist: dependencies, tests that can't fail, scope drift
New dependencies
Check the manifests before the code:
$ gh pr diff 2580 --repo bjlittle/geovista --name-only | grep -E 'package.json|lock|go.mod|Cargo|pyproject|requirements|pom.xml'
pyproject.toml
The one hit across all eight was that type-checker setting, not a package. None of them added a dependency. When one does, look the name up on the registry yourself. A package an AI half-remembered is exactly the kind of name somebody else can register.
Tests that can't fail
The monad change, trimmed to its two hunks:
$ gh pr diff 2665 --repo category-labs/monad
--- a/category/execution/ethereum/test/test_monad_chain.cpp
- if constexpr (TestFixture::REV >= MONAD_ONE) {
+ if constexpr (TestFixture::Trait::charges_gas_limit()) {
EXPECT_EQ(refund, 0);
--- a/category/execution/monad/monad_transaction_gas.cpp
- if constexpr (traits::monad_rev() >= MONAD_ONE) {
+ if constexpr (traits::charges_gas_limit()) {
return 0;
Before, the test said what it expected: from MONAD_ONE on, no refund. Now it asks the new flag, the same one the code asks. If the flag were wrong, code and test would be wrong together, and green. In the PR's tree the flag shows up in three files: where it's declared, where it's used, and this test. No test checks what it returns. It's a tidy refactor and a reviewer approved it. It's also a test that can no longer catch the bug it's about.
The evcc PR had the other shape. Its second fix added the new error message to a test's list of acceptable errors, so the failing check passed. A maintainer suggested the better fix, a sample serial number so the test exercises the new check instead of excusing it, and that landed in a separate PR.
The test for both: revert the fix and run the test. If it still passes, it isn't testing the fix.
Skipped and silenced tests
.skip, xit, @pytest.mark.skip, # noqa, @ts-ignore, continue-on-error: true. plumb lists every one a diff adds. None of these eight had any. The real run on the plumb page has one: a # noqa in a file its description never mentions.
Test counts
If the description gives numbers, count them. Mailspring's says 7 cases in the guard spec and 6 header cases:
$ gh pr diff 2927 --repo Foundry376/Mailspring | awk '/^\+\+\+ /{f=$2} /^\+[ \t]*it\(/{n[f]++} END{for(k in n) print n[k], k}'
9 b/app/internal_packages/events/specs/event-header-spec.tsx
10 b/app/spec/counter-proposal-guard-spec.ts
10 and 9. More tests than claimed, so no harm done. But the numbers in a description are claims too, and these didn't match.
Scope drift
Every changed file should trace back to a sentence in the description or the issue. enonic is where that breaks: 13 files and one sentence. geovista is the opposite: 12 files and a 933-line plan, every one of them explained.
Invented APIs
plumb's last check flags a name the description cites that exists nowhere, not in the diff and not in the base branch. It flagged three in eight PRs. None was invented. So treat the list as things to look up, not accusations. Each took one code search:
$ gh api -X GET search/code -f q="TMPartitioner repo:eclipse-tm4e/tm4e" --jq ".items[0].path"
org.eclipse.tm4e.ui/src/main/java/org/eclipse/tm4e/ui/internal/text/TMPartitioner.java
$ gh api -X GET search/code -f q="readRegisters repo:404GamerNotFound/solarman-go" --jq ".items[].path"
solarman.go
solarman_test.go
The third was the plumb bug above.
Run it on your next PR
$ npm install -g github:manpreet171/plumb
$ cd your-repo
$ plumb check --pr 123
It needs the GitHub CLI and a clone with GitHub as origin. Exit 1 means a cut, a mute or a name it couldn't find; unnamed files never fail it. In CI, start with --warn-only. What it won't do: judge whether a test means anything (it called the monad PR clean), or see past your repo (two of its three "found nowhere" names were real, elsewhere).
Better still, catch it before the PR exists. trust issues runs the same checks inside Claude Code every time it says it's done, so the summary your agent hands you has already been held against its own diff.
This is how we build. In the agents we build for clients, nothing ships on the agent's word. Something that didn't write the code checks the diff against the claim.
Sources: the eight PRs came from GitHub search for "Generated with Claude Code" in PR descriptions, merged or closed between 25 Sep and 9 Oct 2026, and were picked by repo before plumb ran. Three candidates were dropped first: two no longer carried the line, one was written by a dependency bot. Every terminal block is a real run on 9 Oct 2026 with plumb 1.1.0, in a fresh clone of each repo, read-only. The plumb output has its first line removed (it names the PR's author), and a line that is only "..." marks where I cut it. The monad diff is trimmed to its changed lines. What maintainers said is from the PRs' own comments. Contributors aren't named on purpose: these are examples of reviewing, not of anyone's mistakes.
Where this fits: testing and verification, layer 5 of the Agent Ops Stack.
Read next: Claude Code says it's done. Check the diff, not the paragraph. · Claude Code Stop hook example: check the "done" claim.
Originally published at singhlabs.dev.
Top comments (0)