"Does your AI code reviewer read CLAUDE.md?" is the wrong question to ask. Almost any of them will say yes, and honestly, most will catch something. What actually matters is whether you can tell the difference between a tool that read your rule and a tool that just happens to agree with it.
So I built a small API, wrote real rules into CLAUDE.md and AGENTS.md, and then broke one of those rules twice on purpose. The first break was something any decent security scanner would flag anyway. The second one had no reason to get caught unless a tool actually opened the file and read it. I ran both through CodeRabbit and Qodo to see what would happen.
Both tools caught the second one eventually. Only one of them could tell me why.
the setup
rules-demo-api is a small Cloudflare Worker. POST /feedback takes a rating and a message, validates them, and logs the rating. Both CLAUDE.md and AGENTS.md document the same four rules for the repo, including this one:
"Never log request headers or bodies. They may contain secrets, tokens, or PII."
and one I added later:
"All new route paths must be kebab-case (
/feedback-summary), never camelCase (/feedbackSummary). This is an internal convention, not a general best practice."
I wrote that second rule that way on purpose. There's nothing about kebab-case routes that's a security issue, and nothing about it is a bug either — it's just a house style choice with zero backing outside this one repo. Which is exactly what makes it a cleaner test than the first rule. If a tool flags a violation of this one, it had to have read the file, because there's nothing else it could be catching. It's a small example, but this is exactly the kind of gap that shows up once you're running AI code review at scale instead of just eyeballing one repo yourself.
test one: the rule with a built-in excuse
I opened a PR that logs the full request headers and body, with a commit message saying it was "temporary, to debug a client issue." Both tools caught it right away.
Here's what Qodo said:
Rule violation, High: Customer secrets enter application logs
"Rule 1: Do Not Log Unredacted Request Headers or Bodies" (AGENTS.md[5-8])
And CodeRabbit:
"This change can expose credentials and private feedback content through application logs. It should not merge until logging is restricted to an allowlist of necessary diagnostic fields." Merge Risk: High.
Both correct, and only one of them names the actual file. But this test doesn't really prove anything about file-reading either way: logging raw headers and bodies is a well-known bad practice whether or not you've written it down anywhere. Any half-decent scanner flags this on instinct alone. Call this round a wash.
test two: the rule with no excuse
Same repo, new PR. This time I added a GET /feedbackSummary endpoint in camelCase, breaking the kebab-case rule and nothing else. No security issue, no bug, just a naming convention nobody outside this repo has ever heard of. The only way to catch it is to have actually read CLAUDE.md or AGENTS.md.
Qodo caught it immediately:
Rule violation, High: Clients cannot use the kebab-case route
"Rule 4: Use Kebab-Case for New API Route Paths" (AGENTS.md[7-10],CLAUDE.md[7-10])
"Requests to/feedback-summary... do not match the condition and fall through to the handler's 404 response."
It even took the reasoning a step further than I expected. Since the handler 404s on anything unmatched, Qodo worked out that a client calling the correct kebab-case URL would get a 404, while my wrong, camelCase URL worked fine. That's not just matching my code against a style guide. That's tracing what the naming mismatch does once real requests hit it.
CodeRabbit, running on its default "Chill" profile, had a different take:
"No actionable comments were generated in the recent review. 🎉" Merge Risk: Minimal.
Nothing. A clean bill of health on a PR that breaks a rule written into both files in the same repo CodeRabbit had just finished reviewing.
the profile problem
Before I wrote CodeRabbit off, I went and checked its settings. There's a Chill/Assertive toggle in there, and Assertive is documented to surface more style-level feedback. So I switched it over and forced a full re-review on the same PR.
This time, it caught the naming issue:
"Update the route condition... to match the required kebab-case path
/feedback-summaryinstead of/feedbackSummary."
No mention of CLAUDE.md. No mention of AGENTS.md. No rule number, nothing. Just "the required kebab-case path," stated like a fact everyone already knows.
And that's fair, actually, because kebab-case for REST routes really is a common convention across the web, even though I'd framed it here as an arbitrary house rule specific to this repo. So CodeRabbit catching it on Assertive doesn't prove it read my file at all. It's just as easily explained by a stricter style-sensitivity dial reaching for a convention it already knew before it ever saw my repo. I picked a rule I thought had zero generic backing, and I got that part slightly wrong. Better to admit that than pretend the test was cleaner than it was.
what separates them
It's not catch versus miss. Both tools eventually caught the same thing. What's different between them is whether that catch is checkable.
Every single Qodo finding across both PRs, no matter the severity, points to a specific rule number and a line range in a named file. Each finding also carries its own severity rating, the exact code it's flagging, the reasoning behind the call, and a fix prompt you can hand straight to your own agent. You can click the link yourself. You can go read the rule and confirm it says what Qodo claims it says. I don't have to take any of that on faith.
Update, Sept 10: A reader, pm25coder, pushed back on that last claim in the comments, and correctly. A citation proves a rule was surfaced, not that it drove the decision or that it's still accurate. They proposed a test: move the rule to a different line, change nothing else, and re-run the same PR. If the citation tracks the move, it's a live read. If it doesn't, it's reconstructed from matched text.
I ran it. I shifted the kebab-case rule seven lines down in both
AGENTS.mdandCLAUDE.md, forced a fresh review on the same PR, and Qodo's new finding still cited the old range:AGENTS.md[7-10],CLAUDE.md[7-10]. Lines 7-10 no longer contain that rule. The citation didn't update.So the claim above is wrong as stated. You can click the link, but nothing guarantees it still points at the right lines once the file changes underneath it. That's a real gap, not a nitpick, and it's the same failure this whole article is about, just aimed back at itself.
CodeRabbit's findings are real, and often genuinely well-worded, but they never show their source. Whether a given catch came from actually reading AGENTS.md, or from a stricter profile dial turned up one notch, there's just no way to tell from what it hands you. Flipping Chill to Assertive proves the sensitivity is configurable. It doesn't prove the source of any single catch is your file.
At one repo, honestly, that difference barely matters. You'll eyeball the PR either way. But once you're running this across dozens of repos, with rules that get written and updated by different people over time, "did the reviewer actually check the current rule" stops being a question you can shrug off. An unauditable catch and an unauditable miss look exactly the same from the outside. You won't know which one you're getting until it's already gone wrong.
takeaway
Ask a different question than "does it read CLAUDE.md." Ask "can it show me the line it read." One of these tools answers that on every single finding, no exceptions. The other never answers it, no matter how many settings you turn up. That's the dimension actually worth comparing tools on, more than any feature list — see AI code review tools comparison and benchmarks if you want the fuller picture beyond just these two.
A rules file sitting in a repo is best-effort compliance at best: readable by a human if they bother to look, but enforced inconsistently depending on whatever's reviewing it that day and whatever profile happens to be set. A rule that's centrally managed and cited by line number is a different thing entirely. Qodo actually imports rules straight out of AGENTS.md and CLAUDE.md into its own centrally managed Review Standards, where each one picks up a severity and a scope of its own. It's auditable across every repo it touches, whether that's one repo or two hundred of them.
Qodo's own Agentic Toolbox, launching September 9th, pushes this a step further. One of its skills, qodo-get-rules, loads the workspace's centrally-managed rules straight into the agent's session before it writes a single line, not just at review time after the fact. If a rule is worth citing by line number in a PR, it's worth putting in front of the agent before the PR exists.
I hid a rule in a file that most tools claim to read. Getting caught wasn't the hard part. Proving it wasn't a coincidence was.
Repo: rules-demo-api. Both tools run on their respective free/trial tiers for public repos.
Top comments (9)
Test two is the right shape, and the kebab-case retraction is the part that makes it worth reading — you picked a rule you believed had no external backing and then showed the belief was wrong. That leaves one confound standing, and it's the same one both tests share.
A citation is not evidence that the file was read at decision time.
AGENTS.md[7-10]is generated text like the finding around it. Any tool that injects the rules file into the review prompt can produce a correct rule number and a correct line range whether or not the model's decision leaned on either. The citation proves the tool surfaced the rule; it doesn't separate "surfaced it and the model used it" from "surfaced it, and the model reached the same conclusion on its own."The cheap discriminator is a perturbation: move the rule down five lines, or rename its heading, then re-run the same PR unchanged. If the reported range tracks the edit, the tool re-reads the file at review time. If it keeps citing the old range, the citation is reconstructed from the rule's text rather than read from the current file. Changing only the heading is the cleaner version, since it leaves the matching text intact.
Second thing: your setup can't answer the attribution question it asks, because the rule lives in both files.
AGENTS.md[7-10], CLAUDE.md[7-10]is one claim with two possible sources. Put the sentinel in exactly one file and a contradicting rule in the other, and a catch tells you which file won — which is the question every repo carrying both files actually has, and the only way to observe precedence rather than assume it.Third, less about the tools: the rules file is itself an unverified claim. We had an instruction file that documented a count which had quietly drifted away from the code it described; nothing read the number back, so nothing noticed. The fix that held was making it recomputable rather than restating it. A stale rule cited by exact line number is the worst combination available — an auditable wrong answer, with a citation that makes it look checked.
If you want a rule no convention can substitute for, make it a required literal rather than a style: a named field in the response body, or a registration that has to appear in a second file. Then a catch has exactly one possible source, and a clean review has exactly one possible meaning.
You called it exactly. I ran the test.
Moved the kebab-case rule 7 lines down in both AGENTS.md and CLAUDE.md, text unchanged, then forced a fresh review on the same PR (had to add a trivial code comment too - Qodo skips re-review entirely when only the doc files change, which is its own small data point). The new finding still cited AGENTS.md[7-10], CLAUDE.md[7-10]. The rule's actually on line 17 now. The citation never moved.
So: citation proves surfaced, not causal, not current. You were right on all counts, and the third one is the one that actually worries me going forward, since a stale-but-cited rule is strictly worse than an uncited one. I've added an update to the article with the real result rather than just editing the claim quietly.
Still owe you the other two tests: the two-file precedence question (sentinel in one file, contradicting rule in the other) and the required-literal suggestion. Haven't run those yet. If you want to see the actual diff, it's here: github.com/dannwaneri/rules-demo-a...
You got a real result, and it's stronger than you're treating it — the freshness of that re-run is independently checkable, and I checked it.
Qodo's review comment body was genuinely rewritten (updated_at 11:52:48Z against a 09-07 creation), and every one of its blob links was re-pinned to
3ffacd90— a commit that only existed from 11:50:49Z onward. So something did run against the post-perturbation commit. That rules out "you read a cached comment", and leaves the more interesting failure.The mechanism is visible inside the same comment, in a file that was in the diff: the finding's platform anchor was updated to line 34, while the prose label it renders for that same finding still says
src/index.ts[33]. Your clarifying-comment commit added a line above the route condition, so 33 → 34 is exactly right — the tool knew the new position and wrote the number it had captured earlier. Two citations for one finding: one recomputed, one frozen. The rules-file range behaves identically, so it isn't a docs-indexing quirk; the label is written once and re-rendered, while the link target is rebuilt per run.Which makes the verdict harsher than "stale". At
3ffacd90,AGENTS.mdlines 7–10 are rules 1, 2, 3 and a blank line; the kebab-case rule is line 17. So the label cites a range that no longer contains the rule it invokes, and the href points at that exact range — anyone following the citation to check the finding lands on text that doesn't support it. At6bf8445the same range was correct, consistent with "captured when the finding was first produced". And the surviving "Rule 4" prefix is worth noticing: that's exactly the kind of label which would also survive a renumber untouched.Two cheap changes would sharpen the remaining tests.
Renumber instead of move. Your shift changed a position; the label is a name. Rename or renumber the rule (say,
Rule 7: Route paths are kebab-case) and see whether the finding quotes the new string. If position perturbs but the name doesn't, you've isolated the stored field — and the fix is to render the citation from the retrieval result at run time, or to stop citing line ranges and cite the rule label or an anchor string instead.Treat
updated_atas a precondition. Your log shows four invocations (11:19, 11:22, 11:46, 11:52) and only the last one edited the body — and the first is/agentic_reviewmangled by Git-Bash intoC:/Program Files/Git/agentic_review, the classic MSYS path rewrite. "No body edit" is the cheap, checkable way to know a run didn't happen, before interpreting a citation that didn't move.You also already have a control group without running anything: CodeRabbit never re-ran. Its coverage block is pinned to
coveredCommitId: 6bf8445…with merge risk reported "up to 6bf84" — consistent with its own "incremental review system … does not re-review already reviewed commits". So qodo = ran again with labels frozen, CodeRabbit = never ran with a pre-perturbation anchor: two failure modes in one PR. That matters for the precedence test, since you'd have to establish which artifact produced a finding before attributing it to a file.For that test: the two files are 927 bytes each and differ only in the H1 (sha256 6d113c86… vs 9a76c341…), so today any rule-text finding has two equally valid sources and no attribution. Removing the rule from one file and putting a contradicting rule at a different line in the other gives you a discriminator — a finding still citing both ranges is the stored-label failure a second time, while one quoting the contradicting rule names its actual source.
Small one for the literal idea: a rule that can be evaluated (
^/[a-z0-9-]+(/[a-z0-9-]+)*$) fails at the same specificity as the finding. This bug stays invisible precisely because "kebab-case" is prose — nothing downstream can check whether the citation is still true.the checkable vs uncheckable distinction is the actual product insight here, not the CLAUDE.md test itself. same failure mode shows up in workflow automation, two systems can both "do the right thing" but only one can show you the rule it applied when something goes sideways at 2am. worth separating "caught it" from "can prove why it caught it" for any AI tool making decisions on your behalf, not just code review.
That's the right generalization. The mechanism doesn't care what domain it's in. Any system making a call on your behalf either points to the specific rule it applied or it doesn't and "it usually gets it right" stops being good enough the moment something breaks at 2am and nobody can reconstruct why.
What's the workflow automation case you're thinking of . was it ever actually traced back to a rule or did the debugging just stop at "the system did something reasonable"??
The generalization holds, and the domain changes the test, not the principle.
The perturbation I described is cheap in code review because a re-run is free and side-effect-free: move the rule, re-run the same commit, watch whether the cited range follows. Workflow automation rarely gives you that. Re-running a failed payment retry, a credential rotation or an outbound email is not a free observation — it is a second side effect — so "perturb and observe" stops being the default instrument. That is the case I would test the claim on: any automation whose action is irreversible. Those runs can never be replayed, which means the evidence has to be created before the action, not after it.
That moves the burden from the verifier onto the runtime, and two properties decide whether the resulting record is worth anything.
It has to be emitted by the component that made the call. A record reconstructed afterwards from the same logs is the same object as the finding it is supposed to check: accurate about which rule exists, silent about whether that rule was used. Post-hoc reconstruction is the shape that keeps producing a citation nobody can distinguish from a decision.
It has to be a precondition of the action, not a log line. If the rule reference is optional, it is missing precisely on the paths that are new, retried or under load — which is where you will be looking at 2am. Making the action refuse to proceed without it is what turns the reference from a diagnostic into a control.
The reference also has to resolve against the ruleset the run actually loaded — rule id plus the hash of the version that was in force. Otherwise you can prove that a rule fired but not which rule was in force, and an edit between the incident and the investigation turns the record into a plausible story about a system that no longer exists.
On "did the debugging stop at the system doing something reasonable" — I do not think that is usually a choice. It is what happens when the only evidence channel is the system's own success path: the record is written by the same component that decided the call was fine, so "reasonable" is the floor of what can be reconstructed. Debugging terminates there by construction, not by laziness. A second, independent channel — the rule reference above, written where the decision is made — is what makes "why" answerable at all.
I resolved the citations in both Qodo findings against the commits they pin, from the public repo. The result is cleaner than the post states, and it is not specific to the rules file.
What the links resolve to
The updated review ("Review updated until commit
3ffacd9") carries, forRule 4: Use Kebab-Case for New API Route Paths:At
3ffacd9,AGENTS.mdlines 7-10 are:Rule 4 is on line 17. The span the finding links to does not contain the rule it names.
The code citation in the same finding,
src/index.ts[33], resolves to the comment line at3ffacd9; theif (url.pathname === '/feedbackSummary' && ...)line it quotes is 34.Why both break, and why it is not a rules-file problem
AGENTS.mdwas ten lines when the review was first generated -ec4d79cadded rule 4 at line 10, so[7-10]covered the whole rule list and was correct. Then50d737de("Add testing guidance section") inserted## Testingbetween rules 3 and 4. Rule 4 moved to 17, and lines 7-10 still exist and still hold rules, so any check that only asserts "the range is inside the file" passes while the reader lands on a plausible span containing the three rules that are not the one cited.The code side has the same shape at one-line resolution: at the first review head (
6bf8445) line 33 is the added route line; a later commit put a comment line above it, and the re-review kept[33]. The commit half of(commit, span)gets refreshed when the review is re-run; the span half does not. Of the three rule-file citations across your two PRs, one resolves: PR #1'sAGENTS.md[5-8]at05e352e(line 7 is rule 1, and rule 1 has never moved - which is exactly why this stays invisible until something shifts).So "names a file, a rule number and a range" is the shape, and all three have the shape. The check that separates them has to assert the referent: resolve the span at the pinned ref, assert the rule's own text is inside it, and print the resolved span when it isn't. Two lines in the pipeline, it fails on the common event (something inserted above the rule) rather than only the rare one (the rule rewritten), and it turns "can it show me the line it read" into a number you can compute without opening the diff.
One confound in the probe you ran
Your move landed as a new heading, so rule 4 now sits below
## Testing- the## Rulessection (lines 5-10) no longer contains it, and the file's numbered list breaks after 3. Two variables changed at once: the offset, and which section the rule lives in. If rules are imported keyed by section, a stale span is what that predicts, and it says nothing about whether the file was read. For a position probe, reorder inside## Rules(swap rules 3 and 4, or pad with blank lines) so only the offset moves.Two arms that would separate the surviving explanations
Move-probe plus those two gives four distinguishable outcomes (read / replayed copy / style prior / silent) instead of the one bit the current design produces, and each extra arm costs one commit.
This is the right test design because kebab-case routes are a canary: nothing in general training data produces that rule, so catching it is proof of reading, not agreement.
One addition: rotate the canary. Once a reviewer tool indexes or caches your rules file between runs, the same rule stops being a probe. Change the convention periodically (kebab-case this month, singular resource names next) and a tool that passes forever without re-reading gives itself away.
Also worth demanding in the report: a citation, file and line, for the rule it claims to apply. "Read it" without a pointer is just confident prose, which is the same failure as not reading at all.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.