I've been building a tool that checks whether Claude Code actually followed the rules in your CLAUDE.md. Early on I realised I was testing it against one file: my own. So I went and got 559 real ones.
CLAUDE.md, AGENTS.md, .cursorrules, Copilot instructions, Windsurf, Gemini — from PyTorch, Kubernetes, Elasticsearch and a few hundred smaller repos. 21,986 items parsed out of them, no crashes.
Two things came out of it that I didn't expect.
63.4% of a rules file isn't a rule
Directory listings. Reference tables. Import examples. Architecture notes. Glob syntax cheatsheets.
All useful documentation. None of it is telling the agent to do anything. And it's sitting in the same file as the actual rules, competing for the same attention.
So when someone says their CLAUDE.md is 400 lines, it's usually closer to 150 lines of rules with 250 lines of README stapled on.
That reframes the "my CLAUDE.md is too long" problem. It's often not too many rules. It's rules buried in documentation.
Of the real rules, only 44.2% can be checked mechanically
Here's the split that surprised me more.
Checkable:
- "Never commit directly to main" — you can look at what git actually ran
- "No console.log in committed code" — you can look at what was written to files
- "Run the test suite before pushing" — either the command ran or it didn't
Not checkable by any tool, ever:
- "Surface bad news first"
- "Explain the trade-off before choosing an approach"
- "Keep changes reviewable"
There's no command to inspect for the second group. It's not a limitation of my parser — the information isn't in the transcript. A person has to read it and decide.
55.8% of the rules people actually write fall in that second group.
The distribution barely moved as the corpus grew
I started with 40 files: 65.4% non-instructions.
At 559 files: 63.4%.
Between those two points I added formats the parser had never been tuned on. It moved 2.0 percentage points.
That's the part I'd actually stand behind. It suggests "is this a directive?" is a property of language, not of file format — which matters, because the alternative approach (recognising every way people format
rules) is an infinite list.
What I don't know
I don't have a controlled experiment showing longer files perform worse. That's the obvious next question and I haven't run it.
But if a model is budgeting attention across a long file, it seems worth knowing that most of that file was never telling it to do anything.
The tool
It's called RuleReceipt. Free, runs locally, no account:
npx rulereceipt check
I also published the part where it got things badly wrong — it reported 10 violations on a real session of mine and every single one was false. That write-up is at https://rulereceipt.dev/postmortem, including the wrong turn I took while fixing it.
Happy to answer anything about the method. The corpus was public files only, and the classifier is in the repo.
Update, 4 September.
Two bugs in my own parser moved these numbers after publishing, so I'm correcting them in place rather than leaving them wrong. The parser had no concept of fenced code blocks, so a shell comment starting with # or a YAML item starting with - was read as a rule boundary — 588 rules were being cut mid-block, and 1,718 items (7.2% of everything extracted) were lines lifted out of code samples that had never been rules at all.
Separately, command documentation was being counted as rules. Corrected: 21,986 items, 63.4% non-instructions, 44.2% of real rules checkable. The 40-file sample moved 2.0 points rather than 2.5, which if anything strengthens the stability point.
Full history of every figure and what changed it: https://rulereceipt.dev/postmortem#history
Top comments (7)
One question about the method: does the parser separate a rule that's only a sentence from the same rule with a hook behind it? In a file they look identical, and only the hooked one fails loudly when it's ignored - when we cut our own file down we sorted by what was even a rule, not by that, so I still can't tell you what share of what stayed is checkable.
Following up because your question turned into a feature.
You asked whether the parser separates a rule that's only a sentence from the same rule with a hook behind it. It didn't — it read the rules file and nothing else. It's now in rulereceipt rules --coverage,
shipped in 0.1.28.
It lists your rules against the hooks configured in the project and on the machine, and tells you which rules name something a blocking hook also names. Hooks on events that can't refuse anything — SessionStart,
PostToolUse — are counted separately, because they can log or inject context but can't make a rule fail when it's ignored, which was your actual point.
What it deliberately doesn't do is claim a rule is enforced. A hook's command is usually a path to a script the tool doesn't read, so the only evidence available is the event, the matcher, and literal text in the
command string. A hook can guard a rule while sharing no wording with it, and shared wording doesn't mean it guards it. So it reports a possible backing, names the exact token that caused each link so a wrong
one is visible, and says in its own output that this isn't proof.
On my own setup it came back with 1 of 21 rules possibly backed, which is roughly the ratio your comment implied and not a flattering number.
Your framing was better than mine, and I said so at the time — I'd sorted by "is this even a rule", which is upstream of the thing that matters.
No, and that's a real gap. The parser reads the rules file and nothing else. It has no idea whether a hook sits behind any given rule, so a sentence and a sentence-with-a-PreToolUse-guard are identical to it —exactly as you say.
The tool can see hooks, just not connect them: rulereceipt doctor lists every hook and auto-run task configured on the machine and in the project. Nothing correlates that list against the rules. So I can tell you what your hooks are and what your rules are, and not which rules are actually load-bearing.
Your framing is sharper than mine was, though. I sorted by "is this even a rule", which is upstream of the thing that matters. The useful sort is the one you're describing: of the rules that survived, which ones fail loudly when ignored and which just sit there. Same question came up on a Claude Code issue recently — someone with a hook-heavy setup found that every place he'd built enforcement it held, and every place he had only a rule it broke, 13 out of 13 versus 1 out of 1.
I tried to count that split across the corpus and couldn't, which is its own finding: four independent passes over the same 25 rules agreed 28% of the time on which bucket each belonged in. So I can't tell you the share yet. But correlating rules against configured hooks is a thing the tool could actually do, and I hadn't considered it until your question. Thanks for that.
The finding that most of a rules file isn't rules matches what I see when a long CLAUDE.md starts getting ignored in the middle. The agent still loads it, but the actual directives drown under architecture notes and tables that belong in a README. The sharper cut is your checkable split. "Never commit to main" is something a hook can refuse. "Keep changes reviewable" only shows up when a person actually looks at the change, so treating both as the same kind of rule is how people end up trusting a green session that never checked the soft half.
That last sentence is the failure mode I built the reporting around, so it's good to see someone name it unprompted.
The summary line always carries all three counts together, and on a real rules file the judgment count is usually the largest one — on the corpus, 55.8% of genuine rules land there. A run against a small file
just now printed:
2 followed · 0 not followed · 15 need your judgment
Two greens and fifteen things nobody has looked at. It never prints "compliant", "all clear" or "all passed" — I checked, zero occurrences — because a clean-sounding word next to a number like that is exactly how
someone ends up trusting a session where the soft half went unread. Every judgment rule says NEEDS HUMAN REVIEW with the reason, rather than being quietly omitted or guessed at.
The honest limit is that none of this stops a person reading "0 not followed" and stopping there. I can make the number impossible to miss; I can't make anyone act on it.
The checkable vs uncheckable split is the part I hadn't thought to measure but can't unsee now. I've been writing agent instruction files for a while and I think I've had the ratio backwards, spending more words on behavioral guidance like "explain the trade-off" than on things I could actually verify against git or file state. The stable distribution from 40 to 559 files suggests this isn't sloppy teams, it's just how people write instructions naturally.
That was the one that surprised me too. Worth saying I had it backwards in a way that matters more: I'd assumed the uncheckable ones were a minority I could ignore. They're the majority, and they're the ones people are most surprised to see broken.