I keep seeing teams treat their AGENTS.md like a config file: write the rules once, assume the agent obeys. So I went looking for anyone who actually measured which of those lines do anything, and the answer is uncomfortable.
The cleanest number I found comes out of agents-md-evals, a small eval harness built on Anthropic's skill-creator framework. The README opens with the result: most instruction files are 80 to 95% redundant, and in controlled testing 25 of 26 assertions passed identically with and without a 755-line instruction file. One assertion out of 26 moved. The rest of that file was, functionally, decoration.
That is the thing worth sitting with. Everyone in this space argues about whether the reviewer reads your rules. That was the whole industrial complex of AGENTS.md disk-loading posts this year (I wrote one of them). It's the wrong question. The harder one is whether the rules change a single output when they get read.
The codebase already taught the model your rules
The mechanism behind the 25/26 number is what the harness calls the codebase-teaches-patterns effect. A well-structured repo publishes its tech stack in package.json, its conventions in existing imports, its helper functions in the directory tree. The model reads all of it before it ever gets to your instruction file.
So a line like "use TypeScript strict mode" or "follow functional patterns where possible" is not teaching the model anything it couldn't infer from opening two files. The harness reports realistic improvement deltas of +3 to 10% for well-structured projects, and that's the ceiling you're optimizing against once the obvious stuff is already implicit in the code.
The rules that survive are the ones the codebase physically cannot teach:
- coupling knowledge the code doesn't show (why two services must ship together, why a table can't be touched without a migration)
- workflow requirements that aren't visible in any file
- decision-style preferences where two reasonable choices exist
- domain context that lives in people's heads
Everything else is usually the model guessing correctly and your file taking credit for it.
The A/B protocol, because "feels followed" isn't a measurement
What makes the agents-md-evals result reproducible is the isolation. You can't just ask an agent to "ignore the instruction file" for the control run. Claude Code (and others) auto-load those files for every conversation, including subagents. If the file is present, it's in context, and your "control" is a lie.
The harness handles this with a seven-phase isolation protocol and moves the files to /tmp/ rather than renaming them, because a rename on the same path sometimes still resolves. Then it generates eval prompts from your actual git history instead of generic ones, runs a clean A/B, and grades with a visual HTML viewer. Each rule gets categorized, tested, and either kept or flagged as non-discriminating.
If you're rolling your own version, the shape is: same task, two containers, one with the rules file genuinely absent, one with it present, everything else identical. Score both against the same rubric. Any rule whose presence doesn't move the score across runs is dead weight you're paying context tokens for on every single request.
There's a keep test in the repo too. A rule survives only if it hits at least one of: specificity, behavior change, surprise, pain, frequency, or no feedback loop. Fail all six and fail the A/B and it's gone.
Half the failures are the loader, not the rules
The 25/26 result assumes the file actually loaded. A lot of the time it didn't, and the agent gives you no error. The best writeup I've read on this is on how Codex silently skips your AGENTS.md rules, and the specifics are grim:
- Codex includes at most one instruction file per directory, and it checks AGENTS.override.md first. An override replaces the AGENTS.md sitting next to it rather than merging. The docs even annotate a real file in the sample tree as "Ignored because an override exists."
- The search stops at your working directory, so a rules file deeper in the tree is invisible if you launch from the repo root.
- There's a
project_doc_max_bytescap defaulting to 32 KiB, and the two doc pages disagree on whether it's per-file or for the whole chain. The reliable move is to raise it and verify. - The chain is built once per run. In the TUI, once per launched session. Edit the file and the running session won't notice.
None of that throws an error. You get a confident agent doing the wrong thing, and you conclude the model is bad at following instructions when the file never made it into context. This is exactly why the canary test matters: plant a rule the agent can only satisfy if it read the file, and check whether the output reflects it. Loading and complying are two separate failures and you need to see both.
Verify it the way you'd verify any other claim
The through-line here is the same one that runs through measuring what an AI reviewer misses: a green checkmark from a tool that graded itself is not evidence. An instruction file is a claim about model behavior. Treat it like one.
Concretely, that means:
- Prove loading first. Canary rule, one line, output only satisfiable if read. If it doesn't appear, stop and fix the loader.
- A/B the file. Present vs genuinely absent, same tasks from your own history, same rubric. Count how many rules actually move the score.
- Judge independently. Don't let the model that produced the output score it. The judge needs to be a different model or a deterministic check on a typed field, because an LLM judge that changes its mind run to run is a flaky test suite wearing a lab coat.
- Re-check after tool updates. The loaders change. A rule that worked last quarter may be at the bottom of a 32 KiB cutoff now.
For reviewer tools specifically, custom-standards support is the feature to probe here. Some, like Kodus, ship rules files and block-level checks you can point at; the point isn't the toggle, it's whether the rule lands in the model that's actually reviewing. Ask the vendor to show you the loaded context, not the settings page.
The uncomfortable takeaway from the 755-line file: your rules book is probably a wish list. One line in 26 did something measurable. Find that line, delete the other twenty-five, and spend the tokens on something the codebase can't teach.
Top comments (0)