There are dozens of posts about CLAUDE.md best practices. Each one is a list of tips, and none of them explains why. We stopped guessing and went to the primary sources instead: Anthropic's own guidance and the 2026 papers on context engineering.
They mostly agree. Under all the arguments about how many lines CLAUDE.md should have, there are four principles that vendors, academia, and production tools arrived at on their own. This article puts them on one map, so you don't have to read ten papers to find it.
TL;DR. Four principles show up again and again across independent 2026 sources:
- Pointers, not payloads. Keep references in the window and load heavy content when you need it.
- Progressive disclosure. Reveal detail by task depth, not all at once.
- Less, but relevant. Past a threshold, more context lowers accuracy, it does not just raise cost.
- Freshness is a feature. Stale context does not go quiet. It actively misleads the agent.
Context: the bug is in what we feed the model
Model windows grew to 200K-1M tokens, and the reflex is to fill them. That reflex is the bug.
We do not have one repo. We have many services, and each has its own CLAUDE.md. In production that gave us two failure modes at once:
-
Bloat. A root
CLAUDE.mdof about 238 lines loaded in full on every request. You fix a test in one subsystem, and the agent still pulls in schemas, configs, and migration notes from three neighbors it never touches. - Rot. Some engineers forget to update the context files. Some never do. The files drift from the code and start lying to the agent.
What follows is the map we wish we had on day one.
Core: four principles the 2026 sources agree on
Four principles, one idea
Here is the whole field in one table. Each principle rests on independent sources: vendor guidance, academic benchmarks, and production tools that land on the same shape.
| # | Principle | What it says | Evidence (2026) |
|---|---|---|---|
| 1 | Pointers, not payloads | Keep lightweight identifiers in the window (paths, links, a thin index); load heavy data at runtime, on demand | Anthropic's "just-in-time context" guidance; Claude Code skills load only name+description (tens of tokens each), and the body only when triggered |
| 2 | Progressive disclosure | Reveal depth in stages, routed by task complexity | arXiv:2607.17598: one routing level is enough, a second "never helps and sometimes breaks accuracy"; Claude Code skills are a first-party example |
| 3 | Less, but relevant | Past a threshold, extra context lowers answer quality, it does not just raise cost | On a 50-task hotel-expense benchmark (GPT-5, Dynamics 365 tools), recency pruning plus summarization lifted full itemization from 71.0% to 91.6% while cutting tokens by 62.6% (1.48M to 553K); RepoGraph reports a 32.8% average relative gain across 4 SWE-Bench frameworks |
| 4 | Freshness is a feature | Stale content misleads; the agent trusts every token equally | Studies show outdated docs steer generated code, reviews, and repairs, and the agent cannot tell fresh from rotten (arXiv:2605.10990) |
The picture behind the table
The four principles are not four tricks. They are one idea seen from four sides: the window should hold references and route to detail, instead of swallowing the repository.
The apparent counterexample, and why it is not one
You may have seen the headline version: an ETH Zurich study "proves context files do not work." Read the actual paper (arXiv:2602.11988) and it says something sharper, and it does not contradict the consensus above. It fits inside it.
The effect on success rate is not significant in either direction. LLM-generated files moved the resolution rate by -0.5% (SWE-bench) and -2% (CTXbench), both with high p-values (0.87, 0.37). Developer-committed files did +2.4% on average, better than generated ones (p=0.038), but still not a significant absolute gain (p=0.21). The one effect that is rock solid: context files raise inference cost by 20-23% (p<0.001).
So the honest takeaway is not "files hurt." It is this:
A context file rarely makes the agent smarter. It reliably makes the agent cost 20%+ more.
The fix is not "write a better file." It is "stop paying to load a payload you do not need."
That is principle #1, restated by a skeptic. If extra context does not buy accuracy, keep the window thin and load on demand. Curation beats generation, but cheapness beats both.
The shape of "less, but relevant"
The curve is not flat. Accuracy climbs, peaks, then falls as the window keeps growing. The longer the context, the more the model degrades (it even gives up the search early). That is context rot (arXiv:2606.29718), the mechanism behind "less, but relevant."
Solution: how the four principles become one layer
A map is only useful if it tells you where to stand. Here is what the consensus looks like as a real repository layout: pointers on top, payloads on demand.
# Root context = a map, not a manual (loaded every request, keep it thin)
# Expected effect: the agent reads ~80 lines of pointers, not ~238 of payload
<service>/
├── CLAUDE.md # navigation map + task complexity levels (POINTERS)
├── docs/context/
│ ├── integrations.md # loaded only for cross-service work (PAYLOAD)
│ ├── errors.md # loaded only during an incident (PAYLOAD)
│ └── systemPatterns.md# loaded only when adding a component (PAYLOAD)
└── internal/<subsystem>/
└── CLAUDE.md # auto-loaded when editing this directory (PAYLOAD)
Mapped back to the four principles:
| Consensus principle | Where it lives in the layout |
|---|---|
| Pointers, not payloads | The root CLAUDE.md is an index, not a description |
| Progressive disclosure | "Task complexity levels" route which files load |
| Less, but relevant | Detail sits in docs/context/, pulled only when the task needs it |
| Freshness is a feature | A gate keeps the map honest (its own article in this series) |
💡 Context economics: prompt caching vs just-in-time loading.
Pointers save tokens, but prompt caching only rewards a stable prefix (often up to 90% off cached prompt tokens). Split the window into two zones:
- Hot head (cached core): the system prompt, the rules, and a compact index of every pointer or skill sit at the very start of the window and stay cached.
- Cold tail (dynamic): heavy payloads load at runtime, at the end of the window, for the specific request.
You keep the full discount on the base context without bloating it.
Trade-offs, honestly. This is not free:
- Someone has to write the map by hand. ETH found that curated files beat generated ones (p=0.038), and generated dumps only add cost.
- Progressive disclosure adds one more step: someone, the agent or the author, has to decide how hard the task is.
- A thin map only works if the detail files exist and have not rotted. So freshness (principle #4) cannot run on good intentions. It needs automation.
We did not invent any of this. We read four independent sources that say the same thing, and built the boring layout that follows from it.
Insight: it is a consensus, not a bag of tricks
The one thing to remember: this is not a pile of disconnected hacks. Independent sources landed on the same thing. Vendor guidance, academic benchmarks, and production tools arrived at the same four rules: pointers over payloads, disclosure by depth, less but relevant, and freshness as a first-class property.
That changes the work. You are not guessing at CLAUDE.md line counts. You are applying a documented consensus, and the only real question is how well you put it to work in your own repositories.
The rest of this series is that work, one principle at a time: the tool landscape that already exists, the file structure that encodes pointers and disclosure, what "less" does to quality and to your bill, and the machinery that keeps a map from rotting.
The production numbers behind "it worked" come later in the series. First, the map.
Sources
- Anthropic: effective context engineering / just-in-time context (Anthropic Engineering; "Equipping agents with Agent Skills")
- Progressive disclosure: arXiv:2607.17598
- Repo-map benchmarks: RepoGraph (arXiv:2410.14684); Repository Intelligence Graph (arXiv:2601.10112); Aider repo map
- Less-but-relevant: "Less Context, Better Agents" (arXiv:2606.10209); "The Complexity Trap" (arXiv:2508.21433)
- Context rot: arXiv:2606.29718; redis.io/blog/context-rot
- Doc-drift: "Impressive, But Wrong"; SkillGuard (arXiv:2605.10990)
- ETH Zurich: "Evaluating AGENTS.md" (arXiv:2602.11988)




Top comments (0)