Originally published on emeraldleaf.dev.
Anyone who uses a coding agent seriously ends up with a setup: a CLAUDE.md or AGENTS.md, some rules, a few skills, a hook or two, a handful of MCP servers. There are now tools to sync that setup across machines, convert it between agents and share it with a team. I wanted to know what people who do this well actually keep in their setups, how they keep them current, and what the evidence says about any of it.
So I looked at about 25 tools and published setups, and six recent papers. Four things stood out.
1. Agents follow what is in front of them, and skip what they have to fetch
Vercel ran the cleanest test I found. In their agent evals they gave an agent the same documentation two ways: as an index in AGENTS.md, which is always in context, and as skills the agent could load when it judged them relevant. AGENTS.md passed 100% of cases. Skills passed 53%, the same score as giving the agent no docs at all, and in 56% of cases it never loaded the skill. Telling it explicitly to use skills raised the score to 79%.
Scott Spence measured the same effect from the other side. On its own, Claude Code used a skill for 50 to 55% of his 22 test prompts. With a UserPromptSubmit hook that made it evaluate every skill before starting, activation went to 100%, 22 prompts out of 22.
The lesson for any setup: anything the agent has to decide to look up gets skipped about half the time. If something matters, put it in front of the agent before it starts, or let the harness do the deciding.
2. Always-on context costs more than it helps, unless it is the right kind
That does not mean putting everything in AGENTS.md. Gloaguen and colleagues tested repository context files on SWE-bench tasks and on real issues from repositories that already had them. The files "do not generally improve task success rates", whether a model or a developer wrote them, while raising inference cost by over 20% on average. The agents did follow the instructions in those files. What did not help were repository overviews, which the authors note are "popular and recommended by model providers". Their conclusion: context files are useful for specifying non-standard coding practices.
Other results point the same way. IFScale gave models up to 500 instructions at once; the best frontier models followed 68% of them at that density, and favoured the ones that came first. Anthropic's own advice for CLAUDE.md is blunt: for each line, ask "Would removing this cause Claude to make mistakes?" If not, cut it.
The evidence on file size itself is mixed. Damon McMillan ran 1,650 Claude Code sessions, measuring whether the agent followed one simple convention while he varied file size, instruction position, file architecture and contradictions between files. None of the four had a detectable effect. What did show up was time: each additional function the agent wrote in a session was associated with about 5.6% lower odds of following the convention, though he notes it is not a steady per-step decline. Instructions tend to fade as a session goes on, wherever they sit.
Taken together: keep always-on context to specific, actionable instructions, and expect them to fade during a long session.
3. Memory that captures everything mostly does not help. Memory that is checked does
The newest evidence is VibeMemBench, which tested memory systems for coding agents on real repository tasks. Injecting past experience whose usefulness had been verified by running the tests raised task resolution by 1.1 to 4.5 percentage points on four of the five agents tested. But when four existing memory systems had to build and retrieve that experience themselves from the same history, 11 of 12 combinations failed to beat the no-memory baseline.
GitHub's Copilot Memory is the strongest counter-example, and it shows what makes automatic capture work. Its agents save facts about a repository on their own, but each fact carries citations to the code that supports it. Before using a fact, Copilot "checks those citations against the current branch to confirm the information is still accurate. Only validated facts are used." Facts that go unused for 28 days are deleted. From A/B tests, GitHub reports a 90% merge rate for its coding agent's pull requests with memory, against 83% without. That is the vendor's own number.
So the dividing line is not human versus automatic capture. It is whether a memory is checked against the code before the agent relies on it.
ACE, from Zhang and colleagues, adds a point about how memories should change over time. Rewriting a context wholesale leads to "context collapse, where iterative rewriting erodes details over time"; small, structured, incremental updates preserve them. Their approach improved agent benchmarks by 10.6%.
4. Setups drift, and almost nothing notices
Every setup goes stale, and the most careful ones say so. Aristidis Vasilopoulos built a 108,000-line C# system with coding agents, supported by a 660-line constitution, 19 specialized agents and 34 specification documents, and reports that "specification staleness was the primary failure mode". When a subsystem changes and its spec does not, "the AI will generate code based on stale information." His fix is a session-start hook that compares recent commits with a map of subsystems to files, and warns when code changed but its spec did not.
Others do it by hand. One of Sentry's agent skills ships a SOURCES.md that maps each claim to the file and function supporting it, with the date it was captured and the date part of it was updated. Every's compound-engineering plugin, the closest thing I found to a learning loop, records solution notes from sessions and has a refresh command that keeps, updates, merges, replaces or deletes them. Its rule is one I would adopt anywhere: "Age alone is not staleness."
The portability tools solve a different kind of drift. rulesync, Microsoft's APM and similar tools generate each agent's configuration from one source, and can fail CI when a generated file no longer matches it (rulesync generate --check exits 1). That catches a stale copy. It does not ask whether the rule itself is still true of the code.
What I would take from this
- Keep always-on instructions short, specific and actionable. Cut overviews and anything the agent can read from the code.
- Do not rely on the agent to fetch what matters. Put it in front of the agent, or have a hook do it.
- Expect instructions to fade in long sessions, and restate the relevant ones at the moment they matter.
- Only keep memories you can check. A note with no evidence behind it is a guess the agent will believe.
- Retire rules when the evidence says they are wrong, not because they are old.
- Watch for drift in what a rule claims, as well as drift between copies. When the code a rule describes changes, someone should have to look again.
Where my own loop fits
That last point is the gap I have been working on with okl, an open-source tool I have been building. Each lesson can be proven by a check you choose, and CI goes red when the code it governs changes, until someone runs the check again. On my own eval, briefing lessons before each task cut repeated known bugs from 43% to 8%, on a small sample of tasks I wrote.
The research also showed me where okl falls short. Agents other than Claude Code have to call for its lessons, which the Vercel result says they often will not. It flags stale lessons in CI rather than at the moment they are used. And it could brief again when a governed file is about to be edited, since instructions fade within a session. Those are next.
How I did this
I used AI research assistants to survey the tools, published setups and papers, and to check every claim cited here against its source, twice. The checks caught one research summary that misstated a paper's finding and several smaller overstatements; this piece follows the sources' own wording. Tools and numbers in this area change weekly; everything here is as of 1 October 2026.
Sources
Papers
- Fan et al., VibeMemBench: Evaluating Memory Systems for Coding Agents on Real Repository Coding Tasks (2026)
- Gloaguen et al., Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? (2026)
- McMillan, Instruction Adherence in Coding Agent Configuration Files: A Factorial Study of Four File-Structure Variables (2026)
- Jaroslawicz et al., How Many Instructions Can LLMs Follow at Once? (2025)
- Zhang et al., Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models (2025)
- Vasilopoulos, Codified Context: Infrastructure for AI Agents in a Complex Codebase (2026)
Measurements and documentation
- Vercel, AGENTS.md outperforms skills in our agent evals
- Scott Spence, Measuring Claude Code skill activation with sandboxed evals
- GitHub, Building an agentic memory system for GitHub Copilot and the Copilot Memory docs
- Anthropic, Best practices for Claude Code
Setups and tools
- Sentry, an agent skill with SOURCES.md
- Every, compound-engineering refresh guide
-
rulesync and its
--checkoption - Microsoft APM
Top comments (2)
The 5.6% drop in instruction adherence per written function from McMillan matches the failure curve I hit in multi-turn coding runs. When agents derail, it is almost never at step 1; it is around step 8 after two tool errors flood the context with stack traces.
The Vercel finding on skipped skills and the McMillan decay point to the same structural fix: moving enforcement out of the prompt and into the tool boundary. If an agent has to decide to fetch a rule, it skips it; if you put all rules in AGENTS.md, attention dissolves over long trajectories. Injecting constraints inside pre-tool hooks or blocking invalid tool arguments at the harness gate survives long runs because it does not depend on the model maintaining attention over twenty turns of history.
Agreed, and "step 8, after two tool errors flood the context" is a good way to put it. One nuance on McMillan: the 5.6% is an association, and he notes it isn't a steady per-step decline, so I read it as "adherence erodes as the session fills up" rather than a fixed curve. Your failure mode fits that.
I'd split it the way you suggest. Rules a tool can check, like invalid arguments, a protected path, or a commit without passing tests, belong at the gate, where the model can't forget them. Rules that need judgment, like "never trust a client-supplied price", can't be blocked mechanically. For those, the next best thing is re-injecting the rule right before the tool call that matters, when the agent is about to edit the file it governs, and backing it with a test in CI. That re-injection at edit time is the next thing I'm building.