"Context engineering" started turning up in job titles this year. That usually means a thing has become real, and also that it's become vague. The version I hear most is that prompt engineering was about the words in the instruction, and context engineering is about everything else the model sees: the documents, the tool results, the conversation so far, the memory. That's true, and it doesn't tell you what to do.
I find it more useful to treat the context window as a cache. It's fast, it's small next to everything the agent might need, everything in it costs money on every call, and what you're managing is what sits in it at the moment the model has to decide something. Seen that way, every technique that works already has a name in the cache literature, and the ones that don't work are things a cache wouldn't do either.
Bigger did not mean solved
Two years ago the argument went that context management was a temporary problem. Windows would grow, a million tokens would hold the whole codebase and the whole conversation, and we'd stop thinking about it.
The windows grew and the problem stayed, for two reasons that look obvious now.
One is cost. A model call is billed on input tokens, so once a conversation reaches 400,000 tokens, every turn costs 400,000 tokens, and an agent that takes sixty turns to finish a task has paid for that history sixty times.1
The other is that the model doesn't use a long context evenly. An instruction in the first thousand tokens weighs differently from the same instruction at token 300,000 with 299,000 tokens of tool output in front of it. Every benchmark that measures this finds some decay. In practice it looks like the agent forgetting a constraint you gave it at the start, right around when the context got heavy. A bigger window means it forgets later, and pays more when it does.
So the window is a cache with a price per byte per access and a hit rate that gets worse as it grows. Nothing strange about that. Every cache works like that.
What goes in
A cache holds the working set, whatever the next operation is likely to need. For an agent that's the current task description, the constraints on it, the most recent tool results, and whatever the model has learned during this run that it'll need again.
It shouldn't hold everything that happened: the full output of a test run from twenty turns ago, a file the agent read once and moved on from, the six wrong approaches it tried before the right one. That's history. If nothing ever gets evicted you have a log, and a log is what the context window becomes when nobody manages it.
The agents that hold up in long sessions all do the same thing here. Coding agents pin the task and the constraints near the front, keep recent tool results in full, and summarise older ones down to a line each. The summary is the eviction. You keep the fact that the tests ran and which two failed, and drop the 8,000 tokens of output the run produced.
Compaction is write-back
The technique with the most names is the one where the agent, once the context gets full, rewrites its own history into something shorter and carries on from there. Claude Code calls it compaction.2
In cache terms it's write-back. Entries on their way out get written to a compact form before they're dropped, so what they said survives in a size that fits. What matters is what gets kept. A good compaction keeps what the task is, what's been decided, what's done, what's in progress, and any facts found along the way that the model would otherwise have to rediscover with another tool call. A bad one is a paragraph saying "the user asked for X and we worked on it", which is like writing a dirty cache line back as zeros.
I've watched an agent lose a constraint through a bad compaction. In turn three the user said not to modify the migration files. At turn forty the context was compacted and the summary said "working on the billing feature". At turn forty-four the agent modified a migration file. The constraint had been in the cache, the eviction didn't write it back, and the model had no way of knowing it had ever been there.
The fix is in the structure. Constraints and decisions are state, and state shouldn't get compacted along with the history. Give them their own section that survives compaction word for word, or put them in a file the agent re-reads.
Memory files are the L2
That file is the other technique that spread this year: a few files the agent reads at the start of every session and can write to during one.3
In cache terms this is the next level down. It's slower to reach, because the agent has to read it. It's bigger and it persists, and it holds whatever should outlive both compaction and the end of the session: project conventions, things the user has said they always or never want, facts about the environment that were expensive to find out. The agent promotes something to this level when it decides it's worth keeping, and the promotion is a write to disk.
The requirements are the same as for any second level cache. It has to be small enough to load every time, and indexed so the agent can find what it needs without reading all of it. What works is an index file with one line per memory and a separate file for each memory, so the index is always in the context and the memories get pulled in when they're relevant. That's a page table. We didn't invent anything.
Retrieval is the miss path
When the model needs something that isn't in the window, it has to go and get it. That's a cache miss, and the miss path is retrieval: search the codebase, search the documents, query the knowledge base, call the tool.4
A good miss path is what lets you keep the cache small. If the agent can find any file in the repository with one tool call, it doesn't need every file in the context, and it can drop a file as soon as it's done with it. If retrieval is bad, say search returns forty results and the right one is thirty-first, the agent makes up for it by hoarding. It keeps everything it's seen in case it needs it again, and the context fills up with insurance.
Most teams invest in the opposite order. They put the effort into what to load up front and neglect search. Search pays back more, because a cheap, precise miss path makes every other decision easier.
Prefetching, and knowing when not to
The last piece is guessing what the next turn needs and loading it before the model asks. A coding agent asked to change a function can load the function, its callers and its tests before the first model call, because the model is going to want them. That's prefetching, and it saves a turn.
It goes wrong when you prefetch too much: the whole module because the function lives in it, every caller's file in full, the whole test suite. Now the first turn starts with 60,000 tokens of context, the model will use 5,000 of them, and the cache is polluted before any work has happened. The rule is the one hardware uses. Fetch what the access pattern predicts, and leave what merely happens to be nearby.
The discipline, in one place
Pin the task and the constraints, and keep them out of anything that gets summarised.
Keep recent tool results in full and cut older ones down to a line. That's your eviction policy.
When you compact, write back state and skip the narrative: decisions, facts, progress. If you couldn't resume the task from the summary, the write-back failed.
Promote durable facts to files with an index, and load the index every session.
Make search precise, so the agent can afford to forget.
Prefetch what the task predicts and nothing else.
None of this needs a bigger model or a bigger window. It needs someone to look at the context the way they'd look at a cache hit rate graph and ask, of every token in there, whether the next decision needs it. Most of the time it doesn't, and that token is costing you money and attention for nothing.
Originally published at zeybek.dev.
-
The providers' prompt caching takes a lot of the sting out, but it's a discount on re-reading. You still pay, and that cache has its own invalidation rules, which a long agent run trips over all the time. ↩
-
Others call it summarisation, checkpointing or context folding. ↩
-
Claude Code's
CLAUDE.mdand its memory directory are the well known example, and the pattern has been copied everywhere. ↩ -
Retrieval augmented generation is a name for making the miss path good. ↩
Top comments (0)