When we build AI systems, one instinct shows up again and again : if the model can use more information, give it more information.
You have a coding agent with access to an entire repository, so why not give it more files? A RAG pipeline retrieved ten documents, so why throw any of them away? The agent returned a useful tool result, so you append it to the conversation and carry it into the next iteration.
I have found this pattern surprisingly easy to ignore because nothing breaks immediately. The model still responds. The answer may still look reasonable. The context window still has capacity. That is exactly what makes the problem interesting.
Over time, an AI system can accumulate conversation history, retrieved documents, source code, logs, tool output, and intermediate results. Eventually, prompt construction itself becomes part of the system's performance profile.
We already know how to deal with this class of problem elsewhere.
Databases use indexes. Caches use eviction. Operating systems manage working sets. Distributed systems put bounds around queues and buffers.
I think AI systems need the same discipline - context window is not a working set. A context window tells us how much information a model can accept. A working set tells us how much information the model actually needs for the current decision. Those are very different things.
Consider a coding agent working against:
2,400 files
180,000 lines of code
A persistence bug might involve:
OrderController
OrderService
OrderRepository
Order entity
One failing test
Relevant stack trace
Perhaps 6,000–8,000 tokens are enough.
The repository remains the source of truth. The model does not need the repository inside its context.
A better pipeline is:
Repository
↓
Search / Index
↓
Candidate files
↓
Symbol / Dependency analysis
↓
Relevant code slices
↓
Context builder
↓
LLM
This is fundamentally the same idea as database indexing. We don't scan an entire table into application memory just because the database contains the data. We narrow the working set before processing it.
Long-context research gives us another reason to be cautious.
The Lost in the Middle study found that model performance can degrade when relevant information is buried in the middle of a long context. The conclusion is not that long context is bad. It is that capacity is not relevance.
The problem becomes even more obvious when the model is called repeatedly. Suppose an agent performs 100 model calls and each call carries an average of 40,000 input tokens. If better retrieval and context construction reduce the working set to 12,000 tokens. That is a 70% reduction in input-token volume. The exact financial impact depends on the model, provider pricing, caching, and workload. The architectural point is :
Every unnecessary token can be multiplied by every model invocation that carries it.
This matters particularly in agentic workflows, where one user request can result in dozens of model and tool interactions.
There is another boundary I think AI architectures need to make explicit: the boundary between tool output and persistent context.
Suppose an agent executes:
run_tests()
and receives 15,000 tokens of output.
A naïve implementation does:
Tool
↓
15,000 tokens
↓
Conversation history
↓
LLM
A better design is:
Tool
↓
Raw output
↓
Filter / summarize
↓
Relevant findings
↓
Context builder
↓
LLM
The tool response is an intermediate result that may contribute to context. It should not become persistent context by default.
A test runner might produce thousands of lines containing successful tests, timestamps, warnings, stack traces, and environment details. The model may only need:
3 failing tests
2 relevant stack traces
5 changed files
The same principle applies to database queries, APIs, code search, logs, and browser results. Tool interfaces should therefore support filtering, pagination, field selection, structured responses, and bounded results wherever possible.
This is becoming a real design concern for agent platforms. Anthropic, for example, has documented tool-use patterns specifically aimed at avoiding the cost of loading large tool definitions and intermediate tool results into the model's context. Its Tool Search approach allows tools to be discovered and loaded on demand rather than putting every tool definition into context upfront.
A tool that routinely returns 50,000 tokens to an agent is not only a tool-design problem. It is a context-architecture problem.
Once context is treated as a working set, another question becomes important:
What is allowed into it?
I would treat context construction almost like a cache admission policy. For every candidate piece of information:
Is it relevant?
Is it authoritative?
Is it current?
Is it duplicated?
Will the model need it for the next decision?
That gives us:
Candidate information
↓
Relevance
↓
Deduplication
↓
Freshness
↓
Priority
↓
Token budget
↓
Model context
This also gives information a lifecycle.
A system constraint may remain relevant throughout a session. A retrieved document may matter for one decision. A raw database response might only be useful for one tool call. A 10,000-line build log might ultimately produce one useful fact. Treating all of these as permanent conversation history is effectively an eviction policy of never evict which is rarely good for any resource.
So How Do We Reduce Context?
Once the problem is framed this way, the controls become fairly mechanical.
Retrieve Less, but Retrieve Better
Don't retrieve 20 documents and ask the model to decide which three matter. Use metadata filters, hybrid search, reranking, namespace boundaries, and query-specific retrieval before constructing the prompt. The goal isn't simply a smaller top-k. It is a better signal-to-noise ratio.
Slice Code Instead of Retrieving Files
For coding agents, retrieving an entire 2,000-line class because one method is relevant is unnecessary. Use symbol-level retrieval, AST-aware slicing, dependency analysis, or line-range extraction where possible.
2,000-line file
↓
Relevant method
↓
Callers / dependencies
↓
Required context
The model gets the code it needs rather than the file that happens to contain it.
Summarize Tool Output Before It Becomes Context
Treat tool output as an intermediate representation.
Raw output
↓
Parse
↓
Filter
↓
Extract findings
↓
Context
Do not make the LLM perform work that deterministic software can perform first.
Compact Conversation State
Long-running agents accumulate history even after information has served its purpose. Instead of carrying every previous turn forward:
Raw conversation
↓
Decisions
Constraints
Open questions
Current state
↓
Compact state
The objective is to preserve state, not transcript. This is already becoming a standard technique in production agent architectures. OpenAI provides server-side and standalone context compaction for long-running conversations, while Anthropic describes compaction and structured note-taking as mechanisms for maintaining coherence across long-horizon tasks.
Put a Budget Around Context Sources
I would rather have explicit limits than discover context growth through production latency.
For example:
Conversation history: 4K
Retrieved documents: 8K
Tool output: 4K
Source code: 8K
System instructions: 2K
────────────────────────────
Input budget: 26K
The numbers are workload-specific. The principle isn't. Every source competing for context should have a budget and a priority.
The Tradeoff
There is an important tradeoff here. Reducing context is not automatically an optimization.mIf I aggressively trim a prompt and remove the one piece of information that explains a subtle business rule, the model may produce a faster answer that is simply wrong. So I don't think the goal should be:
Minimize context.
It should be:
Maximize useful information within a bounded context budget.
A retrieval system that returns only highly relevant documents but misses critical supporting information has poor coverage. A system that retrieves everything has poor selectivity. The engineering problem is finding the point where the model has enough information to make the decision without carrying everything that might be useful and that is why I would give different context sources different priorities:
System constraints → mandatory
Current task → mandatory
Relevant source code → high
Authoritative documents → high
Recent tool findings → medium/high
Conversation history → medium
Raw tool output → low
Duplicate information → discard
The exact ordering depends on the application but the important part is that the system has a policy rather than relying on the model to make the decision implicitly.
Once context has a budget, something has to enforce it. I would put context construction behind an explicit budget manager:
Candidate Context
↓
Deduplicate
↓
Classify / Prioritize
↓
Per-source budgets
↓
Global budget
↓
┌──────────────────────────┐
│ Fit? │
│ │
│ Yes → Build context │
│ No → Trim / Summarize │
│ / Retrieve again │
└──────────────────────────┘
↓
LLM
There is another detail that is easy to miss - Input tokens and output tokens share the model's total context capacity.
If a model supports a 32K total window, allocating all 32K to input leaves no room for the response. So I would calculate:
output_reservation =
expected_output + output_buffer
input_hard_limit =
model_limit
- output_reservation
- safety_margin
For example:
Model limit: 32K
Expected response: 4K
Output buffer: 1K
Safety margin: 1K
──────────────────────────────
Hard input limit: 26K
The reservation should be dynamic.
A classification request might need only 500 output tokens, while a coding-agent step generating a patch may need 6,000. The same model can therefore have different input budgets for different operations. I would also distinguish between soft and hard limits:
0 ───────────── 20K ───────────── 26K ───── 32K
│ │
Soft limit Hard limit
│ │
Optimize / trim Never cross
The soft limit is an operating target. Crossing it should trigger more aggressive filtering or compaction. The hard limit is an enforcement boundary. The system should never send a request above it.
The hardest case is mandatory context that already exceeds the available input budget. Suppose:
Hard input limit: 26K
Mandatory context:
System constraints: 5K
Current task: 4K
Security policy: 3K
Required task state: 18K
──────────────────────────────
Total: 30K
There is nothing meaningful to evict. The system should not silently truncate it and hope the model can recover. In such cases I would use an explicit escalation path:
Mandatory context > hard limit
↓
Try deterministic compression
↓
Compact structured state
↓
Split / stage the task
↓
Retrieve incrementally
↓
Still too large?
↓
Fail explicitly
For example, 18K of task state might contain verbose conversation history that can be converted into 7K of structured facts, decisions, constraints, and unresolved questions. If that makes the request fit, proceed. If the information is genuinely irreducible, stop. I would rather return:
CONTEXT_BUDGET_EXCEEDED
Required input: 30K
Available input: 26K
Suggested action:
- compact task state
- split the task
- retrieve information incrementally
than silently remove information classified as mandatory.
That gives us a useful rule:
Optional context can be evicted. Mandatory context must be compressed, partitioned, or cause the workflow to stop.
The policy itself can be implemented deterministically:
function buildContext(candidates, modelLimit, expectedOutput):
outputReservation = expectedOutput + OUTPUT_BUFFER
hardInputLimit =
modelLimit - outputReservation - SAFETY_MARGIN
candidates = deduplicate(candidates)
mandatory = candidates.filter(isMandatory)
optional = candidates.filter(notMandatory)
if tokenCount(mandatory) > hardInputLimit:
mandatory = compact(mandatory)
if tokenCount(mandatory) > hardInputLimit:
mandatory = partitionForStagedExecution(mandatory)
if tokenCount(mandatory) > hardInputLimit:
return CONTEXT_OVERFLOW
remaining =
hardInputLimit - tokenCount(mandatory)
optional = prioritize(optional)
optional =
admitUntilBudget(optional, remaining)
context = mandatory + optional
if tokenCount(context) > hardInputLimit:
return CONTEXT_OVERFLOW
return context
The important part is the order.
Mandatory context is protected first. Optional information competes for whatever budget remains.
And once prompt construction sits on the runtime path, I would measure it like any other resource. At minimum:
input_tokens / model_call
retrieved_tokens / request
tool_output_tokens / request
conversation_tokens / request
context_growth / agent step
P50 / P95 / P99 context size
context_truncation_rate
context_budget_utilization
context_overflow_rate
I would particularly watch P95 and P99 context size. An agent that normally operates around 12K tokens but occasionally reaches 120K has a very different architecture from one that consistently stays near 12K, even if their averages look similar. It is also useful to know where those tokens came from:
Context
├── Conversation history 18%
├── Retrieval 32%
├── Tool output 27%
├── Source code 19%
└── System instructions 4%
Now "the prompt is too large" becomes an engineering problem with a location. If tool output dominates, fix the tool boundary. If conversation history dominates, introduce compaction. If retrieval dominates, improve ranking and filtering. If source code dominates, make retrieval more code-aware. The value of these metrics is not in measuring context for its own sake. They tell us where the architecture needs to change and this leads to the architecture I would build: one where information can exist broadly across the system, but only a deliberately selected subset is allowed into the model's active context.
The model sits at the end of that pipeline, not at the center of the storage layer. Repository data can remain in the repository. Documents can remain in the retrieval system. Raw logs can remain in observability infrastructure. Tool results can remain transient until their findings are extracted. The context builder decides what crosses the boundary.
That separation also makes the system easier to reason about. If context grows unexpectedly, we can identify whether the problem is retrieval, tool design, conversation state, code selection, or budget enforcement instead of treating the entire prompt as one opaque string.
What I find interesting is that this isn't just a theoretical architecture anymore. The industry has started to treat context this way now.
OpenAI now exposes context compaction specifically for long-running interactions, reducing accumulated context while preserving state needed for subsequent turns.
Anthropic explicitly describes context engineering as the broader problem of curating the information available to a model, including system instructions, tools, external data, and message history. Its guidance covers just-in-time retrieval, compaction, structured note-taking, and sub-agent architectures.
Anthropic's Tool Search is another concrete example of the same idea. Instead of putting thousands of tool definitions into context, tools can be discovered and loaded when they are actually needed. Anthropic reports that this can substantially reduce context consumption in large tool environments.
Google takes a slightly different approach with long-context models and context caching. Gemini supports very large context windows, while its caching mechanisms allow repeatedly used context to be reused rather than repeatedly processed as ordinary input. Google's documentation also explicitly notes that longer contexts can increase latency and that unnecessary tokens should generally be avoided.
These are different mechanisms solving different problems.
Retrieval controls what enters the working set.
Compaction controls accumulated state.
Caching reduces repeated processing of stable context.
Tool discovery and filtering control what external systems contribute.
The implementation differs, but the architectural direction is remarkably similar - context is becoming something applications need to manage deliberately.
Key Takeaway
The model should not become the database, repository, cache, or log store. It should be the reasoning component operating over a deliberately constructed working set.
My simple rule is:
Store broadly. Retrieve selectively. Admit deliberately. Reason narrowly. Persist the result.
AI systems have given software architects another resource to manage, not because context windows are small, but because they are now large enough for poor context management to hide inside them. And that leads to the principle I keep coming back to:
A bigger context window gives you more capacity. It doesn't give you a reason to fill it.
References
- Liu et al., Lost in the Middle: How Language Models Use Long Contexts — arXiv paper
- OpenAI — Context Compaction
- Anthropic — Effective Context Engineering for AI Agents
- Anthropic — Advanced Tool Use
- Google AI — Context Caching
- Google AI — Long Context

Top comments (0)