I spent six weeks convinced my agents had a reasoning problem.
They contradicted each other, repeated work, and confidently cited facts that nobody had established. I upgraded models. I rewrote prompts. I added a knowledge graph because the research said graphs were the future. Each fix bought me a week of calm before the same failures crept back.
The turning point came when I stopped asking "why are my agents wrong?" and started asking a different question: why does accuracy drop when I give them more context?
That's when I found the answer, and it wasn't the one I expected.
The Failure I Kept Misdiagnosing
I'd built a research pipeline with five agents: scout, analyst, synthesizer, fact-checker, and writer. Each received the accumulated context of everything that had come before. Sources read. Conclusions drawn. Draft sections. Intermediate tool outputs. The entire history, passed forward in one growing blob because that's what the tutorials showed.
The first sign was subtle. The analyst would repeat a conclusion the scout had already established, phrasing it as a new discovery. The second sign was louder. The writer would confidently assert a fact that appeared nowhere in the source material—but which had been hinted at in an intermediate note the fact-checker had already flagged as unverified.
I kept looking for a reasoning bug. What I found was a context bug.
Chroma's 2025 research had already formalized what I was experiencing. They tested 18 frontier models—GPT-4.1, Claude 4, Gemini 2.5, Qwen3—and found that every single one gets worse as input length increases . Not some. Not most. All of them. The phenomenon now has a name: context rot .
The numbers are brutal. Chroma found 20–50% performance drops as input length grew from 10k to 100k+ tokens, with coherent documents paradoxically degrading performance more than shuffled ones . Reasoning-capable models showed up to 80% accuracy drops from contextual distractors. And the degradation isn't a cliff—it's a continuous decline that starts well before any token limit is reached. A model with a 200K token window can exhibit significant degradation at 50K tokens .
I had been treating my context window as a container. Capacity was the metric. Signal-to-noise ratio was what actually determined output quality.
The Research That Changed How I Build
The discipline that replaced my patchwork approach has a name: context engineering. Google's ADK team describes it as treating context as a first-class system with its own architecture, lifecycle, and constraints. The core thesis is that context should be a compiled view over a richer stateful system—not a mutable string buffer that each agent accumulates independently.
I rebuilt my pipeline around four strategies.
Write to external memory, not the context window. Instead of accumulating every intermediate output in the agent's context, write it to a filesystem or a structured store. LangChain's Deep Agents SDK uses filesystem tools—read, write, edit, list, search—to let agents offload context without losing access to it. The context window stays lean. The memory persists.
Select what each agent actually needs. Not everything that came before is relevant to what comes next. Sentex, a context management middleware for multi-agent pipelines, puts every agent output into a shared sentence graph and retrieves exactly the sentences relevant to the next agent's task—traversing semantic KNN edges across agent boundaries—within a token budget you set. Agent 3 doesn't carry the full output of Agents 1 and 2. It carries the eight sentences that matter.
Compress aggressively. Summarize intermediate outputs before they accumulate. LangChain's Deep Agents uses sub-agents to summarize large results into a single compressed output before the main agent sees them . The main agent gets the signal. The noise stays behind.
Isolate context by role. Each agent gets its own view of the shared context, filtered by role. A researcher doesn't need the writer's draft. A fact-checker doesn't need the scout's raw search results.
The Retrieval Trade-Off Most Teams Miss
Here's the part I got wrong for months. When accuracy drops, the instinct is to retrieve less—be more selective, pull fewer chunks, trust the model's context window more.
That instinct is half right and half dangerous.
The LaRA benchmark, published at ICML 2025, evaluated eleven models across 2,326 test cases and found that the optimal choice between RAG and long-context depends on a complex interplay of model capabilities, context length, task type, and retrieval characteristics. Neither RAG nor long-context LLMs are a silver bullet.
The numbers are nuanced. At 32k context length, long-context achieved an average accuracy 2.4% higher than RAG. At 128k context length, the trend reversed, with RAG outperforming long-context by 3.68%.
But the real problem is worse than simple retrieval failure. A 2025 EMNLP paper, "Context Length Alone Hurts LLM Performance Despite Perfect Retrieval," found that even when models can perfectly retrieve all relevant information, their performance still degrades substantially—13.9% to 85%—as input length increases . This failure occurs even when irrelevant tokens are replaced with whitespace, and even when they are masked entirely .
This is the finding that should stop every multi-agent architect cold. The sheer length of the input alone can hurt LLM performance, independent of retrieval quality and without any distraction.
I had been optimizing retrieval for six weeks. The problem wasn't retrieval. The problem was length.
The Silent Killer: Lost in the Middle
The second failure mode compounds the first. A 2026 Bangla long-context study confirmed what the original Stanford "Lost in the Middle" research found: LLMs exhibit a U-shaped performance curve, prioritizing information at the beginning (primacy bias) and end (recency bias) of the input while neglecting the middle .
For English, the middle-position accuracy drop is 20-25 percentage points . For Bangla, the same pattern held . The bias is architectural, not linguistic.
In my pipeline, the most critical findings—the conclusions from the analyst, the flagged issues from the fact-checker—were almost always in the middle of the context. The writer saw them but didn't use them. The model was attending to the edges and skimming the center.
What Production Teams Are Actually Running
The teams shipping multi-agent systems to production aren't solving this with longer context windows. They're solving it with architectural discipline.
Microsoft's Azure SRE Agent started with 100+ specialized tools and 50+ agents. It failed in production—coordination failures, infinite loops, poor generalization . They pivoted to five wide tools (primarily az and kubectl), a handful of generalist agents, and aggressive context management. The impact was transformative: massive context compression by recovering headroom previously consumed by tool definitions, expanded capabilities because the model had access to the entire CLI surface area, and improved reasoning quality because LLMs already possess knowledge of standard CLIs from training data .
The tool count wasn't the problem. The context bloat was.
Slack's security investigation system handles investigations spanning hundreds of inference requests and megabytes of output. Their solution: three specialized context channels—a Director's Journal for structured working memory, a Critic's Review with credibility-scored findings, and a Critic's Timeline for consolidated chronological evidence. Critically, they pass no message history forward between agent invocations. The Critic filters out approximately 26% of findings that don't meet plausibility thresholds .
Lyft's LangGraph-based support system uses router-based multi-agent architecture with state management and handoffs built into the flow . Agent development accelerated from roughly six months to just a few weeks .
Rippling developed three context engineering patterns to reduce context bloat, using Deep Agents middleware for dynamic skill injection.
The Mitigation That Actually Works
The context engineering discipline has four core strategies, and I rebuilt my pipeline around all of them.
First, define the bar explicitly. Not "does it return JSON" but "does it classify the ticket the way the primary model would." The bar has to be measured on your real task distribution, not a generic benchmark.
Second, test the fallback as a production path, not a spare tire. The model swap deserves the same regression testing as a prompt change. Hold the prompt, tools, cases, judge, and inference parameters still, and vary only the model.
Third, gate the fallback with the same validation you'd apply to the primary. A schema check is necessary but nowhere near sufficient. A response can conform perfectly to a schema and still classify incorrectly, choose the wrong tool, overstate confidence, or take a tone that doesn't belong in your product.
Fourth, use the recitation mitigation. The EMNLP paper's simple, model-agnostic mitigation strategy is to transform a long-context task into a short-context one by prompting the model to recite the retrieved evidence before attempting to solve the problem . On RULER, this yielded a consistent improvement of GPT-4o up to 4% on an already strong baseline .
The Trade-Off You're Accepting
Context engineering isn't free. It's more work than dumping everything into the prompt and hoping.
You're designing what each agent sees. You're building the retrieval layer that gets the right context to the right agent at the right time. You're writing the compaction logic that summarizes intermediate outputs. You're maintaining the schemas that make context queryable rather than just readable.
You're also accepting that some context will be lost. Aggressive compression means the writer doesn't see every detail the analyst considered. Progressive disclosure means the CEO doesn't see the raw worker output. You're trading completeness for signal. The research says that's the right trade—context rot degrades performance even when the context is coherent. But it's a trade you have to make deliberately, not accidentally.
And you're accepting that context engineering is a discipline, not a one-time fix. The context that was right for a five-step pipeline isn't right for a fifteen-step one.
The Question I Keep Coming Back To
If you looked at what your agents actually see right now, how much of that context is load-bearing—and how much is rot?
I'd love to hear where you've landed. Filesystem offloading, progressive disclosure, aggressive compaction, or a context budget you've never actually measured—and what finally made you look?
Top comments (0)