DEV Community

Cover image for Giving every agent the full context is why your reasoning degrades — the context engineering pattern that fixes it — Context Engineering
Alex Aslam
Alex Aslam

Posted on

Giving every agent the full context is why your reasoning degrades — the context engineering pattern that fixes it — Context Engineering

I spent two months convinced my agents had a reasoning problem. They contradicted each other, repeated work, and confidently cited facts that nobody had established. I upgraded models. I rewrote prompts. I added a knowledge graph because the research said graphs were the future. Each fix bought me a week of calm before the same failures crept back.

The turning point came when I stopped asking "why are my agents wrong?" and started asking a harder question: "what are they actually seeing?"

That's when I found it. My agents weren't bad at reasoning. They were reasoning over polluted contexts, and I had been the one polluting them.

The Failure I Kept Misdiagnosing

I'd built a research pipeline with five agents: scout, analyst, synthesizer, fact-checker, and writer. Each agent received the full accumulated context of everything that had come before. Sources read. Conclusions drawn. Draft sections. Intermediate tool outputs. The entire history, passed forward in one growing blob because that's what the tutorials showed.

The first sign was subtle. The analyst would repeat a conclusion the scout had already established, phrasing it as a new discovery. The second sign was louder. The writer would confidently assert a fact that appeared nowhere in the source material—but which had been hinted at in an intermediate note the fact-checker had already flagged as unverified.

I kept looking for a reasoning bug. What I found was a context bug.

The research had already named what I was experiencing. Context rot — the measurable degradation in LLM output quality as input length increases — was formalized by Chroma's 2025 research, which tested 18 frontier models and found that every single one gets worse as input grows. Not some. Not most. All of them. Even when the context window isn't close to full.

The numbers are brutal. Chroma found 20–50% performance drops as input length grew from 10k to 100k+ tokens, with coherent documents paradoxically degrading performance more than shuffled ones. Reasoning-capable models showed up to 80% accuracy drops from contextual distractors. And the degradation isn't a cliff—it's a continuous decline that starts well before any token limit is reached. A model with a 200K token window can exhibit significant degradation at 50K tokens.

I had been treating my context window as a container. Capacity was the metric. Signal-to-noise ratio was what actually determined output quality.

The Discipline That Changed How I Build

The shift has a name now: context engineering. Google's ADK team describes it as treating context as a first-class system with its own architecture, lifecycle, and constraints. LangChain's definition is the one I keep taped to my monitor: context engineering is the art and science of filling the context window with just the right information at each step of an agent's trajectory.

Karpathy's framing is the one that finally made it click. LLMs are like a new kind of operating system. The LLM is the CPU. The context window is RAM. And just as an operating system curates what fits into a CPU's RAM, context engineering plays that same role.

I had been treating my context window like a hard drive. Everything went in. Nothing came out. The CPU was drowning in RAM that should never have been there.

What Context Engineering Actually Looks Like

The discipline has four core strategies, and I rebuilt my pipeline around all of them.

Write to external memory, not the context window. Instead of accumulating every intermediate output in the agent's context, write it to a filesystem or a structured store. LangChain's Deep Agents SDK uses filesystem tools—read, write, edit, list, search—to let agents offload context without losing access to it. The context window stays lean. The memory persists.

Select what each agent actually needs. Not everything that came before is relevant to what comes next. Sentex, a context management middleware for multi-agent pipelines, puts every agent output into a shared sentence graph and retrieves exactly the sentences relevant to the next agent's task—traversing semantic KNN edges across agent boundaries—within a token budget you set. Agent 3 doesn't carry the full output of Agents 1 and 2. It carries the eight sentences that matter.

Compress aggressively. Summarize intermediate outputs before they accumulate. LangChain's Deep Agents uses sub-agents to summarize large results into a single compressed output before the main agent sees them. The main agent gets the signal. The noise stays behind.

Isolate context by role. Each agent gets its own view of the shared context, filtered by role. A researcher doesn't need the writer's draft. A fact-checker doesn't need the scout's raw search results. The Atlan framework calls these "role-scoped context views": each agent gets only the context it needs for its specific job.

What the Research Quantifies

The 2026 literature has moved past theory. The numbers are concrete.

A March 2026 paper introduced five production-grade context quality criteria: relevance, sufficiency, isolation, economy, and provenance. Each is a design constraint, not a vibe. Relevance means every token in the context window is load-bearing. Sufficiency means nothing critical is missing. Isolation means one agent's context doesn't leak into another's. Economy means you're not paying for tokens the model doesn't need. Provenance means you can trace every fact back to its source.

The paper's conclusion is the line I keep coming back to: whoever controls the agent's context controls its behavior. Not whoever controls the prompt. Not whoever controls the model. Whoever controls the context.

The Slack security investigation system demonstrates this in production. Their multi-agent system handles investigations spanning hundreds of inference requests and megabytes of output. Their solution: three specialized context channels—a Director's Journal for structured working memory, a Critic's Review with credibility-scored findings, and a Critic's Timeline for consolidated chronological evidence. They pass no message history forward between agent invocations. The agents don't accumulate. They query structured context and produce structured output.

Microsoft's Azure SRE Agent tells the same story from a different angle. They started with 100+ narrow tools and 50+ specialized agents. It failed in production—coordination failures, infinite loops, poor generalization. They pivoted to five wide tools, a handful of generalist agents, and sophisticated context management: code interpreters for computation, progressive disclosure through file-based systems, aggressive context compaction, and planned tool call chaining. The result was dramatically improved reliability and the ability to handle unanticipated scenarios.

The tool count wasn't the problem. The context bloat was. Every tool definition, every schema, every narrow guardrail consumed tokens that should have been available for the actual task. Microsoft's team described their original architecture as "a brittle workflow system with an LLM grafted on top". The pivot to context engineering was what made it an agent.

What Production Teams Are Actually Running

Slack's security investigation system uses three context channels and zero message history passing. The Critic filters out approximately 26% of findings that don't meet plausibility thresholds. The system maintains coherence across hundreds of inference requests by never letting context accumulate in the first place.

monday.com's Sidekick started with one general-purpose agent and a growing list of tools. In production, every new tool made the system more ambiguous, more expensive, and harder to debug. They tore it down and rebuilt around bounded responsibilities, sandboxes, and specialized subagents. Their key insight: a single prompt had to contain instructions for research, content generation, data analysis, board operations, and file processing. The agent wasn't too general. The context was too crowded.

Microsoft's Azure SRE Agent consolidated from 100+ narrow tools to five wide CLI tools, achieving massive context compression and dramatically improved reasoning quality because LLMs already possess knowledge of standard CLIs from their training data.

Rippling developed three context engineering patterns to reduce context bloat, using Deep Agents middleware for dynamic skill injection.

The Trade-Off You're Accepting

Context engineering isn't free. It's more work than dumping everything into the prompt and hoping.

You're designing what each agent sees. You're building the retrieval layer that gets the right context to the right agent at the right time. You're writing the compaction logic that summarizes intermediate outputs. You're maintaining the schemas that make context queryable rather than just readable.

You're also accepting that some context will be lost. Aggressive compression means the writer doesn't see every detail the analyst considered. Progressive disclosure means the CEO doesn't see the raw worker output. You're trading completeness for signal. The research says that's the right trade—context rot degrades performance even when the context is coherent. But it's a trade you have to make deliberately, not accidentally.

And you're accepting that context engineering is a discipline, not a one-time fix. The context that was right for a five-step pipeline isn't right for a fifteen-step one. The schema that worked for two agents doesn't work for ten. You're maintaining an architecture, not writing a prompt.

But here's what I've learned from every pipeline post-mortem I've sat through: the teams that fail aren't the ones with bad models. They're the ones with polluted contexts. The reasoning was always good enough. The signal was just buried under noise.

So here's my question: If you looked at what your agents actually see right now, how much of that context is load-bearing—and how much is rot?

I'd love to hear where you've landed. Filesystem offloading, progressive disclosure, aggressive compaction, or a context budget you've never actually measured and what finally made you look?

Top comments (0)