DEV Community

Cover image for Your multi-agent system doesn’t have a memory problem — it has a shared-context problem — Shared Context Layer
Alex Aslam
Alex Aslam

Posted on

Your multi-agent system doesn’t have a memory problem — it has a shared-context problem — Shared Context Layer

I spent six weeks convinced my multi-agent system had a memory problem. I added a vector store. Then a second one. Then a knowledge graph, because the research said graphs were the future. Each addition bought me a week of calm and then the same failures crept back: agents contradicting each other, repeating work, and confidently citing facts that no one had actually established.

The turning point came when I stopped asking "how do I give my agents better memory?" and started asking "why are they each building their own version of reality?"

That's when I saw it. My agents didn't have a memory problem. They had a shared-context problem.

The Failure Nobody Names Correctly

I'd built a system where each agent accumulated its own context. The researcher remembered sources it had read. The analyst remembered conclusions it had drawn. The writer remembered the draft it had produced. Each agent's context was internally consistent and completely isolated from every other agent's.

The first time I noticed the cost, a support agent told a customer that a refund was "being processed" while the billing agent had already flagged the request as ineligible. Both agents were reasoning correctly. Neither was wrong. They were operating on different pictures of the same situation, and no one had told them to compare notes.

The research had already quantified what I was feeling. In large-scale deployments of multi-agent frameworks, the task failure rate can climb to 40–80% without coordination through shared memory. Roughly 36.9% of those failures are attributed to misalignment issues—inconsistent states or goal deviations between agents. The paper's conclusion was blunt: operating independently without a common memory repository, different agents develop conflicting understandings of the environment, leading to contradictory decisions that undermine the stability of the overall system.

I had been solving the wrong problem. Vector stores and knowledge graphs help a single agent remember more. They do nothing for the failure mode where two agents disagree because they're reasoning over different fragments of the truth.

What Shared Context Actually Means

The discipline that replaced my patchwork approach has a name now: context engineering. Google's ADK team describes it as treating context as a first-class system with its own architecture, lifecycle, and constraints. The core thesis is that context should be a compiled view over a richer stateful system—not a mutable string buffer that each agent accumulates independently.

The mental model shift is this. In a single-agent system, you ask: "What context does this agent need to answer well?" In a multi-agent system, you have to ask a harder question: "How do several agents share the same business understanding while each agent sees only the context its role requires?"

That second question is what I had never asked. I'd been optimizing retrieval per agent. I hadn't designed the shared substrate that makes those agents' retrievals compatible.

The Architecture That Fixed It

The fix has three layers. Each one solves a specific failure mode I'd been papering over with prompt engineering.

First, a shared context repository that agents read from and write to. Not a message queue. Not point-to-point handoffs. A persistent, versioned store of facts, decisions, and artifacts that every agent can query. The Atlan guide calls this the "single trusted source" layer—every agent uses the same approved definitions, metrics, and policies. When my researcher established a fact, it didn't just pass that fact forward. It wrote it to the shared layer. When my analyst needed context, it didn't ask a peer. It queried the shared layer.

Second, role-scoped views instead of a shared prompt. This is the part I got wrong the first time. I tried sharing everything with everyone. The result was context saturation and signal degradation. The fix is that each agent gets its own view of the shared context—a projection filtered by role. Atlan's framework calls these "role-scoped context views": each agent gets only the context it needs for its specific job. Tencent's Team Memory does this with visibility tiers—Private, Team, Restricted—so that a Scout agent researching a market doesn't get the code graph a Builder agent needs.

Third, governed retrieval paths. Agents don't pull context from random documents or their own private memory. They pull from approved systems through governed paths. This is what prevents one agent's hallucination from spreading to others through shared state. A fact enters the shared layer only if it came from a governed source. A claim is only reusable if it was validated before it was written.

What the Research Backs Up

The shared-context pattern isn't just a hunch. It's a research-backed architecture that shows up across three distinct lines of work.

ContextDB is an open-source unified context layer that replaces the patchwork of vector databases, session stores, and glue code with a single memory operating system. Its multi-agent memory sharing protocol includes conflict resolution and role-aware routing—exactly the layer I was missing. The evaluation reports up to 90% token savings compared to full-context baselines while maintaining production-grade latency under 100ms p95.

SE-Blackboard applies the classical blackboard architecture—a shared workspace accessible to all agents—to software engineering pipelines. The paper's key metric is "knowledge drift": the cumulative loss or distortion of technical entities as task-relevant content is paraphrased through successive agent handoffs. Their empirical result: the blackboard architecture improves information fidelity by 62% and raises resolve rates by 4 percentage points compared to message-passing architectures.

The Context-Graph Shared Memory pattern takes a different approach: storing cross-agent state as typed entity-relationship triples rather than vector chunks. The benchmark data is nuanced—graph memory beats vector RAG on multi-hop join queries (80% vs 20% accuracy) but falls behind on general long-conversation recall by roughly 25 accuracy points at ten times the cost. The pattern's own guidance is refreshingly honest: benchmark all three approaches on the queries your system actually runs before defaulting to any one of them.

What Production Teams Are Building

Tencent's Team Memory is the most complete production implementation I've found. It distinguishes itself from RAG in a way that matters: "RAG answers 'what can be found?' Team Memory also answers 'who can use it, which version is valid, and which Agent should receive it.'" The system registers four kinds of reusable assets—Chat Memory, Skills, LLM-Wiki, and Code-Graph—and equips each agent with an "Agent Loadout" of only what it needs. The repo hit number one on GitHub's TypeScript trending list within days of launch.

CrewContext takes a database-first approach. PostgreSQL is the source of truth—an append-only event log with versioned entity snapshots and causal link tables. A Policy Router evaluates events against composable rules before they enter the shared state. Neo4j is an optional projection for graph queries and lineage visualization.

memX, open-sourced by a Microsoft AutoGen contributor, treats the shared memory layer as a semantic store rather than just a coordination bus. The architecture separates two concerns: a pub/sub layer (Redis) for real-time signals like task status and inter-agent pings, and a vector store (LanceDB) for accumulated knowledge with embedding-based retrieval. The result: a memory written by the research agent can be recalled by the marketing agent with a thematic query, with no direct messaging and no schema coupling.

Google's ADK builds context engineering into the framework itself. Sessions, memory, and artifacts are the sources; flows and processors are the compiler pipeline; the working context is the compiled view shipped to the LLM for a single invocation. The design principle is separating storage from presentation so that schemas and prompt formats can evolve independently.

Blackboard-Core implements the classical blackboard pattern as a Python SDK: a centralized typed state model, a supervisor LLM that decides which worker runs next, and workers that read from and write to the shared state rather than messaging each other directly.

The Trade-Off You're Accepting

A shared context layer isn't free. It adds a dependency that every agent's reasoning now flows through, and that dependency can become a bottleneck if you don't design it carefully.

The failure mode is what one team called "agreement manufactured by copying rather than earned by observation." Once an agent's context is shared across a whole team, a wrong fact doesn't cost one person a repeated explanation. It costs the whole team. One agent records a wrong value, others copy it into shared memory, and the system displays consensus that was never real.

The fix is governance. Every write to the shared layer needs provenance—which agent, which source, which timestamp. Every read needs a scope. Conflict resolution can't be an afterthought; it has to be a first-class operation on the shared layer. And someone has to own the schema.

You're also accepting more infrastructure. A shared context layer is another system to deploy, monitor, and scale. The payoff is that agents stop duplicating each other's work, stop contradicting each other, and stop burning tokens reconstructing state that already exists elsewhere.

Where This Fits in Your Architecture

Use a shared context layer when:

  • Your agents routinely need facts another agent established. If you're passing raw outputs between agents and hoping meaning survives the handoff, you have a shared-context problem.
  • You've seen contradictory outputs from agents that were each individually correct. That's the signature of fragmented context.
  • You're scaling past three or four agents. The coordination overhead of ad-hoc context sharing grows faster than the agent count.

Don't use it when:

  • Your agents are genuinely independent. If each agent operates on its own data and produces its own output with no shared reasoning, a shared context layer is unnecessary overhead.
  • Your sessions are short and single-purpose. A shared context layer earns its keep over sessions with many agents and many interactions. For a single-shot pipeline, it's premature.
  • You haven't defined what "correct" looks like. The shared layer doesn't fix bad reasoning. It surfaces conflicts. If you can't adjudicate a conflict, sharing context will just make the contradictions more visible—which is progress, but it's uncomfortable progress.

The Question I Keep Coming Back To

If two of your agents answered the same question right now, would they agree—not because one copied the other, but because they're both reading from the same version of the truth?

I'd love to hear where you've landed. A shared blackboard, a context graph, a governed memory hub, or a patchwork that mostly works and what finally made you stop treating it as a memory problem?

Top comments (0)