DEV Community

Cover image for AI agent memory vs RAG — what's the difference?
Statewave
Statewave

Posted on Originally published at statewave.ai

AI agent memory vs RAG — what's the difference?

Most teams building on LLMs end up with two patterns in the same codebase: RAG for looking things up in a corpus, and some hand-rolled memory for remembering what the agent has done or what the user has said. The two are often confused, and the confusion costs real engineering time when one is used in place of the other.

RAG retrieves content the agent doesn't already know. Memory retrieves context the agent has already participated in.

They share a vector store but they answer different questions, store different shapes of data, and have different correctness requirements.

The shared substrate

Both patterns sit on top of an embedding store — usually pgvector or a dedicated vector database — and both use approximate-nearest-neighbour search at retrieval time. That's where the overlap ends.

  • RAG stores chunks of source documents (markdown pages, PDFs, support articles). Retrieval returns the most semantically similar chunks to the user's question, which are spliced into the prompt as grounding.
  • Memory stores typed records of what happened: episodes (raw conversation turns, tool calls, decisions) and compiled memories (typed facts derived from episodes, with provenance back to the source). Retrieval returns the most relevant prior context for this subject, ranked by recency, kind, validity, and similarity.

If you only need to ground answers in static documentation, you need RAG. If you need the agent to remember the user across sessions, you need memory.

Where RAG breaks down

RAG was designed to answer "what does our documentation say about X?" — not "what did this user tell us last month?" When teams stretch RAG to cover memory, three failure modes show up:

  1. Embedding-nearest is not decision-relevant. A user message "I'm allergic to peanuts" and a later question "What should I order for lunch?" don't have close embeddings. Cosine similarity will surface every restaurant chunk before the allergy note. Memory ranking needs more than similarity — it needs kind priority, temporal validity, and explicit provenance.
  2. There's no compaction. RAG indexes the corpus as-is. Memory has to summarize — turning 200 conversation turns into "user is a senior engineer at a fintech, prefers terse responses" — because the prompt budget is finite. RAG doesn't do that.
  3. No invalidation model. If the user's job changes, the old "works at a fintech" record needs to be marked superseded, not retrieved alongside the new one. RAG's append-only chunk store has no concept of validity windows or supersession.

You can paper over each of these in your application code. Teams that do end up with a memory runtime in everything but name.

Where memory goes beyond retrieval

A memory runtime adds three things on top of the vector store:

  • Compilation — a pass over raw episodes that produces typed memories — profile facts, preferences, procedures, episode summaries — with confidence scores and validity windows. This is what shrinks 200 turns into one fact.
  • Deterministic ranking — scoring that mixes similarity with kind priority (a procedure beats a casual mention), recency, temporal validity, and an explicit token budget. Same query → same bundle. No silent re-ordering.
  • Provenance — every compiled memory carries the IDs of the episodes it was derived from. When an agent answers from memory, the answer is auditable back to the raw event that produced it.

None of those are properties of "RAG" in the literature sense. They're what makes memory infrastructure rather than retrieval over chat logs.

When to reach for which

RAG Memory
Question shape "What does our content say about X?" "What does this user / agent / project need to know right now?"
Data shape Document chunks Episodes → compiled memories with provenance
Ranking signal Cosine similarity Similarity + kind priority + recency + validity + token budget
Mutability Append-only chunks; reindex on doc update Episodes append-only; memories supersede; compaction is idempotent
Output Top-K chunks Token-bounded bundle ready to drop into a prompt

Most production agents need both. The grounding corpus (docs, knowledge base) lives in RAG. The user / account / project context lives in memory.

Trying to make either pattern do the other's job is the common architecture mistake — and it's the one we built Statewave to stop people from making.

What Statewave is in this picture

Statewave is the memory layer — episodes in, ranked context bundles out, with deterministic ranking and provenance. It uses pgvector under the hood (no separate vector DB to operate) but it's not a RAG framework: there's no document loader, no chunker, no retriever interface for grounding-over-corpora. If you want RAG, plug Statewave alongside your existing RAG stack — Statewave handles the who you're talking to layer, your RAG framework handles the what does the knowledge base say layer.

The architecture page on the docs site goes deeper on the ranking signals and the compile-vs-retrieve split. The getting-started guide is a five-minute Docker Compose path if you want to try it side-by-side with your current RAG setup.

FAQ

1. What's the difference between RAG and AI agent memory?

RAG retrieves content the agent doesn't already know — document chunks ranked by cosine similarity. Memory retrieves context the agent has already participated in — episodes and compiled facts ranked by recency, kind, validity, and similarity.

2. Can I use RAG in place of a memory layer for agent memory?

You can stretch it, but three failure modes show up: embedding-nearest isn't decision-relevant (an allergy note won't be the closest embedding to a lunch question), there's no compaction of history into durable facts, and there's no invalidation model for facts that get superseded.

3. Do I need both RAG and a memory runtime?

Most production agents do. The grounding corpus — docs, knowledge base — lives in RAG. The user, account, or project context lives in memory. Trying to make either pattern do the other's job is the common architecture mistake.

4. What does a memory runtime add that a vector store alone doesn't?

Three things: compilation (turning raw episodes into typed facts with confidence and validity), deterministic ranking (the same query always returns the same bundle), and provenance (every compiled memory carries the IDs of the episodes it came from).

5. Is Statewave a RAG framework?

No. It uses pgvector under the hood but ships no document loader, chunker, or retriever for grounding over a corpus. It's the who-you're-talking-to layer, meant to run alongside your existing RAG stack rather than replace it.

6. How do I decide which one to reach for?

Look at the shape of the question. "What does our content say about X?" is RAG. "What does this user, agent, or project need to know right now?" is memory.


Originally published on the Statewave blog. Statewave is an open-source, self-hosted memory runtime for AI agents — GitHub.

Top comments (10)

Collapse
 
raknaos profile image
Raknaos •

Clean split, and the three failure modes match what we stumbled into. The invalidation one bit us hardest: we kept both the old and the new fact in the store, and retrieval happily surfaced "user works at X" months after they moved to Y because cosine similarity does not care about dates.

We ended up adding a superseded_by pointer plus a recency boost at ranking time, which is a hand-rolled version of what you describe. How do you handle conflicts between an episode-derived fact and a compiled memory that contradict each other — does provenance win, or do you re-derive?

Collapse
 
statewave profile image
Statewave •

Thanks, and "user works at X" is exactly the case that pushed us toward validity windows. Cosine doesn't care about dates; a validity window does.

To your question: neither, strictly. Provenance doesn't decide anything; it's the audit trail. And we don't re-derive old memories, because compile only processes new episodes. The new episode compiles into its own fact, and conflict resolution runs at that point. If both facts carry the same registered single-valued key, the newer value supersedes the older one when their validity windows overlap. The old fact isn't deleted: it's marked superseded, its valid_to is set to the new one's valid_from, and it drops out of the read path. If the windows don't overlap, both stay, since "worked at X until 2025" and "works at Y" are both true. Facts without a key fall back to word overlap.

The honest gap: resolution only happens at compile time. Between compiles, the new raw episode can sit in the bundle as a recent interaction next to the still-active old fact. So compile cadence is the knob.

Collapse
 
hannune profile image
Tae Kim •

The compaction problem is the one that bit us hardest: we started by shoving conversation turns into a vector store and couldn't figure out why retrieval kept pulling up stale context that contradicted what the user had just told us. The "allergy to peanuts" example is a perfect illustration of why cosine similarity alone can't model salience: importance and semantic distance really are orthogonal. What we ended up building was a two-tier system where raw episodes get compacted into typed facts with a validity window, which is basically what you're describing as the compiled memory layer. Curious how you handle the compaction trigger, time-based or turn-count or something more semantic?

Collapse
 
statewave profile image
Statewave •

Thanks, Tae — "importance and semantic distance are orthogonal" is the sentence I'd put on the wall. It's why kind priority sits next to similarity in our ranking instead of under it.

On the trigger, honestly: none of the three. Compilation is an explicit call. The server doesn't decide when; most setups compile every few episodes or nightly, and there's an episode.created webhook if you'd rather drive it from ingest.

The reason that matters less than it sounds: the most recent raw episodes are ranked directly alongside compiled facts, with a boost for the active session, so what the user just said reaches the bundle before any compile runs.

The trade-off we haven't solved: supersession only happens at compile time. Between runs, the bundle can hold the old fact and the new turn that contradicts it, side by side. That's a milder version of the exact problem you describe, and the trigger frequency decides how long that window stays open.

Did a semantic trigger ever work for you (say, compiling when an incoming turn contradicts an existing fact), or did you end up on a schedule too?

Collapse
 
izgorodin profile image
Edward Izgorodin •

Kind priority next to similarity gets the peanut note into the race, but the note still has to win a race it arguably should not be running. It rises together with every other fact of its kind, its recency only falls as it ages, and if the session has been about restaurants, those recent turns come in with the active session boost and bid for the same token budget. Whether the note reaches the bundle then depends on how crowded the conversation is, which is an odd thing for a food allergy to depend on.

A preference for terse answers can lose to a better candidate and nothing breaks. An allergy cannot be allowed to lose, so it may not belong in the competition at all. It binds to the kind of action the agent is about to take, suggesting food, rather than to the wording of the question, and it enters the context whenever that action is in play, before the budget goes to anything else. Retrieved by similarity to the question, it will surface often enough to pass a demo, which is how a missing constraint stays unnoticed until the conversation where it mattered.

Collapse
 
statewave profile image
Statewave •

You're right, and the numbers back you up more than I'd like. Kind priority is the largest fixed term a fact gets, 10 against 3 for a raw episode. But the active-session boost, +6, only applies to episodes, never to compiled facts, and recency is min-max across the candidate pool, so the oldest candidate scores zero by construction. A long-standing allergy note is structurally the oldest thing in the room.

Run your restaurant session through that and the note sits around 10 plus whatever similarity it earns, while the last few turns sit at 3 plus recency plus 6. "Depends on how crowded the conversation is" is not a figure of speech here. It is arithmetic.

The distinction you draw is the one that matters: a preference can lose and nothing breaks, a constraint cannot. That makes it a different admission path, not a bigger weight. Any weight I pick still loses to enough competitors, and the one conversation where it loses is the one that counts.

The part I keep landing on is upstream of ranking, and it is the same wall as the failed-attempt thread: something has to mark the note as a constraint when it is written. Similarity cannot infer "this must never lose", and neither can a kind that also holds "prefers terse answers". Bind it to the action, admit it before the budget is divided - but only after someone has said, at write time, that it is that kind of statement.

And your last line is the one I'd put on the wall. Surfacing often enough to pass a demo is exactly how a missing constraint stays invisible until it matters.

Collapse
 
jo-do profile image
Jo Do •

The "different correctness requirements" line is the one teams underweight. RAG can be stale and still be right (the corpus didn't change); memory can be fresh and still be wrong (it records what happened, not what was true). That split drives the schema: memory entries need provenance and confidence; RAG chunks need source and version.

One more thing worth a paragraph: the shared vector store is a liability for memory. Retrieval that surfaces a user's old message because it embeds near the query is how private context leaks across sessions. The two patterns sharing an index is an accident waiting for a permission model.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.