DEV Community

Cover image for How AI Agents Remember: A Practical Guide to Agent Memory
shadrach temitayo
shadrach temitayo

Posted on Originally published at taytimes.substack.com AI-assisted

How AI Agents Remember: A Practical Guide to Agent Memory

Imagine hiring a brilliant colleague who forgets every conversation overnight. Each morning you re-explain the project, your preferences, and what you decided yesterday. They are capable, but they never get better at working with you.

That is an AI agent without memory. The underlying language model starts every conversation from a blank page, so anything the agent should "know" has to be stored somewhere else and handed back to it.

In this guide, you'll learn:

  • What agent memory is, and why a bigger context window doesn't replace it
  • The kinds of memory an agent can have, each with an analogy and a diagram
  • The two main ways memory gets updated, and what can go wrong
  • How LangGraph, Claude, and CrewAI implement memory today
  • How the pieces fit together in one architecture

What is agent memory?

Agent memory is a mechanism that lets an agent store, recall, and use information across multiple interactions. Before each model call, the system looks up relevant memories and places them into the model's context window (the text the model can see at one time).

The key point is that a large language model (LLM) cannot remember on its own. Memory is a component you add around it.

Think of the context window as your desk. Everything on the desk is easy to reach, but the desk is small and gets cleared at the end of the day. Memory is the filing cabinet: bigger, permanent, and only useful if someone fetches the right folder and puts it on the desk.

Diagram of the context window as a desk and the memory store as a filing cabinet. The agent recalls relevant items from the store into the context window and saves new items back.

Why do agents need memory?

Consider a simple contrast. A basic thermostat reads the temperature and reacts. A smart thermostat with memory learns your usage patterns and anticipates what you'll want. Both are agents, but only one improves with experience.

Memory gives an agent four things:

  • Context that carries across a conversation
  • Personalization that lasts across sessions
  • Pattern recognition over past events
  • Adaptation based on what worked before

There is a real cost, though. Store too much and retrieval slows down and gets noisy. The hard part of memory design is keeping what is relevant and fast to fetch, not keeping everything.

The kinds of memory

You will see agent memory counted as three, four, or five kinds depending on who is counting. They all start from the same split: short-term versus long-term, with long-term divided into semantic, episodic, and procedural memory. The names come from human psychology and were brought to agents by a framework called CoALA. This guide uses that shared structure.

Table of the kinds of agent memory: short-term, semantic, episodic, and procedural, with what each holds, an analogy, and an example.

Short-term memory

Short-term memory (also called working memory) holds what the agent needs right now: the current conversation, recent tool results, and the state of the task. It is limited in size and cleared when the session ends.

Analogy: the sticky notes on your monitor while you work on one task. Useful today, thrown away tomorrow.

Example: a chat history that keeps the last few turns so the agent understands "do that again."

Diagram of short-term memory as a rolling window. Turns 1 to 3 enter the window, the LLM sees only the window, and the oldest turns drop off and are forgotten.

Long-term memory

Long-term memory survives across sessions. It lives outside the model, in a database, a set of files, or a vector store (a database that finds items by meaning, not exact words). It comes in three flavors.

Semantic memory

Semantic memory stores facts: a user's preferences, rules of a domain, what an entity is. It is not tied to a specific moment.

Analogy: your general knowledge, like knowing Paris is the capital of France without remembering when you learned it.

Example: "this customer prefers email follow-ups," stored as a fact the agent can look up.

Diagram of semantic memory. Facts are extracted from a conversation into a semantic store, retrieved by meaning, and added to the prompt.

Episodic memory

Episodic memory records specific past experiences, with when and what happened. It lets an agent reason from similar situations it has seen before.

Analogy: a diary. You can flip back and find the time something similar happened.

Example: a timestamped log of a past task, including what the agent tried and whether it worked. These past episodes are often used as few-shot examples (sample question-and-answer pairs placed in the prompt to show the model what good looks like).

Diagram of episodic memory. Episodes of task, actions, and outcome go into an event log. A new similar task finds similar episodes, which become few-shot examples in the prompt.

Procedural memory

Procedural memory is knowing how to do things: skills, rules, and routines. The agent can act without reasoning from scratch each time.

Analogy: riding a bike. You don't recall instructions, you just do it.

Example: an instruction block or workflow template the agent follows. People define this slightly differently. Some treat it as the combination of the model's weights and the agent's code, and note that agents rarely rewrite their own system prompt automatically. Others treat it as updated system prompts or reusable workflow templates.

Diagram of procedural memory. Experience and feedback, rarely and carefully, update the agent's instructions or workflow, which shape agent behavior.

Ways of updating memory

Knowing the kinds of memory is half the story. The other half is deciding when and how memory gets written.

Hot path versus background

There are two timings:

  1. In the hot path: the agent decides to save something while it is answering, usually by calling a memory tool. Memory is available immediately, but each save adds delay and mixes memory logic into the agent's own logic. ChatGPT works this way.
  2. In the background: a separate process updates memory during or after the conversation. There is no added latency and memory logic stays separate, but updates arrive late and you need a rule for when to run them.

Two lanes compared. In the hot path the agent saves memory while answering. In the background a separate process updates memory after the conversation.

Management strategies

There are five common strategies, and you can combine them:

  • Sliding window: keep only the most recent N messages
  • Summarization: compress older history into a short summary
  • Fact extraction: pull key facts into structured records
  • Retrieval-augmented recall: store chunks externally and fetch only the relevant ones by similarity
  • Hybrid: mix the above, often with background management

Letting the agent manage its own memory

Researchers have also explored letting the agent learn to manage its own memory. In an approach called AgeMem, the agent gets six operations as tools: ADD, UPDATE, and DELETE for long-term memory, and RETRIEVE, SUMMARY, and FILTER for short-term context. The authors train the agent with reinforcement learning and report improved results across five benchmarks. It is research, not a product you install, but it shows where the field is heading.

What can go wrong

These failure modes are worth testing for:

  • The agent forgets to look up memory at the start
  • Duplicate or conflicting entries get saved
  • A stale memory gets retrieved and produces a wrong answer
  • An important detail never gets saved at all

How common frameworks implement memory

LangGraph

LangGraph splits memory along the short-term and long-term line. Short-term memory is the graph's state, saved as checkpoints within a thread (one conversation) through a checkpointer. Long-term memory lives in a Store: JSON documents organized under namespaces, searchable by content and, with an embedding function configured, by meaning. LangGraph also frames long-term memory as semantic, episodic, and procedural, and provides two templates: memory-agent for the hot path and memory-service for the background.

Claude

Anthropic's memory tool lets Claude create, read, update, and delete files in a /memories directory. The important design choice is that it runs client-side: Claude asks for a file operation, and your application performs it, so you decide where the data lives. The commands are view, create, str_replace, insert, delete, and rename. When the tool is enabled, the API adds instructions telling Claude to check its memory directory first and to save progress as it goes, because its context could be reset at any time. It pairs with context editing and compaction for long-running work. Because paths come from the model, your handler must validate them to prevent directory traversal.

CrewAI

CrewAI now exposes a single Memory class that replaces the older separate short-term, long-term, entity, and external memory types. Memories are organized in hierarchical scopes such as /project/alpha. When saving, an LLM infers scope and importance and extracts discrete facts; a consolidation step compares new items with similar existing ones and decides whether to keep, update, or delete. Recall scores results by semantic similarity, recency, and importance. The default store is LanceDB. You turn it on with memory=True on a crew. (CrewAI's memory API has changed over time, so check its current docs.)

Comparison table of LangGraph, the Claude memory tool, and CrewAI across short-term memory, long-term memory, who writes, and where it lives.

Putting it together

Here is one architecture that combines everything above. The agent loop reads short-term state from checkpoints, recalls from the three long-term stores through a retriever, and writes through a single memory writer that is fed by the hot path and by a background job.

Architecture diagram. A user talks to the agent loop, which reads short-term thread state, recalls from semantic, episodic, and procedural stores through a retriever, and writes through a memory writer fed by the hot path and a background extract-and-consolidate job.

What I built

I've been putting this into practice with a memory MCP server. MCP (Model Context Protocol) is a standard way for agents to call external tools, which makes it a natural home for shared memory. I built it following industry standards and deployed it with authentication.

Conclusion

Agent memory comes down to a few decisions, and it helps to group them by purpose:

  • What the agent holds right now: short-term memory, kept small on purpose
  • What it knows: semantic memory, the facts
  • What happened: episodic memory, the diary
  • How it acts: procedural memory, the instructions it follows
  • When it writes: the hot path for immediacy, the background for speed
  • Where it lives: a store you control, whichever framework you pick

Start small. Add one kind of memory, write it in the background, and test for the failure modes above before adding more. If you're building agents, try the Claude memory tool or a LangGraph Store on a toy project and watch what the agent chooses to remember.

Summary card: four kinds of memory to hold (short-term, semantic, episodic, procedural) and two ways to write (hot path and background).

If you want to talk shop on agent memory, you can find me on LinkedIn · GitHub · X.

References

Top comments (1)

Collapse
 
mansio profile image
Mikhail •

Your four failure modes match what we measured, with two key corrections from 4,087 server-audited calls (zero-shot, temp=0, v2 EN):

1. File fragments vs. bare tokens (The recall jump)

Bare token strings gave Qwen 3.6 / Qwen 3.7 a recall of 0.08 / 0.20 on 25 true facts. Presenting the exact same facts inside a ±12-line file fragment pushed recall to 0.88 / 0.88 (22/25 each), at FA 0.04/0.02, 0/16 absent, and 0/3 silent in this setup.

  • Wilson CI: [0.70, 0.96] vs. baseline [0.02, 0.25] — anchor bias, not paranoia (report.md:357-364).

2. Graph is a subject check, not an additive boost

Qwen hybrid accuracy (0.900) was worse than file-only (0.940) — the fragment dominates, and graph context only helps when the file fragment is missing.

  • On subject-validated 20 false / 10 true traps:
  • File content FA: Qwen 2/20, DeepSeek 15/20, GLM 13/20
  • Graph-first FA: Qwen 2/20, DeepSeek 8/20, GLM 14/20

  • Graph context halves DeepSeek trap vulnerability (75% → 40%), does nothing for GLM, while Qwen pays with a recall penalty (2–3/10).

Two Methodological Lessons

  • Label noise hides miss_true: 4 out of 6 of our v4_rep "false" traps were actually true (R43/R45/R46/R47 — value was present AND used in the subject file; only R42 was genuinely false). Real trap-FA was 0 on corrected labels, but this defect hid trap miss_true — fail-closed models rejecting 4 out of 5 true usage claims.
  • N=1 trap metrics are noise: On honest 20/10 labels, trap FA spans 10% to 75% by model, proving it is a primary behavior rather than a residual anomaly. trap_facts_generator.py now enforces grep(subject) == 0, backed by 5 unit tests.