Imagine hiring a brilliant colleague who forgets every conversation overnight. Each morning you re-explain the project, your preferences, and what you decided yesterday. They are capable, but they never get better at working with you.
That is an AI agent without memory. The underlying language model starts every conversation from a blank page, so anything the agent should "know" has to be stored somewhere else and handed back to it.
In this guide, you'll learn:
- What agent memory is, and why a bigger context window doesn't replace it
- The kinds of memory an agent can have, each with an analogy and a diagram
- The two main ways memory gets updated, and what can go wrong
- How LangGraph, Claude, and CrewAI implement memory today
- How the pieces fit together in one architecture
What is agent memory?
Agent memory is a mechanism that lets an agent store, recall, and use information across multiple interactions. Before each model call, the system looks up relevant memories and places them into the model's context window (the text the model can see at one time).
The key point is that a large language model (LLM) cannot remember on its own. Memory is a component you add around it.
Think of the context window as your desk. Everything on the desk is easy to reach, but the desk is small and gets cleared at the end of the day. Memory is the filing cabinet: bigger, permanent, and only useful if someone fetches the right folder and puts it on the desk.
Why do agents need memory?
Consider a simple contrast. A basic thermostat reads the temperature and reacts. A smart thermostat with memory learns your usage patterns and anticipates what you'll want. Both are agents, but only one improves with experience.
Memory gives an agent four things:
- Context that carries across a conversation
- Personalization that lasts across sessions
- Pattern recognition over past events
- Adaptation based on what worked before
There is a real cost, though. Store too much and retrieval slows down and gets noisy. The hard part of memory design is keeping what is relevant and fast to fetch, not keeping everything.
The kinds of memory
You will see agent memory counted as three, four, or five kinds depending on who is counting. They all start from the same split: short-term versus long-term, with long-term divided into semantic, episodic, and procedural memory. The names come from human psychology and were brought to agents by a framework called CoALA. This guide uses that shared structure.
Short-term memory
Short-term memory (also called working memory) holds what the agent needs right now: the current conversation, recent tool results, and the state of the task. It is limited in size and cleared when the session ends.
Analogy: the sticky notes on your monitor while you work on one task. Useful today, thrown away tomorrow.
Example: a chat history that keeps the last few turns so the agent understands "do that again."
Long-term memory
Long-term memory survives across sessions. It lives outside the model, in a database, a set of files, or a vector store (a database that finds items by meaning, not exact words). It comes in three flavors.
Semantic memory
Semantic memory stores facts: a user's preferences, rules of a domain, what an entity is. It is not tied to a specific moment.
Analogy: your general knowledge, like knowing Paris is the capital of France without remembering when you learned it.
Example: "this customer prefers email follow-ups," stored as a fact the agent can look up.
Episodic memory
Episodic memory records specific past experiences, with when and what happened. It lets an agent reason from similar situations it has seen before.
Analogy: a diary. You can flip back and find the time something similar happened.
Example: a timestamped log of a past task, including what the agent tried and whether it worked. These past episodes are often used as few-shot examples (sample question-and-answer pairs placed in the prompt to show the model what good looks like).
Procedural memory
Procedural memory is knowing how to do things: skills, rules, and routines. The agent can act without reasoning from scratch each time.
Analogy: riding a bike. You don't recall instructions, you just do it.
Example: an instruction block or workflow template the agent follows. People define this slightly differently. Some treat it as the combination of the model's weights and the agent's code, and note that agents rarely rewrite their own system prompt automatically. Others treat it as updated system prompts or reusable workflow templates.
Ways of updating memory
Knowing the kinds of memory is half the story. The other half is deciding when and how memory gets written.
Hot path versus background
There are two timings:
- In the hot path: the agent decides to save something while it is answering, usually by calling a memory tool. Memory is available immediately, but each save adds delay and mixes memory logic into the agent's own logic. ChatGPT works this way.
- In the background: a separate process updates memory during or after the conversation. There is no added latency and memory logic stays separate, but updates arrive late and you need a rule for when to run them.
Management strategies
There are five common strategies, and you can combine them:
- Sliding window: keep only the most recent N messages
- Summarization: compress older history into a short summary
- Fact extraction: pull key facts into structured records
- Retrieval-augmented recall: store chunks externally and fetch only the relevant ones by similarity
- Hybrid: mix the above, often with background management
Letting the agent manage its own memory
Researchers have also explored letting the agent learn to manage its own memory. In an approach called AgeMem, the agent gets six operations as tools: ADD, UPDATE, and DELETE for long-term memory, and RETRIEVE, SUMMARY, and FILTER for short-term context. The authors train the agent with reinforcement learning and report improved results across five benchmarks. It is research, not a product you install, but it shows where the field is heading.
What can go wrong
These failure modes are worth testing for:
- The agent forgets to look up memory at the start
- Duplicate or conflicting entries get saved
- A stale memory gets retrieved and produces a wrong answer
- An important detail never gets saved at all
How common frameworks implement memory
LangGraph
LangGraph splits memory along the short-term and long-term line. Short-term memory is the graph's state, saved as checkpoints within a thread (one conversation) through a checkpointer. Long-term memory lives in a Store: JSON documents organized under namespaces, searchable by content and, with an embedding function configured, by meaning. LangGraph also frames long-term memory as semantic, episodic, and procedural, and provides two templates: memory-agent for the hot path and memory-service for the background.
Claude
Anthropic's memory tool lets Claude create, read, update, and delete files in a /memories directory. The important design choice is that it runs client-side: Claude asks for a file operation, and your application performs it, so you decide where the data lives. The commands are view, create, str_replace, insert, delete, and rename. When the tool is enabled, the API adds instructions telling Claude to check its memory directory first and to save progress as it goes, because its context could be reset at any time. It pairs with context editing and compaction for long-running work. Because paths come from the model, your handler must validate them to prevent directory traversal.
CrewAI
CrewAI now exposes a single Memory class that replaces the older separate short-term, long-term, entity, and external memory types. Memories are organized in hierarchical scopes such as /project/alpha. When saving, an LLM infers scope and importance and extracts discrete facts; a consolidation step compares new items with similar existing ones and decides whether to keep, update, or delete. Recall scores results by semantic similarity, recency, and importance. The default store is LanceDB. You turn it on with memory=True on a crew. (CrewAI's memory API has changed over time, so check its current docs.)
Putting it together
Here is one architecture that combines everything above. The agent loop reads short-term state from checkpoints, recalls from the three long-term stores through a retriever, and writes through a single memory writer that is fed by the hot path and by a background job.
What I built
I've been putting this into practice with a memory MCP server. MCP (Model Context Protocol) is a standard way for agents to call external tools, which makes it a natural home for shared memory. I built it following industry standards and deployed it with authentication.
Conclusion
Agent memory comes down to a few decisions, and it helps to group them by purpose:
- What the agent holds right now: short-term memory, kept small on purpose
- What it knows: semantic memory, the facts
- What happened: episodic memory, the diary
- How it acts: procedural memory, the instructions it follows
- When it writes: the hot path for immediacy, the background for speed
- Where it lives: a store you control, whichever framework you pick
Start small. Add one kind of memory, write it in the background, and test for the failure modes above before adding more. If you're building agents, try the Claude memory tool or a LangGraph Store on a toy project and watch what the agent chooses to remember.
If you want to talk shop on agent memory, you can find me on LinkedIn · GitHub · X.
References
- What is AI agent memory? (IBM)
- Memory Types in Agentic AI: A Breakdown (Gökçe Belgusen, Medium)
- Memory for agents (LangChain)
- Agentic memory (Patronus AI)
- Agentic Memory: Learning Unified Long-Term and Short-Term Memory Management for LLM Agents (arXiv:2601.01885)
- LangGraph memory docs
- Anthropic memory tool docs
- CrewAI memory docs










Top comments (1)
Your four failure modes match what we measured, with two key corrections from 4,087 server-audited calls (zero-shot, temp=0, v2 EN):
1. File fragments vs. bare tokens (The recall jump)
Bare token strings gave Qwen 3.6 / Qwen 3.7 a recall of 0.08 / 0.20 on 25 true facts. Presenting the exact same facts inside a ±12-line file fragment pushed recall to 0.88 / 0.88 (22/25 each), at FA 0.04/0.02, 0/16 absent, and 0/3 silent in this setup.
[0.70, 0.96]vs. baseline[0.02, 0.25]— anchor bias, not paranoia (report.md:357-364).2. Graph is a subject check, not an additive boost
Qwen hybrid accuracy (
0.900) was worse than file-only (0.940) — the fragment dominates, and graph context only helps when the file fragment is missing.2/20, DeepSeek15/20, GLM13/20Graph-first FA: Qwen
2/20, DeepSeek8/20, GLM14/20Graph context halves DeepSeek trap vulnerability (75% → 40%), does nothing for GLM, while Qwen pays with a recall penalty (2–3/10).
Two Methodological Lessons
miss_true: 4 out of 6 of ourv4_rep"false" traps were actually true (R43/R45/R46/R47— value was present AND used in the subject file; onlyR42was genuinely false). Realtrap-FAwas 0 on corrected labels, but this defect hidtrap miss_true— fail-closed models rejecting 4 out of 5 true usage claims.trap_facts_generator.pynow enforcesgrep(subject) == 0, backed by 5 unit tests.