DEV Community

Abdeljabbar Elassali
Abdeljabbar Elassali

Posted on

How Do AI Agents Store Long-Term Memory? The 4 Layers Between Runs

How Do AI Agents Store Long-Term Memory? The 4 Layers Between Runs

Here is the uncomfortable fact underneath every "give your agent memory" tutorial: the language model stores nothing. An LLM is stateless during generation. Every message you send is, technically, a brand new independent call, as if it were a new session, because the model itself keeps no state from one call to the next. That is why an agent can hold a brilliant conversation on Monday and, on Tuesday's scheduled run, have no idea who it talked to.

So when people say "the agent remembers," where does the remembering actually happen? It is not one thing. It is four storage layers, each holding a different kind of information, each with its own cost and failure mode. If you run scheduled agents in n8n, Make, Zapier, or a cron script, knowing which layer holds what is the difference between an agent that gets smarter every week and one that re-learns the same lesson every morning.

Layer 1: The transcript stack (working memory)

The simplest layer is also the most fragile. A conversation is faked as a growing stack: every exchange gets appended, and the entire stack is re-sent to the model on every single call. That is how "memory" works inside one session — the model re-reads everything, every time, because nothing from the previous call survives on its own.

For a scheduled agent, this layer is gone between runs. The context window closes, the stack dies, and the next run starts with an empty window. Everything the agent "knew" at 9 AM is ashes by the 9 PM run unless something outside the call saved it.

Layer 2: The rolling summary

When the transcript gets too long, the standard fix is a summary buffer: keep the last N turns verbatim, compress everything older into a rolling summary, and prepend it on every call. The BeeAI open-source docs sketch this pattern plainly — a summarizer that takes the existing summary plus new transcript turns and produces an updated one, over and over.

Summaries are cheap and fast, and they are lossy in exactly the way that bites scheduled agents. "The Shopify API returned 429 rate-limit errors on Monday's run; the 9:40 retry succeeded" becomes "there were some API issues earlier in the week." The incident survives; the detail that would prevent the next incident does not.

Layer 3: The fact store (long-term memory)

The third layer lives outside the conversation entirely: a persistent store of durable facts — preferences, decisions, client rules, things that should still be true tomorrow. The BeeAI docs show the minimal version as a JSON fact store on disk: save a key-value pair like a timezone or a preference, then inject all the facts into the system prompt at session start. For larger collections, the same idea scales into a vector database: embed each memory, retrieve the most relevant ones per query, exactly like RAG.

This is the layer most "give your agent memory" tutorials actually build. It works — until the agent needs to remember what happened, not what is. A fact store knows "refund requests over $500 route to billing." It cannot tell you "this refund was already issued on Monday's run." Facts without episodes produce an agent that knows the rules but has no past.

Layer 4: The working-memory scratchpad

The fourth layer is the one frameworks are still naming. Mastra's docs call it "working memory": a persistent scratchpad of structured information about users or tasks that survives across sessions — distinct from message history and from semantic recall.

For a scheduled agent, the scratchpad is where task state lives: which items were already processed, where the last run stopped, what the retry count is. This is not knowledge. It is the agent's to-do list, and without it every run re-processes the same queue from the top.

The loop that connects them: read at start, write at end

Here is the part most DIY setups get wrong. The four layers do not connect themselves. A scheduled agent needs an explicit loop: read at the start of the run (inject the facts, load the scratchpad, pull the summary of the last run), write at the end of the run (update the stores with what just happened, or the next run starts blind again), and set an expiry rule (a memory from eighteen months ago about a customer's preference might be actively wrong now).

Skip step 2 once and the whole system silently degrades. The agent does not error. It just starts every run from scratch, and because language models are confident by default, nobody notices for weeks.

The token bill per layer

Each layer has a per-run cost. The transcript stack costs everything it contains, every call. The summary costs a fixed, bounded amount — that is the point of summarizing. The fact store costs retrieval plus injection: small if you pull only relevant facts, large if you dump the whole store into the prompt. The scratchpad costs almost nothing.

The common failure is layer-3 bloat: the fact store grows, nobody prunes it, and every run injects thousands of tokens of stale facts into the prompt. Memory that is never read is not memory. It is an expensive decoration.

The lazy version: one memory layer that keeps all four

Building this yourself means a transcript store, a summarizer, a fact store with retrieval, a scratchpad, a write-back loop, and expiry rules — all hosted and maintained. That is a real engineering project, and for most automation operators it is not the project they wanted.

This is the gap Vilix AI is built for. It is a cloud-hosted memory layer, so there is no database to run and nothing to maintain. The same memory follows your agent across every tool over MCP: the n8n workflow, the cron script, Claude Code, the phone app, all reading and writing one shared store. It keeps full conversation history, not just extracted facts or lossy summaries, so nothing a future run might need is thrown away.

The free plan is free forever, the Pro trial is 7 days with no credit card, and your data is portable. Export everything or delete it anytime, in a portable format.

FAQ

Does the LLM itself store anything between runs?
No. A language model is stateless during generation: every call is independent, and nothing the model "learned" in one call exists in the next. All persistence is something you build around it.

Is long-term agent memory the same as RAG?
They share a retrieval mechanism — semantic recall over embeddings is RAG applied to past messages — but RAG answers "what does the documentation say" while agent memory answers "what did we do and learn."

How often should a scheduled agent write back to memory?
At the end of every run, unconditionally. Write the summary, update the facts, commit the scratchpad. A run that finishes without writing is a run the next one cannot learn from.

Four storage layers, and most setups are missing at least one connection between them. Close the loop — read at start, write at end — and the 7 AM run stops re-learning what the midnight run already knew.

Top comments (0)