DEV Community

Cover image for Your AI forgets what you told it. A memory layer fixes most of it
qianqiuwanzi
qianqiuwanzi

Posted on

Your AI forgets what you told it. A memory layer fixes most of it

You 官网 hm.qianshi.cooll your agent something on Monday. On Tuesday it asks you again. Not because it did not listen. The transcript is still there. The problem is that nothing which survived the session is addressable.

cover

A recent post on Dev.to put a number on that feeling: your AI ignores what you told it most of the time. I do not have a better measurement and I am not going to invent one. What I can do is point at public issue trackers, where the same failure was filed by people who hit it in production.

One report: assistant text emitted before a tool call is dropped from the UI, even though the text is persisted in the transcript. Another: conversations become unreachable, plus ghost and duplicate project entries. A third is a request rather than a bug: when resuming a large and stale session, suggest starting a new conversation instead.

Read those three together and the shape is clear. The memory is often already written down. What is missing is a way back to it.

The fix people reach for first

Rewriting the prompt. "Remember that I prefer X." It works inside one session and evaporates with the context window, because a prompt is an input, not a store. You cannot index it, version it, or cite it three sessions later.

The second instinct is a bigger context window. That moves the wall, it does not remove it. Compaction still fires, and it fires on its own schedule rather than on yours.

Four things a memory layer has to do instead

I build HyperMarrow, a local-first memory layer for agents. All four of these are write-time decisions, not query-time tricks.

1. Write before you compact

Preferences, decisions and conclusions go to the local store first, each with a source and a timestamp. Only then is the working context allowed to be summarized or dropped. If the summarizer throws something away, the durable copy is already on disk.

2. Separate the kinds of memory at write time

Not every sentence deserves the same treatment. A stated preference, like "we use pnpm here", has to survive verbatim. A conclusion is durable. Raw chatter is noise. Sorting them at query time is too late, because by then the original phrasing is gone.

3. Recall returns your words, not a summary of your words

This is where the reported bugs bite hardest. If what got stored was already a lossy summary, no ranker can recover the sentence you needed. Keep the exact phrasing and recall can hand it back. That is the part which makes an agent stop re-asking.

4. The boundary is a setting, not a promise

On a local-first design the store lives on your machine, and sending context to a hosted service is an explicit action rather than the default. "You own your data" and "your data never leaves your machine" are different statements, and only the second one is checkable from your side.

What changes in practice

The agent stops starting from zero. It knows you said pnpm. It knows which approach you already ruled out. It can point at the paragraph where you said it. The difference is not that the model got smarter. The difference is that it can read something which outlived the session.

If you run agents

The useful question is not whether your agent has memory. It is where that memory lives, and whether it is still addressable tomorrow. If the answer is "somewhere in the context window", you are one compaction away from starting over.

I am building HyperMarrow, the local-first memory layer described above. Docs and the client are here: HyperMarrow.

If you run agents: what is the thing you have to repeat most often?

(Disclosure: I build HyperMarrow, the local-first memory layer described above.)

Top comments (0)