The public issue trackers of the major agent CLIs are the best place to read what "my agent got dumber" actually looks like in the wild. One of the most concrete reports is titled, verbatim:
[FEATURE] suggest starting a new conversation when resuming a large/stale session
The request is small and specific. When you resume a session that has grown large or has gone stale, the CLI should suggest starting a new conversation, because resuming it drags back a pile of context that no longer describes the work in front of you.
Two sibling reports sit in the same class, and their titles are also verbatim:
- A bug report: "Subagents persist in memory after stopping without manual removal". Stopped subagents keep occupying memory until they are removed by hand.
- A bug report from another tracker: "memory_search hybrid ranking drops the only chunk that contains the whole query". A sentence quoted verbatim from an indexed note does not return that note.
None of these are model-quality problems. They are one problem wearing three hats: memory with no boundary and no plan. What gets kept, when it gets dropped and who decides are all left to whatever the host happens to do at the moment.
Ruling out the two reflexes
A better prompt does not fix it. A prompt is an input, not a store. It has no index, no source pointer and no address three sessions later. If the thing you needed was never written down somewhere addressable, no amount of instruction text brings it back.
A bigger context window does not fix it either. It moves the wall. Compaction still fires on its own schedule, and whatever was not durably stored at that moment is gone. Resuming a stale session is the visible edge of that: the context is technically still there, and it is precisely the part that no longer applies.
The common denominator is a missing layer between the transcript and the model: something that decides at write time what is worth keeping, keeps it verbatim, and drops the rest on a schedule instead of never.
Five write-time moves
These are the rules I ended up enforcing in HyperMarrow, a local-first memory layer for agents on Windows. They are ordered by how early they act, because every failure above is decided at write time, not at query time.
1. Decisions are written down, not re-derived. A preference, a decision or a conclusion has to land in durable local storage with a source and a timestamp. The stale-session request is a user asking, in effect, for the opposite: that the things already decided do not have to be re-established every time context is rebuilt.
2. Write before you compact. Nothing is allowed to summarise or drop the working context until that durable copy exists and has been acknowledged. Once the durable copy comes first, a summariser drops from single point of failure to convenience. The subagents-persist report is the same rule seen from the other side: memory that is never reclaimed is as bad as memory that is lost.
3. Forgetting is bounded and scheduled, and pinned records never decay. An unbounded memory is not a feature. Sessions that stopped being relevant should age out on a schedule instead of being carried forward forever, while records you explicitly pin are exempt. This is what "suggest a new conversation" asks for one level up: a deliberate boundary on what a session still carries.
4. Recall returns your words, not a summary of your words. If what got written was already lossy, no ranker recovers the sentence you actually needed. The hybrid-ranking report is a reminder that recall quality is capped by write quality, and ranking is the second-most important knob, not the first.
5. The privacy boundary is a setting, not a promise. Where records live, and whether a given record may leave the machine, is a configuration decision rather than a paragraph. Memory that holds decisions has to make that boundary explicit, because the alternative is a policy statement you cannot audit.
The modules are record, recall, consolidation and file-bridge, exposed over MCP, so the agent does not need a bespoke integration per tool.
Mapping the rules back to the reported symptoms
- Resuming a large or stale session. Retrieval is scoped and ranked, so resuming does not mean re-injecting everything. Old records decay on schedule; pinned ones do not.
- Subagents persisting after stopping. Reclamation is part of the write contract: a record with no live referent is collected, not left to accumulate.
- Quoted sentences losing to paraphrases. Short queries against long records still fail, which is why the fix starts at write time with tighter, verbatim records rather than at the ranker.
- "The agent forgot what we decided." The decision has an address, a source and a timestamp, so it can be re-found from a different session instead of being re-explained.
What is still wrong
The reports are fair criticism, so the limits are worth stating plainly:
- Hybrid ranking is genuinely hard. A quoted sentence does not always outrank a paraphrase, and short queries against long records still fail more often than I would like.
- Scheduled decay needs a usable notion of relevance. Get it wrong in the aggressive direction and you drop something that mattered; get it wrong in the patient direction and you are back to unbounded memory.
- Anything that depends on a model to classify memory kinds will misclassify some fraction of the time. It fails more quietly than losing the data, but it still fails.
Over to you
If you run agents across long-lived sessions: what is the thing you keep having to re-establish, and what would you want forgotten instead?
HyperMarrow is a Windows desktop memory layer. The local-first build is here: HyperMarrow
Disclosure: I build HyperMarrow, so read the above with that in mind. The issue reports quoted here are real and public, and no numbers, user counts or testimonials have been invented for this post.

Top comments (0)