Open a long-running agent session, walk away, come back. In at least one widely used coding agent, compaction can fire while the session is idle and drop context you still needed, with no way to opt out. The issue title says it better than I can: idle compaction silently discards working context in long-running sessions.
That report is not an outlier. Reading the public issue trackers of the major agent CLIs, the same handful of failures keeps showing up. They are not exotic edge cases. They are what happens when "memory" is really just a context window with a summarizer bolted on.
I work on HyperMarrow, a local-first memory layer for agents. I treat these reports as a backlog rather than as talking points. Here are four of them, quoted from the trackers, and the write-time rule each one implies.
Bug 1. "Idle compaction silently discards working context in long-running sessions; no opt-out"
Source: anthropics/claude-code issue 98747
What happens: compaction is a background job. It runs when the session is idle, not when it is safe. Anything the summarizer judges low-value at that moment is gone, and the user never got a vote. A mirror-image report in another repo describes subagent-completion turns skipping preflight compaction, so a session over its threshold never gets compacted at all.
The rule: a write decision has to happen before compaction, not after it.
In our design, compaction is not allowed to be the first thing that touches recent state. Records, decisions and conclusions are written to the local store first, each with a source tag and a timestamp. Only then may the working context be summarized or dropped. If the summarizer throws something away, the durable copy is already on disk. That turns compaction from a data-loss problem into a view problem.
Bug 2. "memory_search hybrid ranking drops the only chunk that contains the whole query"
Source: openclaw/openclaw issue 162764
What happens: retrieval quality is treated as a ranking problem. It is usually a writing problem. If what you stored was already a lossy summary, no ranker can recover the sentence you needed.
The rule: decide at write time what is a fact, what is a decision, and what is noise.
Our layer separates the three. Facts and decisions are stored verbatim and immutably. Noise is stored as a pointer, or not stored at all. Retrieval then has something exact to hit. When a recall misses, the first question is not "which embedding model", it is "what did we fail to write down".
Bug 3. "Subagents persist in memory after stopping without manual removal"
Source: anthropics/claude-code issue 98804
What happens: everything is remembered forever, so stale agents, dead sessions and one-off experiments accumulate and start to pollute later decisions.
The rule: forgetting has to be bounded and scheduled, not manual.
Our layer runs a decay pass. Low-value records fade along a curve, pinned records never fade, and nothing is hard-deleted while it is still referenced. The user should not have to be the garbage collector. A memory system that needs manual cleanup is not a memory system, it is a leak.
Bug 4. "Paragraph anchors with cross-session references"
Source: anthropics/claude-code issue 98768 (a feature request)
What happens: users want to point at one paragraph from an earlier session. Today they can point at a whole conversation, or at nothing.
The rule: continuity has to be addressable.
This one is filed as a feature, and that is the interesting part. The distance between "I remember your last chat" and "I can cite the exact paragraph from three sessions ago" is where continuity actually lives.
The rules, in one place
- Write before you compact. The durable copy exists before the summarizer runs.
- Separate facts, decisions and noise at write time, not at query time.
- Forgetting is scheduled and bounded; pinned records never decay.
- Continuity is addressable across sessions, down to a paragraph.
- All of it stays local by default, and the privacy boundary is a first-class setting rather than an afterthought.
None of this is exotic. It is the boring part of storage that most agent stacks skip, because a context window is good enough in a demo and the bill arrives later.
If you run agents
The useful question is not whether your agent has memory. It is where that memory lives, and what happens to it when the session ends. If the answer is "in the context window", these four bugs are already on your roadmap, whether or not you filed them.
I am building HyperMarrow, the local-first memory layer described above. Docs and the client are here: HyperMarrow.
If you maintain an agent runtime: which of these four have you hit, and which one do you consider unfixable?
(Disclosure: I build HyperMarrow, the local-first memory layer described above.)

Top comments (0)