DEV Community

Cover image for Session discovery: conversations become unreachable, plus ghost and duplicate project entries
qianqiuwanzi
qianqiuwanzi

Posted on

Session discovery: conversations become unreachable, plus ghost and duplicate project entries

The public issue trackers of the major agent CLIs are the best place to read what "my agent forgot" actually looks like in the wild. One of the most detailed reports is titled, verbatim:

cover

Session discovery: conversations become unreachable, plus ghost and duplicate project entries

The reporter works with roughly ten active conversations. The complaint is not that the model got weaker. It is that the conversations stopped being addressable: some of them cannot be opened at all, and the project list has grown ghost entries plus duplicates of the same project, so there is no obvious place to go back to.

Two sibling reports describe the same class of failure from different angles:

  • A feature request asks the CLI to suggest starting a new conversation when resuming a large or stale session, because resuming it drags in context that no longer describes the work.
  • A second bug report asks for conversation-scoped visibility, so an agent serving one channel can recall sibling threads in that channel without reaching into the other channels it serves.

None of these are recall-quality problems in the narrow sense. They are storage and addressing problems, and they sit underneath everything else.

Ruling out the two reflexes

Rewriting the prompt does not fix it. A prompt is an input, not a store. It has no index, no source pointer and no address three sessions later. If the thing you need was never written down somewhere addressable, no amount of system-prompt engineering brings it back.

A larger context window does not fix it either. It moves the wall. Compaction still fires on its own schedule, and whatever was not durably stored at that moment is gone. This is not hypothetical: another report in the same tracker describes an idle session being compacted before the prompt cache expired, with no opt-out, and logged as if it had been a manual action.

The common denominator is a missing layer between the transcript and the model: something that writes facts down, gives each one an address, and can find it again from a different session.

Four write-time rules

These are the rules I ended up enforcing in HyperMarrow, a local-first memory layer for agents on Windows. They are ordered by how early they act, because the failures above are decided at write time, not at query time.

1. Write before you compact. A preference, a decision or a conclusion has to land in durable local storage with a source and a timestamp before anything is allowed to summarise or drop the working context. The idle-compaction report argues for this rule by counterexample. Once the durable copy exists first, a summariser drops from single point of failure to convenience.

2. Decide at write time what kind of memory it is. A stated preference has to survive verbatim; a conclusion is durable but rewritable; routine chatter is noise. Sorting after the fact is too late, because the original phrasing has already been discarded. In practice that means three buckets with three retention policies, not one undifferentiated pool of vectors.

3. Recall returns your words, not a summary of your words. If what got stored was already lossy, the best ranker available cannot recover the sentence you actually needed. This is why rule 2 matters more than retrieval tuning: recall quality is capped by write quality.

4. Forgetting is scheduled and bounded, and pinned records never decay. An unbounded memory is not a feature. Sessions that stopped being relevant should age out on a schedule instead of being carried forward forever, while records you explicitly pin are exempt. The stale-session request above is a user asking for exactly this, one level up.

Two constraints run across the whole set: memory is local by default, and the privacy boundary is a setting rather than a promise. Where records live, and whether a given record may leave the machine, is a configuration decision, not a policy paragraph.

Mapping the rules back to the reported symptoms

  • Conversations become unreachable. Each session gets a stable local record with an address that does not depend on scrollback or on the CLI's own session list. Recall queries the store directly.
  • Ghost and duplicate project entries. The project entry is derived from the stored record instead of being accumulated separa官网 hm.qianshi.cooly, so the list cannot drift away from what actually exists.
  • Resuming a large or stale session. Retrieval is scoped and ranked, so resuming does not mean re-injecting everything. Old records decay on a schedule; pinned ones do not.
  • Conversation-scoped visibility. Scope is a property of the record, so recall can be limited to one channel's threads without exposing the rest.

The modules are record, recall, consolidation and file-bridge, exposed over MCP, so the agent does not need a bespoke integration per tool.

What is still wrong

The reports above are fair criticism, so the limits are worth stating plainly:

  • Hybrid ranking is genuinely hard. Short queries against long records still fail more often than I would like, and a quoted sentence does not always outrank a paraphrase of it.
  • Session addressing is only as good as the session identifiers the host gives us. If the CLI renames or recycles them, the memory layer has to reconcile, and reconciliation can be wrong.
  • Anything that depends on a model to classify memory kinds will misclassify some fraction of the time. It fails more quietly than losing the data, but it still fails.

Over to you

If you run agents across several sessions: what is the thing you keep having to re-establish? And if you have filed one of these reports, I would rather hear where this model still does not match what you saw.

HyperMarrow is a Windows desktop memory layer. The local-first build is here: HyperMarrow

Disclosure: I build HyperMarrow, so read the above with that in mind. The issue reports quoted here are real and public, and no numbers, user counts or testimonials have been invented for this post.

Top comments (0)