<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: qianqiuwanzi</title>
    <description>The latest articles on DEV Community by qianqiuwanzi (@qianqiuwanzi).</description>
    <link>https://dev.to/qianqiuwanzi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4127114%2Fa8e08787-440e-4370-93e8-c8f9c64ad35c.jpg</url>
      <title>DEV Community: qianqiuwanzi</title>
      <link>https://dev.to/qianqiuwanzi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/qianqiuwanzi"/>
    <language>en</language>
    <item>
      <title>[FEATURE] suggest starting a new conversation when resuming a large/stale session</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Wed, 07 Oct 2026 01:53:49 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/feature-suggest-starting-a-new-conversation-when-resuming-a-largestale-session-48mm</link>
      <guid>https://dev.to/qianqiuwanzi/feature-suggest-starting-a-new-conversation-when-resuming-a-largestale-session-48mm</guid>
      <description>&lt;p&gt;The public issue trackers of the major agent CLIs are the best place to read what "my agent got dumber" actually looks like in the wild. One of the most concrete reports is titled, verbatim:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqrkxom5785ibkoshxsx.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhqrkxom5785ibkoshxsx.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;[FEATURE] suggest starting a new conversation when resuming a large/stale session&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The request is small and specific. When you resume a session that has grown large or has gone stale, the CLI should suggest starting a new conversation, because resuming it drags back a pile of context that no longer describes the work in front of you.&lt;/p&gt;

&lt;p&gt;Two sibling reports sit in the same class, and their titles are also verbatim:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A bug report: "Subagents persist in memory after stopping without manual removal". Stopped subagents keep occupying memory until they are removed by hand.&lt;/li&gt;
&lt;li&gt;A bug report from another tracker: "memory_search hybrid ranking drops the only chunk that contains the whole query". A sentence quoted verbatim from an indexed note does not return that note.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are model-quality problems. They are one problem wearing three hats: memory with no boundary and no plan. What gets kept, when it gets dropped and who decides are all left to whatever the host happens to do at the moment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ruling out the two reflexes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A better prompt does not fix it.&lt;/strong&gt; A prompt is an input, not a store. It has no index, no source pointer and no address three sessions later. If the thing you needed was never written down somewhere addressable, no amount of instruction text brings it back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A bigger context window does not fix it either.&lt;/strong&gt; It moves the wall. Compaction still fires on its own schedule, and whatever was not durably stored at that moment is gone. Resuming a stale session is the visible edge of that: the context is technically still there, and it is precisely the part that no longer applies.&lt;/p&gt;

&lt;p&gt;The common denominator is a missing layer between the transcript and the model: something that decides at write time what is worth keeping, keeps it verbatim, and drops the rest on a schedule instead of never.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five write-time moves
&lt;/h2&gt;

&lt;p&gt;These are the rules I ended up enforcing in HyperMarrow, a local-first memory layer for agents on Windows. They are ordered by how early they act, because every failure above is decided at write time, not at query time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Decisions are written down, not re-derived.&lt;/strong&gt; A preference, a decision or a conclusion has to land in durable local storage with a source and a timestamp. The stale-session request is a user asking, in effect, for the opposite: that the things already decided do not have to be re-established every time context is rebuilt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Write before you compact.&lt;/strong&gt; Nothing is allowed to summarise or drop the working context until that durable copy exists and has been acknowledged. Once the durable copy comes first, a summariser drops from single point of failure to convenience. The subagents-persist report is the same rule seen from the other side: memory that is never reclaimed is as bad as memory that is lost.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Forgetting is bounded and scheduled, and pinned records never decay.&lt;/strong&gt; An unbounded memory is not a feature. Sessions that stopped being relevant should age out on a schedule instead of being carried forward forever, while records you explicitly pin are exempt. This is what "suggest a new conversation" asks for one level up: a deliberate boundary on what a session still carries.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Recall returns your words, not a summary of your words.&lt;/strong&gt; If what got written was already lossy, no ranker recovers the sentence you actually needed. The hybrid-ranking report is a reminder that recall quality is capped by write quality, and ranking is the second-most important knob, not the first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. The privacy boundary is a setting, not a promise.&lt;/strong&gt; Where records live, and whether a given record may leave the machine, is a configuration decision rather than a paragraph. Memory that holds decisions has to make that boundary explicit, because the alternative is a policy statement you cannot audit.&lt;/p&gt;

&lt;p&gt;The modules are record, recall, consolidation and file-bridge, exposed over MCP, so the agent does not need a bespoke integration per tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping the rules back to the reported symptoms
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Resuming a large or stale session.&lt;/strong&gt; Retrieval is scoped and ranked, so resuming does not mean re-injecting everything. Old records decay on schedule; pinned ones do not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Subagents persisting after stopping.&lt;/strong&gt; Reclamation is part of the write contract: a record with no live referent is collected, not left to accumulate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quoted sentences losing to paraphrases.&lt;/strong&gt; Short queries against long records still fail, which is why the fix starts at write time with tighter, verbatim records rather than at the ranker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;"The agent forgot what we decided."&lt;/strong&gt; The decision has an address, a source and a timestamp, so it can be re-found from a different session instead of being re-explained.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What is still wrong
&lt;/h2&gt;

&lt;p&gt;The reports are fair criticism, so the limits are worth stating plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Hybrid ranking is genuinely hard. A quoted sentence does not always outrank a paraphrase, and short queries against long records still fail more often than I would like.&lt;/li&gt;
&lt;li&gt;Scheduled decay needs a usable notion of relevance. Get it wrong in the aggressive direction and you drop something that mattered; get it wrong in the patient direction and you are back to unbounded memory.&lt;/li&gt;
&lt;li&gt;Anything that depends on a model to classify memory kinds will misclassify some fraction of the time. It fails more quietly than losing the data, but it still fails.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;If you run agents across long-lived sessions: what is the thing you keep having to re-establish, and what would you want forgotten instead?&lt;/p&gt;

&lt;p&gt;HyperMarrow is a Windows desktop memory layer. The local-first build is here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;HyperMarrow&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Disclosure: I build HyperMarrow, so read the above with that in mind. The issue reports quoted here are real and public, and no numbers, user counts or testimonials have been invented for this post.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>memory</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Session discovery: conversations become unreachable, plus ghost and duplicate project entries</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Tue, 06 Oct 2026 03:30:11 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/session-discovery-conversations-become-unreachable-plus-ghost-and-duplicate-project-entries-29j1</link>
      <guid>https://dev.to/qianqiuwanzi/session-discovery-conversations-become-unreachable-plus-ghost-and-duplicate-project-entries-29j1</guid>
      <description>&lt;p&gt;The public issue trackers of the major agent CLIs are the best place to read what "my agent forgot" actually looks like in the wild. One of the most detailed reports is titled, verbatim:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkdzl5sq15fhpl7rhmyjw.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkdzl5sq15fhpl7rhmyjw.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Session discovery: conversations become unreachable, plus ghost and duplicate project entries&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The reporter works with roughly ten active conversations. The complaint is not that the model got weaker. It is that the conversations stopped being addressable: some of them cannot be opened at all, and the project list has grown ghost entries plus duplicates of the same project, so there is no obvious place to go back to.&lt;/p&gt;

&lt;p&gt;Two sibling reports describe the same class of failure from different angles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A feature request asks the CLI to suggest starting a new conversation when resuming a large or stale session, because resuming it drags in context that no longer describes the work.&lt;/li&gt;
&lt;li&gt;A second bug report asks for conversation-scoped visibility, so an agent serving one channel can recall sibling threads in that channel without reaching into the other channels it serves.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these are recall-quality problems in the narrow sense. They are storage and addressing problems, and they sit underneath everything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ruling out the two reflexes
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Rewriting the prompt does not fix it.&lt;/strong&gt; A prompt is an input, not a store. It has no index, no source pointer and no address three sessions later. If the thing you need was never written down somewhere addressable, no amount of system-prompt engineering brings it back.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A larger context window does not fix it either.&lt;/strong&gt; It moves the wall. Compaction still fires on its own schedule, and whatever was not durably stored at that moment is gone. This is not hypothetical: another report in the same tracker describes an idle session being compacted before the prompt cache expired, with no opt-out, and logged as if it had been a manual action.&lt;/p&gt;

&lt;p&gt;The common denominator is a missing layer between the transcript and the model: something that writes facts down, gives each one an address, and can find it again from a different session.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four write-time rules
&lt;/h2&gt;

&lt;p&gt;These are the rules I ended up enforcing in HyperMarrow, a local-first memory layer for agents on Windows. They are ordered by how early they act, because the failures above are decided at write time, not at query time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Write before you compact.&lt;/strong&gt; A preference, a decision or a conclusion has to land in durable local storage with a source and a timestamp before anything is allowed to summarise or drop the working context. The idle-compaction report argues for this rule by counterexample. Once the durable copy exists first, a summariser drops from single point of failure to convenience.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Decide at write time what kind of memory it is.&lt;/strong&gt; A stated preference has to survive verbatim; a conclusion is durable but rewritable; routine chatter is noise. Sorting after the fact is too late, because the original phrasing has already been discarded. In practice that means three buckets with three retention policies, not one undifferentiated pool of vectors.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Recall returns your words, not a summary of your words.&lt;/strong&gt; If what got stored was already lossy, the best ranker available cannot recover the sentence you actually needed. This is why rule 2 matters more than retrieval tuning: recall quality is capped by write quality.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Forgetting is scheduled and bounded, and pinned records never decay.&lt;/strong&gt; An unbounded memory is not a feature. Sessions that stopped being relevant should age out on a schedule instead of being carried forward forever, while records you explicitly pin are exempt. The stale-session request above is a user asking for exactly this, one level up.&lt;/p&gt;

&lt;p&gt;Two constraints run across the whole set: memory is local by default, and the privacy boundary is a setting rather than a promise. Where records live, and whether a given record may leave the machine, is a configuration decision, not a policy paragraph.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping the rules back to the reported symptoms
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Conversations become unreachable.&lt;/strong&gt; Each session gets a stable local record with an address that does not depend on scrollback or on the CLI's own session list. Recall queries the store directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ghost and duplicate project entries.&lt;/strong&gt; The project entry is derived from the stored record instead of being accumulated separa官网 hm.qianshi.cooly, so the list cannot drift away from what actually exists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resuming a large or stale session.&lt;/strong&gt; Retrieval is scoped and ranked, so resuming does not mean re-injecting everything. Old records decay on a schedule; pinned ones do not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversation-scoped visibility.&lt;/strong&gt; Scope is a property of the record, so recall can be limited to one channel's threads without exposing the rest.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The modules are record, recall, consolidation and file-bridge, exposed over MCP, so the agent does not need a bespoke integration per tool.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is still wrong
&lt;/h2&gt;

&lt;p&gt;The reports above are fair criticism, so the limits are worth stating plainly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Hybrid ranking is genuinely hard. Short queries against long records still fail more often than I would like, and a quoted sentence does not always outrank a paraphrase of it.&lt;/li&gt;
&lt;li&gt;Session addressing is only as good as the session identifiers the host gives us. If the CLI renames or recycles them, the memory layer has to reconcile, and reconciliation can be wrong.&lt;/li&gt;
&lt;li&gt;Anything that depends on a model to classify memory kinds will misclassify some fraction of the time. It fails more quietly than losing the data, but it still fails.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Over to you
&lt;/h2&gt;

&lt;p&gt;If you run agents across several sessions: what is the thing you keep having to re-establish? And if you have filed one of these reports, I would rather hear where this model still does not match what you saw.&lt;/p&gt;

&lt;p&gt;HyperMarrow is a Windows desktop memory layer. The local-first build is here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;HyperMarrow&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Disclosure: I build HyperMarrow, so read the above with that in mind. The issue reports quoted here are real and public, and no numbers, user counts or testimonials have been invented for this post.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>memory</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Your AI forgets what you told it. A memory layer fixes most of it</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Sat, 03 Oct 2026 01:53:51 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/your-ai-forgets-what-you-told-it-a-memory-layer-fixes-most-of-it-1djg</link>
      <guid>https://dev.to/qianqiuwanzi/your-ai-forgets-what-you-told-it-a-memory-layer-fixes-most-of-it-1djg</guid>
      <description>&lt;p&gt;You 官网 hm.qianshi.cooll your agent something on Monday. On Tuesday it asks you again. Not because it did not listen. The transcript is still there. The problem is that nothing which survived the session is addressable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9r3m3fuky3zzu7rkbmu.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx9r3m3fuky3zzu7rkbmu.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A recent post on Dev.to put a number on that feeling: your AI ignores what you told it most of the time. I do not have a better measurement and I am not going to invent one. What I can do is point at public issue trackers, where the same failure was filed by people who hit it in production.&lt;/p&gt;

&lt;p&gt;One report: assistant text emitted before a tool call is dropped from the UI, even though the text is persisted in the transcript. Another: conversations become unreachable, plus ghost and duplicate project entries. A third is a request rather than a bug: when resuming a large and stale session, suggest starting a new conversation instead.&lt;/p&gt;

&lt;p&gt;Read those three together and the shape is clear. The memory is often already written down. What is missing is a way back to it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix people reach for first
&lt;/h2&gt;

&lt;p&gt;Rewriting the prompt. "Remember that I prefer X." It works inside one session and evaporates with the context window, because a prompt is an input, not a store. You cannot index it, version it, or cite it three sessions later.&lt;/p&gt;

&lt;p&gt;The second instinct is a bigger context window. That moves the wall, it does not remove it. Compaction still fires, and it fires on its own schedule rather than on yours.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things a memory layer has to do instead
&lt;/h2&gt;

&lt;p&gt;I build HyperMarrow, a local-first memory layer for agents. All four of these are write-time decisions, not query-time tricks.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Write before you compact
&lt;/h3&gt;

&lt;p&gt;Preferences, decisions and conclusions go to the local store first, each with a source and a timestamp. Only then is the working context allowed to be summarized or dropped. If the summarizer throws something away, the durable copy is already on disk.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Separate the kinds of memory at write time
&lt;/h3&gt;

&lt;p&gt;Not every sentence deserves the same treatment. A stated preference, like "we use pnpm here", has to survive verbatim. A conclusion is durable. Raw chatter is noise. Sorting them at query time is too late, because by then the original phrasing is gone.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Recall returns your words, not a summary of your words
&lt;/h3&gt;

&lt;p&gt;This is where the reported bugs bite hardest. If what got stored was already a lossy summary, no ranker can recover the sentence you needed. Keep the exact phrasing and recall can hand it back. That is the part which makes an agent stop re-asking.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. The boundary is a setting, not a promise
&lt;/h3&gt;

&lt;p&gt;On a local-first design the store lives on your machine, and sending context to a hosted service is an explicit action rather than the default. "You own your data" and "your data never leaves your machine" are different statements, and only the second one is checkable from your side.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes in practice
&lt;/h2&gt;

&lt;p&gt;The agent stops starting from zero. It knows you said pnpm. It knows which approach you already ruled out. It can point at the paragraph where you said it. The difference is not that the model got smarter. The difference is that it can read something which outlived the session.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you run agents
&lt;/h2&gt;

&lt;p&gt;The useful question is not whether your agent has memory. It is where that memory lives, and whether it is still addressable tomorrow. If the answer is "somewhere in the context window", you are one compaction away from starting over.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow, the local-first memory layer described above. Docs and the client are here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;HyperMarrow&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you run agents: what is the thing you have to repeat most often?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Disclosure: I build HyperMarrow, the local-first memory layer described above.)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>memory</category>
      <category>agents</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Four agent-memory bugs from real repos: idle compaction silently discards working context</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Fri, 02 Oct 2026 02:00:55 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/four-agent-memory-bugs-from-real-repos-idle-compaction-silently-discards-working-context-562b</link>
      <guid>https://dev.to/qianqiuwanzi/four-agent-memory-bugs-from-real-repos-idle-compaction-silently-discards-working-context-562b</guid>
      <description>&lt;p&gt;Open a long-running agent session, walk away, come back. In at least one widely used coding agent, compaction can fire while the session is idle and drop context you still needed, with no way to opt out. The issue title says it better than I can: idle compaction silently discards working context in long-running sessions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2geep3oj2gbdnk2kcnp.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fy2geep3oj2gbdnk2kcnp.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That report is not an outlier. Reading the public issue trackers of the major agent CLIs, the same handful of failures keeps showing up. They are not exotic edge cases. They are what happens when "memory" is really just a context window with a summarizer bolted on.&lt;/p&gt;

&lt;p&gt;I work on HyperMarrow, a local-first memory layer for agents. I treat these reports as a backlog rather than as talking points. Here are four of them, quoted from the trackers, and the write-time rule each one implies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug 1. "Idle compaction silently discards working context in long-running sessions; no opt-out"
&lt;/h2&gt;

&lt;p&gt;Source: anthropics/claude-code issue 98747&lt;/p&gt;

&lt;p&gt;What happens: compaction is a background job. It runs when the session is idle, not when it is safe. Anything the summarizer judges low-value at that moment is gone, and the user never got a vote. A mirror-image report in another repo describes subagent-completion turns skipping preflight compaction, so a session over its threshold never gets compacted at all.&lt;/p&gt;

&lt;p&gt;The rule: a write decision has to happen before compaction, not after it.&lt;/p&gt;

&lt;p&gt;In our design, compaction is not allowed to be the first thing that touches recent state. Records, decisions and conclusions are written to the local store first, each with a source tag and a timestamp. Only then may the working context be summarized or dropped. If the summarizer throws something away, the durable copy is already on disk. That turns compaction from a data-loss problem into a view problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug 2. "memory_search hybrid ranking drops the only chunk that contains the whole query"
&lt;/h2&gt;

&lt;p&gt;Source: openclaw/openclaw issue 162764&lt;/p&gt;

&lt;p&gt;What happens: retrieval quality is treated as a ranking problem. It is usually a writing problem. If what you stored was already a lossy summary, no ranker can recover the sentence you needed.&lt;/p&gt;

&lt;p&gt;The rule: decide at write time what is a fact, what is a decision, and what is noise.&lt;/p&gt;

&lt;p&gt;Our layer separates the three. Facts and decisions are stored verbatim and immutably. Noise is stored as a pointer, or not stored at all. Retrieval then has something exact to hit. When a recall misses, the first question is not "which embedding model", it is "what did we fail to write down".&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug 3. "Subagents persist in memory after stopping without manual removal"
&lt;/h2&gt;

&lt;p&gt;Source: anthropics/claude-code issue 98804&lt;/p&gt;

&lt;p&gt;What happens: everything is remembered forever, so stale agents, dead sessions and one-off experiments accumulate and start to pollute later decisions.&lt;/p&gt;

&lt;p&gt;The rule: forgetting has to be bounded and scheduled, not manual.&lt;/p&gt;

&lt;p&gt;Our layer runs a decay pass. Low-value records fade along a curve, pinned records never fade, and nothing is hard-deleted while it is still referenced. The user should not have to be the garbage collector. A memory system that needs manual cleanup is not a memory system, it is a leak.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug 4. "Paragraph anchors with cross-session references"
&lt;/h2&gt;

&lt;p&gt;Source: anthropics/claude-code issue 98768 (a feature request)&lt;/p&gt;

&lt;p&gt;What happens: users want to point at one paragraph from an earlier session. Today they can point at a whole conversation, or at nothing.&lt;/p&gt;

&lt;p&gt;The rule: continuity has to be addressable.&lt;/p&gt;

&lt;p&gt;This one is filed as a feature, and that is the interesting part. The distance between "I remember your last chat" and "I can cite the exact paragraph from three sessions ago" is where continuity actually lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rules, in one place
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Write before you compact. The durable copy exists before the summarizer runs.&lt;/li&gt;
&lt;li&gt;Separate facts, decisions and noise at write time, not at query time.&lt;/li&gt;
&lt;li&gt;Forgetting is scheduled and bounded; pinned records never decay.&lt;/li&gt;
&lt;li&gt;Continuity is addressable across sessions, down to a paragraph.&lt;/li&gt;
&lt;li&gt;All of it stays local by default, and the privacy boundary is a first-class setting rather than an afterthought.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is exotic. It is the boring part of storage that most agent stacks skip, because a context window is good enough in a demo and the bill arrives later.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you run agents
&lt;/h2&gt;

&lt;p&gt;The useful question is not whether your agent has memory. It is where that memory lives, and what happens to it when the session ends. If the answer is "in the context window", these four bugs are already on your roadmap, whether or not you filed them.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow, the local-first memory layer described above. Docs and the client are here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;HyperMarrow&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you maintain an agent runtime: which of these four have you hit, and which one do you consider unfixable?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Disclosure: I build HyperMarrow, the local-first memory layer described above.)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>mcp</category>
      <category>llm</category>
    </item>
    <item>
      <title>The 3 Quiet Tricks That Make AI Memory Actually Pay Off Over a Year</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Thu, 01 Oct 2026 01:45:02 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/the-3-quiet-tricks-that-make-ai-memory-actually-pay-off-over-a-year-186h</link>
      <guid>https://dev.to/qianqiuwanzi/the-3-quiet-tricks-that-make-ai-memory-actually-pay-off-over-a-year-186h</guid>
      <description>&lt;p&gt;Everyone demos AI memory with a fresh session. The real test is what happens after a year of use.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffbivwafmhlbc5s7wn1dn.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffbivwafmhlbc5s7wn1dn.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A memory layer that only stores is a junk drawer. The systems that pay off long-term do three quiet things:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Consolidate&lt;/strong&gt; — turn raw session logs into durable facts, so the useful signal survives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Forget on a schedule&lt;/strong&gt; — let low-value noise decay instead of piling up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stay local&lt;/strong&gt; — the longer the history, the more sensitive it gets; keeping it on your disk keeps the risk bounded.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We built HyperMarrow around exactly these three. The four building blocks (recall / intercept-record / consolidation / file-bridge) are what make a year of memory feel like an asset instead of a liability.&lt;/p&gt;

&lt;p&gt;The download and docs are here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;https://hm.qianshi.cool/api/v2/dl?from=devto&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;What would change for you if your AI's memory actually compounded instead of rotting?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Disclosure: I build HyperMarrow, the local-first memory system described above.)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>memory</category>
      <category>llm</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Local vs Cloud AI Memory: What You Actually Trade When You Pick One</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Wed, 30 Sep 2026 02:27:42 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/local-vs-cloud-ai-memory-what-you-actually-trade-when-you-pick-one-29jl</link>
      <guid>https://dev.to/qianqiuwanzi/local-vs-cloud-ai-memory-what-you-actually-trade-when-you-pick-one-29jl</guid>
      <description>&lt;p&gt;Most teams pick an AI memory product by the model it uses. The more important question is where the memory physically lives.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjy3tkrc5rm2mtwrprvs9.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjy3tkrc5rm2mtwrprvs9.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A cloud memory layer is convenient — until you realize every decision, customer note, and internal context your team produces is now sitting in someone else's database. You traded ownership for convenience.&lt;/p&gt;

&lt;p&gt;A local-first layer keeps the store on your own hardware:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Raw data never leaves the intranet unless you explicitly choose to share it.&lt;/li&gt;
&lt;li&gt;Consolidation and forgetting run on machines you control.&lt;/li&gt;
&lt;li&gt;Sharing is opt-in, per document, per person.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You still get retrieval-augmented memory. You just stop handing the crown jewels to a vendor by default.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow, a local-first memory system for AI agents and coding assistants. The docs and download are here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;https://hm.qianshi.cool/api/v2/dl?from=devto&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you had to pick, where would your team's memory store physically live?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Disclosure: I build HyperMarrow, the local-first memory system described above.)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>privacy</category>
      <category>local</category>
      <category>llm</category>
    </item>
    <item>
      <title>Looking for Tech Co-builders: Let's Make Local AI Memory an Open Standard</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:26:10 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/looking-for-tech-co-builders-lets-make-local-ai-memory-an-open-standard-2k3n</link>
      <guid>https://dev.to/qianqiuwanzi/looking-for-tech-co-builders-lets-make-local-ai-memory-an-open-standard-2k3n</guid>
      <description>&lt;p&gt;Local, private AI memory is where this whole space is headed — but the interfaces are still a mess. Everyone builds their own silo, and switching costs lock users in.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zjjoqrgom02u1h565h5.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F0zjjoqrgom02u1h565h5.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I have been building HyperMarrow, a local-first memory system with four building blocks (recall / intercept-record / consolidation / file-bridge), exposed over MCP so the data stays on the user's machine.&lt;/p&gt;

&lt;p&gt;I am looking for a few technical co-builders who care about this problem: people who want to help shape how local memory is structured, shared, and made interoperable — not just ship another closed product.&lt;/p&gt;

&lt;p&gt;What is on the table: architecture decisions, the open spec, and the hard parts (privacy boundaries, forgetting curves, cross-agent sharing). If that sounds like your kind of problem, the docs and contact path are here: &lt;a href="https://hm.qianshi.cool/api/v2/dl?from=devto" rel="noopener noreferrer"&gt;https://hm.qianshi.cool/api/v2/dl?from=devto&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;What would make a local-memory standard actually stick for you?&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Disclosure: I build HyperMarrow, the local-first memory system described above — that is also why I am recruiting co-builders.)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>mcp</category>
      <category>llm</category>
    </item>
    <item>
      <title>What Your AI Should Remember — and What It Must Forget: A Privacy Boundary</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Mon, 28 Sep 2026 06:20:54 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/what-your-ai-should-remember-and-what-it-must-forget-a-privacy-boundary-2h6m</link>
      <guid>https://dev.to/qianqiuwanzi/what-your-ai-should-remember-and-what-it-must-forget-a-privacy-boundary-2h6m</guid>
      <description>&lt;p&gt;Giving an AI a memory sounds great until you ask: what exactly is it keeping? A memory system with no boundary is just a surveillance machine you trained yourself.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjyzm5omk5wxvf933e596.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fjyzm5omk5wxvf933e596.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The boundary we draw is simple and local:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Record the &lt;em&gt;work&lt;/em&gt;, not the &lt;em&gt;person&lt;/em&gt;. Decisions, code, fixes — yes. Private credentials, health, unrelated personal context — never.&lt;/li&gt;
&lt;li&gt;The store lives on your disk. Nothing leaves unless you explicitly share it.&lt;/li&gt;
&lt;li&gt;You can delete any record, and the deletion is real (local file, not a vendor's 'we deleted it' promise).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The test: if you would be uncomfortable reading it back aloud in a meeting, it should not be in the store.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow around exactly this boundary. The docs are in my profile (?from=devto).&lt;/p&gt;

&lt;p&gt;Where would you draw the line on what your AI is allowed to remember?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>privacy</category>
      <category>llm</category>
      <category>local</category>
    </item>
    <item>
      <title>Wiring Local Memory Into Your Stack via MCP: A Practical Walkthrough</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Sun, 27 Sep 2026 04:38:19 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/wiring-local-memory-into-your-stack-via-mcp-a-practical-walkthrough-4g8j</link>
      <guid>https://dev.to/qianqiuwanzi/wiring-local-memory-into-your-stack-via-mcp-a-practical-walkthrough-4g8j</guid>
      <description>&lt;p&gt;Most 'AI memory' demos hardcode the store. That breaks the moment you change models or add a second agent. The fix is to treat memory as a capability your tools expose, not a database you wire by hand.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs5lspm4bfbjkk8wy4idq.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs5lspm4bfbjkk8wy4idq.jpg" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Enter MCP. A local memory layer can expose recall / record / consolidation as MCP resources and tools. Any MCP-aware client — Claude, Codex, your own agent — can then read and write memory through one standard interface, with the actual data staying on your disk.&lt;/p&gt;

&lt;p&gt;What that buys you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Swap models without rewriting memory code.&lt;/li&gt;
&lt;li&gt;One agent's record is visible to the next.&lt;/li&gt;
&lt;li&gt;The store stays local; only the interface is shared.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I am building HyperMarrow, whose four building blocks are exposed this way. The docs are in my profile (?from=devto).&lt;/p&gt;

&lt;p&gt;If you have wired MCP before, what was the hardest part to get right?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>llm</category>
      <category>local</category>
    </item>
    <item>
      <title>Build a Team Knowledge Base That Never Quits: Local Memory for AI Agents</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 01:45:27 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/build-a-team-knowledge-base-that-never-quits-local-memory-for-ai-agents-leb</link>
      <guid>https://dev.to/qianqiuwanzi/build-a-team-knowledge-base-that-never-quits-local-memory-for-ai-agents-leb</guid>
      <description>&lt;p&gt;The most expensive loss in a team is not a missed deadline. It is the knowledge that walks out the door when someone leaves — the 'why we built it this way', the fix nobody wrote down, the tribal context.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs4k9h1g44s1qaaeksgve.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs4k9h1g44s1qaaeksgve.png" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A local memory layer changes that. Every decision, fix, and discussion your agents and people produce gets consolidated into a shared, structured store that stays on your machines. New teammates — human or agent — start with the team's history already loaded.&lt;/p&gt;

&lt;p&gt;It is not a wiki nobody updates. It is memory that writes itself as work happens, and forgets on a schedule so the noise fades.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow, a local-first memory system for exactly this. The download and docs are in my profile (?from=devto).&lt;/p&gt;

&lt;p&gt;What would you preserve if your team's memory could outlast any single person?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>knowledge</category>
      <category>mcp</category>
      <category>llm</category>
    </item>
    <item>
      <title>Why Enterprise AI Memory Should Never Leave the Intranet</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Sat, 26 Sep 2026 01:11:02 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/why-enterprise-ai-memory-should-never-leave-the-intranet-2e5l</link>
      <guid>https://dev.to/qianqiuwanzi/why-enterprise-ai-memory-should-never-leave-the-intranet-2e5l</guid>
      <description>&lt;p&gt;When a company adopts AI memory, the first question should not be &lt;em&gt;which model&lt;/em&gt;. It should be &lt;em&gt;where does the memory live&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8w77683pavwfbrg510sx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8w77683pavwfbrg510sx.png" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;For an enterprise, memory is the most sensitive asset there is — it is a concentrated record of decisions, customers, and internal context. Shipping that to a third-party cloud by default is a liability most teams sign up for without thinking.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;local-first&lt;/strong&gt; memory architecture keeps the store inside the intranet:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Raw data is never uploaded unless explicitly chosen.&lt;/li&gt;
&lt;li&gt;Consolidation and forgetting run on local hardware.&lt;/li&gt;
&lt;li&gt;Sharing is opt-in, per document, per person.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You still get the power of retrieval-augmented memory. You just do not hand the crown jewels to a vendor.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow, a local-first memory system designed for exactly this boundary. The download and docs are in my profile (?from=devto).&lt;/p&gt;

&lt;p&gt;For teams evaluating AI memory: what is your policy on where the memory store physically lives?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>privacy</category>
      <category>enterprise</category>
      <category>local</category>
    </item>
    <item>
      <title>Turn Your Local Files Into Your AI's Second Brain (Without the Cloud)</title>
      <dc:creator>qianqiuwanzi</dc:creator>
      <pubDate>Fri, 25 Sep 2026 16:16:46 +0000</pubDate>
      <link>https://dev.to/qianqiuwanzi/turn-your-local-files-into-your-ais-second-brain-without-the-cloud-do4</link>
      <guid>https://dev.to/qianqiuwanzi/turn-your-local-files-into-your-ais-second-brain-without-the-cloud-do4</guid>
      <description>&lt;p&gt;Your best knowledge is not in chat history. It is in the folders on your disk — docs, notes, specs, past reports. The problem is your AI cannot see any of it without you pasting it in by hand.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs4k9h1g44s1qaaeksgve.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs4k9h1g44s1qaaeksgve.png" alt="cover" width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;file-bridge&lt;/strong&gt; fixes that. It maps your local files into the same memory layer your agent already uses, so a document becomes a retrievable memory — without uploading a single byte to a cloud.&lt;/p&gt;

&lt;p&gt;How it works in practice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You point the bridge at a folder.&lt;/li&gt;
&lt;li&gt;Files are indexed locally (never sent out).&lt;/li&gt;
&lt;li&gt;When you ask a question, the agent pulls the relevant passage the same way it pulls a stored memory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The key word is &lt;em&gt;local&lt;/em&gt;. The index lives on your machine. Your private docs stay private.&lt;/p&gt;

&lt;p&gt;I am building HyperMarrow, whose file-bridge is one of four building blocks (recall / intercept-record / consolidation / file-bridge). The download and docs are in my profile (?from=devto).&lt;/p&gt;

&lt;p&gt;Which folder on your machine would you most want your AI to actually understand?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>productivity</category>
      <category>local</category>
    </item>
  </channel>
</rss>
