<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: lanbass869-cell</title>
    <description>The latest articles on DEV Community by lanbass869-cell (@lanbass869cell).</description>
    <link>https://dev.to/lanbass869cell</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4156498%2Faaddd45e-5cad-46a1-92e0-f315b155167b.png</url>
      <title>DEV Community: lanbass869-cell</title>
      <link>https://dev.to/lanbass869cell</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lanbass869cell"/>
    <language>en</language>
    <item>
      <title>I Built a Cross-Client Memory Hub for AI Agents — Here's What I Learned</title>
      <dc:creator>lanbass869-cell</dc:creator>
      <pubDate>Fri, 02 Oct 2026 06:22:36 +0000</pubDate>
      <link>https://dev.to/lanbass869cell/i-built-a-cross-client-memory-hub-for-ai-agents-heres-what-i-learned-418l</link>
      <guid>https://dev.to/lanbass869cell/i-built-a-cross-client-memory-hub-for-ai-agents-heres-what-i-learned-418l</guid>
      <description>&lt;p&gt;I use Claude Code for coding, Cursor for refactoring, and Windsurf for exploration. Each has its own memory. Switch tools and my AI forgets everything.&lt;/p&gt;

&lt;p&gt;So I built &lt;a href="https://github.com/MemTether/MemTether" rel="noopener noreferrer"&gt;MemTether&lt;/a&gt; — a local-first memory hub that lets 23+ AI clients share one physical SQLite database.&lt;/p&gt;

&lt;p&gt;Here's what I learned building it.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem Is Simpler Than You Think
&lt;/h2&gt;

&lt;p&gt;Existing memory solutions (mem0, cognee, zep) all treat memory as a &lt;strong&gt;service&lt;/strong&gt;. You send memories to their cloud API, they store and retrieve them. This works, but it means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your memories live on someone else's server&lt;/li&gt;
&lt;li&gt;You pay per API call&lt;/li&gt;
&lt;li&gt;You need an API key even for a local project&lt;/li&gt;
&lt;li&gt;You can't inspect the raw data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;My insight: if all your AI tools run on the same machine, you don't need a &lt;strong&gt;service&lt;/strong&gt; — you need a &lt;strong&gt;file&lt;/strong&gt;. Just make them all point to the same SQLite database.&lt;/p&gt;

&lt;p&gt;No cloud. No API fees. No abstraction layer. One &lt;code&gt;memory.db&lt;/code&gt;, 23 clients reading and writing to it.&lt;/p&gt;




&lt;h2&gt;
  
  
  Design Decisions That Matter
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Supersession, Not Deletion
&lt;/h3&gt;

&lt;p&gt;When a memory needs updating, I don't delete the old one. I mark it &lt;code&gt;superseded&lt;/code&gt; and create a new version. This means you can always trace "what did we believe before we learned X?"&lt;/p&gt;

&lt;p&gt;This sounds simple, but it changes everything about how the system works. Search must filter out superseded entries. FTS5 needs triggers to auto-sync. The projection system needs to prefer the latest version.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Bi-Temporal: Two Clocks, Not One
&lt;/h3&gt;

&lt;p&gt;Every memory has two timestamps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;T (valid time)&lt;/strong&gt;: when the fact was true in the real world&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;T′ (recorded time)&lt;/strong&gt;: when the system learned about it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example: an API key expired on September 15th, but I didn't notice until September 18th. Querying "what did we know on September 15th?" returns "the key is valid" (which is what the system believed). Querying "what was actually true?" returns "expired."&lt;/p&gt;

&lt;p&gt;Without this distinction, you get retroactive truth bias — using today's knowledge to judge yesterday's decisions.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Q-Value: Memories That Get Used Should Rank Higher
&lt;/h3&gt;

&lt;p&gt;Most memory systems rank by recency or semantic similarity. But a memory from 6 months ago that gets used every day is more valuable than one from yesterday that's never been retrieved.&lt;/p&gt;

&lt;p&gt;So I added a &lt;strong&gt;Q-Value&lt;/strong&gt; (inspired by reinforcement learning): every time a memory is searched and actually useful, its Q-Value increases. Next search, it ranks higher. Simple, effective, and nobody else does it.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Four-Factor Re-Ranking
&lt;/h3&gt;

&lt;p&gt;Raw semantic similarity isn't enough. I blend four factors:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Semantic similarity (45%)&lt;/li&gt;
&lt;li&gt;Recency (25%)&lt;/li&gt;
&lt;li&gt;Usage frequency (5%)&lt;/li&gt;
&lt;li&gt;Type importance (10%)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then normalize with z-score + sigmoid and blend 70/30 with the RRF score. This fixed asset lookup queries that pure semantic search missed.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. SQLite Triggers for FTS Sync (Not Python)
&lt;/h3&gt;

&lt;p&gt;My first version had Python code to sync the FTS5 full-text index. It had &lt;code&gt;except Exception: pass&lt;/code&gt; around the sync calls. Result: 254 stale entries and 260 missing entries.&lt;/p&gt;

&lt;p&gt;The fix: &lt;strong&gt;SQLite triggers&lt;/strong&gt;. Three triggers (INSERT, DELETE, UPDATE) at the SQL level. No Python code can accidentally skip them. The inconsistency went from 514 entries to zero.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Lesson: if SQLite can do it at the SQL level, don't do it in Python.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  The Hard Part Isn't Code — It's Packaging
&lt;/h2&gt;

&lt;p&gt;Writing the memory engine took 3 weeks. Making it installable took 2 months:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wheel building with correct &lt;code&gt;py-modules&lt;/code&gt; (flat layout, not packages)&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;check_packaging.py&lt;/code&gt; to catch version drift and missing modules&lt;/li&gt;
&lt;li&gt;FTS5 triggers in the right place (inside &lt;code&gt;init_db&lt;/code&gt;, not floating in the schema string)&lt;/li&gt;
&lt;li&gt;Console entry points (&lt;code&gt;memtether&lt;/code&gt;, &lt;code&gt;memtether-connect&lt;/code&gt;)&lt;/li&gt;
&lt;li&gt;Optional dependencies (&lt;code&gt;[vector]&lt;/code&gt; for chromadb, &lt;code&gt;[server]&lt;/code&gt; for fastapi)&lt;/li&gt;
&lt;li&gt;A Quick Start that actually works (I initially wrote commands that didn't exist)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I'm a solo developer. Every hour spent on packaging is an hour not spent on features. But without packaging, nobody can use your features.&lt;/p&gt;




&lt;h2&gt;
  
  
  Honest Benchmark Numbers
&lt;/h2&gt;

&lt;p&gt;I ran LongMemEval (500 questions, full run):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Score&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Strict match&lt;/td&gt;
&lt;td&gt;62.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LLM judge&lt;/td&gt;
&lt;td&gt;54.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-session strict&lt;/td&gt;
&lt;td&gt;48.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multi-session LLM judge&lt;/td&gt;
&lt;td&gt;60.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The multi-session gap (strict vs judge) is the most interesting finding. When an answer is computed (like "3 weeks" from multiple data points), strict substring matching fails because the computed answer doesn't appear verbatim in any single memory. The LLM judge is more forgiving.&lt;/p&gt;

&lt;p&gt;I also built an E-Hybrid method (session summaries + flat evidence) that improved multi-session strict from 48.8% to 71.4% on a 15-question test. Small sample, but promising.&lt;/p&gt;

&lt;p&gt;I'm not going to claim these are better than mem0's 94.4% — different harness, different methodology. But 48.8% strict is above the industry average of 27.9%.&lt;/p&gt;




&lt;h2&gt;
  
  
  What's Next
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;MCP Registry submission (done: &lt;a href="https://github.com/modelcontextprotocol/servers/pull/4946" rel="noopener noreferrer"&gt;PR #4946&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;Community building (this blog post is part of that)&lt;/li&gt;
&lt;li&gt;NLPCC 2027 paper submission (CCF C, deadline ~April 2027)&lt;/li&gt;
&lt;li&gt;More examples and better docs&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Try It
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;memtether
memtether init
memtether connect &lt;span class="nt"&gt;--all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;GitHub: &lt;a href="https://github.com/MemTether/MemTether" rel="noopener noreferrer"&gt;MemTether/MemTether&lt;/a&gt;&lt;br&gt;
PyPI: &lt;a href="https://pypi.org/project/memtether/" rel="noopener noreferrer"&gt;memtether&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>python</category>
      <category>mcp</category>
    </item>
  </channel>
</rss>
