<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Andrea Redegalli</title>
    <description>The latest articles on DEV Community by Andrea Redegalli (@aendrix03).</description>
    <link>https://dev.to/aendrix03</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4127096%2F14027752-0719-4833-9b60-3cbe33818bdb.jpg</url>
      <title>DEV Community: Andrea Redegalli</title>
      <link>https://dev.to/aendrix03</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/aendrix03"/>
    <language>en</language>
    <item>
      <title>Why vector similarity alone is not enough for coding-agent memory</title>
      <dc:creator>Andrea Redegalli</dc:creator>
      <pubDate>Wed, 16 Sep 2026 21:48:20 +0000</pubDate>
      <link>https://dev.to/aendrix03/why-vector-similarity-alone-is-not-enough-for-coding-agent-memory-26bb</link>
      <guid>https://dev.to/aendrix03/why-vector-similarity-alone-is-not-enough-for-coding-agent-memory-26bb</guid>
      <description>&lt;p&gt;A coding agent fixes a tricky deployment failure. Two weeks later, another session sees the same failure described in different words. The expensive part is not finding a nearby document chunk. It is deciding whether a past conclusion is safe enough to reuse.&lt;/p&gt;

&lt;p&gt;That distinction is why I think agent memory needs more than vector similarity.&lt;/p&gt;

&lt;h2&gt;
  
  
  Similar is not the same as reusable
&lt;/h2&gt;

&lt;p&gt;Vector search is excellent at finding semantic neighbours. Ask for help with a Docker healthcheck and it may find memories about container startup, service discovery, or a previous networking issue. Those can be useful leads.&lt;/p&gt;

&lt;p&gt;But a ranked list is not a decision. The top result may be merely adjacent: it may apply to a different framework version; describe an abandoned approach; be true only under an unstated deployment constraint; or have been superseded by a later architectural decision.&lt;/p&gt;

&lt;p&gt;For a human, these are normal caveats. For an agent working under time and context pressure, treating the nearest neighbour as an answer can turn retrieval into a subtle source of confident mistakes. The question is not only “what is semantically close?” It is also “have we seen this problem closely enough before to make that prior learning part of the current reasoning?”&lt;/p&gt;

&lt;h2&gt;
  
  
  Separate recall from exploration
&lt;/h2&gt;

&lt;p&gt;A useful memory system can expose different operations for different levels of certainty.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verified recall&lt;/strong&gt; is the narrowest path. It returns one candidate only when multiple signals support the match; otherwise it says that the match is weak or absent. The important output is not just a snippet, but a gate: STRONG, WEAK, or MISS.&lt;/p&gt;

&lt;p&gt;A STRONG result is still evidence, not an instruction. The agent should inspect the returned context and decide whether it fits. A WEAK result should encourage investigation rather than shortcut it. A MISS tells the agent to solve the problem normally instead of inventing continuity.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hybrid retrieval&lt;/strong&gt; is broader. Semantic similarity is valuable, but it should be combined with lexical signals. Exact terms often carry disproportionate meaning in engineering work: a configuration key, exception type, table name, package version, or command flag. Dense retrieval captures paraphrase; BM25 preserves those precise anchors. Reciprocal-rank fusion is a practical way to blend independent rankings without pretending either is universally correct.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Graph exploration&lt;/strong&gt; answers a third question: what else is connected to this? A deployment failure may touch a prior decision about ingress, a known local-development exception, and a later supersession. Those relationships are often more useful than another ten chunks with similar embeddings. Exploration should favour diversity and keep the relationship visible, rather than flattening everything into one similarity score.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write memories as learnings, not documents
&lt;/h2&gt;

&lt;p&gt;This changes what we store. A codebase index asks: “What documents might answer this query?” Agent memory asks: “What did the agent learn that would be expensive to rediscover?”&lt;/p&gt;

&lt;p&gt;Good memory entries are compact and explicit: the root cause and fix; the architectural choice and why it was made; a framework or infrastructure gotcha; a constraint that rules out an appealing approach; or a failed experiment worth not repeating.&lt;/p&gt;

&lt;p&gt;Lifecycle matters too. If a decision changes, deleting the old memory loses context; leaving it untreated risks a stale answer. A better model marks the old learning as superseded and links it to the newer one. History remains inspectable while retrieval can favour what is current.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retrieval quality is a product decision
&lt;/h2&gt;

&lt;p&gt;There is no universal threshold at which a match becomes safe. A recall system has to decide its false-positive and false-negative trade-off. For an agent that proposes code changes, false positives are costly: a plausible but stale memory can send work down the wrong path. For an agent doing background research, a false negative may be less harmful: it can simply search more broadly. Confidence should therefore be exposed as part of the contract rather than hidden inside a score.&lt;/p&gt;

&lt;p&gt;A practical evaluation loop should include failures, not just happy-path hits: paraphrases of the same historical problem; near misses involving a different version, environment, or dependency; contradictions and superseded decisions; exact-identifier queries that vectors may dilute; and small noisy stores where every result appears relevant. The goal is not a leaderboard score. It is to learn when the system should stay silent.&lt;/p&gt;

&lt;h2&gt;
  
  
  A local-first implementation
&lt;/h2&gt;

&lt;p&gt;I am building these ideas into &lt;a href="https://github.com/AEndrix03/Graft" rel="noopener noreferrer"&gt;Graft&lt;/a&gt;, an Apache-2.0 local memory layer for coding agents. Its core is a C11 daemon backed by SQLite, FTS5, sqlite-vec, llama.cpp, and local BGE-M3 embeddings. It does not require a SaaS account, external embedding API, or external LLM call for its storage and retrieval loop.&lt;/p&gt;

&lt;p&gt;Graft keeps the three paths separate: confidence-gated query, hybrid retrieve using vector search plus BM25 title/body search with reciprocal-rank fusion, and explore for semantic and keyword relationships. It is an active alpha, and the cross-encoder reranker is scaffolded rather than active today; verification currently uses vector similarity plus lexical signals.&lt;/p&gt;

&lt;p&gt;If you are building long-running coding-agent workflows, I would value feedback on the hard part: what evidence would make you willing—or unwilling—to reuse a prior agent learning?&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>Your coding agent already learned this: local-first memory with Graft</title>
      <dc:creator>Andrea Redegalli</dc:creator>
      <pubDate>Wed, 16 Sep 2026 02:23:18 +0000</pubDate>
      <link>https://dev.to/aendrix03/your-coding-agent-already-learned-this-local-first-memory-with-graft-4dgb</link>
      <guid>https://dev.to/aendrix03/your-coding-agent-already-learned-this-local-first-memory-with-graft-4dgb</guid>
      <description>&lt;p&gt;Many coding-agent sessions end with an expensive loss: the bug is understood, the fix is shipped, and the useful reasoning disappears with the session.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/AEndrix03/Graft" rel="noopener noreferrer"&gt;Graft&lt;/a&gt; is an open-source project for keeping that kind of knowledge available locally across later agent work. Its focus is not document ingestion or a hosted chatbot. It is persistent memory for the fixes, decisions, constraints, and project-specific gotchas that agents uncover while solving real tasks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The workflow is intentionally small
&lt;/h2&gt;

&lt;p&gt;The core loop is three questions:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Have we seen this before?&lt;/strong&gt; Use &lt;code&gt;graft query&lt;/code&gt; for a confidence-gated top result: &lt;code&gt;STRONG&lt;/code&gt;, &lt;code&gt;WEAK&lt;/code&gt;, or &lt;code&gt;MISS&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What could help here?&lt;/strong&gt; Use &lt;code&gt;graft retrieve&lt;/code&gt; for ranked memories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What else is connected?&lt;/strong&gt; Use &lt;code&gt;graft explore&lt;/code&gt; to walk semantic and keyword relationships.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When a solution is worth preserving, it can be added as a memory node with a title, body, and keywords:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;graft insert &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--title&lt;/span&gt; &lt;span class="s1"&gt;'Spring @Valid must also be applied to nested DTO fields'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--body&lt;/span&gt; &lt;span class="s1"&gt;'Without @Valid on the nested field, validation does not cascade into it.'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--keyword&lt;/span&gt; spring-boot &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--keyword&lt;/span&gt; validation &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--keyword&lt;/span&gt; gotcha
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The project documents the intended pattern clearly: search before a non-trivial task, solve normally when nothing useful exists, and save the reusable learning afterward.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local runtime, not a hosted memory service
&lt;/h2&gt;

&lt;p&gt;Graft’s core runtime is local: a CLI communicates with a local daemon, which uses SQLite storage alongside FTS5 and sqlite-vec. Embeddings run locally through llama.cpp with BGE-M3. The repository states that the default setup needs no SaaS account, external embedding API, or API key.&lt;/p&gt;

&lt;p&gt;The retrieval modes are deliberately different:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;query&lt;/code&gt; uses embedding-based candidates plus lexical verification and confidence gating.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;retrieve&lt;/code&gt; combines BGE-M3 vectors, BM25 over titles, and BM25 over bodies with Reciprocal Rank Fusion.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;explore&lt;/code&gt; traverses semantic and keyword relationships using beam search, score decay, and MMR diversity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why Graft is not positioned as a replacement for a vector database. The repository distinguishes bulk document indexing from remembering agent learnings produced during work.&lt;/p&gt;

&lt;h2&gt;
  
  
  Integrations and current scope
&lt;/h2&gt;

&lt;p&gt;Graft ships as a C11 project with a CLI contract, so agents that can launch a subprocess can use it. The repository lists integrations for Claude Code, Codex, Open Code, Gemini CLI, Claude Desktop, ChatGPT, and custom agents through CLI, subprocess, REST, or MCP.&lt;/p&gt;

&lt;p&gt;It also documents profiles for separate memory spaces, plus optional REST, MCP, graph-viewer, and analytics tooling. Knowledge can evolve through supersession: older memories remain inspectable while a newer memory becomes the useful one.&lt;/p&gt;

&lt;p&gt;The project is currently an active alpha in the v0.1.x line. Its README lists the local daemon and CLI, SQLite storage, BGE-M3 embeddings, verified recall, hybrid retrieval, graph exploration, profiles, coding-agent skills, and MCP bridge as working today. It also explicitly calls out areas still evolving, including the API surface before 1.0, packaging and platform coverage, remote/shared memory, team workflows, and neural reranking.&lt;/p&gt;

&lt;p&gt;If your problem is indexing millions of documents, use the right retrieval stack for that. If the problem is that your coding agent keeps rediscovering the same hard-won lesson, Graft is built for that narrower job.&lt;/p&gt;

&lt;p&gt;Repository: &lt;a href="https://github.com/AEndrix03/Graft" rel="noopener noreferrer"&gt;https://github.com/AEndrix03/Graft&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;This post was prepared with AI assistance and reviewed against the repository README.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
    </item>
  </channel>
</rss>
