<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ujwal Malladi</title>
    <description>The latest articles on DEV Community by Ujwal Malladi (@ujwal_malladi).</description>
    <link>https://dev.to/ujwal_malladi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4150104%2Ff50cdcb6-3d8e-4a9c-a88f-f946ceb7dc3f.png</url>
      <title>DEV Community: Ujwal Malladi</title>
      <link>https://dev.to/ujwal_malladi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ujwal_malladi"/>
    <language>en</language>
    <item>
      <title>The Hard Part of Giving an AI Agent Memory Was Deciding What to Remember</title>
      <dc:creator>Ujwal Malladi</dc:creator>
      <pubDate>Tue, 29 Sep 2026 15:43:33 +0000</pubDate>
      <link>https://dev.to/ujwal_malladi/the-hard-part-of-giving-an-ai-agent-memory-was-deciding-what-to-remember-8lb</link>
      <guid>https://dev.to/ujwal_malladi/the-hard-part-of-giving-an-ai-agent-memory-was-deciding-what-to-remember-8lb</guid>
      <description>&lt;p&gt;By M Gnan Ujwal&lt;br&gt;
Software Engineer | Continuum Project&lt;/p&gt;

&lt;p&gt;I focused on the production design of Continuum's memory architecture — deciding what should be retained, separating private customer memory from shared knowledge, and designing the Hindsight integration behind a testable service boundary.&lt;/p&gt;

&lt;p&gt;I expected the difficult part of giving a support agent memory to be connecting an LLM to a memory system. It wasn't. The harder problem was deciding what information should survive a conversation, who should be allowed to benefit from it, and where that information should live.&lt;/p&gt;

&lt;p&gt;That became the central engineering problem behind Continuum, a customer-support agent for a fictional SaaS product called Nimbus Sync. My role focused on the production-design side of the system: memory boundaries, the Hindsight integration, data responsibilities, and the abstractions that keep those pieces testable and replaceable.&lt;/p&gt;

&lt;p&gt;Continuum uses &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight for agent memory&lt;/a&gt; in two different ways. Each customer has a private memory bank containing their environment, previous troubleshooting attempts, unresolved issues, and relevant interaction context. Separately, the system maintains a shared known-issues bank containing distilled technical knowledge about recurring problems across customers.&lt;/p&gt;

&lt;p&gt;The important part is not that we stored more information. It is that we gave different information different lifetimes and different scopes.&lt;/p&gt;
&lt;h2&gt;
  
  
  Memory is not a transcript
&lt;/h2&gt;

&lt;p&gt;My first design constraint was simple: the support transcript and the agent's memory should not be the same thing.&lt;/p&gt;

&lt;p&gt;A transcript answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What exactly happened?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Memory answers:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What should the agent remember because it may matter later?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Continuum therefore keeps the literal customer, ticket, and message records in SQLite. The database is the system of record for the conversation that the UI renders.&lt;/p&gt;

&lt;p&gt;Hindsight has a different job. It stores semantic knowledge that can be recalled when it is useful: previous troubleshooting attempts, resolutions, customer context, and recurring technical issues.&lt;/p&gt;

&lt;p&gt;That separation is deliberate. If I treated every chat message as memory, the memory bank would become noisy. If I treated the memory store as the transcript, I would lose the clean, deterministic record of what actually happened.&lt;/p&gt;

&lt;p&gt;The architecture is essentially:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer message
      |
      +---------------------&amp;gt; SQLite
      |                        raw transcript
      |
      v
  Hindsight
  distilled knowledge
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This distinction ended up being one of the most reusable lessons from the project: &lt;strong&gt;storage and memory are different responsibilities, even when both contain text.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The more interesting problem: private versus shared knowledge
&lt;/h2&gt;

&lt;p&gt;A support agent needs to remember things about a customer, but a support organization also needs to remember things about the product.&lt;/p&gt;

&lt;p&gt;Those are not the same memory.&lt;/p&gt;

&lt;p&gt;Continuum gives every customer a bank such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customer::{id}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and keeps cross-customer technical knowledge in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;product::known-issues
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The customer bank is allowed to contain information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the customer's environment and plan&lt;/li&gt;
&lt;li&gt;troubleshooting steps they already tried&lt;/li&gt;
&lt;li&gt;previous resolutions&lt;/li&gt;
&lt;li&gt;relevant interaction context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The shared bank has a stricter rule: it contains the technical shape of recurring problems and their resolutions, but not the customer's personal details.&lt;/p&gt;

&lt;p&gt;That boundary matters.&lt;/p&gt;

&lt;p&gt;Imagine three customers report the same sync failure. The useful organizational memory is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Recurring issue in sync-engine.
Root cause: sleep-wake-stall.
Resolution: reconnect the session.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Marcus Chen reported this from his laptop on Tuesday.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first is reusable product knowledge. The second belongs to a customer's private history.&lt;/p&gt;

&lt;p&gt;This is one reason I think "give the agent memory" is too vague as a design requirement. The real questions are: &lt;strong&gt;memory of what, scoped to whom, and reusable by whom?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  One &lt;code&gt;reflect()&lt;/code&gt; call closes the learning loop
&lt;/h2&gt;

&lt;p&gt;The most important part of the Hindsight integration is the closed loop around &lt;code&gt;reflect()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;For each memory-enabled message, Continuum does four things:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Retain the customer's message
2. Recall relevant shared known issues
3. Reflect using customer memory + issue context
4. Write the outcome back into memory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fourth step is what makes the system accumulate knowledge.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;reflect()&lt;/code&gt; call returns not just the response text but structured signals. The response schema includes fields such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reply&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sentiment&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;frustrated&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;resolved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;False&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;root_cause_tag&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sync-timeout-vpn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;is_known_issue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;module&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sync-engine&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application can then turn those signals into memory.&lt;/p&gt;

&lt;p&gt;A resolution gets retained in the customer's bank. If the outcome is identified as a known issue, the technical signature is also retained in the shared bank.&lt;/p&gt;

&lt;p&gt;That means the loop looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customer message
      |
      v
recall shared known issues
      |
      v
reflect(customer memory + issue context)
      |
      +----&amp;gt; response to customer
      |
      +----&amp;gt; sentiment / resolution / root cause
                    |
                    v
              retain outcome
                    |
             +------+------+
             |             |
        customer bank   shared bank
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the part I would describe as learning in Continuum. The underlying model is not being retrained after every ticket. Instead, the system changes the context available to future reasoning by retaining useful outcomes.&lt;/p&gt;

&lt;p&gt;The next ticket therefore starts with more knowledge than the previous one.&lt;/p&gt;

&lt;h2&gt;
  
  
  I kept Hindsight behind one boundary
&lt;/h2&gt;

&lt;p&gt;I did not want Hindsight calls scattered throughout the application.&lt;/p&gt;

&lt;p&gt;The rest of the application talks to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MemoryService
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and &lt;code&gt;MemoryService&lt;/code&gt; talks to the Hindsight SDK.&lt;/p&gt;

&lt;p&gt;The adapter exposes operations that make sense to Continuum:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain_customer_message&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall_known_issues&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reflect_reply&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain_resolution&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;memory&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain_known_issue&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is more useful than exposing generic SDK calls everywhere.&lt;/p&gt;

&lt;p&gt;The rest of the application does not need to know how banks, recall parameters, structured outputs, or SDK objects are represented. That knowledge stays in one place.&lt;/p&gt;

&lt;p&gt;It also makes testing practical. The memory layer depends on a small &lt;code&gt;HindsightLike&lt;/code&gt; protocol rather than requiring the concrete SDK everywhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;HindsightLike&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Protocol&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;create_bank&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Any&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;reflect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Our tests provide a fake implementation of that surface.&lt;/p&gt;

&lt;p&gt;The result is a useful boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AgentService
     |
     v
MemoryService
     |
     v
Hindsight SDK
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the memory provider changes, or the SDK surface changes, I have one integration boundary to revisit rather than an application-wide rewrite.&lt;/p&gt;

&lt;p&gt;For developers evaluating agent-memory infrastructure, the &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight GitHub repository&lt;/a&gt; and &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; are useful starting points. The broader idea is also captured well in &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Vectorize's overview of agent memory&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  SQLite still has a job
&lt;/h2&gt;

&lt;p&gt;It would have been tempting to put everything into Hindsight and call the problem solved.&lt;/p&gt;

&lt;p&gt;I think that would have made the architecture worse.&lt;/p&gt;

&lt;p&gt;SQLite stores:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;customers
tickets
messages
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The repositories expose operations such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;ticket_repo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_message&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;ticket_repo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_messages&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;ticket_repo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_status&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Hindsight stores the semantic memory that the agent needs for future reasoning.&lt;/p&gt;

&lt;p&gt;This is a classic separation of concerns:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SQLite
= what happened

Hindsight
= what the agent should remember
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That distinction also makes debugging easier. If an agent gives an unexpected answer, I can inspect the literal transcript independently from the semantic memory that was retrieved.&lt;/p&gt;

&lt;h2&gt;
  
  
  Repository pattern keeps SQL out of the agent
&lt;/h2&gt;

&lt;p&gt;I used repositories for the SQLite side rather than letting the orchestration layer contain SQL.&lt;/p&gt;

&lt;p&gt;There are separate &lt;code&gt;CustomerRepository&lt;/code&gt; and &lt;code&gt;TicketRepository&lt;/code&gt; classes. They receive a database object and expose typed application-level operations.&lt;/p&gt;

&lt;p&gt;That means &lt;code&gt;AgentService&lt;/code&gt; can focus on the support workflow:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tickets&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add_message&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;reply&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;strategy&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;respond&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tickets&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set_status&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;instead of mixing ticket logic with SQL statements.&lt;/p&gt;

&lt;p&gt;It sounds like a small architectural decision, but it pays off when the system grows. The service layer describes business behavior; the repository layer describes persistence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Strategy made the memory comparison honest
&lt;/h2&gt;

&lt;p&gt;Continuum also has a memory-off baseline.&lt;/p&gt;

&lt;p&gt;Rather than scattering conditionals throughout the code, I used a common &lt;code&gt;AgentStrategy&lt;/code&gt; interface with two implementations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AgentStrategy
   |
   +-- WithMemoryStrategy
   |
   +-- NoMemoryStrategy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The memory-enabled strategy performs the retain → recall → reflect → write-back loop.&lt;/p&gt;

&lt;p&gt;The no-memory strategy makes a stateless LLM call and retains nothing.&lt;/p&gt;

&lt;p&gt;That distinction matters because the UI can send the exact same customer message through both paths. The comparison is therefore structural rather than a prompt trick.&lt;/p&gt;

&lt;p&gt;It also makes the code easier to reason about: the question "what happens when memory is off?" has a concrete implementation instead of being hidden behind branches inside the main agent logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  Dependency injection made the boundaries testable
&lt;/h2&gt;

&lt;p&gt;FastAPI's dependency injection is what ties the pieces together without hard-coding them into routes.&lt;/p&gt;

&lt;p&gt;The route can request an &lt;code&gt;AgentService&lt;/code&gt;, while the service receives its ticket repository, memory service, and baseline LLM.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FastAPI route
      |
      v
AgentService
  |     |     |
  v     v     v
Repo  Memory  LLM
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For tests, those dependencies can be substituted.&lt;/p&gt;

&lt;p&gt;That matters especially for Hindsight. The test suite uses a &lt;code&gt;FakeHindsight&lt;/code&gt; implementation, so the memory behavior can be exercised without requiring a live API key or network connection.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually difficult
&lt;/h2&gt;

&lt;p&gt;The difficult part was not writing the API endpoint that sends a message.&lt;/p&gt;

&lt;p&gt;It was deciding what should happen after the message.&lt;/p&gt;

&lt;p&gt;If the customer says something that sounds important, should I retain the entire message? Which parts are durable facts? Is an issue specific to this customer or systemic? If it is systemic, what can be safely shared? How do I avoid turning the shared bank into a dump of unrelated ticket details?&lt;/p&gt;

&lt;p&gt;Those questions pushed the architecture toward explicit missions, directives, structured output, tags, separate banks, and a dedicated memory service.&lt;/p&gt;

&lt;p&gt;We also had to make the memory behavior observable. The application exposes metadata such as memories used, directives applied, and known-issue matches. That makes it possible to inspect why a response had the context it did instead of treating memory as an invisible black box.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I would reuse in another agent
&lt;/h2&gt;

&lt;p&gt;Three lessons stand out.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Design memory boundaries before choosing storage
&lt;/h3&gt;

&lt;p&gt;Don't start with "where can I store the chat?"&lt;/p&gt;

&lt;p&gt;Start with "what should the agent remember, for how long, and who can use it?"&lt;/p&gt;

&lt;p&gt;That naturally leads to different memory scopes and retention rules.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Keep semantic memory separate from the system of record
&lt;/h3&gt;

&lt;p&gt;A memory system is optimized for helping future reasoning. A database transcript is optimized for accurately recording what happened.&lt;/p&gt;

&lt;p&gt;Trying to make one system perform both jobs can make both less useful.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Put the memory provider behind an application-level interface
&lt;/h3&gt;

&lt;p&gt;Agent code should ask for things like "recall known issues" and "retain this resolution," not manipulate provider-specific primitives everywhere.&lt;/p&gt;

&lt;p&gt;That makes the integration easier to test, replace, and evolve.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Make the learning signal structured
&lt;/h3&gt;

&lt;p&gt;Returning a response alone makes it difficult to build a reliable write-back loop.&lt;/p&gt;

&lt;p&gt;Returning the response plus explicit fields such as &lt;code&gt;resolved&lt;/code&gt;, &lt;code&gt;root_cause_tag&lt;/code&gt;, and &lt;code&gt;is_known_issue&lt;/code&gt; gives the application something concrete to retain and act on.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Treat privacy boundaries as architecture
&lt;/h3&gt;

&lt;p&gt;Private customer memory and shared organizational memory should not be separated only by a prompt instruction. They should be separated by the data model itself.&lt;/p&gt;

&lt;p&gt;In Continuum, that separation is represented directly by different Hindsight banks.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;The biggest lesson I took from building Continuum is that agent memory is not primarily a storage feature.&lt;/p&gt;

&lt;p&gt;It is a design decision about &lt;strong&gt;what survives a conversation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Once I treated memory that way, the rest of the architecture became clearer: SQLite records the conversation, Hindsight stores distilled knowledge, customer banks preserve private context, the shared bank captures systemic issues, &lt;code&gt;MemoryService&lt;/code&gt; isolates the provider, repositories isolate persistence, and strategies make memory-on versus memory-off behavior explicit.&lt;/p&gt;

&lt;p&gt;The result is a support agent that does more than remember previous text. It accumulates operational knowledge while keeping that knowledge scoped to the people and problems it belongs to.&lt;/p&gt;

&lt;p&gt;That is the part of agent memory I think is worth designing carefully: not how much an agent can remember, but whether it remembers the &lt;strong&gt;right things&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>python</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
