<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: TejaAndhoju</title>
    <description>The latest articles on DEV Community by TejaAndhoju (@tejaandhoju).</description>
    <link>https://dev.to/tejaandhoju</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147583%2F8b7b86cc-82de-4aef-b99e-b93d0d467105.png</url>
      <title>DEV Community: TejaAndhoju</title>
      <link>https://dev.to/tejaandhoju</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/tejaandhoju"/>
    <language>en</language>
    <item>
      <title>Hindsight Recall Was Easy. Deciding What to Retain Wasn't.</title>
      <dc:creator>TejaAndhoju</dc:creator>
      <pubDate>Mon, 28 Sep 2026 16:24:21 +0000</pubDate>
      <link>https://dev.to/tejaandhoju/hindsight-recall-was-easy-deciding-what-to-retainwasnt-52mc</link>
      <guid>https://dev.to/tejaandhoju/hindsight-recall-was-easy-deciding-what-to-retainwasnt-52mc</guid>
      <description>&lt;p&gt;My sales agent told a stalled CFO deal to try milestone-based payments. It got that idea from a different company's closed deal, weeks earlier, which nobody had mentioned in the conversation. Recall worked on the first try. What took me a lot longer was deciding what the agent should be allowed to remember.&lt;/p&gt;

&lt;p&gt;This is a write-up of how I wired &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Hindsight agent memory&lt;/a&gt; into a FastAPI sales assistant, and the design mistake I only noticed when I read my own save endpoint out loud.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;DealMind is a small assistant for B2B sales reps. A rep opens a deal that has stalled, asks "what should we do?", and gets a strategy back. The pieces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;SQLite&lt;/strong&gt; holds the CRM: 20 sample deals with stage, value, contact role, blocker, and a &lt;code&gt;success_reason&lt;/code&gt; for the ones that closed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;FastAPI&lt;/strong&gt; exposes &lt;code&gt;/api/chat&lt;/code&gt; as a server-sent-events stream, plus endpoints to save a strategy and update a deal's outcome.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hindsight&lt;/strong&gt; is the memory layer. I use its Python client to &lt;code&gt;recall&lt;/code&gt; past strategies and &lt;code&gt;retain&lt;/code&gt; new ones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;An LLM on Groq&lt;/strong&gt; (&lt;code&gt;openai/gpt-oss-120b&lt;/code&gt;) writes the final answer from the CRM context plus whatever memory came back.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The CRM answers "what is this deal?" Hindsight answers "have we solved this before?" Those are different questions, and keeping them in different systems turned out to matter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just put the closed deals in the prompt?
&lt;/h2&gt;

&lt;p&gt;I could have pasted every &lt;code&gt;success_reason&lt;/code&gt; into the system prompt. With 20 deals that works. With 2,000 it doesn't, and even at 20 it hides the real problem: the model gets no signal about which past win is relevant to &lt;em&gt;this&lt;/em&gt; blocker.&lt;/p&gt;

&lt;p&gt;Hindsight does the matching. I hand it a query and it returns the memories that are semantically closest. I never write the "if blocker mentions payment terms, look up the payment playbook" logic myself, which is the part I would have gotten wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recall: the blocker is the query
&lt;/h2&gt;

&lt;p&gt;The first design decision was what to search on. The user's question ("what should we do?") is nearly useless as a query. The useful text is the deal's blocker, so I append it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;search_q&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;query&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;active_deal&lt;/span&gt; &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;active_deal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;blocker&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;search_q&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;active_deal&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;blocker&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;

&lt;span class="n"&gt;memories&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;search_hindsight&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;search_q&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;limit&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Stark Industries the blocker in the CRM is "Pushing back hard on upfront costs. Wants Net-90." The call underneath is short:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;api&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_hindsight_api&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RecallRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;api&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall_memories&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;recall_request&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;authorization&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;HINDSIGHT_API_KEY&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The memory bank contains a strategy written in completely different words, from a company that isn't Stark:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;When a prospect pushes back hard on upfront costs and wants Net-90 terms (like Oscorp), structure a milestone-based payment plan tied to deployment phases.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;No keyword overlap on the company name, no shared deal ID. It matched on meaning. The recalled text goes into the system prompt under a "what worked in the past" heading, and the model tailors it to the current deal.&lt;/p&gt;

&lt;p&gt;I also stream a visible "memory match" message to the UI before the LLM starts, so the rep can see the recalled playbook itself and not only the model's rewrite of it. In a sales tool, being able to check the source matters more than a polished paragraph.&lt;/p&gt;

&lt;h2&gt;
  
  
  The before and after
&lt;/h2&gt;

&lt;p&gt;Without memory, the prompt falls back to rule 2 in my system prompt: "generate a fresh, highly tactical market-based sales solution." You get the advice any LLM gives about discounts and negotiation. It's fine and interchangeable.&lt;/p&gt;

&lt;p&gt;With memory, the model has a concrete precedent, including the mechanism (payments tied to deployment phases) and why it worked (it removed the upfront risk). The output stops being generic and starts referencing something the company has actually done.&lt;/p&gt;

&lt;p&gt;I haven't run a controlled comparison, so I'm not going to put a number on "better." What I can say is that the two answers are visibly different, and the second one is traceable to a specific memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Retain: where I got it wrong
&lt;/h2&gt;

&lt;p&gt;Retaining is symmetrical to recalling, and that symmetry made it look easy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;content_text&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Novel solution successfully applied for &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;deal_company&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;solution&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;MemoryItem&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;Content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;actual_instance&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;content_text&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="n"&gt;req&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;RetainRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;api&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain_memories&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;BANK_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;retain_request&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;authorization&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The UI shows an "Accept Strategy &amp;amp; Save to Hindsight" button after the model proposes a fix for a stalled deal. Clicking it retains the text and moves the deal to a "Strategy Applied" stage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;conn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;execute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;UPDATE deals SET stage = &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Strategy Applied&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;, success_reason = ? WHERE company = ?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;solution_text&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;deal_company&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that back. The memory says "successfully applied." The moment of saving is when a rep &lt;em&gt;agrees with&lt;/em&gt; the advice, not when the deal closes. I'm storing a hypothesis under a label that says it's a result.&lt;/p&gt;

&lt;p&gt;That is a real bug in how the memory bank behaves over time. If a rep accepts three plausible-sounding strategies and two of the deals die anyway, the next rep who asks a similar question gets those strategies recalled with the same weight as the ones that actually closed. Recall doesn't know the difference, because I never told the bank.&lt;/p&gt;

&lt;p&gt;There's already an &lt;code&gt;/api/update_outcome&lt;/code&gt; endpoint that marks a deal Closed Won or Closed Lost. It updates SQLite and never touches Hindsight. The fix I'm working toward is to split retention into two events:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;On acceptance, retain the strategy as &lt;em&gt;proposed&lt;/em&gt;, clearly labeled.&lt;/li&gt;
&lt;li&gt;On outcome, retain a second memory that links the result to the strategy: worked, or didn't, and for which deal.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Then a recall for a payment-terms blocker can return the playbook along with its track record.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cold start
&lt;/h2&gt;

&lt;p&gt;An empty memory bank gives you nothing to recall, and a brand-new deployment is empty. I wrote a small script that retains four playbooks drawn from the sample deals that had closed, covering upfront costs, migration downtime, vendor lock-in, and on-prem requirements. It's honest seeding: those strategies came from the CRM data, not from thin air. But anyone reading this should know that the first recall in a fresh install works because I planted the memory, not because the agent learned it. Real learning starts after that first batch of accepted, resolved deals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Things that bit me
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Unclosed HTTP sessions.&lt;/strong&gt; The generated client uses &lt;code&gt;aiohttp&lt;/code&gt;, and I was creating a new client per request without closing it. Sockets leaked until I added an explicit &lt;code&gt;await api.api_client.close()&lt;/code&gt; on both the success and error paths. If you build on any async SDK, check this first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent failures hide missing memory.&lt;/strong&gt; My first integration swallowed an import error, so the agent was answering without memory and looked fine. Now &lt;code&gt;search_hindsight&lt;/code&gt; logs a warning when the key is missing, and the UI shows whether memories were found. An agent that quietly loses its memory is the worst failure mode because the output still sounds confident.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sample data resets on restart.&lt;/strong&gt; My database initializer reloads the CRM every time the app starts, so local stage changes disappear. Hindsight persists, the SQLite copy doesn't. That inconsistency is next on my list.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Lessons
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Query with the problem, not the question.&lt;/strong&gt; The user's words were the weakest search text I had. The blocker from the CRM was the strongest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show the recalled memory, not just the answer.&lt;/strong&gt; Reps trust a strategy more when they can read the precedent behind it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retain outcomes, not agreement.&lt;/strong&gt; "The user liked this advice" and "this advice worked" are different facts. Store them separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be upfront about seeding.&lt;/strong&gt; A pre-loaded bank is fine for cold start, but say it, and don't present it as learning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make missing memory loud.&lt;/strong&gt; Log it and surface it, so you notice when the agent is running blind.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The code is on &lt;a href="https://github.com/TejaAndhoju/DealMind" rel="noopener noreferrer"&gt;GitHub&lt;/a&gt;. If you're deciding whether to add memory to an agent, start with &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt; and read the &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;documentation&lt;/a&gt;. Then spend more time than you expect on what gets written into the bank, because the retrieval side will mostly take care of itself.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
