<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sreeja</title>
    <description>The latest articles on DEV Community by Sreeja (@ysreeja_12_881a6fe10bd391).</description>
    <link>https://dev.to/ysreeja_12_881a6fe10bd391</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147238%2F4feb79af-8f08-4bba-b842-03d4fd5afb83.png</url>
      <title>DEV Community: Sreeja</title>
      <link>https://dev.to/ysreeja_12_881a6fe10bd391</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ysreeja_12_881a6fe10bd391"/>
    <language>en</language>
    <item>
      <title>Swapping Hardcoded Customer Context for Hindsight Recall</title>
      <dc:creator>Sreeja</dc:creator>
      <pubDate>Mon, 28 Sep 2026 13:57:38 +0000</pubDate>
      <link>https://dev.to/ysreeja_12_881a6fe10bd391/swapping-hardcoded-customer-context-for-hindsight-recall-3b33</link>
      <guid>https://dev.to/ysreeja_12_881a6fe10bd391/swapping-hardcoded-customer-context-for-hindsight-recall-3b33</guid>
      <description>&lt;p&gt;Swapping Hardcoded Customer Context for Hindsight Recall&lt;/p&gt;

&lt;p&gt;The first version of my support agent had a great memory, as long as I typed it in by hand. Every request carried a &lt;code&gt;customer_context&lt;/code&gt; field, and every test passed because I was the one writing the context. Then I asked the obvious question: who fills that field in production?&lt;/p&gt;

&lt;p&gt;This is the story of replacing that field with real memory, using &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight&lt;/a&gt;, and of the design decisions that mattered more than the swap itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the system does
&lt;/h2&gt;

&lt;p&gt;The project is a customer support agent. A customer sends a message, and the agent works out what kind of problem it is, decides whether a human needs to get involved, pulls in what it knows about that customer, and produces a response. The stack is deliberately plain: Python, FastAPI, and a local Llama 3.2 model served by Ollama. There is no hosted LLM API and no API key to manage. Everything runs on hardware I control, which matters when the input is support tickets full of account details.&lt;/p&gt;

&lt;p&gt;Every message moves through a fixed pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer Message
        ↓
Intent Detection
        ↓
Escalation Check
        ↓
Customer Context
        ↓
Ollama AI Agent
        ↓
Support Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Intent detection sorts messages into six categories: &lt;code&gt;PAYMENT&lt;/code&gt;, &lt;code&gt;LOGIN&lt;/code&gt;, &lt;code&gt;ACCOUNT&lt;/code&gt;, &lt;code&gt;SUBSCRIPTION&lt;/code&gt;, &lt;code&gt;TECHNICAL&lt;/code&gt;, and &lt;code&gt;OTHER&lt;/code&gt;. The escalation check decides whether the model should be involved at all. Only after those two steps does the customer context get assembled and handed to Llama 3.2.&lt;/p&gt;

&lt;p&gt;Here is the same flow with Hindsight in place. Recall sits in the request path, retain sits at the end, and both talk to a per-customer bank:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgawa6k2glfxphvuup7a7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fgawa6k2glfxphvuup7a7.png" alt="Architecture diagram" width="800" height="367"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The model side is unremarkable, and that's the point. Llama 3.2 runs locally through Ollama, and a quick &lt;code&gt;ollama run llama3.2&lt;/code&gt; is all it takes to confirm it's answering before I involve any of my own code:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5d8hcvesnmpu2blmacz5.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5d8hcvesnmpu2blmacz5.png" alt="Ollama Running in terminal" width="798" height="170"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The API surface is a single endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;POST&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;/chat&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"customer_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"C001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"My payment failed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"customer_context"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That &lt;code&gt;customer_context&lt;/code&gt; field is the subject of this article.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with context you pass in
&lt;/h2&gt;

&lt;p&gt;In the early version, the caller supplied the context:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"previous_issue"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Payment failed before"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"previous_solution"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Retrying the payment resolved the issue"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This was useful for proving the prompt worked. With that context present, the agent stopped suggesting generic fixes and started referring to what had already been tried. The tests I wrote (payment issues, login issues, escalation, customer context) all passed.&lt;/p&gt;

&lt;p&gt;But the design had a flaw that no test could catch. The frontend, or whatever sits in front of the agent, would have to know what is relevant about a customer, fetch it, format it, and send it on every request. That pushes the hardest part of the problem onto the caller. It also means the agent itself has no memory. It only has whatever the caller remembered to send.&lt;/p&gt;

&lt;p&gt;I wanted the agent to own this. The request should need only a &lt;code&gt;customer_id&lt;/code&gt; and a &lt;code&gt;message&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I used a memory layer instead of a table
&lt;/h2&gt;

&lt;p&gt;My first instinct was a Postgres table: &lt;code&gt;customer_id&lt;/code&gt;, &lt;code&gt;previous_issue&lt;/code&gt;, &lt;code&gt;previous_solution&lt;/code&gt;. That works until you notice what the fields imply. Support history isn't a fixed schema. One customer's history is "card declined, retry fixed it." Another's is "locked out after a phone change, verified by email, asked twice about downgrading." Forcing that into two columns means deciding up front what matters, and I wouldn't know until I saw real tickets.&lt;/p&gt;

&lt;p&gt;What I actually needed was to store conversation outcomes as they happened, and later retrieve the ones relevant to the current message. That is what &lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;agent memory&lt;/a&gt; is for, and it's why I looked at Hindsight instead of building retrieval on top of a database myself. It's open source, so I could read how it works and run it next to the rest of the stack. The &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight documentation&lt;/a&gt; covers the model well, and there is a Python client, &lt;code&gt;hindsight-client&lt;/code&gt;, which is already in my &lt;code&gt;requirements.txt&lt;/code&gt; alongside &lt;code&gt;fastapi&lt;/code&gt;, &lt;code&gt;uvicorn&lt;/code&gt;, &lt;code&gt;python-dotenv&lt;/code&gt;, and &lt;code&gt;requests&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The through-line: one memory bank per customer
&lt;/h2&gt;

&lt;p&gt;The decision that shaped everything else was scoping. Hindsight organizes memory into banks. I create one bank per customer, keyed by the same &lt;code&gt;customer_id&lt;/code&gt; the API already receives.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;hindsight_client&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Hindsight&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Hindsight&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HINDSIGHT_URL&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;bank_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;customer-&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives me isolation for free. When the agent recalls memory for &lt;code&gt;C001&lt;/code&gt;, it queries only &lt;code&gt;C001&lt;/code&gt;'s bank. There is no filtering step that could be forgotten, and no shared index where one customer's payment failure could surface in another customer's reply. For a support system, a cross-customer leak is the worst class of bug, and I preferred to make it structurally impossible rather than something I had to test for.&lt;/p&gt;

&lt;p&gt;Retrieval replaces the field that used to arrive in the request body:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_customer_context&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;recall&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;bank_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;format_memories&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The query is the customer's current message. If someone writes "my payment failed again," the recall is driven by that text, so payment-related history ranks above an old login problem. That is the behavior I was faking by hand before, except now it selects from real history instead of a string I typed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where recall sits in the pipeline
&lt;/h2&gt;

&lt;p&gt;Placement mattered. Recall happens after the escalation check, not before.&lt;/p&gt;

&lt;p&gt;The reasoning: escalation is a decision about the message and the category, and it should not depend on a network call to a memory service. If a message needs a human, I want that path to be fast and to work even if the memory service is down. Recall only runs on the path where the model will actually use it. That keeps the failure modes separate. A memory outage degrades the agent into a competent but generic support bot. It doesn't take escalation with it.&lt;/p&gt;

&lt;p&gt;The empty case needed the same care. A new customer has an empty bank, and recall returns nothing. That is a normal condition, not an error. The context block becomes a plain statement that there is no prior history, and the prompt is written so the model doesn't invent any. This was the first thing I tested after the swap, because a model handed an empty "history" section will sometimes fill the gap with confident nonsense.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing memory back
&lt;/h2&gt;

&lt;p&gt;Recall is only half of it. After the agent replies and the outcome is known, I write it back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;remember_outcome&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;resolution&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;retain&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;bank_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;bank_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;customer_id&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Customer reported: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;. &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Outcome: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;resolution&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I write the outcome, not the raw transcript. The useful sentence is "payment failed, retrying resolved it," not forty lines of greetings. That is the same information as the old hardcoded &lt;code&gt;previous_issue&lt;/code&gt; and &lt;code&gt;previous_solution&lt;/code&gt; fields, except it now accumulates on its own, and I don't have to decide the schema ahead of time.&lt;/p&gt;

&lt;p&gt;There's a judgment call here about escalated tickets. When a case goes to a human, the agent doesn't know how it ended. I record that it was escalated and why, so the next conversation can acknowledge that a person is already handling it rather than starting over.&lt;/p&gt;

&lt;h2&gt;
  
  
  Behavior in practice
&lt;/h2&gt;

&lt;p&gt;The clearest way I found to see what memory changes is to send nearly the same message twice. These two captures are from the FastAPI &lt;code&gt;/docs&lt;/code&gt; page, from the stage where I still passed the context by hand. That is exactly the payload recall now produces on its own, so they show the effect cleanly.&lt;/p&gt;

&lt;p&gt;First, a customer with history. The context says a previous payment failure was fixed by retrying, and the message is "My payment failed again":&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvnlpxkntv9ob0t4uap4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsvnlpxkntv9ob0t4uap4.png" alt="FastAPI docs showing a POST /chat request with previous_issue and previous_solution in customer_context, and a response that references the earlier failure and the retry that fixed it" width="800" height="724"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The response acknowledges that this has happened before, mentions that retrying resolved it last time, and then asks what the customer has already tried and whether their payment method or account details changed. It reads like someone who has seen this customer's file.&lt;/p&gt;

&lt;p&gt;Now a customer with nothing to recall. Same endpoint, same customer ID, empty context, message "My payment failed":&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftrl3aknyalwo903qwyhj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftrl3aknyalwo903qwyhj.png" alt="FastAPI docs showing a POST /chat request with an empty customer_context and a generic response asking which payment method was used and why the payment might have failed" width="800" height="750"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The intent is still classified as &lt;code&gt;PAYMENT&lt;/code&gt; and the action is still &lt;code&gt;ASSIST&lt;/code&gt;, but the reply is a generic intake: which payment method, was the card expired, were there insufficient funds. Nothing in it claims a history that doesn't exist.&lt;/p&gt;

&lt;p&gt;Same model, same prompt template, same intent. The only difference is the recalled context. With Hindsight in the loop, the request body shrinks to &lt;code&gt;customer_id&lt;/code&gt; and &lt;code&gt;message&lt;/code&gt;, and the first response is what the agent produces when the bank has a relevant memory. The second is what it produces when the bank is empty.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell someone doing the same thing
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Move context ownership into the agent.&lt;/strong&gt; If the caller has to assemble memory, you've built a prompt template, not an agent. The request should carry identity and intent; everything else is the agent's job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scope memory by the identifier you already have.&lt;/strong&gt; One bank per &lt;code&gt;customer_id&lt;/code&gt; meant isolation was a property of the data layout rather than a filter I had to get right on every query.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Decide where a dependency sits in the failure path.&lt;/strong&gt; Putting recall after escalation means a memory outage can only degrade answer quality. It can't block the paths that need to work no matter what.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat empty memory as a first-class case.&lt;/strong&gt; The new-customer path is the one most likely to produce fabricated history. Write the prompt for it explicitly and test it before anything else.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Store outcomes, not transcripts.&lt;/strong&gt; What a future conversation needs is what happened and what fixed it. Short, factual entries recall better than raw chat logs, and they read cleanly when injected into a prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it ended up
&lt;/h2&gt;

&lt;p&gt;The pipeline diagram barely changed. It gained one arrow: between the escalation check and the model, customer context now comes from Hindsight recall instead of the request body. The &lt;code&gt;customer_context&lt;/code&gt; field went from load-bearing to unnecessary, and the frontend got simpler because it stopped needing to know anything about customer history.&lt;/p&gt;

&lt;p&gt;If you're building a similar agent, the &lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight repository&lt;/a&gt; is the place to start, and the &lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;docs&lt;/a&gt; explain banks, retain, and recall in more detail than I can here. The most useful thing I got out of this wasn't a feature. It was the ability to stop hand-writing what my agent was supposed to remember.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>python</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
