<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Bhavani Prasanna</title>
    <description>The latest articles on DEV Community by Bhavani Prasanna (@bhavani_prasanna_b123d32f).</description>
    <link>https://dev.to/bhavani_prasanna_b123d32f</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4149929%2F50b02adf-d901-43cc-af82-6205f93e9dbc.png</url>
      <title>DEV Community: Bhavani Prasanna</title>
      <link>https://dev.to/bhavani_prasanna_b123d32f</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/bhavani_prasanna_b123d32f"/>
    <language>en</language>
    <item>
      <title>I Built an AI Support Agent That Remembers What Worked</title>
      <dc:creator>Bhavani Prasanna</dc:creator>
      <pubDate>Tue, 29 Sep 2026 14:24:56 +0000</pubDate>
      <link>https://dev.to/bhavani_prasanna_b123d32f/i-built-an-ai-support-agent-that-remembers-what-worked-4o1c</link>
      <guid>https://dev.to/bhavani_prasanna_b123d32f/i-built-an-ai-support-agent-that-remembers-what-worked-4o1c</guid>
      <description>&lt;p&gt;Customer support agents usually know what the customer is asking right now, but they often don't know what worked the last time the same customer had a similar problem.&lt;/p&gt;

&lt;p&gt;I wanted to build a support agent where previous interactions could actually influence future responses.&lt;/p&gt;

&lt;p&gt;That led me to MemoryCare, an AI customer-support system that combines screenshot understanding with persistent memory using Hindsight.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fro39et3xl3tl9tfpsr54.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fro39et3xl3tl9tfpsr54.jpeg" alt=" " width="800" height="376"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The interesting part isn't simply remembering a conversation. The useful part is remembering which previous experience is relevant to the current problem.&lt;/p&gt;

&lt;p&gt;The Problem: Every Support Conversation Starts Too Close to Zero&lt;/p&gt;

&lt;p&gt;Consider a customer who previously experienced a payment error.&lt;/p&gt;

&lt;p&gt;The support agent suggested clearing the browser cache, and the payment succeeded.&lt;/p&gt;

&lt;p&gt;Later, the same customer encounters the same error again.&lt;/p&gt;

&lt;p&gt;A normal chatbot can understand the current message, but unless the previous interaction is available in its context, it has no reason to know that clearing the cache worked previously.&lt;/p&gt;

&lt;p&gt;That can lead to repetitive troubleshooting.&lt;/p&gt;

&lt;p&gt;MemoryCare approaches this differently:&lt;/p&gt;

&lt;p&gt;Current Customer Issue&lt;br&gt;
        ↓&lt;br&gt;
Understand the Issue&lt;br&gt;
        ↓&lt;br&gt;
Recall Relevant Customer Memory&lt;br&gt;
        ↓&lt;br&gt;
Generate Personalized Support&lt;br&gt;
        ↓&lt;br&gt;
Resolve the Case&lt;br&gt;
        ↓&lt;br&gt;
Retain the New Experience&lt;br&gt;
        ↓&lt;br&gt;
Use It During Future Support&lt;/p&gt;

&lt;p&gt;The persistent memory layer is powered by Hindsight.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;Why Hindsight Became the Important Part&lt;br&gt;
*&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feore847fkm9ntku8yclz.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Feore847fkm9ntku8yclz.jpeg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
I didn't want to simply attach a database containing old conversations to the chatbot.&lt;/p&gt;

&lt;p&gt;The useful question is:&lt;/p&gt;

&lt;p&gt;“What previous experience is actually relevant to this problem?”&lt;/p&gt;

&lt;p&gt;MemoryCare uses Hindsight for that Recall → Response → Retain loop.&lt;/p&gt;

&lt;p&gt;For a current customer issue, the system sends a meaningful query to the customer's memory bank and retrieves relevant memories.&lt;/p&gt;

&lt;p&gt;A simplified version of the Recall implementation looks like this:&lt;/p&gt;

&lt;p&gt;def recall(query):&lt;br&gt;
    url = f"{BASE_URL}/v1/default/banks/{BANK_ID}/memories/recall"&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payload = {
    "query": query,
    "max_tokens": 1000,
    "budget": "low"
}

response = requests.post(
    url,
    headers=HEADERS,
    json=payload,
    timeout=60
)

response.raise_for_status()

data = response.json()

return data.get("results", [])
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;The important part here is that the agent doesn't need to know the answer in advance. It asks the memory layer for context relevant to the current issue.&lt;/p&gt;

&lt;p&gt;Recall Before Responding&lt;/p&gt;

&lt;p&gt;The support agent follows a simple sequence:&lt;/p&gt;

&lt;p&gt;from memory import recall&lt;br&gt;
from llm import generate_response&lt;/p&gt;

&lt;p&gt;def support_agent(user_message):&lt;br&gt;
    memories = recall(user_message)&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;response = generate_response(
    user_message=user_message,
    memories=memories
)

return response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;This creates a clear separation:&lt;/p&gt;

&lt;p&gt;The customer provides the current problem.&lt;br&gt;
Hindsight retrieves relevant previous experience.&lt;br&gt;
The response generator combines the current issue with that context.&lt;br&gt;
The customer receives a response based on more than the current message.&lt;/p&gt;

&lt;p&gt;This is where persistent memory changes the behavior of the support agent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Concrete Example&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz3i8jbedwgf398kwp0eb.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz3i8jbedwgf398kwp0eb.jpeg" alt=" " width="800" height="438"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Suppose the customer reports:&lt;/p&gt;

&lt;p&gt;My payment is failing again. Can you check what happened last time?&lt;/p&gt;

&lt;p&gt;The previous support case contained:&lt;/p&gt;

&lt;p&gt;Customer: Bhavani&lt;/p&gt;

&lt;p&gt;Issue: Payment failed&lt;br&gt;
Error: 4032&lt;br&gt;
Product: MemoryCare Mobile App&lt;/p&gt;

&lt;p&gt;Previous troubleshooting:&lt;br&gt;
Clear browser cache&lt;/p&gt;

&lt;p&gt;Outcome:&lt;br&gt;
Payment succeeded after clearing the browser cache.&lt;/p&gt;

&lt;p&gt;When the current issue matches that previous experience, Hindsight can return the relevant memory.&lt;/p&gt;

&lt;p&gt;The response can then use that context:&lt;/p&gt;

&lt;p&gt;You previously experienced this error, and clearing&lt;br&gt;
the browser cache resolved it.&lt;/p&gt;

&lt;p&gt;You can try clearing the browser cache first and&lt;br&gt;
then attempt the payment again.&lt;/p&gt;

&lt;p&gt;The important behavior is not that the agent knows a predefined answer.&lt;/p&gt;

&lt;p&gt;It is that the previous successful experience becomes available when the current issue is similar.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Before Memory vs After Memory&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff5eq3rjgv9q8hrhlwcu8.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ff5eq3rjgv9q8hrhlwcu8.jpeg" alt=" " width="658" height="583"&gt;&lt;/a&gt;&lt;br&gt;
Before Memory&lt;/p&gt;

&lt;p&gt;Customer:&lt;/p&gt;

&lt;p&gt;My payment is failing again.&lt;/p&gt;

&lt;p&gt;Agent:&lt;/p&gt;

&lt;p&gt;Please check your internet connection and try again.&lt;/p&gt;

&lt;p&gt;This is generic troubleshooting.&lt;/p&gt;

&lt;p&gt;After Relevant Memory&lt;/p&gt;

&lt;p&gt;Customer:&lt;/p&gt;

&lt;p&gt;My payment is failing again.&lt;/p&gt;

&lt;p&gt;Agent:&lt;/p&gt;

&lt;p&gt;You previously experienced this error, and clearing the browser cache resolved it. You can try that first and then attempt the payment again.&lt;/p&gt;

&lt;p&gt;The second response has more context because the agent can use relevant previous experience.&lt;/p&gt;

&lt;p&gt;That is the behavior I wanted persistent memory to enable.&lt;/p&gt;

&lt;p&gt;*&lt;em&gt;Retaining the New Experience&lt;br&gt;
*&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ob6pqh8psdk129dtwix.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ob6pqh8psdk129dtwix.jpeg" alt=" " width="593" height="395"&gt;&lt;/a&gt;&lt;br&gt;
Recall is only half of the system.&lt;/p&gt;

&lt;p&gt;If the support agent learns something useful from the current interaction, that experience should become available later.&lt;/p&gt;

&lt;p&gt;MemoryCare therefore also supports retention.&lt;/p&gt;

&lt;p&gt;The simplified Retain implementation is:&lt;/p&gt;

&lt;p&gt;def remember(content):&lt;br&gt;
    url = f"{BASE_URL}/v1/default/banks/{BANK_ID}/memories"&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;payload = {
    "items": [
        {
            "content": content
        }
    ]
}

response = requests.post(
    url,
    headers=HEADERS,
    json=payload,
    timeout=60
)

response.raise_for_status()

return response.json()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;After a case is resolved, the system can store the actual customer issue, troubleshooting step, and outcome.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Customer: Bhavani&lt;br&gt;
Issue: Payment failure&lt;br&gt;
Error: 4032&lt;br&gt;
Solution: Cleared browser cache&lt;br&gt;
Outcome: Payment succeeded&lt;/p&gt;

&lt;p&gt;A future support interaction can then retrieve that experience.&lt;/p&gt;

&lt;p&gt;This creates the loop:&lt;/p&gt;

&lt;p&gt;RECALL&lt;br&gt;
  ↓&lt;br&gt;
RESPOND&lt;br&gt;
  ↓&lt;br&gt;
RESOLVE&lt;br&gt;
  ↓&lt;br&gt;
RETAIN&lt;br&gt;
  ↓&lt;br&gt;
RECALL AGAIN&lt;br&gt;
What I Learned About Agent Memory&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory is useful only when it is relevant&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Storing everything isn't enough.&lt;/p&gt;

&lt;p&gt;A support agent shouldn't use an old delivery problem when the customer is currently reporting a login failure.&lt;/p&gt;

&lt;p&gt;The current issue should determine what memory is useful.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Current evidence still matters&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Memory should complement the current customer message or screenshot, not replace it.&lt;/p&gt;

&lt;p&gt;The current issue is the primary evidence.&lt;/p&gt;

&lt;p&gt;Previous experience provides additional context.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A memory failure should not break the support agent&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Hindsight is an important component, but the application should still be able to respond using the current issue if memory retrieval fails or returns nothing useful.&lt;/p&gt;

&lt;p&gt;This makes the support flow more resilient.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retain makes the system different from a simple chatbot&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Recall gives the agent access to previous experience.&lt;/p&gt;

&lt;p&gt;Retain allows useful new experience to become part of future interactions.&lt;/p&gt;

&lt;p&gt;Without both sides, the memory loop is incomplete.&lt;/p&gt;

&lt;p&gt;What Was Difficult&lt;/p&gt;

&lt;p&gt;One of the practical challenges was making the memory layer behave like part of the support workflow instead of treating it as an independent feature.&lt;/p&gt;

&lt;p&gt;The application needs to decide:&lt;/p&gt;

&lt;p&gt;What should be recalled?&lt;br&gt;
Which memories are relevant?&lt;br&gt;
How should they influence the response?&lt;br&gt;
What should be retained after resolution?&lt;br&gt;
What should happen when no useful memory exists?&lt;/p&gt;

&lt;p&gt;These decisions matter more than simply adding a memory API call.&lt;/p&gt;

&lt;p&gt;Another important lesson was avoiding hard-coded assumptions around one support scenario. Payment failure is useful for demonstrating the workflow, but the architecture should also support other customer-support issues such as login failures, OTP problems, order issues, refunds, subscriptions, and application errors.&lt;/p&gt;

&lt;p&gt;What I Would Improve Next&lt;/p&gt;

&lt;p&gt;The current system demonstrates the persistent-memory workflow, but there are several directions I would take it further.&lt;/p&gt;

&lt;p&gt;I would expand evaluation across a larger set of support scenarios, improve memory relevance evaluation, add stronger observability around Recall and Retain, and integrate the system with real support-ticket workflows.&lt;/p&gt;

&lt;p&gt;The goal is not to make the agent remember everything.&lt;/p&gt;

&lt;p&gt;The goal is to make it remember the right things at the right time.&lt;/p&gt;

&lt;p&gt;Conclusion&lt;/p&gt;

&lt;p&gt;Building MemoryCare changed how I think about AI support agents.&lt;/p&gt;

&lt;p&gt;A model can generate a useful answer from the current conversation. But persistent memory creates another possibility: the agent can use previous customer experiences to make future interactions more contextual.&lt;/p&gt;

&lt;p&gt;The core loop is simple:&lt;/p&gt;

&lt;p&gt;Current Issue&lt;br&gt;
     ↓&lt;br&gt;
Hindsight Recall&lt;br&gt;
     ↓&lt;br&gt;
Personalized Response&lt;br&gt;
     ↓&lt;br&gt;
Resolution&lt;br&gt;
     ↓&lt;br&gt;
Hindsight Retain&lt;br&gt;
     ↓&lt;br&gt;
Better Context for the Next Interaction&lt;/p&gt;

&lt;p&gt;For me, the most interesting part isn't making an agent remember everything.&lt;/p&gt;

&lt;p&gt;It is making previous experience useful when it actually matters.&lt;/p&gt;

&lt;p&gt;Learn More&lt;/p&gt;

&lt;p&gt;Hindsight:&lt;/p&gt;

&lt;p&gt;Hindsight GitHub&lt;br&gt;
Hindsight Documentation&lt;br&gt;
Vectorize Agent Memory&lt;/p&gt;

&lt;p&gt;Hindsight links&lt;/p&gt;

&lt;h2&gt;
  
  
  Learn More
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/vectorize-io/hindsight" rel="noopener noreferrer"&gt;Hindsight GitHub&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://hindsight.vectorize.io/" rel="noopener noreferrer"&gt;Hindsight Documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://vectorize.io/what-is-agent-memory" rel="noopener noreferrer"&gt;Vectorize Agent Memory&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
