<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ivan Kamal Thuppari</title>
    <description>The latest articles on DEV Community by Ivan Kamal Thuppari (@ivan_kamalthuppari_a5a0d).</description>
    <link>https://dev.to/ivan_kamalthuppari_a5a0d</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147826%2Ffb9f8af4-f13a-41b8-a0e1-773b4f84e3bd.png</url>
      <title>DEV Community: Ivan Kamal Thuppari</title>
      <link>https://dev.to/ivan_kamalthuppari_a5a0d</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ivan_kamalthuppari_a5a0d"/>
    <language>en</language>
    <item>
      <title>How Hindsight Changed the Way My Agent Handles Support</title>
      <dc:creator>Ivan Kamal Thuppari</dc:creator>
      <pubDate>Mon, 28 Sep 2026 18:25:26 +0000</pubDate>
      <link>https://dev.to/ivan_kamalthuppari_a5a0d/how-hindsight-changed-the-way-my-agent-handles-support-292a</link>
      <guid>https://dev.to/ivan_kamalthuppari_a5a0d/how-hindsight-changed-the-way-my-agent-handles-support-292a</guid>
      <description>&lt;p&gt;Every time a customer came back with the same problem, I wanted my support agent to know what had already happened.&lt;/p&gt;

&lt;p&gt;Instead, a typical conversational agent starts with the latest message and tries to solve it from scratch. That works for isolated questions, but customer support is rarely isolated. Troubleshooting is a process, and the history of that process matters.&lt;/p&gt;

&lt;p&gt;That was the problem I wanted to solve with RecallDesk: give a customer support agent persistent memory so that previous interactions become useful context for future ones.&lt;/p&gt;

&lt;p&gt;The problem with starting from zero&lt;/p&gt;

&lt;p&gt;Consider a customer named Sarah.&lt;/p&gt;

&lt;p&gt;She previously reported that her dashboard was freezing when uploading large CSV files. The support process already tried clearing the browser cache and reducing the file size.&lt;/p&gt;

&lt;p&gt;Later, Sarah returns:&lt;/p&gt;

&lt;p&gt;"Hi, I'm still having the dashboard freezing problem."&lt;/p&gt;

&lt;p&gt;That message alone doesn't tell the agent what happened previously.&lt;/p&gt;

&lt;p&gt;A stateless agent could easily suggest clearing the cache again. A human support representative, however, would ideally look at the previous interaction first.&lt;/p&gt;

&lt;p&gt;The difference is simple:&lt;/p&gt;

&lt;p&gt;Without memory:&lt;/p&gt;

&lt;p&gt;Customer message&lt;br&gt;
       ↓&lt;br&gt;
      LLM&lt;br&gt;
       ↓&lt;br&gt;
New troubleshooting&lt;/p&gt;

&lt;p&gt;With memory:&lt;/p&gt;

&lt;p&gt;Customer message&lt;br&gt;
       ↓&lt;br&gt;
Retrieve relevant history&lt;br&gt;
       ↓&lt;br&gt;
      LLM&lt;br&gt;
       ↓&lt;br&gt;
Next troubleshooting step&lt;/p&gt;

&lt;p&gt;I wanted RecallDesk to follow the second path.&lt;/p&gt;

&lt;p&gt;The architecture&lt;/p&gt;

&lt;p&gt;RecallDesk is intentionally small. The application consists of a Streamlit interface, a memory layer using Hindsight, and an AI agent using Groq.&lt;/p&gt;

&lt;p&gt;The basic flow is:&lt;/p&gt;

&lt;p&gt;Customer&lt;br&gt;
   │&lt;br&gt;
   ▼&lt;br&gt;
Streamlit UI&lt;br&gt;
   │&lt;br&gt;
   ▼&lt;br&gt;
Hindsight Recall&lt;br&gt;
   │&lt;br&gt;
   ▼&lt;br&gt;
Relevant customer history&lt;br&gt;
   │&lt;br&gt;
   ▼&lt;br&gt;
Groq AI Agent&lt;br&gt;
   │&lt;br&gt;
   ▼&lt;br&gt;
Context-aware response&lt;br&gt;
   │&lt;br&gt;
   ▼&lt;br&gt;
Hindsight Retain&lt;/p&gt;

&lt;p&gt;The important part isn't just retrieving information.&lt;/p&gt;

&lt;p&gt;The system also stores the new interaction after generating the response.&lt;/p&gt;

&lt;p&gt;That creates a continuous loop:&lt;/p&gt;

&lt;p&gt;RECALL → RESPOND → REMEMBER&lt;/p&gt;

&lt;p&gt;The next customer interaction can then use what happened this time.&lt;/p&gt;

&lt;p&gt;Using Hindsight as the memory layer&lt;/p&gt;

&lt;p&gt;I kept the Hindsight integration isolated in memory.py.&lt;/p&gt;

&lt;p&gt;The application creates a memory bank for RecallDesk:&lt;/p&gt;

&lt;p&gt;BANK_ID = "recall-desk"&lt;/p&gt;

&lt;p&gt;client = Hindsight(&lt;br&gt;
    api_key=API_KEY,&lt;br&gt;
    base_url="&lt;a href="https://api.hindsight.vectorize.io" rel="noopener noreferrer"&gt;https://api.hindsight.vectorize.io&lt;/a&gt;"&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;When an interaction needs to be stored, RecallDesk calls remember():&lt;/p&gt;

&lt;p&gt;def remember(conversation):&lt;br&gt;
    client.retain(&lt;br&gt;
        bank_id=BANK_ID,&lt;br&gt;
        content=conversation,&lt;br&gt;
        context="Customer support conversation"&lt;br&gt;
    )&lt;/p&gt;

&lt;p&gt;The important distinction is that I'm not manually maintaining a giant conversation history inside the application.&lt;/p&gt;

&lt;p&gt;Instead, the interaction is retained in Hindsight and can later be recalled when it becomes relevant.&lt;/p&gt;

&lt;p&gt;For retrieval, RecallDesk uses:&lt;/p&gt;

&lt;p&gt;def recall(query):&lt;br&gt;
    results = client.recall(&lt;br&gt;
        bank_id=BANK_ID,&lt;br&gt;
        query=query&lt;br&gt;
    )&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;memories = []

for result in results:
    memories.append(str(result))

return memories
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;This gives the agent a way to search its previous support context rather than treating every request as completely independent.&lt;/p&gt;

&lt;p&gt;Turning memory into useful context&lt;/p&gt;

&lt;p&gt;Retrieving memories isn't enough by itself.&lt;/p&gt;

&lt;p&gt;The agent needs to understand what those memories mean for the current request.&lt;/p&gt;

&lt;p&gt;That's handled in agent.py.&lt;/p&gt;

&lt;p&gt;First, RecallDesk creates a query containing the customer and current issue:&lt;/p&gt;

&lt;p&gt;query = f"""&lt;br&gt;
Customer: {customer_name}&lt;/p&gt;

&lt;p&gt;Current issue:&lt;br&gt;
{customer_message}&lt;br&gt;
"""&lt;/p&gt;

&lt;p&gt;Then it recalls relevant history:&lt;/p&gt;

&lt;p&gt;memories = recall(query)&lt;br&gt;
memory_text = "\n".join(str(memory) for memory in memories[:4])&lt;/p&gt;

&lt;p&gt;I deliberately keep the retrieved context bounded before sending it to the language model. This matters because the amount of historical information can otherwise make the model request unnecessarily large.&lt;/p&gt;

&lt;p&gt;The retrieved history is then supplied to the model alongside the current customer message.&lt;/p&gt;

&lt;p&gt;The agent's system instructions also explicitly tell it how to use that information:&lt;/p&gt;

&lt;p&gt;SYSTEM_PROMPT = """&lt;br&gt;
You are RecallDesk, an AI customer support agent.&lt;/p&gt;

&lt;p&gt;Use relevant customer history from Hindsight to help answer the customer.&lt;/p&gt;

&lt;p&gt;Rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use remembered information when relevant.&lt;/li&gt;
&lt;li&gt;Do not invent memories.&lt;/li&gt;
&lt;li&gt;Do not ask for information already known.&lt;/li&gt;
&lt;li&gt;Acknowledge previous failed troubleshooting.&lt;/li&gt;
&lt;li&gt;Suggest the next reasonable troubleshooting step.&lt;/li&gt;
&lt;li&gt;Be concise and professional.
"""&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last part is important.&lt;/p&gt;

&lt;p&gt;Memory shouldn't mean that the model blindly repeats everything it finds. The retrieved information is context, and the agent still has to decide what is relevant to the current problem.&lt;/p&gt;

&lt;p&gt;The retain → recall loop&lt;/p&gt;

&lt;p&gt;After the response is generated, RecallDesk stores the interaction:&lt;/p&gt;

&lt;p&gt;remember(&lt;br&gt;
    f"""&lt;br&gt;
Customer: {customer_name}&lt;/p&gt;

&lt;p&gt;Customer message:&lt;br&gt;
{customer_message}&lt;/p&gt;

&lt;p&gt;RecallDesk response:&lt;br&gt;
{answer}&lt;br&gt;
"""&lt;br&gt;
)&lt;/p&gt;

&lt;p&gt;This is what makes the system different from simply attaching a static knowledge base to an LLM.&lt;/p&gt;

&lt;p&gt;The agent's own interactions become future context.&lt;/p&gt;

&lt;p&gt;A simplified version of the lifecycle is:&lt;/p&gt;

&lt;p&gt;Interaction #1&lt;/p&gt;

&lt;p&gt;Customer → "Dashboard freezes with large CSV"&lt;br&gt;
                ↓&lt;br&gt;
          Troubleshooting&lt;br&gt;
                ↓&lt;br&gt;
             RETAIN&lt;/p&gt;

&lt;p&gt;Interaction #2&lt;/p&gt;

&lt;p&gt;Customer → "It is still freezing"&lt;br&gt;
                ↓&lt;br&gt;
             RECALL&lt;br&gt;
                ↓&lt;br&gt;
    Previous troubleshooting&lt;br&gt;
                ↓&lt;br&gt;
          Groq AI Agent&lt;br&gt;
                ↓&lt;br&gt;
      Next troubleshooting step&lt;br&gt;
                ↓&lt;br&gt;
             RETAIN&lt;/p&gt;

&lt;p&gt;The second interaction can therefore build on the first.&lt;/p&gt;

&lt;p&gt;What the user actually sees&lt;/p&gt;

&lt;p&gt;I wanted the memory mechanism to be visible rather than hiding everything behind the model.&lt;/p&gt;

&lt;p&gt;The RecallDesk interface has three main areas:&lt;/p&gt;

&lt;p&gt;Customer information and message&lt;br&gt;
Retrieved memory&lt;br&gt;
Generated response&lt;/p&gt;

&lt;p&gt;A support representative can enter:&lt;/p&gt;

&lt;p&gt;Customer:&lt;br&gt;
Sarah Mitchell&lt;/p&gt;

&lt;p&gt;Message:&lt;br&gt;
Hi, I'm still having the dashboard freezing problem.&lt;/p&gt;

&lt;p&gt;RecallDesk searches Hindsight and displays the retrieved history.&lt;/p&gt;

&lt;p&gt;The agent then receives that history before generating its response.&lt;/p&gt;

&lt;p&gt;In our example, the resulting response can acknowledge that previous troubleshooting was already attempted and move toward another diagnostic step instead of automatically restarting the same checklist.&lt;/p&gt;

&lt;p&gt;That's the behavior I was looking for.&lt;/p&gt;

&lt;p&gt;The useful part of memory isn't simply that the system can say "I remember Sarah."&lt;/p&gt;

&lt;p&gt;The useful part is:&lt;/p&gt;

&lt;p&gt;"I remember what we already tried, so I can decide what to do next."&lt;/p&gt;

&lt;p&gt;What I learned&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Memory is only useful when it changes behavior&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;It would be easy to add a memory database and call the project memory-enabled.&lt;/p&gt;

&lt;p&gt;That isn't enough.&lt;/p&gt;

&lt;p&gt;The important question is what the agent does differently because the memory exists.&lt;/p&gt;

&lt;p&gt;For RecallDesk, the desired behavior is straightforward: previous troubleshooting should influence the next response.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retrieval belongs between the user and the model&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The language model doesn't need every historical interaction.&lt;/p&gt;

&lt;p&gt;It needs the history that is relevant to the current problem.&lt;/p&gt;

&lt;p&gt;That's why the architecture separates retrieval from generation:&lt;/p&gt;

&lt;p&gt;Current message&lt;br&gt;
      ↓&lt;br&gt;
   Retrieve&lt;br&gt;
      ↓&lt;br&gt;
Relevant context&lt;br&gt;
      ↓&lt;br&gt;
    Generate&lt;/p&gt;

&lt;p&gt;This keeps the model's input focused on the current task.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Failed troubleshooting is valuable information&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A support interaction isn't valuable only when it contains a successful solution.&lt;/p&gt;

&lt;p&gt;Knowing that a particular troubleshooting step didn't work is also useful.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Clear cache → failed&lt;br&gt;
Reduce file size → failed&lt;/p&gt;

&lt;p&gt;That history can prevent the next response from simply repeating the same suggestions.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The memory loop should be simple&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;RecallDesk doesn't require a complicated orchestration layer.&lt;/p&gt;

&lt;p&gt;The fundamental loop is:&lt;/p&gt;

&lt;p&gt;recall()&lt;br&gt;
   ↓&lt;br&gt;
generate()&lt;br&gt;
   ↓&lt;br&gt;
remember()&lt;/p&gt;

&lt;p&gt;That simplicity makes the behavior easier to reason about and debug.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Agent memory changes what "conversation history" means&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Before building RecallDesk, I mostly thought about conversation history as messages that need to remain inside a chat session.&lt;/p&gt;

&lt;p&gt;Persistent agent memory changes that model.&lt;/p&gt;

&lt;p&gt;A useful interaction doesn't necessarily have to remain in the same session to remain useful.&lt;/p&gt;

&lt;p&gt;It can become information that the agent retrieves later when another interaction makes it relevant.&lt;/p&gt;

&lt;p&gt;Where I would take RecallDesk next&lt;/p&gt;

&lt;p&gt;The current architecture provides a foundation for a larger support intelligence system.&lt;/p&gt;

&lt;p&gt;For example, the same memory layer could eventually support incident histories, recurring customer issues, previously attempted solutions, unresolved problems, and patterns across support interactions.&lt;/p&gt;

&lt;p&gt;That could turn RecallDesk from a conversational support interface into a system that helps support teams understand what has already happened before deciding what should happen next.&lt;/p&gt;

&lt;p&gt;But the core idea would remain the same.&lt;/p&gt;

&lt;p&gt;Recall before responding. Remember after responding.&lt;/p&gt;

&lt;p&gt;That's the change Hindsight introduced to my approach to building the agent.&lt;/p&gt;

&lt;p&gt;Instead of building an assistant that only answers the current message, I built one that can use what happened before.&lt;/p&gt;

&lt;h1&gt;
  
  
  ai
&lt;/h1&gt;

&lt;h1&gt;
  
  
  python
&lt;/h1&gt;

&lt;h1&gt;
  
  
  webdev
&lt;/h1&gt;

&lt;h1&gt;
  
  
  opensource
&lt;/h1&gt;

&lt;h1&gt;
  
  
  agents
&lt;/h1&gt;

&lt;h1&gt;
  
  
  python
&lt;/h1&gt;

&lt;h1&gt;
  
  
  streamlit
&lt;/h1&gt;

&lt;h1&gt;
  
  
  hindsight
&lt;/h1&gt;

&lt;p&gt;&lt;a class="mentioned-user" href="https://dev.to/vectorize"&gt;@vectorize&lt;/a&gt; &lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flbe0py5mmysvgbjcxzlx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flbe0py5mmysvgbjcxzlx.png" alt=" " width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnqbp22bqcun7wb0lv0qy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnqbp22bqcun7wb0lv0qy.png" alt=" " width="800" height="437"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9l62e1otzakgsgt866tb.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9l62e1otzakgsgt866tb.png" alt=" " width="800" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>saas</category>
    </item>
  </channel>
</rss>
