DEV Community

Ivan Kamal Thuppari
Ivan Kamal Thuppari

Posted on

How Hindsight Changed the Way My Agent Handles Support

Every time a customer came back with the same problem, I wanted my support agent to know what had already happened.

Instead, a typical conversational agent starts with the latest message and tries to solve it from scratch. That works for isolated questions, but customer support is rarely isolated. Troubleshooting is a process, and the history of that process matters.

That was the problem I wanted to solve with RecallDesk: give a customer support agent persistent memory so that previous interactions become useful context for future ones.

The problem with starting from zero

Consider a customer named Sarah.

She previously reported that her dashboard was freezing when uploading large CSV files. The support process already tried clearing the browser cache and reducing the file size.

Later, Sarah returns:

"Hi, I'm still having the dashboard freezing problem."

That message alone doesn't tell the agent what happened previously.

A stateless agent could easily suggest clearing the cache again. A human support representative, however, would ideally look at the previous interaction first.

The difference is simple:

Without memory:

Customer message
↓
LLM
↓
New troubleshooting

With memory:

Customer message
↓
Retrieve relevant history
↓
LLM
↓
Next troubleshooting step

I wanted RecallDesk to follow the second path.

The architecture

RecallDesk is intentionally small. The application consists of a Streamlit interface, a memory layer using Hindsight, and an AI agent using Groq.

The basic flow is:

Customer
│
▼
Streamlit UI
│
▼
Hindsight Recall
│
▼
Relevant customer history
│
▼
Groq AI Agent
│
▼
Context-aware response
│
▼
Hindsight Retain

The important part isn't just retrieving information.

The system also stores the new interaction after generating the response.

That creates a continuous loop:

RECALL → RESPOND → REMEMBER

The next customer interaction can then use what happened this time.

Using Hindsight as the memory layer

I kept the Hindsight integration isolated in memory.py.

The application creates a memory bank for RecallDesk:

BANK_ID = "recall-desk"

client = Hindsight(
api_key=API_KEY,
base_url="https://api.hindsight.vectorize.io"
)

When an interaction needs to be stored, RecallDesk calls remember():

def remember(conversation):
client.retain(
bank_id=BANK_ID,
content=conversation,
context="Customer support conversation"
)

The important distinction is that I'm not manually maintaining a giant conversation history inside the application.

Instead, the interaction is retained in Hindsight and can later be recalled when it becomes relevant.

For retrieval, RecallDesk uses:

def recall(query):
results = client.recall(
bank_id=BANK_ID,
query=query
)

memories = []

for result in results:
    memories.append(str(result))

return memories
Enter fullscreen mode Exit fullscreen mode

This gives the agent a way to search its previous support context rather than treating every request as completely independent.

Turning memory into useful context

Retrieving memories isn't enough by itself.

The agent needs to understand what those memories mean for the current request.

That's handled in agent.py.

First, RecallDesk creates a query containing the customer and current issue:

query = f"""
Customer: {customer_name}

Current issue:
{customer_message}
"""

Then it recalls relevant history:

memories = recall(query)
memory_text = "\n".join(str(memory) for memory in memories[:4])

I deliberately keep the retrieved context bounded before sending it to the language model. This matters because the amount of historical information can otherwise make the model request unnecessarily large.

The retrieved history is then supplied to the model alongside the current customer message.

The agent's system instructions also explicitly tell it how to use that information:

SYSTEM_PROMPT = """
You are RecallDesk, an AI customer support agent.

Use relevant customer history from Hindsight to help answer the customer.

Rules:

  • Use remembered information when relevant.
  • Do not invent memories.
  • Do not ask for information already known.
  • Acknowledge previous failed troubleshooting.
  • Suggest the next reasonable troubleshooting step.
  • Be concise and professional. """

That last part is important.

Memory shouldn't mean that the model blindly repeats everything it finds. The retrieved information is context, and the agent still has to decide what is relevant to the current problem.

The retain → recall loop

After the response is generated, RecallDesk stores the interaction:

remember(
f"""
Customer: {customer_name}

Customer message:
{customer_message}

RecallDesk response:
{answer}
"""
)

This is what makes the system different from simply attaching a static knowledge base to an LLM.

The agent's own interactions become future context.

A simplified version of the lifecycle is:

Interaction #1

Customer → "Dashboard freezes with large CSV"
↓
Troubleshooting
↓
RETAIN

Interaction #2

Customer → "It is still freezing"
↓
RECALL
↓
Previous troubleshooting
↓
Groq AI Agent
↓
Next troubleshooting step
↓
RETAIN

The second interaction can therefore build on the first.

What the user actually sees

I wanted the memory mechanism to be visible rather than hiding everything behind the model.

The RecallDesk interface has three main areas:

Customer information and message
Retrieved memory
Generated response

A support representative can enter:

Customer:
Sarah Mitchell

Message:
Hi, I'm still having the dashboard freezing problem.

RecallDesk searches Hindsight and displays the retrieved history.

The agent then receives that history before generating its response.

In our example, the resulting response can acknowledge that previous troubleshooting was already attempted and move toward another diagnostic step instead of automatically restarting the same checklist.

That's the behavior I was looking for.

The useful part of memory isn't simply that the system can say "I remember Sarah."

The useful part is:

"I remember what we already tried, so I can decide what to do next."

What I learned

  1. Memory is only useful when it changes behavior

It would be easy to add a memory database and call the project memory-enabled.

That isn't enough.

The important question is what the agent does differently because the memory exists.

For RecallDesk, the desired behavior is straightforward: previous troubleshooting should influence the next response.

  1. Retrieval belongs between the user and the model

The language model doesn't need every historical interaction.

It needs the history that is relevant to the current problem.

That's why the architecture separates retrieval from generation:

Current message
↓
Retrieve
↓
Relevant context
↓
Generate

This keeps the model's input focused on the current task.

  1. Failed troubleshooting is valuable information

A support interaction isn't valuable only when it contains a successful solution.

Knowing that a particular troubleshooting step didn't work is also useful.

For example:

Clear cache → failed
Reduce file size → failed

That history can prevent the next response from simply repeating the same suggestions.

  1. The memory loop should be simple

RecallDesk doesn't require a complicated orchestration layer.

The fundamental loop is:

recall()
↓
generate()
↓
remember()

That simplicity makes the behavior easier to reason about and debug.

  1. Agent memory changes what "conversation history" means

Before building RecallDesk, I mostly thought about conversation history as messages that need to remain inside a chat session.

Persistent agent memory changes that model.

A useful interaction doesn't necessarily have to remain in the same session to remain useful.

It can become information that the agent retrieves later when another interaction makes it relevant.

Where I would take RecallDesk next

The current architecture provides a foundation for a larger support intelligence system.

For example, the same memory layer could eventually support incident histories, recurring customer issues, previously attempted solutions, unresolved problems, and patterns across support interactions.

That could turn RecallDesk from a conversational support interface into a system that helps support teams understand what has already happened before deciding what should happen next.

But the core idea would remain the same.

Recall before responding. Remember after responding.

That's the change Hindsight introduced to my approach to building the agent.

Instead of building an assistant that only answers the current message, I built one that can use what happened before.

ai

python

webdev

opensource

agents

python

streamlit

hindsight

@vectorize

Top comments (0)