DEV Community

Bhavani Prasanna
Bhavani Prasanna

Posted on

I Built an AI Support Agent That Remembers What Worked

Customer support agents usually know what the customer is asking right now, but they often don't know what worked the last time the same customer had a similar problem.

I wanted to build a support agent where previous interactions could actually influence future responses.

That led me to MemoryCare, an AI customer-support system that combines screenshot understanding with persistent memory using Hindsight.

The interesting part isn't simply remembering a conversation. The useful part is remembering which previous experience is relevant to the current problem.

The Problem: Every Support Conversation Starts Too Close to Zero

Consider a customer who previously experienced a payment error.

The support agent suggested clearing the browser cache, and the payment succeeded.

Later, the same customer encounters the same error again.

A normal chatbot can understand the current message, but unless the previous interaction is available in its context, it has no reason to know that clearing the cache worked previously.

That can lead to repetitive troubleshooting.

MemoryCare approaches this differently:

Current Customer Issue
↓
Understand the Issue
↓
Recall Relevant Customer Memory
↓
Generate Personalized Support
↓
Resolve the Case
↓
Retain the New Experience
↓
Use It During Future Support

The persistent memory layer is powered by Hindsight.

*Why Hindsight Became the Important Part
*


I didn't want to simply attach a database containing old conversations to the chatbot.

The useful question is:

“What previous experience is actually relevant to this problem?”

MemoryCare uses Hindsight for that Recall → Response → Retain loop.

For a current customer issue, the system sends a meaningful query to the customer's memory bank and retrieves relevant memories.

A simplified version of the Recall implementation looks like this:

def recall(query):
url = f"{BASE_URL}/v1/default/banks/{BANK_ID}/memories/recall"

payload = {
    "query": query,
    "max_tokens": 1000,
    "budget": "low"
}

response = requests.post(
    url,
    headers=HEADERS,
    json=payload,
    timeout=60
)

response.raise_for_status()

data = response.json()

return data.get("results", [])
Enter fullscreen mode Exit fullscreen mode

The important part here is that the agent doesn't need to know the answer in advance. It asks the memory layer for context relevant to the current issue.

Recall Before Responding

The support agent follows a simple sequence:

from memory import recall
from llm import generate_response

def support_agent(user_message):
memories = recall(user_message)

response = generate_response(
    user_message=user_message,
    memories=memories
)

return response
Enter fullscreen mode Exit fullscreen mode

This creates a clear separation:

The customer provides the current problem.
Hindsight retrieves relevant previous experience.
The response generator combines the current issue with that context.
The customer receives a response based on more than the current message.

This is where persistent memory changes the behavior of the support agent.

A Concrete Example

Suppose the customer reports:

My payment is failing again. Can you check what happened last time?

The previous support case contained:

Customer: Bhavani

Issue: Payment failed
Error: 4032
Product: MemoryCare Mobile App

Previous troubleshooting:
Clear browser cache

Outcome:
Payment succeeded after clearing the browser cache.

When the current issue matches that previous experience, Hindsight can return the relevant memory.

The response can then use that context:

You previously experienced this error, and clearing
the browser cache resolved it.

You can try clearing the browser cache first and
then attempt the payment again.

The important behavior is not that the agent knows a predefined answer.

It is that the previous successful experience becomes available when the current issue is similar.

Before Memory vs After Memory


Before Memory

Customer:

My payment is failing again.

Agent:

Please check your internet connection and try again.

This is generic troubleshooting.

After Relevant Memory

Customer:

My payment is failing again.

Agent:

You previously experienced this error, and clearing the browser cache resolved it. You can try that first and then attempt the payment again.

The second response has more context because the agent can use relevant previous experience.

That is the behavior I wanted persistent memory to enable.

*Retaining the New Experience
*


Recall is only half of the system.

If the support agent learns something useful from the current interaction, that experience should become available later.

MemoryCare therefore also supports retention.

The simplified Retain implementation is:

def remember(content):
url = f"{BASE_URL}/v1/default/banks/{BANK_ID}/memories"

payload = {
    "items": [
        {
            "content": content
        }
    ]
}

response = requests.post(
    url,
    headers=HEADERS,
    json=payload,
    timeout=60
)

response.raise_for_status()

return response.json()
Enter fullscreen mode Exit fullscreen mode

After a case is resolved, the system can store the actual customer issue, troubleshooting step, and outcome.

For example:

Customer: Bhavani
Issue: Payment failure
Error: 4032
Solution: Cleared browser cache
Outcome: Payment succeeded

A future support interaction can then retrieve that experience.

This creates the loop:

RECALL
↓
RESPOND
↓
RESOLVE
↓
RETAIN
↓
RECALL AGAIN
What I Learned About Agent Memory

  1. Memory is useful only when it is relevant

Storing everything isn't enough.

A support agent shouldn't use an old delivery problem when the customer is currently reporting a login failure.

The current issue should determine what memory is useful.

  1. Current evidence still matters

Memory should complement the current customer message or screenshot, not replace it.

The current issue is the primary evidence.

Previous experience provides additional context.

  1. A memory failure should not break the support agent

Hindsight is an important component, but the application should still be able to respond using the current issue if memory retrieval fails or returns nothing useful.

This makes the support flow more resilient.

  1. Retain makes the system different from a simple chatbot

Recall gives the agent access to previous experience.

Retain allows useful new experience to become part of future interactions.

Without both sides, the memory loop is incomplete.

What Was Difficult

One of the practical challenges was making the memory layer behave like part of the support workflow instead of treating it as an independent feature.

The application needs to decide:

What should be recalled?
Which memories are relevant?
How should they influence the response?
What should be retained after resolution?
What should happen when no useful memory exists?

These decisions matter more than simply adding a memory API call.

Another important lesson was avoiding hard-coded assumptions around one support scenario. Payment failure is useful for demonstrating the workflow, but the architecture should also support other customer-support issues such as login failures, OTP problems, order issues, refunds, subscriptions, and application errors.

What I Would Improve Next

The current system demonstrates the persistent-memory workflow, but there are several directions I would take it further.

I would expand evaluation across a larger set of support scenarios, improve memory relevance evaluation, add stronger observability around Recall and Retain, and integrate the system with real support-ticket workflows.

The goal is not to make the agent remember everything.

The goal is to make it remember the right things at the right time.

Conclusion

Building MemoryCare changed how I think about AI support agents.

A model can generate a useful answer from the current conversation. But persistent memory creates another possibility: the agent can use previous customer experiences to make future interactions more contextual.

The core loop is simple:

Current Issue
↓
Hindsight Recall
↓
Personalized Response
↓
Resolution
↓
Hindsight Retain
↓
Better Context for the Next Interaction

For me, the most interesting part isn't making an agent remember everything.

It is making previous experience useful when it actually matters.

Learn More

Hindsight:

Hindsight GitHub
Hindsight Documentation
Vectorize Agent Memory

Hindsight links

Learn More

Top comments (0)