DEV Community

Gowtham Madagoni
Gowtham Madagoni

Posted on

Building a Customer-Support Agent That Actually Remembers: Adding Persistent Memory with Hindsight

Imagine contacting customer support about a problem you've already explained twice.

The support agent asks:

“Could you please provide your order number?”

You already provided it in your previous conversation.

Then you explain the issue again.

And again.

This is one of the biggest limitations of many AI customer-support agents: they can remember the current conversation, but they don't truly remember the customer.

For our project, we wanted to solve exactly this problem.

We built SupportMemory, a customer-support agent that uses Hindsight persistent memory to recall previous customer interactions and retain new information for future conversations.

The goal was simple:

Customers shouldn't have to repeat themselves.


The Problem: Stateless Customer-Support Agents

Most conversational AI systems work well within a single conversation.

For example:

Customer:
Hi, my order #4821 arrived damaged.

Agent:
I'm sorry about that. Can you provide a photo?

Customer:
Sure. I've uploaded it.
Enter fullscreen mode Exit fullscreen mode

During this conversation, the agent has access to the previous messages.

But what happens when the customer comes back tomorrow?

Customer:
Hi, I need help with my damaged order.

Agent:
Sure! What is the issue with your order?
Enter fullscreen mode Exit fullscreen mode

The customer has to start over.

This creates several problems:

  • Customers repeat information.
  • Support conversations take longer.
  • Agents have less context.
  • Customer experience becomes frustrating.
  • Personalization becomes difficult.

We wanted our agent to behave differently.


The Idea: Give the Agent Persistent Memory

Instead of treating every conversation as a completely new interaction, we introduced a persistent memory layer.

Our architecture looks roughly like this:

                 ┌────────────────────┐
                 │      Customer      │
                 └─────────┬──────────┘
                           │
                           ▼
                 ┌────────────────────┐
                 │ SupportMemory Agent│
                 └─────────┬──────────┘
                           │
                  ┌────────┴────────┐
                  │                 │
                  ▼                 ▼
          ┌──────────────┐   ┌──────────────┐
          │ Recall Memory│   │ New Message  │
          └──────┬───────┘   └──────┬───────┘
                 │                  │
                 └────────┬─────────┘
                          ▼
                  ┌───────────────┐
                  │ Hindsight     │
                  │ Memory Layer  │
                  └───────┬───────┘
                          │
                          ▼
                  ┌───────────────┐
                  │ Agent Response│
                  └───────┬───────┘
                          │
                          ▼
                  ┌───────────────┐
                  │ Retain New    │
                  │ Information   │
                  └───────────────┘
Enter fullscreen mode Exit fullscreen mode

The important idea is that memory isn't just a database containing old conversations.

The agent needs two capabilities:

Recall → Retrieve useful information from previous interactions.

Retain → Store useful information from the current interaction for future conversations.

That's where Hindsight becomes important.


How Hindsight Fits Into the Agent

We use Hindsight as the persistent memory layer for our support agent.

The basic flow is:

Customer Message
       │
       ▼
   Recall relevant
   customer memory
       │
       ▼
Combine memory +
current conversation
       │
       ▼
    LLM Agent
       │
       ▼
Generate response
       │
       ▼
Retain new information
Enter fullscreen mode Exit fullscreen mode

This means the agent doesn't need to blindly load the customer's entire conversation history every time.

Instead, it can retrieve information that is relevant to the current request.

For example, imagine a customer named Ananya.

During a previous conversation:

Ananya:
My preferred language is English,
and I usually want email updates.
Enter fullscreen mode Exit fullscreen mode

The agent can retain that information.

Later, Ananya starts another conversation:

Ananya:
Can you give me an update on my refund?
Enter fullscreen mode Exit fullscreen mode

The agent can recall the relevant customer information and respond with more context.

Instead of treating Ananya as a completely new customer, the agent can use what it has learned from previous interactions.


Before Memory vs. With Hindsight

This was the most important part of our demo.

We created a simple comparison between a stateless agent and our memory-enabled agent.

Without Memory

Imagine Steven contacts support.

Steven:
My laptop replacement hasn't arrived yet.
Enter fullscreen mode Exit fullscreen mode

The agent asks:

Agent:
Could you provide your order number?
Enter fullscreen mode Exit fullscreen mode

Steven provides it.

Later, he contacts support again:

Steven:
I'm checking on my laptop replacement.
Enter fullscreen mode Exit fullscreen mode

The stateless agent responds:

Agent:
Sure. Could you provide your order number?
Enter fullscreen mode Exit fullscreen mode

Steven has to repeat himself.


With Hindsight

Now the same interaction happens with persistent memory.

First conversation:

Steven:
My laptop replacement hasn't arrived yet.

Agent:
I can help with that. Your order number is
#7392, correct?

Steven:
Yes.
Enter fullscreen mode Exit fullscreen mode

The interaction is retained.

Later:

Steven:
I'm checking on my laptop replacement.
Enter fullscreen mode Exit fullscreen mode

The agent can recall the previous context:

Agent:
I remember you were waiting for the replacement
for order #7392. Let me check the latest status
for you.
Enter fullscreen mode Exit fullscreen mode

The difference is small from a technical perspective.

But from the customer's perspective, it is significant.

The agent feels like it actually knows them.


Recall and Retain

One of the most useful concepts we learned while building SupportMemory was separating recall from retain.

1. Recall

When a new message arrives, we first ask:

“What information from this customer's history could help answer this message?”

For example:

Customer:
I still haven't received my replacement.
Enter fullscreen mode Exit fullscreen mode

Relevant memories might include:

Customer: Steven

Previous issue:
Laptop replacement

Order:
#7392

Previous conversation:
Customer was waiting for replacement
Enter fullscreen mode Exit fullscreen mode

The agent can then use those memories as additional context.


2. Retain

After the agent responds, the new interaction can be stored.

For example:

Steven confirmed that order #7392
is still missing its replacement.
Enter fullscreen mode Exit fullscreen mode

This becomes part of the customer's persistent history.

The next conversation can use this information.

So the cycle becomes:

       ┌───────────────┐
       │ New Customer  │
       │   Message     │
       └───────┬───────┘
               │
               ▼
          ┌─────────┐
          │ Recall  │
          └────┬────┘
               │
               ▼
        ┌──────────────┐
        │ Agent + LLM  │
        └──────┬───────┘
               │
               ▼
          ┌─────────┐
          │ Respond │
          └────┬────┘
               │
               ▼
          ┌─────────┐
          │ Retain  │
          └────┬────┘
               │
               ▼
        Persistent Memory
Enter fullscreen mode Exit fullscreen mode

This creates a continuous learning loop for the agent.


Our Demo: Ananya, Steven, and Max

For the demonstration, we used three fictional customers:

Ananya

Ananya had previously discussed a refund issue.

When she returned later, the memory-enabled agent could use the previous interaction instead of asking her to explain everything again.

Steven

Steven had an unresolved replacement-order issue.

The agent could recall the previous order information and continue from where the previous conversation ended.

Max

Max had previous preferences and support interactions stored in memory.

When he returned, the agent could use those previous details to provide a more personalized response.

The demo showed the difference between:

WITHOUT MEMORY
      ↓
New conversation
      ↓
No previous context
      ↓
Ask customer again
Enter fullscreen mode Exit fullscreen mode

and:

WITH HINDSIGHT
      ↓
New conversation
      ↓
Recall relevant history
      ↓
Personalized response
      ↓
Retain new interaction
Enter fullscreen mode Exit fullscreen mode

Why Not Just Use the Conversation History?

This was an important design question.

One simple solution would be:

“Just send the customer's entire conversation history to the LLM.”

But that approach doesn't scale well.

Imagine a customer who has contacted support 50 times.

Sending everything every time can result in:

  • Larger prompts
  • Higher token usage
  • More irrelevant information
  • Slower responses
  • More difficult context management

Most importantly, not every historical conversation is relevant to the current question.

Persistent memory gives us a way to retrieve useful information instead of blindly replaying everything.

The agent can focus on relevant memories.


Personalization Becomes Possible

Once an agent can remember useful information, personalization becomes much easier.

For example:

Customer:
I need help with my subscription again.
Enter fullscreen mode Exit fullscreen mode

A stateless agent sees only that message.

A memory-enabled agent might know:

Previous subscription issue
Previous support interaction
Customer preferences
Previously discussed resolution
Relevant account context
Enter fullscreen mode Exit fullscreen mode

The response can therefore be more contextual.

Instead of:

“Please provide more information.”

The agent can potentially say:

“I remember you had a similar subscription issue previously. Let's continue from there.”

This is the difference between an AI assistant that simply responds and one that can maintain continuity.


Architecture Decisions

While building the project, we focused on keeping the architecture simple.

The main components were:

Frontend
   │
   ▼
Support Agent
   │
   ├── Current Conversation
   │
   ├── Hindsight Memory
   │      ├── Recall
   │      └── Retain
   │
   └── LLM
          │
          ▼
       Response
Enter fullscreen mode Exit fullscreen mode

The important part is that memory is treated as a separate layer.

This makes the architecture easier to reason about.

The LLM doesn't have to be responsible for remembering everything.

Instead:

LLM
 ↓
Reasoning + Response Generation

Memory
 ↓
Long-term Customer Context
Enter fullscreen mode Exit fullscreen mode

Each component has a clearer responsibility.


What We Learned

Building SupportMemory taught us a few practical lessons.

1. Memory is more than chat history

A long conversation history doesn't automatically create useful long-term memory.

The important question is:

What information should the agent remember?


2. Retrieval matters

Having thousands of memories isn't useful if the agent cannot retrieve the right ones.

Relevant memory is more valuable than simply having more memory.


3. Retention should be intentional

Not every sentence from a conversation needs to become permanent memory.

For a customer-support system, useful information might include:

  • Previous issues
  • Preferences
  • Important interactions
  • Unresolved problems
  • Relevant decisions
  • Customer-specific context

This makes memory more useful and manageable.


4. Memory changes the user experience

The biggest improvement isn't necessarily a technical metric.

It's the feeling that:

“This system remembers me.”

That continuity can make a support interaction feel much more natural.


What's Next?

SupportMemory is a starting point.

There are several directions we could explore next:

Better memory selection

Determine which interactions are worth retaining and which should be ignored.

Memory updates

Customer information can change.

For example:

Old:
Preferred contact method = Email

New:
Preferred contact method = SMS
Enter fullscreen mode Exit fullscreen mode

The memory system should handle updates instead of blindly accumulating conflicting information.

Memory evaluation

We also need to measure whether the agent actually recalls the right information.

Useful metrics could include:

  • Recall accuracy
  • Relevant-memory retrieval
  • Response quality
  • Customer resolution time
  • Repeated-question rate

Privacy and security

Customer memory contains potentially sensitive information.

A production system needs strong controls around:

  • Data access
  • Retention
  • Deletion
  • Authentication
  • Authorization
  • Sensitive information handling

Persistent memory is powerful, but it also makes responsible data handling even more important.


Final Thoughts

The biggest lesson from building SupportMemory is simple:

An AI agent becomes much more useful when it can maintain continuity across conversations.

Without persistent memory:

Conversation → Response → Forget
Enter fullscreen mode Exit fullscreen mode

With persistent memory:

Conversation
     ↓
Recall previous context
     ↓
Generate response
     ↓
Retain useful information
     ↓
Future conversation
Enter fullscreen mode Exit fullscreen mode

Hindsight gave us a way to build that recall-and-retain loop into our customer-support agent.

Instead of asking customers to start from zero every time, the agent can carry useful context forward.

And that's the direction we're excited to explore further:

AI agents that don't just answer questions, but remember the people they're helping.


Built With

  • Hindsight — persistent memory
  • LLM-based agent — reasoning and response generation
  • SupportMemory — customer-support application layer

If you're building AI agents, persistent memory is worth experimenting with. The difference between “I can answer your question” and “I remember what we discussed last time” can completely change the experience.

Top comments (0)