Imagine contacting customer support about a problem you've already explained twice.
The support agent asks:
“Could you please provide your order number?”
You already provided it in your previous conversation.
Then you explain the issue again.
And again.
This is one of the biggest limitations of many AI customer-support agents: they can remember the current conversation, but they don't truly remember the customer.
For our project, we wanted to solve exactly this problem.
We built SupportMemory, a customer-support agent that uses Hindsight persistent memory to recall previous customer interactions and retain new information for future conversations.
The goal was simple:
Customers shouldn't have to repeat themselves.
The Problem: Stateless Customer-Support Agents
Most conversational AI systems work well within a single conversation.
For example:
Customer:
Hi, my order #4821 arrived damaged.
Agent:
I'm sorry about that. Can you provide a photo?
Customer:
Sure. I've uploaded it.
During this conversation, the agent has access to the previous messages.
But what happens when the customer comes back tomorrow?
Customer:
Hi, I need help with my damaged order.
Agent:
Sure! What is the issue with your order?
The customer has to start over.
This creates several problems:
- Customers repeat information.
- Support conversations take longer.
- Agents have less context.
- Customer experience becomes frustrating.
- Personalization becomes difficult.
We wanted our agent to behave differently.
The Idea: Give the Agent Persistent Memory
Instead of treating every conversation as a completely new interaction, we introduced a persistent memory layer.
Our architecture looks roughly like this:
┌────────────────────┐
│ Customer │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ SupportMemory Agent│
└─────────┬──────────┘
│
┌────────┴────────┐
│ │
▼ ▼
┌──────────────┐ ┌──────────────┐
│ Recall Memory│ │ New Message │
└──────┬───────┘ └──────┬───────┘
│ │
└────────┬─────────┘
▼
┌───────────────┐
│ Hindsight │
│ Memory Layer │
└───────┬───────┘
│
▼
┌───────────────┐
│ Agent Response│
└───────┬───────┘
│
▼
┌───────────────┐
│ Retain New │
│ Information │
└───────────────┘
The important idea is that memory isn't just a database containing old conversations.
The agent needs two capabilities:
Recall → Retrieve useful information from previous interactions.
Retain → Store useful information from the current interaction for future conversations.
That's where Hindsight becomes important.
How Hindsight Fits Into the Agent
We use Hindsight as the persistent memory layer for our support agent.
The basic flow is:
Customer Message
│
▼
Recall relevant
customer memory
│
▼
Combine memory +
current conversation
│
▼
LLM Agent
│
▼
Generate response
│
▼
Retain new information
This means the agent doesn't need to blindly load the customer's entire conversation history every time.
Instead, it can retrieve information that is relevant to the current request.
For example, imagine a customer named Ananya.
During a previous conversation:
Ananya:
My preferred language is English,
and I usually want email updates.
The agent can retain that information.
Later, Ananya starts another conversation:
Ananya:
Can you give me an update on my refund?
The agent can recall the relevant customer information and respond with more context.
Instead of treating Ananya as a completely new customer, the agent can use what it has learned from previous interactions.
Before Memory vs. With Hindsight
This was the most important part of our demo.
We created a simple comparison between a stateless agent and our memory-enabled agent.
Without Memory
Imagine Steven contacts support.
Steven:
My laptop replacement hasn't arrived yet.
The agent asks:
Agent:
Could you provide your order number?
Steven provides it.
Later, he contacts support again:
Steven:
I'm checking on my laptop replacement.
The stateless agent responds:
Agent:
Sure. Could you provide your order number?
Steven has to repeat himself.
With Hindsight
Now the same interaction happens with persistent memory.
First conversation:
Steven:
My laptop replacement hasn't arrived yet.
Agent:
I can help with that. Your order number is
#7392, correct?
Steven:
Yes.
The interaction is retained.
Later:
Steven:
I'm checking on my laptop replacement.
The agent can recall the previous context:
Agent:
I remember you were waiting for the replacement
for order #7392. Let me check the latest status
for you.
The difference is small from a technical perspective.
But from the customer's perspective, it is significant.
The agent feels like it actually knows them.
Recall and Retain
One of the most useful concepts we learned while building SupportMemory was separating recall from retain.
1. Recall
When a new message arrives, we first ask:
“What information from this customer's history could help answer this message?”
For example:
Customer:
I still haven't received my replacement.
Relevant memories might include:
Customer: Steven
Previous issue:
Laptop replacement
Order:
#7392
Previous conversation:
Customer was waiting for replacement
The agent can then use those memories as additional context.
2. Retain
After the agent responds, the new interaction can be stored.
For example:
Steven confirmed that order #7392
is still missing its replacement.
This becomes part of the customer's persistent history.
The next conversation can use this information.
So the cycle becomes:
┌───────────────┐
│ New Customer │
│ Message │
└───────┬───────┘
│
▼
┌─────────┐
│ Recall │
└────┬────┘
│
▼
┌──────────────┐
│ Agent + LLM │
└──────┬───────┘
│
▼
┌─────────┐
│ Respond │
└────┬────┘
│
▼
┌─────────┐
│ Retain │
└────┬────┘
│
▼
Persistent Memory
This creates a continuous learning loop for the agent.
Our Demo: Ananya, Steven, and Max
For the demonstration, we used three fictional customers:
Ananya
Ananya had previously discussed a refund issue.
When she returned later, the memory-enabled agent could use the previous interaction instead of asking her to explain everything again.
Steven
Steven had an unresolved replacement-order issue.
The agent could recall the previous order information and continue from where the previous conversation ended.
Max
Max had previous preferences and support interactions stored in memory.
When he returned, the agent could use those previous details to provide a more personalized response.
The demo showed the difference between:
WITHOUT MEMORY
↓
New conversation
↓
No previous context
↓
Ask customer again
and:
WITH HINDSIGHT
↓
New conversation
↓
Recall relevant history
↓
Personalized response
↓
Retain new interaction
Why Not Just Use the Conversation History?
This was an important design question.
One simple solution would be:
“Just send the customer's entire conversation history to the LLM.”
But that approach doesn't scale well.
Imagine a customer who has contacted support 50 times.
Sending everything every time can result in:
- Larger prompts
- Higher token usage
- More irrelevant information
- Slower responses
- More difficult context management
Most importantly, not every historical conversation is relevant to the current question.
Persistent memory gives us a way to retrieve useful information instead of blindly replaying everything.
The agent can focus on relevant memories.
Personalization Becomes Possible
Once an agent can remember useful information, personalization becomes much easier.
For example:
Customer:
I need help with my subscription again.
A stateless agent sees only that message.
A memory-enabled agent might know:
Previous subscription issue
Previous support interaction
Customer preferences
Previously discussed resolution
Relevant account context
The response can therefore be more contextual.
Instead of:
“Please provide more information.”
The agent can potentially say:
“I remember you had a similar subscription issue previously. Let's continue from there.”
This is the difference between an AI assistant that simply responds and one that can maintain continuity.
Architecture Decisions
While building the project, we focused on keeping the architecture simple.
The main components were:
Frontend
│
▼
Support Agent
│
├── Current Conversation
│
├── Hindsight Memory
│ ├── Recall
│ └── Retain
│
└── LLM
│
▼
Response
The important part is that memory is treated as a separate layer.
This makes the architecture easier to reason about.
The LLM doesn't have to be responsible for remembering everything.
Instead:
LLM
↓
Reasoning + Response Generation
Memory
↓
Long-term Customer Context
Each component has a clearer responsibility.
What We Learned
Building SupportMemory taught us a few practical lessons.
1. Memory is more than chat history
A long conversation history doesn't automatically create useful long-term memory.
The important question is:
What information should the agent remember?
2. Retrieval matters
Having thousands of memories isn't useful if the agent cannot retrieve the right ones.
Relevant memory is more valuable than simply having more memory.
3. Retention should be intentional
Not every sentence from a conversation needs to become permanent memory.
For a customer-support system, useful information might include:
- Previous issues
- Preferences
- Important interactions
- Unresolved problems
- Relevant decisions
- Customer-specific context
This makes memory more useful and manageable.
4. Memory changes the user experience
The biggest improvement isn't necessarily a technical metric.
It's the feeling that:
“This system remembers me.”
That continuity can make a support interaction feel much more natural.
What's Next?
SupportMemory is a starting point.
There are several directions we could explore next:
Better memory selection
Determine which interactions are worth retaining and which should be ignored.
Memory updates
Customer information can change.
For example:
Old:
Preferred contact method = Email
New:
Preferred contact method = SMS
The memory system should handle updates instead of blindly accumulating conflicting information.
Memory evaluation
We also need to measure whether the agent actually recalls the right information.
Useful metrics could include:
- Recall accuracy
- Relevant-memory retrieval
- Response quality
- Customer resolution time
- Repeated-question rate
Privacy and security
Customer memory contains potentially sensitive information.
A production system needs strong controls around:
- Data access
- Retention
- Deletion
- Authentication
- Authorization
- Sensitive information handling
Persistent memory is powerful, but it also makes responsible data handling even more important.
Final Thoughts
The biggest lesson from building SupportMemory is simple:
An AI agent becomes much more useful when it can maintain continuity across conversations.
Without persistent memory:
Conversation → Response → Forget
With persistent memory:
Conversation
↓
Recall previous context
↓
Generate response
↓
Retain useful information
↓
Future conversation
Hindsight gave us a way to build that recall-and-retain loop into our customer-support agent.
Instead of asking customers to start from zero every time, the agent can carry useful context forward.
And that's the direction we're excited to explore further:
AI agents that don't just answer questions, but remember the people they're helping.
Built With
- Hindsight — persistent memory
- LLM-based agent — reasoning and response generation
- SupportMemory — customer-support application layer
If you're building AI agents, persistent memory is worth experimenting with. The difference between “I can answer your question” and “I remember what we discussed last time” can completely change the experience.
Top comments (0)