DEV Community

Tejaswini Tingirkar
Tejaswini Tingirkar

Posted on

How I Gave a Customer Support Agent Memory with Hindsight

A support agent that forgets every conversation has a simple problem: it makes the customer do the remembering.
I built MemoraSupport around a different idea. Instead of treating every support conversation as an isolated chat, I wanted the agent to carry useful information from one conversation into the next. Hindsight became the persistent memory layer that makes that possible.
The Problem I Wanted to Solve
Customer support conversations often contain information that matters later: an order number, a previous problem, a preferred communication channel, or an issue that was never fully resolved.
Consider this conversation:

“My order #4521 was delayed. I prefer email communication.”
A conventional chatbot can answer the immediate question, but when the customer starts another conversation, that context may be gone. The customer might later write:
“My refund for that order hasn't arrived.”
Without persistent memory, the agent has to ask which order the customer means. The customer has to repeat information that has already been provided.
That repetition was the problem I wanted MemoraSupport to address.
The goal is not simply to give an LLM access to a longer chat history. The more interesting problem is deciding what information should survive a conversation and retrieving the relevant part when it becomes useful again.

How MemoraSupport Fits Together

The application is split into a customer-facing frontend, an API/backend layer, an agent layer, a persistent memory service, an LLM, and an application database.
The basic flow is:
Customer → React UI → FastAPI → Support Agent → Hindsight Recall → Gemini → Response
When a customer sends a message, the support agent identifies the customer and retrieves relevant memories from that customer's Hindsight memory bank. Those memories are added to the context used to generate the response.
After processing the new message, useful information can be retained as long-term memory. If the conversation requires escalation, the agent can also create a support ticket in the application database.
That gives the workflow four important stages:
Understand → Recall → Respond → Take Action
The LLM is still responsible for generating the natural-language response, but it is no longer the only component responsible for the quality of the interaction.

Why Hindsight Became the Interesting Part


The most important architectural decision was separating application data from long-term customer memory.
MemoraSupport uses SQLite for application information such as users, conversations, messages, and support tickets. Hindsight handles the persistent memory that the support agent needs across conversations.
I use a separate memory bank for each customer. For example:

customer\_123
Enter fullscreen mode Exit fullscreen mode

This gives each customer's memories an explicit boundary instead of putting every customer's information into one shared memory space.
The two operations I rely on most are Retain and Recall.
Retain stores useful information in the customer's memory bank. Recall searches that memory when a new message arrives.
That separation matters because storing everything is not the same as having useful memory. A support system needs to retrieve information that is relevant to the current problem rather than blindly replaying an entire conversation history.
I used the Hindsight GitHub repository and the Hindsight documentation as references while working with the memory layer. The Vectorize guide to agent memory also provides useful context for thinking about memory as a distinct capability in an agent system.
The Code Path: Retain and Recall
The Hindsight service in MemoraSupport is intentionally kept behind a small service layer. That means the rest of the support agent does not need to know the details of the Hindsight HTTP API.
A simplified version of the retain path looks like this:

def retain\_memory(
    self,
    customer\_id: str,
    content: str,
    category: str = "general"
):
    bank\_id = self.\_get\_bank\_id(customer\_id)

    payload = {
        "bank\_id": bank\_id,
        "items": \[{
            "content": content,
            "category": category
        }]
    }

    response = client.post(
        f"{self.base\_url}/v1/retain",
        json=payload
    )

    return response
Enter fullscreen mode Exit fullscreen mode

The important detail is that the customer ID determines the memory bank. The content is then sent to Hindsight for retention.
Recall follows the same customer-specific pattern:

def recall\_memories(
    self,
    customer\_id: str,
    query: str,
    limit: int = 5
):
    bank\_id = self.\_get\_bank\_id(customer\_id)

    payload = {
        "bank\_id": bank\_id,
        "query": query,
        "limit": limit
    }

    response = client.post(
        f"{self.base\_url}/v1/recall",
        json=payload
    )

    return response.json()
Enter fullscreen mode Exit fullscreen mode

The support agent calls Recall using the customer's new message as the query:

recalled\_memories = hindsight\_service.recall\_memories(
    customer\_id,
    query=message\_text,
    limit=5
)
Enter fullscreen mode Exit fullscreen mode

Those recalled memories can then become part of the context supplied to the LLM.
This is the part of the architecture I found most useful: the model does not have to rediscover the customer's history from scratch every time. The memory layer handles the retrieval step.
A Concrete Before-and-After


The easiest way to see the difference is with two conversations.
Conversation 1
The customer says:

“My order #4521 was delayed. I prefer email communication.”
The system can identify two useful pieces of information:
Order #4521 had a delivery delay.
The customer prefers email communication.
Those facts are retained in the customer's Hindsight memory.
Conversation 2
Later, the customer says:
“My refund for that order hasn't arrived.”
The new message is sent to Recall. Because the query is associated with the same customer, relevant historical information can be retrieved.
Instead of treating “that order” as an unknown reference, the support agent can use the recalled information about order #4521. It can also take the communication preference into account when constructing the response.
The important change is not that the LLM suddenly became better at conversation. The system gave it access to the right historical context at the right time.
That is the difference between a stateless chatbot and a support agent with persistent memory.

Memory Is Useful Only When It Leads Somewhere

I also did not want MemoraSupport to stop at generating text.
Some support conversations require an action rather than another paragraph. For example, refund-related issues can lead to the creation of a support ticket when the application's rules determine that escalation is appropriate.
That gives the support workflow a simple pattern:
Understand → Recall → Respond → Take Action
The application database stores the operational information needed by the rest of the system, while Hindsight provides the long-term customer context.
This separation also makes the architecture easier to reason about. A support ticket is an application record. A customer's preference or previous issue is useful memory. They are related, but they do not have to live in the same storage layer.

Technology Stack

The main technologies in MemoraSupport are:
React + Vite — customer-facing frontend
FastAPI — backend API
Python — support-agent and memory logic
Hindsight — persistent customer memory
Gemini — LLM provider
SQLite — application data
Axios — frontend/backend communication
Docker — running Hindsight and supporting services
The backend exposes the application API, while the Hindsight service runs separately as the persistent memory layer. The LLM is configurable through the application's environment settings, and the current setup uses Gemini.


The architecture is intentionally straightforward:

Customer
   |
React + Vite
   |
FastAPI
   |
Support Agent
   |------------ Recall ----------> Hindsight
   |<----------- Memories ---------|
   |
   |------------ Prompt -----------> Gemini
   |<----------- Response ----------|
   |
   +---------- Store data ---------> SQLite
Enter fullscreen mode Exit fullscreen mode

Keeping these responsibilities separate makes it easier to change one layer without rewriting the entire application.
What I Learned

  1. An LLM is not the same thing as memory A model can generate a good response without knowing what happened in a previous conversation. Giving an agent persistent memory is a separate engineering problem. Hindsight provides that missing layer by giving the application a way to retain and retrieve information beyond the immediate conversation.
  2. More context is not automatically better It is tempting to send the entire customer history to the model. That approach becomes harder to control as conversations grow. Relevant retrieval is more useful than blindly adding more text. The agent needs the history that matters to the current message.
  3. Memory needs boundaries Customer information should not be mixed accidentally. Using a customer-specific memory bank makes the ownership boundary explicit and gives the retrieval layer a clear scope.
  4. Memory should support an action The most useful memory is not memory for its own sake. A previous issue, preference, or order reference becomes valuable when it helps the agent understand the current request or decide what should happen next.
  5. The difficult part comes after the first successful recall A simple demo can prove that a fact can be stored and retrieved. A real support system has harder questions: when should information become long-term memory, how should outdated information be updated, and how should conflicting information be handled? Those are the areas I would continue improving as the system grows. ## Limitations and Next Steps MemoraSupport currently demonstrates the core persistent-memory workflow, but there is still room to make the system more robust. Memory extraction can be expanded to recognize more categories of customer information automatically. Memory management also needs stronger handling for information that becomes outdated or needs to be removed. The support workflow could eventually connect to real e-commerce, CRM, email, and ticketing systems. That would allow the agent to move from a self-contained support application to a system that can actually retrieve order status, update customer records, and coordinate actions across existing business tools. Another area I would improve is evaluation. Instead of only checking whether a single memory can be recalled, I would test retrieval quality across many customer histories, including irrelevant memories, conflicting information, and long gaps between conversations. Conclusion The idea behind MemoraSupport is simple: Customer support should remember the customer instead of making the customer remember everything. The interesting part was not adding another chatbot interface. It was designing the memory boundary around the customer and connecting that memory to the support workflow.


With Hindsight as the persistent memory layer, the agent can carry useful information from one conversation into another, retrieve relevant history when a new message arrives, and use that context when responding or deciding whether a support action is needed.
For me, that changed the way I think about support agents. The question is no longer only, “How good is the model's answer?”
It is also:
“Does the agent remember the right thing when it matters?”
Resources

Top comments (0)