DEV Community

Yashashree yellapu
Yashashree yellapu

Posted on

I Built a Customer Support Agent That Remembers With Hindsight

The most frustrating thing about customer support is not always the problem itself. It is having to explain the same problem again after someone already knows the history.

Imagine a customer saying, “My Wi-Fi is dropping again,” and the support agent already knows which router they use, which fixes have already failed, how many times the issue has returned, and whether the customer is getting frustrated. That is the kind of support agent I wanted to build.

I built RecallAI, a customer support agent that uses Hindsight for long-term customer memory and Groq for response generation. The important part is not simply storing old conversations. The system needs to retrieve the right facts at the right time, keep customers isolated from one another, and use previous outcomes to avoid repeating failed troubleshooting steps.

The problem with a stateless support agent

A normal LLM conversation has access to the current conversation context, but that context is not the same thing as a customer's long-term history.

Consider a returning customer:

“My Wi-Fi is dropping again.”

A stateless agent might respond with the usual troubleshooting checklist:

Restart the router.
Restart the device.
Check the Wi-Fi connection.
Update the firmware.

The problem is that the customer may have already done all of that.

For my example customer Priya, the stored support history says she uses a Netgear Nighthawk R7000 router and has previously reported repeated Wi-Fi drops. Restarting the router only helped temporarily. A firmware update did not fix the problem either. Switching Wi-Fi channels reduced the frequency of the drops, but did not solve them completely.

That history changes what a useful answer looks like.

Instead of starting from zero, the agent should recognize that this is a recurring problem and move toward the next appropriate action.

That became the central design question for RecallAI:

How can the agent remember the customer's history without putting the entire conversation history into every LLM request?

I separated conversation context from customer memory

The architecture is deliberately simple.

The application uses Streamlit for the interface, Groq for generating responses, and Hindsight as the long-term memory layer.

The flow for every new message is:

Customer message
↓
Hindsight recall
↓
Relevant customer memories
↓
Groq + current message + recent chat
↓
Support response
↓
Hindsight retain

There are two different kinds of context here.

The first is short-term conversation context. RecallAI keeps only the most recent four chat turns when constructing the LLM request.

The second is long-term customer memory. Hindsight stores information that can remain useful across conversations and even after the application is restarted.

This distinction matters because I don't want to keep sending an ever-growing transcript to the model.

The LLM gets the current problem, a small recent conversation window, and the relevant memories retrieved from Hindsight.

Every customer gets an isolated memory bank

One of the most important implementation decisions was customer isolation.

I did not want Priya's router information accidentally appearing in Rahul's conversation.

RecallAI creates a deterministic Hindsight bank for each customer:

def bank_id_for(customer_id: str) -> str:
slug = re.sub(
r"[^a-z0-9_-]",
"-",
customer_id.strip().lower()
).strip("-")[:60]

if not slug:
    raise ValueError(
        "Customer ID must contain letters or numbers."
    )

return f"{config.BANK_PREFIX}-{slug}"
Enter fullscreen mode Exit fullscreen mode

So a customer such as priya gets a bank like:

recallai-priya

while another customer gets a different bank.

The memory service then performs both retention and recall against that customer's bank.

This is more than a UI feature. I also wrote tests specifically for isolation.

The test creates two different customers, stores different router information for each, and verifies that recalling one customer's memory does not return the other's information.

That gives the memory layer a clear boundary:

one customer → one memory bank → one isolated history.

Hindsight stores facts, not just transcripts

Another design choice was what exactly to retain.

For every support exchange, RecallAI stores both the customer's message and the agent's response:

items = [
{
"content": f"[{today}] {who} wrote to support: {customer_message}",
"context": "customer message in a support chat",
},
{
"content": f"[{today}] Support agent's reply to {who}: {agent_reply}",
"context": (
"what the support agent suggested; the outcome is unknown "
"unless the customer later confirms it"
),
},
]

I intentionally don't try to manually build a giant customer profile every time a message arrives.

Instead, Hindsight processes the retained information and extracts useful facts from it.

The memory bank's mission tells Hindsight what matters for this application:

customer identity and plan
devices and software
support issues and tickets
solutions that worked
solutions that only worked temporarily
solutions that failed
recurring problems
communication preferences
frustration when explicitly expressed

That gives the memory system a purpose rather than treating every piece of conversation as equally important.

Retrieval happens before the LLM answers

When a customer sends a new message, RecallAI first asks Hindsight for relevant memories.

with st.spinner("Recalling from Hindsight..."):
rec = memory.recall(customer_id, text)

memories = rec.memories if rec.ok else []

Those memories are then passed to the LLM together with the current message and the short recent conversation window.

The system prompt also gives the model rules for using memory:

  • Use the memory to personalize your answer.
  • Never repeat a fix that memory says already failed or only helped temporarily.
  • Follow the customer's communication preferences.
  • Only state history that appears in CUSTOMER MEMORY.
  • If the customer sounds frustrated, acknowledge it briefly.

This is an important distinction.

Memory should not simply make the agent say:

“I remember you.”

It should change the agent's decision-making context.

If a firmware update already failed, the agent should not recommend the same firmware update as if nothing happened.

If a customer prefers concise instructions, the response should be concise.

If the customer has already reported the same issue several times, the response should recognize that history rather than treating it as a brand-new ticket.

The before-and-after behavior

For Priya, the stored history includes multiple Wi-Fi tickets.

Without long-term memory, the message:

“My Wi-Fi is dropping again.”

could produce a generic troubleshooting response.

With RecallAI, Hindsight can retrieve the customer's router, previous tickets, failed fixes, and the fact that the issue has returned.

The agent can therefore respond based on what has already happened instead of asking the customer to start over.

The same idea works for a different kind of support problem.

Rahul's history contains a recurring payment error, PAY-402. Re-entering his card details worked previously, but the problem returned at the next renewal. His memory also says that he prefers short, concise instructions.

When Rahul reports that the payment failed again, the agent can use both pieces of information:

technical history + communication preference.

That is much more useful than simply remembering the customer's name.

Memory survives a new conversation

I wanted to verify that the memory was actually persistent rather than just another form of Streamlit session state.

RecallAI has a “Start new conversation” action that clears the on-screen chat:

if st.button("🆕 Start new conversation"):
st.session_state.chats[customer_id] = []
st.session_state.recalled.pop(customer_id, None)
st.session_state.status.pop(customer_id, None)
st.rerun()

Notice what this does not clear: the Hindsight memory bank.

That means I can clear the visible conversation and ask:

“What have we already tried?”

The previous customer history can still be recalled.

The project also contains a test specifically designed around this behavior. It creates one CustomerMemory instance, stores a customer exchange, destroys that session, creates a new CustomerMemory instance, and then checks whether the relevant memory can still be recalled.

That is the behavior I wanted from persistent memory: the conversation can end without the customer's history disappearing with it.

Memory also learns from new customers

The system is not limited to preloaded demo customers.

RecallAI supports a custom customer ID. A new customer can say something like:

“I use a Mac with Firefox and the checkout page keeps crashing.”

The exchange is retained in that customer's Hindsight bank.

After starting a new conversation, the agent can retrieve the browser and operating-system information from memory.

So the memory isn't just a static database of facts I prepared beforehand. The support agent can build customer context through interactions.

One implementation detail I had to respect

Persistent memory is useful, but it isn't instantaneous.

Hindsight performs fact extraction when information is stored, so the project documentation notes that saves can take several seconds. The tests also account for the time needed for newly retained information to become available for recall.

That changed how I think about memory in an agent system.

A memory layer is not simply a faster dictionary lookup. There is processing involved in deciding what information should become useful long-term memory.

For a production system, that latency would be something I would monitor and optimize rather than hiding from the user.

What I learned

  1. Memory is useful only when it changes behavior

Simply retrieving an old ticket isn't enough.

The valuable part is using that history to avoid repeating failed fixes, respect preferences, and understand recurring problems.

  1. Short-term context and long-term memory should have different jobs

The recent conversation helps the model understand what is happening right now.

Hindsight handles information that needs to survive beyond the current conversation.

Keeping those responsibilities separate prevents the context window from becoming the storage layer.

  1. Customer isolation needs to be designed explicitly

A support agent handling multiple customers cannot treat memory as one shared pool.

Giving each customer an isolated memory bank makes that boundary explicit and testable.

  1. Retrieval quality matters as much as storage

Saving thousands of conversation messages is not the goal.

The goal is retrieving the few pieces of history that actually matter for the current question.

  1. The best support agent is not the one that remembers everything

It is the one that remembers the right things at the right time.

That is what I found most interesting while building RecallAI. Long-term memory turns customer support from a sequence of isolated conversations into a continuous relationship with context.

For the memory layer, I used Hindsight because it provides the retain-and-recall workflow I needed instead of forcing the application to manually construct and maintain a complete customer-memory system.

You can also learn more about agent memory from Vectorize.

RecallAI is still a compact system, but the architecture points toward something much larger: support agents that don't make returning customers start from zero every time they open a new conversation.

Top comments (0)