I Stopped My Customer Support Agent From Forgetting Every Conversation
I built a customer support agent that handled complaints perfectly except it had the memory of a goldfish. Every time the same customer came back, it asked them to repeat their entire story. So I gave it one.
What the System Does
The project is a Customer Resolution Engine a LangGraph based agentic workflow that triages electronics customer complaints. When a customer submits a complaint, the agent:
- Classifies the issue into a category and priority
- Queries a PostgreSQL knowledge base (read only, via a SQL agent) to check if it's a known issue with a known fix
- Decides whether to return an instant resolution, create a support ticket, or pause for human review (if the complaint involves legal threats or repeat failures)
- Responds to the customer
The stack is deliberately boring: FastAPI, PostgreSQL, LangGraph, and Gemini. Nothing exotic. The interesting part is what I bolted on afterward.
The Problem I Was Trying to Solve
After shipping v1, I ran a demo for a friend who runs a small electronics repair shop. Halfway through, he asked: "What happens when the same customer calls back three times about the same laptop?"
I ran the test. Here's what happened:
Call 1 : "My laptop screen is flickering." → Agent creates ticket TKT 2026 00011, assigns a hardware technician.
Call 2 (two days later): "My laptop is acting up again." → Agent responds: *"I'm sorry to hear that. Can you describe the issue in more detail?"*
My friend stared at the screen. "You just made the customer repeat themselves. That's the exact thing that makes people hate calling support."
He was right. The agent had access to the ticket database, but it had no concept of this specific customer's history their frustration level, what had already been tried, what had worked. It was treating every interaction as day one.
The Fix: A Real Memory Layer
I didn't want to solve this by dumping raw chat logs into PostgreSQL. Unstructured conversation history doesn't belong in a relational database it belongs in a semantic memory layer that can retrieve relevant context, not just recent rows.
That's when I found Hindsight, an open source memory system for AI agents by Vectorize. It lets agents retain information after an interaction and recall semantically relevant memories before the next one. You can read more about the design philosophy in their agent memory overview.
The integration took about 45 minutes and changed the agent from a stateless chatbot into something that actually knows its customers.
How I Wired It In
The architecture is simple. Hindsight sits at the absolute entry and exit points of the LangGraph workflow:
User Request → recall_memory → understand → classify → search_db → decide → respond → retain_memory → END
# The Recall Node (runs first)
At the start of every complaint, the agent queries Hindsight for the customer's past interactions:
from hindsight import HindsightClient
import os
hindsight_client = HindsightClient(api_key=os.getenv("HINDSIGHT_API_KEY"))
def recall_customer_memory(state: CustomerCareState):
customer_id = state.get("customer_id", "unknown")
try:
memories = hindsight_client.recall(
query=f"Past issues, frustration level, and resolved solutions for customer {customer_id}",
user_id=customer_id,
limit=3
)
state["memory_context"] = str(memories)
except Exception as e:
state["memory_context"] = "No prior memory found."
return state
# The Retain Node (runs last)
After the agent resolves or tickets the issue, it saves the outcome back to Hindsight with metadata:
def retain_resolution_memory(state: CustomerCareState):
customer_id = state.get("customer_id", "unknown")
complaint = state.get("complaint", "Unknown complaint")
resolution = state.get("resolution") or state.get("ticket_reference") or state.get("status")
hindsight_client.retain(
text=f"Customer {customer_id} complained about: {complaint}. Resolution: {resolution}.",
user_id=customer_id,
metadata={
"status": state.get("status", "resolved"),
"category": state.get("category", "general")
}
)
return state
# The Prompt Injection (where the magic happens)
The memory_context is injected directly into the Gemini prompt. This is the single most important line of code in the entire project:
prompt = f"""
You are an expert customer support agent.
IMPORTANT CONTEXT FROM PAST INTERACTIONS:
{memory_context}
CURRENT CUSTOMER COMPLAINT:
{state.get('complaint')}
INSTRUCTIONS:
1. If the customer is experiencing a recurring issue, explicitly acknowledge their past frustration and previous resolution attempts.
2. Do NOT ask them to repeat information already present in the memory context.
3. Formulate a personalized, context aware resolution.
"""
The Before/After: CUST 999
Here's the same customer, same vague follow up complaint, with and without memory.
Without Hindsight:
Customer : "My laptop is acting up again."
Agent : "I'm sorry to hear that. Can you describe the issue in more detail?"
With Hindsight:
Customer : "My laptop is acting up again."
Agent : "I see from your history that your screen was flickering last week and we replaced the display cable under ticket TKT 2026 00011. I apologize it's acting up again. Is this the exact same flickering issue, or something new? I'll escalate this to a senior hardware technician either way."
The second response is what my friend wanted to see. The customer doesn't repeat themselves. The agent demonstrates it remembers. The frustration level drops. The ticket gets routed correctly on the first try.
You can explore the full implementation in the Hindsight documentation if you want to replicate this pattern.
Five Lessons I Learned
Memory is a first class architectural concern, not a feature flag.
Don't bolt memory on at the end. Design your workflow so thatrecallruns at the entry point andretainruns at the exit. If you treat memory as an afterthought, your agent will feel like one.Don't store chat logs in PostgreSQL.
Relational databases are great for structured data (tickets, customers, products). They are terrible for semantic retrieval of unstructured conversation history. Use a purpose built memory layer.Prompt injection is the whole game.
Retrieving memory is only half the battle. If you don't explicitly instruct the LLM to use the memory and to not ask the user to repeat themselves it will ignore the context. Negative constraints ("Do NOT ask...") are as important as positive ones.Fail gracefully on memory outages.
If Hindsight is down, the agent should still work just without memory. Thetry/exceptblock inrecall_customer_memoryensures the workflow never crashes because of a memory API failure.The "Aha!" moment has to happen in 60 seconds.
Judges (and users) don't have time for a 10 minute demo. Show the generic response first, then show the personalized response. The contrast is the entire pitch.
What's Next
The current implementation uses an in memory LangGraph checkpointer, which means intra session state is lost on restart. The next iteration will swap that for a Postgres backed checkpointer so human in the loop interrupts survive server restarts. But the long term memory the part that actually matters for customer experience is already handled by Hindsight.
If you're building agents that interact with
the same users repeatedly, persistent memory isn't optional. It's the difference between a chatbot and a colleague.
Top comments (0)