DEV Community

Cover image for I Stopped My Customer Support Agent From Forgetting Every Conversation
syed abdul khader
syed abdul khader

Posted on

I Stopped My Customer Support Agent From Forgetting Every Conversation

I Stopped My Customer Support Agent From Forgetting Every Conversation

I built a customer support agent that handled complaints perfectly except it had the memory of a goldfish. Every time the same customer came back, it asked them to repeat their entire story. So I gave it one.

What the System Does

The project is a Customer Resolution Engine a LangGraph based agentic workflow that triages electronics customer complaints. When a customer submits a complaint, the agent:

  1. Classifies the issue into a category and priority
  2. Queries a PostgreSQL knowledge base (read only, via a SQL agent) to check if it's a known issue with a known fix
  3. Decides whether to return an instant resolution, create a support ticket, or pause for human review (if the complaint involves legal threats or repeat failures)
  4. Responds to the customer

The stack is deliberately boring: FastAPI, PostgreSQL, LangGraph, and Gemini. Nothing exotic. The interesting part is what I bolted on afterward.

The Problem I Was Trying to Solve

After shipping v1, I ran a demo for a friend who runs a small electronics repair shop. Halfway through, he asked: "What happens when the same customer calls back three times about the same laptop?"

I ran the test. Here's what happened:

Call 1 : "My laptop screen is flickering." → Agent creates ticket TKT  2026  00011, assigns a hardware technician.
Call 2  (two days later): "My laptop is acting up again." → Agent responds: *"I'm sorry to hear that. Can you describe the issue in more detail?"*
Enter fullscreen mode Exit fullscreen mode

My friend stared at the screen. "You just made the customer repeat themselves. That's the exact thing that makes people hate calling support."

He was right. The agent had access to the ticket database, but it had no concept of this specific customer's history their frustration level, what had already been tried, what had worked. It was treating every interaction as day one.

The Fix: A Real Memory Layer

I didn't want to solve this by dumping raw chat logs into PostgreSQL. Unstructured conversation history doesn't belong in a relational database it belongs in a semantic memory layer that can retrieve relevant context, not just recent rows.

That's when I found Hindsight, an open source memory system for AI agents by Vectorize. It lets agents retain information after an interaction and recall semantically relevant memories before the next one. You can read more about the design philosophy in their agent memory overview.

The integration took about 45 minutes and changed the agent from a stateless chatbot into something that actually knows its customers.

How I Wired It In

The architecture is simple. Hindsight sits at the absolute entry and exit points of the LangGraph workflow:

User Request → recall_memory → understand → classify → search_db → decide → respond → retain_memory → END
Enter fullscreen mode Exit fullscreen mode

# The Recall Node (runs first)

At the start of every complaint, the agent queries Hindsight for the customer's past interactions:

from hindsight import HindsightClient
import os

hindsight_client = HindsightClient(api_key=os.getenv("HINDSIGHT_API_KEY"))

def recall_customer_memory(state: CustomerCareState):
    customer_id = state.get("customer_id", "unknown")
    try:
        memories = hindsight_client.recall(
            query=f"Past issues, frustration level, and resolved solutions for customer {customer_id}",
            user_id=customer_id,
            limit=3
        )
        state["memory_context"] = str(memories)
    except Exception as e:
        state["memory_context"] = "No prior memory found."
    return state
Enter fullscreen mode Exit fullscreen mode

# The Retain Node (runs last)

After the agent resolves or tickets the issue, it saves the outcome back to Hindsight with metadata:

def retain_resolution_memory(state: CustomerCareState):
    customer_id = state.get("customer_id", "unknown")
    complaint = state.get("complaint", "Unknown complaint")
    resolution = state.get("resolution") or state.get("ticket_reference") or state.get("status")

    hindsight_client.retain(
        text=f"Customer {customer_id} complained about: {complaint}. Resolution: {resolution}.",
        user_id=customer_id,
        metadata={
            "status": state.get("status", "resolved"),
            "category": state.get("category", "general")
        }
    )
    return state
Enter fullscreen mode Exit fullscreen mode

# The Prompt Injection (where the magic happens)

The memory_context is injected directly into the Gemini prompt. This is the single most important line of code in the entire project:

prompt = f"""
You are an expert customer support agent.

IMPORTANT CONTEXT FROM PAST INTERACTIONS:
{memory_context}

CURRENT CUSTOMER COMPLAINT:
{state.get('complaint')}

INSTRUCTIONS:
1. If the customer is experiencing a recurring issue, explicitly acknowledge their past frustration and previous resolution attempts.
2. Do NOT ask them to repeat information already present in the memory context.
3. Formulate a personalized, context  aware resolution.
"""
Enter fullscreen mode Exit fullscreen mode

The Before/After: CUST 999

Here's the same customer, same vague follow up complaint, with and without memory.

Without Hindsight:

Customer : "My laptop is acting up again."
Agent : "I'm sorry to hear that. Can you describe the issue in more detail?"

With Hindsight:

Customer : "My laptop is acting up again."
Agent : "I see from your history that your screen was flickering last week and we replaced the display cable under ticket TKT 2026 00011. I apologize it's acting up again. Is this the exact same flickering issue, or something new? I'll escalate this to a senior hardware technician either way."

The second response is what my friend wanted to see. The customer doesn't repeat themselves. The agent demonstrates it remembers. The frustration level drops. The ticket gets routed correctly on the first try.

You can explore the full implementation in the Hindsight documentation if you want to replicate this pattern.

Five Lessons I Learned

  1. Memory is a first class architectural concern, not a feature flag.
    Don't bolt memory on at the end. Design your workflow so that recall runs at the entry point and retain runs at the exit. If you treat memory as an afterthought, your agent will feel like one.

  2. Don't store chat logs in PostgreSQL.
    Relational databases are great for structured data (tickets, customers, products). They are terrible for semantic retrieval of unstructured conversation history. Use a purpose built memory layer.

  3. Prompt injection is the whole game.
    Retrieving memory is only half the battle. If you don't explicitly instruct the LLM to use the memory and to not ask the user to repeat themselves it will ignore the context. Negative constraints ("Do NOT ask...") are as important as positive ones.

  4. Fail gracefully on memory outages.
    If Hindsight is down, the agent should still work just without memory. The try/except block in recall_customer_memory ensures the workflow never crashes because of a memory API failure.

  5. The "Aha!" moment has to happen in 60 seconds.
    Judges (and users) don't have time for a 10 minute demo. Show the generic response first, then show the personalized response. The contrast is the entire pitch.

What's Next

The current implementation uses an in memory LangGraph checkpointer, which means intra session state is lost on restart. The next iteration will swap that for a Postgres backed checkpointer so human in the loop interrupts survive server restarts. But the long term memory the part that actually matters for customer experience is already handled by Hindsight.

If you're building agents that interact with
 the same users repeatedly, persistent memory isn't optional. It's the difference between a chatbot and a colleague.

Top comments (0)