DEV Community

Cover image for How I Built a Customer Support Agent That Remembers Every Ticket
Adithi Sagar
Adithi Sagar

Posted on

How I Built a Customer Support Agent That Remembers Every Ticket

Building a Stateful Customer Support Agent with Persistent Memory

  1. Introduction

Modern customer-support agents are increasingly expected to do more than answer isolated questions. A useful support agent should understand the customer's current problem, remember previous interactions, identify relevant tickets and technical entities, and maintain continuity across multiple conversations.

Traditional chatbot systems are often stateless. Each request is processed independently, meaning the agent may ask customers to repeat information that was already provided in an earlier conversation. This creates unnecessary friction and increases the time required to resolve support issues.

To address this limitation, this project introduces a stateful customer-support agent with persistent memory. The system integrates an AI orchestration layer with Hindsight memory so that customer information, previous support interactions, ticket details, and important entities can be retained and recalled when required.

The central idea is simple:

«The agent should not only understand the current message; it should also understand the relevant history behind that message.»

The architecture therefore separates the system into three major memory-aware operations:

  1. Recall – retrieve relevant historical information before generating a response.
  2. Reasoning / Generation – provide the retrieved context to the LLM along with the current customer request.
  3. Retain – store important information from the completed interaction for future conversations.

Hindsight provides these core memory capabilities through its "retain", "recall", and "reflect" operations, with memory stored in logical memory banks. Its recall mechanism combines multiple retrieval strategies, including semantic, keyword, graph, and temporal retrieval.

  1. Problem Statement

A conventional customer-support chatbot generally follows the following pattern:

Customer Message → Chat API → LLM → Response

Although this architecture works for simple questions, it has a major limitation: the LLM does not automatically possess a reliable, persistent record of previous customer interactions.

For example, a customer may report a webhook failure on Monday and provide the endpoint environment, error code, ticket number, and previous troubleshooting steps. If the customer returns on Tuesday and says:

«"My webhook delivery is still failing."»

A stateless chatbot may respond by asking for the endpoint URL, error code, and environment again.

This creates three problems.

First, the customer has to repeat information.

Second, the support agent spends additional time reconstructing the history of the problem.

Third, the LLM receives insufficient context to determine whether the new message represents a continuation of an existing issue or an entirely new problem.

A persistent-memory architecture solves this by retrieving relevant historical information before the response is generated.

  1. Proposed Solution

The proposed system introduces a persistent memory layer between the customer-support interface and the LLM.

Instead of sending only the latest customer message to the model, the system performs the following sequence:

Customer Message → Request Routing → Memory Recall → Context Assembly → LLM Generation → Support Response → Memory Retention

The memory system maintains information about:

  • Previous customer conversations
  • Support tickets
  • Ticket identifiers
  • Error messages
  • Technical environments
  • Previously attempted solutions
  • Customer-specific preferences
  • Important entities and relationships
  • Relevant timestamps and historical events

Hindsight's "retain" operation processes incoming information to extract useful facts, entities, and relationships, while "recall" retrieves relevant memories using multiple retrieval strategies.

  1. High-Level System Architecture

The complete customer-support architecture can be represented as follows:

                     ┌─────────────────────────────┐
                     │     CUSTOMER PORTAL         │
                     │     Message / Ticket        │
                     └──────────────┬──────────────┘
                                    │
                                    ▼
                     ┌─────────────────────────────┐
                     │      CHAT API / ROUTER      │
                     │  Authentication + Routing   │
                     └──────────────┬──────────────┘
                                    │
                                    ▼
                ┌─────────────────────────────────────┐
                │       MEMORY ORCHESTRATION           │
                │                                     │
                │  ┌────────────────┐  ┌────────────┐ │
                │  │ Recall Memory  │  │ Retrieve   │ │
                │  │ Historical     │  │ Entities   │ │
                │  │ Context        │  │ & Tickets  │ │
                │  └───────┬────────┘  └─────┬──────┘ │
                └──────────┼─────────────────┼────────┘
                           │                 │
                           └────────┬────────┘
                                    ▼
                     ┌─────────────────────────────┐
                     │      CONTEXT ASSEMBLER      │
                     │ Current Query + Memory      │
                     │ + Ticket + Entity Context  │
                     └──────────────┬──────────────┘
                                    │
                                    ▼
                     ┌─────────────────────────────┐
                     │         LLM ENGINE          │
                     │ Reasoning + Response        │
                     └──────────────┬──────────────┘
                                    │
                                    ▼
                     ┌─────────────────────────────┐
                     │       SUPPORT RESPONSE      │
                     └──────────────┬──────────────┘
                                    │
                                    ▼
                     ┌─────────────────────────────┐
                     │    ASYNC MEMORY RETENTION   │
                     │ Ticket + Facts + Entities   │
                     │ + Interaction Summary       │
                     └─────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

This architecture ensures that memory retrieval occurs before generation, while memory retention occurs after the response, preferably without delaying the customer's response.

  1. Memory Architecture

The memory layer consists of two primary operations: Recall and Retain.

5.1 Recall

Recall is responsible for answering the question:

«"What previous information is relevant to the customer's current request?"»

When a customer sends a new message, the system creates a retrieval query using the customer's current request and relevant identifiers.

For example:

Customer ID: CUST-1042

Current Query:
"My webhook delivery is still failing."

Relevant identifiers:
Ticket: #2451
Environment: Production
Service: Webhook Gateway

The memory layer then searches for relevant historical information.

Hindsight's recall process uses multiple retrieval strategies, including semantic similarity, keyword search, graph traversal, and temporal retrieval, and combines the results into a ranked set of memories.

5.2 Retain

Retain is responsible for answering:

«"What information from this interaction should become part of the customer's future memory?"»

After the support interaction completes, the system can retain information such as:

Ticket #2451
Customer reported repeated webhook failures.
Production environment: AWS ECS.
Previous issue involved TLS handshake timeouts.
TLS policy was updated to TLS 1.3.
Customer is now reporting continued 504 Gateway Timeout responses.

Hindsight processes retained content to extract meaningful facts, identify entities, and establish relationships rather than simply treating the entire conversation as an undifferentiated text block.

  1. Entity-Aware Customer Memory

One of the most important improvements in a persistent-memory support system is the use of entities.

An entity represents an important object or concept within the conversation.

Examples include:

Entity Type| Example
Customer| CUST-1042
Ticket| #2451
Service| Webhook Gateway
Environment| Production
Infrastructure| AWS ECS
Error| 504 Gateway Timeout
Configuration| TLS 1.3
Endpoint| Customer webhook URL

Instead of storing a conversation as a single large text block, the system can identify important entities and relationships.

For example:

This structure allows the agent to reason about relationships between customer, ticket, service, environment, error, and previous actions.

Hindsight's retention pipeline is designed to extract facts, resolve entities, and construct connected representations that can subsequently support recall.

  1. Code-Backed Implementation

The persistent-memory architecture is implemented by adding explicit recall and retain hooks to the customer-support orchestration loop.

The current Hindsight Python SDK uses the "Hindsight" client and a "bank_id" to scope memory. In an asynchronous application, the SDK provides async methods such as "arecall" and "aretain".

7.1 Initializing the Hindsight Client

API credentials should never be hard-coded inside application source code.

Instead, they should be loaded from environment variables.

import os
from hindsight_client import Hindsight

hindsight = Hindsight(
base_url=os.environ["HINDSIGHT_BASE_URL"],
api_key=os.environ["HINDSIGHT_API_KEY"]
)

The memory bank can then be associated with a customer or another appropriate isolation boundary.

def get_memory_bank(customer_id: str) -> str:
return f"customer-{customer_id}"

This creates a logical memory boundary such as:

customer-CUST-1042
customer-CUST-1043
customer-CUST-1044

This separation helps prevent unrelated customer information from being mixed together.

  1. Recalling Customer Memory Before Generation

Before generating a response, the system performs a recall operation using the customer's current request.

async def get_customer_context(
customer_id: str,
query: str
) -> str:

bank_id = get_memory_bank(customer_id)

response = await hindsight.arecall(
    bank_id=bank_id,
    query=query
)

if not response.results:
    return "No relevant prior support history was found."

context = "\n".join(
    f"- {memory.text}"
    for memory in response.results[:5]
)

return context
Enter fullscreen mode Exit fullscreen mode

The resulting memories are then passed into the LLM prompt together with the current customer message.

Conceptually:

Current Customer Message
+
Retrieved Customer Memory
+
Relevant Ticket Information
+
System Instructions
↓
LLM Prompt
↓
Support Response

This is an important architectural principle:

«Memory should be retrieved before reasoning, not after the response has already been generated.»

  1. Context Assembly Before LLM Generation

The retrieved memory should not simply be appended blindly to the prompt.

A context-assembly layer should combine the most relevant information.

async def build_support_context(
customer_id: str,
customer_message: str
) -> str:

memory_context = await get_customer_context(
    customer_id,
    customer_message
)

return f"""
Enter fullscreen mode Exit fullscreen mode

Relevant customer history:
{memory_context}

Current customer message:
{customer_message}
"""

The LLM can then receive a structured prompt:

You are a customer-support AI agent.

Use the relevant customer history below to maintain
continuity across conversations.

Relevant History:
[retrieved memories]

Current Customer Message:
[user message]

Instructions:

  • Do not invent historical information.
  • Use retrieved information only when relevant.
  • Ask for clarification when the retrieved context is insufficient.
  • Clearly distinguish previous issues from the current issue.

This reduces the risk of blindly trusting irrelevant historical information.

  1. Asynchronously Retaining Ticket Information

Once the response has been generated, important information from the interaction should be retained for future conversations.

A simplified implementation can look like this:

async def store_ticket_memory(
customer_id: str,
ticket_id: str,
user_msg: str,
agent_res: str
):
bank_id = get_memory_bank(customer_id)

content = f"""
Enter fullscreen mode Exit fullscreen mode

Ticket: {ticket_id}

Customer:
{user_msg}

Agent:
{agent_res}
"""

await hindsight.aretain(
    bank_id=bank_id,
    content=content,
    context="customer support interaction",
    metadata={
        "ticket_id": ticket_id,
        "source": "support_chat"
    }
)
Enter fullscreen mode Exit fullscreen mode

For larger workloads, Hindsight also supports asynchronous retention/batch ingestion so that memory processing can occur without blocking the main application flow.

  1. Complete Orchestration Flow

The complete interaction can therefore be represented as:

This feedback cycle transforms a conventional chatbot into a stateful support agent.

  1. Memory in Action: Before vs. After

12.1 Without Agent Memory — Stateless Interaction

Customer:
"My webhook delivery is still failing."

Agent:
"I'd be happy to help with that. Could you please share
your endpoint URL, error code, and environment setup?"

Customer:
"I already provided those details yesterday."

The problem is not necessarily that the LLM cannot answer the question. The problem is that it lacks the historical context required to understand the customer's statement.

The result is repeated questioning, slower troubleshooting, and a less continuous customer experience.

12.2 With Persistent Agent Memory — Stateful Interaction

Customer:
"My webhook delivery is still failing."

Memory:
Ticket #2451
Production environment: AWS ECS
Previous issue: TLS handshake timeout
Previous action: TLS policy updated to 1.3
Latest known error: 504 Gateway Timeout

Agent:
"I found your previous support history for Ticket #2451.
The earlier issue involved TLS handshake timeouts on your
production AWS ECS endpoint, and the TLS policy was updated
to 1.3.

Are you still receiving 504 Gateway Timeout responses,
or has the error code changed?"

The agent can now continue the troubleshooting process instead of restarting it.

The difference can be summarized as:

STATELESS

Message
↓
LLM
↓
Generic Response

STATEFUL

Message
↓
Recall
↓
Customer History
↓
Ticket + Entity Context
↓
LLM
↓
Context-Aware Response
↓
Retain New Information

This is the core value of persistent memory in customer-support systems.

  1. Why Structured Memory Matters

A common approach to conversational memory is to store complete conversation transcripts and pass them directly into the LLM.

Although this approach is simple, it becomes inefficient as conversations grow.

A long support history may contain:

  • Repeated greetings
  • Unimportant messages
  • Duplicate information
  • Debugging attempts that are no longer relevant
  • Old error messages
  • Irrelevant conversational text

Instead, the memory layer should identify the information that is actually useful.

For example:

RAW CONVERSATION

Customer:
"Hi."

Agent:
"Hello! How can I help?"

Customer:
"My webhook isn't working."

Agent:
"Can you share your environment?"

Customer:
"We're running on AWS ECS."

Agent:
"What error are you seeing?"

Customer:
"504 Gateway Timeout."

    ↓
Enter fullscreen mode Exit fullscreen mode

STRUCTURED MEMORY

Ticket: #2451
Service: Webhook
Environment: AWS ECS
Error: 504 Gateway Timeout
Status: Investigation

This makes future retrieval more focused.

Hindsight's retention pipeline is specifically designed to process raw content into extracted facts, entities, and connected memory representations.

  1. Deterministic Identifiers + Semantic Retrieval

Pure semantic retrieval is powerful, but customer-support systems also contain identifiers that should be handled carefully.

Examples include:

Ticket #2451
Order #893421
Customer ID CUST-1042
Incident INC-7821
Build 2026.09.28
HTTP 504

A semantic search system may understand the meaning of these identifiers, but exact identifiers should also be preserved as metadata wherever possible.

A robust retrieval architecture therefore combines:

            MEMORY RETRIEVAL
                   │
      ┌────────────┴────────────┐
      │                         │
      ▼                         ▼
Enter fullscreen mode Exit fullscreen mode

Semantic Search Deterministic Filters
│ │
│ Ticket ID
│ Customer ID
│ Service ID
│ │
└────────────┬────────────┘
▼
Ranked Context
│
▼
LLM

Hindsight's recall architecture itself combines semantic, keyword, graph, and temporal retrieval, which is useful for support scenarios where both meaning and structured identifiers matter.

  1. Engineering Lessons Learned

15.1 Structured Facts Are More Useful Than Raw Logs

Large conversation transcripts contain a significant amount of irrelevant information.

Extracting useful facts, entities, ticket identifiers, an

Top comments (0)