DEV Community

Manogna Korrapati
Manogna Korrapati

Posted on

What Happened When I Gave My Support Agent Hindsight Memory...!!

# Building a Persistent-Memory Support Copilot with Hindsight

A temporal-memory architecture for support agents that remembers customer history, tracks effort trajectories, and detects escalation risk before frustration becomes a complaint.


## The Problem: Every Conversation Starts From Zero

Stateless language models create a frustrating loop in enterprise customer support.

Every time a customer reaches out, the system treats them like a total stranger.

Support representatives are forced to:

  • Ask for order numbers again
  • Re-verify tracking details
  • Request the same issue description
  • Repeat troubleshooting steps
  • Reconstruct previous conversations manually

For customers, this creates a simple but costly experience:

β€œI already told you this three times.”

To solve this, I built a support copilot with persistent temporal memory across separate customer interactions.

By integrating Hindsight into a FastAPI + React architecture, the copilot can:

  • Recall previous customer experiences
  • Preserve context across separate support sessions
  • Calculate customer-effort trajectories
  • Detect repeated contacts
  • Identify declining sentiment
  • Flag escalation risks before the customer explicitly asks for a manager

This article explains how the architecture works, why per-customer memory isolation became the hardest technical requirement, and what I learned while building persistent context for support workflows.


# πŸ—οΈ Architecture Overview

The system is designed as a rep-facing support copilot.

It sits between the incoming customer message and the support representative, enriching every ticket with historical context and suggested actions before the representative drafts a response.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                        React + Vite UI                            β”‚
β”‚                                                                   β”‚
β”‚  Ticket Queue     Conversation Thread     Hindsight Panel        β”‚
β”‚  Risk Badges      Agent Composer          Memory ON/OFF Toggle    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                               β”‚
                               β”‚ HTTP / REST API
                               β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                       FastAPI Backend                             β”‚
β”‚                                                                   β”‚
β”‚  Ticket Routes      Fallback Engine       Core Memory             β”‚
β”‚  /api/tickets       Local Cache           Pinned Facts            β”‚
β”‚                                                                   β”‚
β”‚                    Risk & Trajectory Classifier                   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                        β”‚                        β”‚
                 async recall/retain          inference
                        β”‚                        β”‚
                        β–Ό                        β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚       Hindsight Cloud           β”‚    β”‚       Groq LPU Engine       β”‚
β”‚                                β”‚    β”‚                             β”‚
β”‚  β€’ Retain Experiences          β”‚    β”‚  β€’ gpt-oss-120b             β”‚
β”‚  β€’ Scoped Tag Recall           β”‚    β”‚    Primary Model             β”‚
β”‚  β€’ Temporal Memory             β”‚    β”‚  β€’ qwen3-32b                β”‚
β”‚  β€’ Reflect & Form Opinions     β”‚    β”‚    Fallback Model            β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Enter fullscreen mode Exit fullscreen mode

Core Components

Component Responsibility
React + Vite Agent-facing support interface
FastAPI REST API, orchestration, prompt construction and fallback logic
Groq LPU Engine High-throughput LLM inference
Hindsight Persistent temporal memory and semantic recall
Local Cache Offline/degraded-mode fallback
Risk Classifier Contact-frequency and sentiment-based escalation detection

### πŸ€– LLM Layer

The system uses:

Primary: openai/gpt-oss-120b

Fallback: qwen/qwen3-32b

If function calling or generation fails with the primary model, the backend automatically falls back to the secondary model.


🧠 Why Traditional RAG Isn't Enough

A traditional Retrieval-Augmented Generation pipeline generally retrieves historical information based on semantic similarity.

That works well for many knowledge-retrieval tasks, but customer support introduces another dimension:

Time.

Imagine a customer has contacted support five times about three different orders.

A naive vector search might retrieve:

  • An old delivery complaint
  • A previously resolved replacement
  • A different product's return
  • A recent billing question

All of these may be semantically similar.

But similarity alone doesn't tell the agent:

Which experience belongs to the current issue?

Which issues are already resolved?

How many times has the customer contacted support?

Is the customer's effort increasing?

This is where persistent temporal memory becomes valuable.

Instead of treating every retrieved chunk as equivalent, the system separates:

Raw Experiences

Individual historical interactions and events.

Consolidated Opinions

Higher-level observations derived from those experiences.

For example:

Experiences
     β”‚
     β”œβ”€β”€ Contact #1 β†’ Delivery delayed
     β”œβ”€β”€ Contact #2 β†’ Carrier failed delivery
     └── Contact #3 β†’ Still waiting
             β”‚
             β–Ό
       Reflection
             β”‚
             β–Ό
   Customer Effort Trajectory
             β”‚
             β–Ό
      Escalation Risk
Enter fullscreen mode Exit fullscreen mode

This allows the copilot to reason about the trajectory of a support relationship, rather than simply retrieving similar text.


# πŸ” The Hardest Problem: Strict Customer Memory Isolation

Persistent memory introduces a serious multi-customer data problem.

Suppose Customer A's memory accidentally appears in Customer B's support session.

The system could expose:

  • Order numbers
  • Shipping information
  • Previous complaints
  • Account details
  • Resolution history

Even if the language model generates the response correctly, the underlying retrieval architecture would already have violated the required isolation boundary.

*### The Rule
*

Every memory operation must be explicitly scoped to the customer.

The implementation therefore enforces:

customer_id + brand
        ↓
Metadata
        ↓
Tags
        ↓
Scoped Recall
Enter fullscreen mode Exit fullscreen mode

For example:

Customer A
    └── customer_id = 4471

Customer B
    └── customer_id = 8234
Enter fullscreen mode Exit fullscreen mode

A recall for Customer A must never retrieve Customer B's memories simply because both customers experienced delayed deliveries.

Semantic similarity is not an authorization boundary.

That became one of the most important architectural lessons from this project.


# πŸ’» Code-Backed Implementation

1. Scoped Memory Retention

Every retained experience is automatically associated with the customer's identity and brand.

async def retain(
    self,
    customer_id: str,
    brand: str,
    text: str,
    metadata: Optional[dict] = None
) -> bool:
    """Retains an experience in Hindsight asynchronously
    with customer isolation."""

    meta = metadata or {}

    meta["customer_id"] = str(customer_id)
    meta["brand"] = str(brand)

    client = self._get_client()

    if not client:
        return True

    try:
        res = await client.aretain(
            bank_id=self.bank_id,
            content=text,
            metadata={
                "customer_id": str(customer_id),
                "brand": str(brand)
            },
            tags=[customer_id, brand]
        )

        return getattr(res, "success", True)

    except Exception as e:
        logger.error(f"Hindsight retain exception: {e}")
        return False
Enter fullscreen mode Exit fullscreen mode

*### Why this matters
*

The important part isn't simply storing the conversation.

The important part is storing it with explicit ownership metadata.

Memory
 β”œβ”€β”€ customer_id
 β”œβ”€β”€ brand
 └── experience
Enter fullscreen mode Exit fullscreen mode

This establishes the isolation boundary before retrieval even begins.


# 2. Isolated Memory Recall

The recall operation applies the customer's unique identifier as a strict filter.

async def recall(
    self,
    customer_id: str,
    query: str = ""
) -> List[RecalledItem]:

    """Recalls memories strictly filtered
    by customer_id tag."""

    client = self._get_client()

    if not client:
        return self._simulated_recall(customer_id)

    try:
        res = await asyncio.wait_for(
            client.arecall(
                bank_id=self.bank_id,
                query=query or (
                    "previous support issues orders "
                    "replacements and resolutions"
                ),
                tags=[customer_id]
            ),
            timeout=4.0
        )

        items = []

        for r in getattr(res, "results", []) or []:
            text_content = (
                getattr(r, "text", "")
                or getattr(r, "content", "")
                or str(r)
            )

            raw_type = str(
                getattr(r, "type", "")
            ).lower()

            item_type = (
                "opinion"
                if "opinion" in raw_type
                else "experience"
            )

            items.append(
                RecalledItem(
                    type=item_type,
                    summary=text_content,
                    confidence=0.88
                )
            )

        return items

    except Exception as e:
        return self._simulated_recall(customer_id)
Enter fullscreen mode Exit fullscreen mode

Two design decisions are particularly important here:
**

πŸ”’ Customer-scoped filtering**

tags=[customer_id]
Enter fullscreen mode Exit fullscreen mode

⏱️ Bounded latency

timeout=4.0
Enter fullscreen mode Exit fullscreen mode

If Hindsight becomes slow or unavailable, the system doesn't leave the support representative waiting indefinitely.

Instead, it falls back to a simulated/local memory path.


# 3. Post-Resolution Reflection

Memory retrieval and memory consolidation are deliberately separated.

When an agent resolves a ticket, the system asynchronously triggers reflection.

async def reflect(
    self,
    customer_id: str
) -> Optional[dict]:

    """Triggers Hindsight reflection
    after issue resolution."""

    client = self._get_client()

    if not client:
        return self._local_opinions.get(customer_id)

    try:
        res = await client.areflect(
            bank_id=self.bank_id,
            query=(
                f"Assess customer effort trajectory "
                f"and frustration risk for {customer_id}"
            ),
            tags=[customer_id]
        )

        return {
            "customer_id": customer_id,
            "confidence": 0.94,
            "reflection_text": getattr(
                res,
                "text",
                ""
            )
        }

    except Exception:
        return None
Enter fullscreen mode Exit fullscreen mode

This creates a two-stage architecture:

LIVE MESSAGE
     β”‚
     β–Ό
Fast Recall
     β”‚
     β–Ό
LLM Response
     β”‚
     β–Ό
Agent Resolves Ticket
     β”‚
     β–Ό
Async Reflection
     β”‚
     β–Ό
Updated Customer Opinion
Enter fullscreen mode Exit fullscreen mode

The key idea is:

Don't make the customer wait while the system consolidates memory.


βš™οΈ Memory ON vs Memory OFF

The copilot supports two modes.

Memory OFF

The model behaves like a traditional stateless support assistant.

Memory ON

The model receives historical experiences and consolidated customer facts.

The prompt builder controls this behavior dynamically:

def _build_system_prompt(
    self,
    customer_label: str,
    memory_enabled: bool,
    recalled_items: Optional[List[RecalledItem]],
    core_memory: Optional[List[str]] = None
) -> str:

    core_facts_text = (
        f"\n[CORE MEMORY FACTS]:\n"
        + "\n".join(
            [f"- {f}" for f in core_memory]
        )
        if core_memory
        else ""
    )

    if not memory_enabled or not recalled_items:
        return (
            "You are an AmazonHelp support copilot. "
            "Memory is DISABLED for this session. "
            "Respond strictly based on the user's latest message. "
            "Ask standard clarification questions "
            "(order numbers, tracking) as if hearing "
            "about the issue for the first time."
        )

    mem_text = "\n".join(
        [
            f"- [{item.type.upper()}] {item.summary}"
            for item in recalled_items
        ]
    )

    return (
        f"You are an AmazonHelp support copilot "
        f"for {customer_label}. Memory is ENABLED.\n"
        f"Hindsight Memories:\n{mem_text}\n"
        f"{core_facts_text}\n"
        "Instructions: Reference past details so "
        "the customer never repeats themselves. "
        "If repeat contacts are noted, acknowledge "
        "frustration directly and propose immediate solution."
    )
Enter fullscreen mode Exit fullscreen mode

This gives the UI a meaningful demonstration of the difference between:

Stateless generation

and

Memory-grounded generation.


# πŸ“Š Behavioral Comparison

Consider Customer #4471, who has contacted support twice during the previous week about a delayed package.

The customer opens a third ticket:

β€œI am still waiting for an update on my package.”

Memory OFF Memory ON
Treats interaction as new Retrieves previous interactions
Requests order information again Reuses previously known context
Repeats standard questions Acknowledges repeated contact
No historical trajectory Calculates effort trajectory
No memory-based risk signal Surfaces escalation risk

## πŸ”΄ Memory OFF β€” Stateless Mode

A typical response might be:

β€œHello Customer #4471, thanks for reaching out to AmazonHelp! Could you please provide your 17-digit Order ID and confirm which item you are waiting for so I can check tracking status?”

The problem is not that the response is grammatically incorrect.

The problem is that it ignores the customer's history.

The customer has already provided this information twice.


## 🟒 Memory ON β€” Hindsight-Grounded Mode

With memory enabled, the system can surface the previous interactions and associated trajectory.

A generated response could look like:

β€œHello Customer #4471, I am very sorry to see this is your 3rd contact regarding your Echo Dot shipment (Order #302-8220-4471). I see carrier attempt failures occurred earlier this week. I have escalated this directly to carrier dispatch for morning redelivery and applied a $15 courtesy credit to your account.”

The important difference is not simply that the response is more personalized.

It is that the response is grounded in the customer's previous support journey.


🚨 Proactive Escalation Detection

The agent interface also surfaces a proactive escalation banner when the system detects a qualifying pattern.

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ 🚨 PROACTIVE ESCALATION ALERT                                β”‚
β”‚                                                              β”‚
β”‚ Confidence: 94%                                              β”‚
β”‚                                                              β”‚
β”‚ This customer has contacted support 3 times about this       β”‚
β”‚ recurring issue with a declining sentiment trajectory.       β”‚
β”‚                                                              β”‚
β”‚ Recommended action: Review for manager intervention or      β”‚
β”‚ immediate goodwill resolution.                               β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Enter fullscreen mode Exit fullscreen mode

The escalation classifier uses an explicit invariant rather than triggering on a single negative message:

Contacts β‰₯ 3
      +
Declining sentiment
      ↓
Escalation Tier
Enter fullscreen mode Exit fullscreen mode

This reduces the possibility of escalating every isolated negative interaction.


🧩 Key Engineering Lessons

1. Metadata Scoping Is Mandatory

Semantic similarity should never be treated as a sufficient isolation mechanism for customer-specific memory.

Every memory operation should carry explicit customer-scoping information.

❌ Semantic similarity only

Customer A ─┐
Customer B ─┼──► Vector Search ──► Similar memories
Customer C β”€β”˜


βœ… Explicit customer scope

Customer A ──► customer_id=A ──► A's memories
Customer B ──► customer_id=B ──► B's memories
Customer C ──► customer_id=C ──► C's memories
Enter fullscreen mode Exit fullscreen mode

2. Separate Reflection From Generation

Running memory consolidation during every live response adds unnecessary latency.

Instead:

Message
  ↓
Recall
  ↓
Generate
  ↓
Respond

             Ticket Resolution
                    ↓
              Async Reflection
Enter fullscreen mode Exit fullscreen mode

This keeps the interactive path fast while still allowing the memory system to improve after the interaction.


3. Risk Detection Needs Stable Invariants

Escalation systems can become noisy when risk levels change based on a single message.

The implementation therefore requires both:

Repeated contact

and

Declining sentiment

before assigning the escalation tier.

This provides a more stable signal than simply detecting words such as angry, manager, or complaint.


4. Always Have an Offline Fallback

External memory services can experience:

  • Latency spikes
  • Temporary failures
  • Network problems
  • API errors

A support agent shouldn't lose access to the interface because a memory service is temporarily unavailable.

The system therefore includes:

Hindsight
   β”‚
   β”œβ”€β”€ Available β†’ Persistent Memory
   β”‚
   └── Unavailable β†’ Local Fallback
Enter fullscreen mode Exit fullscreen mode

The goal is graceful degradation rather than complete workflow failure.


πŸ”„ End-to-End Flow

The complete workflow can be summarized as:

Customer Message
       β”‚
       β–Ό
FastAPI Ticket API
       β”‚
       β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β–Ί Hindsight Recall
       β”‚                       β”‚
       β”‚                       β–Ό
       β”‚                Customer History
       β”‚                       β”‚
       β–Ό                       β”‚
Risk & Trajectory β—„β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚
       β–Ό
Prompt Assembly
       β”‚
       β–Ό
Groq LPU
       β”‚
       β”œβ”€β”€ gpt-oss-120b
       β”‚
       └── qwen3-32b fallback
       β”‚
       β–Ό
Suggested Agent Response
       β”‚
       β–Ό
Agent Resolves Ticket
       β”‚
       β–Ό
Async Hindsight Reflection
       β”‚
       β–Ό
Updated Customer Memory
Enter fullscreen mode Exit fullscreen mode

This creates a closed loop:

Recall β†’ Respond β†’ Resolve β†’ Reflect β†’ Remember


πŸš€ What This Architecture Changes

The central idea behind this project is simple:

A support copilot shouldn't only remember what was said. It should understand what happened over time.

Traditional support automation focuses primarily on the current ticket.

A persistent-memory copilot can additionally consider:

  • What happened before?
  • How many times has the customer contacted support?
  • Which problems were already resolved?
  • Is the customer's effort increasing?
  • Is sentiment changing over time?
  • Does this interaction represent a recurring failure?
  • Should the agent consider escalation?

That transforms memory from a passive retrieval mechanism into an operational signal for the support workflow.


🎯 Final Takeaway

Building persistent memory for customer support isn't simply a matter of connecting an LLM to a vector database.

The difficult engineering problems are around:

Isolation.

Temporal context.

Latency.

Reflection.

Fallback behavior.

Risk consistency.

Hindsight provided the persistent memory layer, while FastAPI handled orchestration and React exposed the resulting context directly to the support representative.

The resulting architecture follows a simple principle:

Don't make the customer repeat what the system already knows.

Instead of treating every ticket as an isolated conversation, the copilot builds an evolving picture of the customer's support journey and gives the representative the context needed to act on it.


πŸ”— Resources

  • Hindsight GitHub Repository β€” Explore the underlying memory architecture and implementation.
  • Hindsight Documentation β€” Learn how retention, recall, and reflection work.
  • Vectorize Agent Memory Guide β€” Explore approaches for building persistent memory into AI agents.
  • Project Source Code β€” Explore the complete FastAPI + React implementation and experiment with the architecture yourself. @Code.in

Top comments (0)