DEV Community

Cover image for RecallDesk: Building Long-Term Memory for Customer Support

RecallDesk: Building Long-Term Memory for Customer Support

How We Architected RecallDesk Around Hindsight Persistent Agent Memory

When an enterprise customer opens a high-priority incident, the hardest challenge is rarely the immediate technical defect. In complex cloud infrastructure, the primary friction is context fragmentation.

Consider an engineer investigating an edge proxy mutual TLS failure. If the customer spent three days debugging certificate rotation paths two months ago, or if an earlier engineer discovered that Envoy drops intermediate authorities when leaf certificates are mounted standalone, that context is critical. Yet in traditional support architectures, tickets arrive as isolated threads. Specialists often start from scratch, asking repetitive questions, re-executing failed steps, and re-learning environment constraints the customer already explained.

When designing RecallDesk, we set out to address this state fragmentation. We wanted support specialists to operate with long-term memory across sessions, recalling customer environments, previous troubleshooting attempts, verified resolutions, and known dead ends. Rather than building an ad-hoc vector pipeline from scratch, we architected RecallDesk directly around Hindsight.

In this article, we walk through the architecture of RecallDesk, how we integrated Hindsight as a dedicated agent memory service across FastAPI and React, and the practical engineering lessons we learned along the way.


The Stateless Agent Bottleneck in Technical Support

Support interactions do not occur in vacuum-sealed prompt windows; incidents span multiple days, diagnostic attempts, and environment constraints. Treating memory as ephemeral frontend state or simple session history causes three failure modes:

  • Context Amputation: Once a conversation closes, technical details—such as Kubernetes cluster versions or ingress configurations—remain buried in closed ticket rows that specialists rarely inspect in real time.
  • Repetition of Known Dead Ends: If a specialist previously confirmed that a TLS 1.2 downgrade failed, another specialist working on the same account might spend hours re-testing that same failed workaround.
  • Information Fatigue: Customers are forced to act as the vendor's human memory, repeatedly detailing their clusters, operating systems, and network topologies.

Addressing these issues requires a dedicated memory layer capable of semantic extraction, customer-scoped querying, and continuous retention across ticket lifecycles.


RecallDesk: High-Level System Architecture

RecallDesk high-level system architecture

RecallDesk pairs a modern frontend workspace with an asynchronous FastAPI backend and an external persistent memory layer powered by Hindsight:

  • Frontend (React 19, Vite, Tailwind CSS): A three-pane workspace featuring a ticket queue, an active conversation thread, and a Memory Intelligence panel separating operational ticket state from retrieved memory context.
  • Backend (FastAPI, Pydantic, Python 3.10+): Exposes REST endpoints for customers, conversations, and health checks, managing validation, input sanitization, and asynchronous Hindsight communication.
  • Data Storage Separation: Operational customer profiles and tickets are maintained in an in-memory data store (backend/app/services/mock_store.py), while long-term cognitive context is managed externally in Hindsight.
  • Memory Subsystem (Hindsight): Serves as our persistent memory engine, indexing dialogue turns, troubleshooting attempts, and ticket outcomes into dedicated memory banks for customer-scoped semantic recall.
+-------------------+        +--------------------+        +-----------------------+
|  React 19 Client  |  --->  |  FastAPI Backend   |  --->  | Hindsight Memory Bank |
|  - Ticket Queue   |  REST  |  - conversations   | Async  | - arecall()           |
|  - Memory Panel   |  <---  |  - customers       | Client | - aretain()           |
|  - Chat Thread    |        |  - hindsight_serv  |  <---  | - persistent bank     |
+-------------------+        +--------------------+        +-----------------------+
Enter fullscreen mode Exit fullscreen mode

Decoupling Memory: Why Hindsight Over Frontend State

In agent workflows, it is tempting to treat memory as an extension of application state—caching messages in React context, saving transcripts in relational tables, or dumping raw JSON into an ad-hoc database table.

Support memory is not simple chat history; it is an evolving corpus of facts, environment constraints, verified solutions, and failed hypotheses. Treating memory as frontend state breaks down when customers open tickets from different sessions or new specialists take over.

Instead, we adopted Vectorize agent memory via Hindsight. Hindsight provides a persistent memory bank accessible via its official Python SDK (hindsight_client), allowing our backend to store structured text interactions and execute semantic queries over previous customer history.

Our integration lives in backend/app/services/hindsight_service.py. Below is our client initialization:

from typing import Optional
from hindsight_client import Hindsight
from app.core.config import settings

class HindsightMemoryService:
    def __init__(self):
        self.base_url = settings.HINDSIGHT_BASE_URL
        self.api_key = settings.HINDSIGHT_API_KEY
        self.bank_id = settings.HINDSIGHT_BANK_ID
        self._client: Optional[Hindsight] = None
        self._init_client()

    def _init_client(self):
        try:
            self._client = Hindsight(
                base_url=self.base_url,
                api_key=self.api_key if self.api_key else None,
                timeout=15.0,
                user_agent="RecallDesk-Support/0.1.0"
            )
        except Exception:
            self._client = None
Enter fullscreen mode Exit fullscreen mode

We configure a dedicated bank ID (recalldesk-support) with an explicit mission: "Persistent support memory and troubleshooting history for enterprise customer incidents."


The Memory Lifecycle: Retain, Sanitize, and Recall

The interaction lifecycle between our FastAPI routes and Hindsight follows a disciplined sequence: sanitize, recall, append, and retain.

1. Pre-Ingestion Sanitization

Before support dialogue touches the memory layer, it passes through a regex redaction pipeline in sanitize_content(). Support tickets frequently contain bearer tokens, temporary passwords, API credentials, and private keys. Our service strips these patterns prior to invoking Hindsight retention APIs.

2. Scoped Semantic Recall

When a specialist inspects a conversation, RecallDesk queries Hindsight before the specialist drafts a reply. We scope the recall using customer metadata tags:

async def recall_customer_memories(
    self, customer_id: str, query: str, max_tokens: int = 2048, budget: str = "mid"
) -> Dict[str, Any]:
    if not self._client:
        return {"success": False, "memories": [], "bank_id": self.bank_id}

    sanitized_query = sanitize_content(query)
    recall_res = await asyncio.wait_for(
        self._client.arecall(
            bank_id=self.bank_id,
            query=sanitized_query,
            tags=[f"customer:{customer_id}"],
            tags_match="any",
            max_tokens=max_tokens,
            budget=budget
        ),
        timeout=8.0
    )
    return {"success": True, "count": len(recall_res.results), "memories": recall_res.results}
Enter fullscreen mode Exit fullscreen mode

Filtering with tags=[f"customer:{customer_id}"] helps scope recall to memories associated with the active customer. Because the current call uses tags_match="any", the supplied customer tag is treated as an OR-style tag filter.

3. Contextual Retention

When a message is sent or an issue is marked as resolved, RecallDesk compiles a structured record containing account context, technical environment details, dialogue, and resolution outcomes.

We construct deterministic document IDs (cust_{customer_id}_conv_{conversation_id}) and issue an asynchronous retain call:

document_id = f"cust_{customer_id}_conv_{conversation_id}"
tags = [
    f"customer:{customer_id}",
    f"conv:{conversation_id}",
    f"status:{status or 'open'}",
    f"tier:{tier.lower()}"
]

retain_res = await asyncio.wait_for(
    self._client.aretain(
        bank_id=self.bank_id,
        content=content,
        document_id=document_id,
        tags=tags,
        metadata=metadata,
        context=f"Support ticket interaction for {customer_name} at {company}"
    ),
    timeout=8.0
)
Enter fullscreen mode Exit fullscreen mode

4. Route Orchestration in FastAPI

In backend/app/api/routes/conversations.py, the messaging route links operational ticket storage with the Hindsight memory lifecycle:

@router.post("/{conversation_id}/messages", response_model=Message)
async def send_message(conversation_id: str, payload: SendMessageRequest):
    conversation = data_store.get_conversation(conversation_id)
    if not conversation:
        raise HTTPException(status_code=404, detail="Conversation not found.")

    # 1. Recall prior customer context to assist the specialist
    query = f"{conversation.subject} - {payload.text}"
    recall_res = await hindsight_service.recall_customer_memories(
        customer_id=conversation.customer_id, query=query
    )

    # 2. Append support reply to operational thread
    message = data_store.add_message(
        conversation_id=conversation_id, text=payload.text, sender_type="agent"
    )

    # 3. Retain updated dialogue in Hindsight memory bank
    await hindsight_service.retain_customer_conversation(
        customer_id=conversation.customer_id,
        conversation_id=conversation.id,
        messages=[m.model_dump() for m in conversation.messages],
        subject=conversation.subject,
        status=conversation.status,
        customer_context=customer_context
    )
    return message
Enter fullscreen mode Exit fullscreen mode

The support specialist reviews the recalled context in the workspace UI to guide their recommendations, rather than relying on an autonomous generation step.


Concrete Behavior: What the Agent Recalls

To test and demonstrate this flow during development, we seeded representative technical support cases in mock_store.py (such as Elena Rostova at Acme Cloud Infrastructure troubleshooting an Envoy mTLS issue).

When an issue regarding TLS handshake failures is selected, RecallDesk issues an arecall() query to Hindsight using the ticket subject and customer messages. In the React frontend, the CustomerContextPanel displays these recollections in the Memory Intelligence tab:

  • What Worked: Seeded records indicating that Envoy v1.28+ requires fullchain.pem in secret mounts rather than leaf-only cert.pem.
  • What Failed: Earlier troubleshooting records demonstrating that a TLS 1.2 protocol downgrade failed to bypass intermediate CA validation.
  • Customer Environment: Environmental metadata confirming that the account runs AWS EKS v1.29 with Envoy ingress edge proxies.

This provides the human specialist with immediate situational awareness before drafting a reply.


Implementation Experience: Dual Environments and Liveness

During development, managing service availability between local environments and cloud infrastructure was an instructive challenge.

Our configuration supports two operational targets:

  1. A self-hosted local Hindsight instance (http://localhost:8888), common in local container workflows.
  2. The managed cloud endpoint (https://api.hindsight.vectorize.io).

Because external services can experience transient network latency, RecallDesk avoids fabricating synthetic memory items when the service is unreachable. Instead, we implemented active health checks using aget_version(), bounded by an 8.0-second timeout.

When Hindsight is offline, the backend reports its status as standby or unavailable, and the React UI indicates standby mode. When live, real semantic recollections populate immediately. This separation ensured debugging reflected actual client-server communication.


What Works in the Current Implementation

The following capabilities are currently implemented and demonstrated:

  • Asynchronous Hindsight Connectivity: Real connection verification against the Hindsight server via aget_version(), surfaced through /health endpoints to the frontend.
  • Scoped Retention: Support conversations and resolution updates serialize into Hindsight memory banks using customer, conversation, and status tags.
  • Semantic Recall: Natural language queries execute against the Hindsight bank, returning recalled memory facts and document IDs.
  • Frontend Memory Display: The React 19 interface triggers FastAPI endpoints, rendering ticket threads alongside recalled memory panels.
  • Credential Redaction: Regex-based sanitization runs prior to payload retention to protect sensitive credentials.

Practical Engineering Lessons

Building this persistent memory architecture provided four practical takeaways:

  1. Scope Queries With Entity Tags: Combining semantic search with explicit metadata tags (customer:{id}) bounds retrieval to the appropriate customer context, avoiding multi-tenant cross-talk.
  2. Index Failures Alongside Fixes: Retaining diagnostic dead ends prevents specialists from recommending workarounds that have already proven ineffective.
  3. Redact Prior to Ingestion: Scrubbing credentials before invoking retention APIs prevents memory banks from storing unintended secrets.
  4. Decouple Memory From Operational State: Tickets have rigid lifecycles (open, pending, resolved), while memories are continuous. Treating Hindsight as an independent memory service rather than an operational database kept both layers cleanly isolated.

Current Limitations

While the persistent memory pipeline with Hindsight is functional, our repository currently stores operational ticket and customer entities in an in-memory data store (mock_store.py) rather than an external relational database.

Furthermore, autonomous LLM reply generation is currently in standby mode within the UI; the current implementation uses recalled context to assist the human specialist rather than generating unreviewed automated replies.

Importantly, the Hindsight memory layer persists separately from local server restarts when connected to an external Hindsight instance. Memories stored in the Hindsight bank remain queryable across backend reboots, demonstrating the structural separation between operational tickets and persistent memory.


Conclusion

Building persistent memory into support workspaces changes how engineering teams handle recurring customer incidents. By architecting RecallDesk around Hindsight, we decoupled operational support tickets from long-term memory, gave specialists visibility into past resolutions, and reduced the need to ask customers repetitive questions during critical outages.

To explore Hindsight or integrate persistent memory into your own software, check out the Hindsight documentation, explore the Hindsight GitHub repository, and read more about Vectorize agent memory.

Top comments (0)