How We Architected RecallDesk Around Hindsight Persistent Agent Memory
When an enterprise customer opens a high-priority incident, the hardest challenge is rarely the immediate technical defect. In complex cloud infrastructure, the primary friction is context fragmentation.
Consider an engineer investigating an edge proxy mutual TLS failure. If the customer spent three days debugging certificate rotation paths two months ago, or if an earlier engineer discovered that Envoy drops intermediate authorities when leaf certificates are mounted standalone, that context is critical. Yet in traditional support architectures, tickets arrive as isolated threads. Specialists often start from scratch, asking repetitive questions, re-executing failed steps, and re-learning environment constraints the customer already explained.
When designing RecallDesk, we set out to address this state fragmentation. We wanted support specialists to operate with long-term memory across sessions, recalling customer environments, previous troubleshooting attempts, verified resolutions, and known dead ends. Rather than building an ad-hoc vector pipeline from scratch, we architected RecallDesk directly around Hindsight.
In this article, we walk through the architecture of RecallDesk, how we integrated Hindsight as a dedicated agent memory service across FastAPI and React, and the practical engineering lessons we learned along the way.
The Stateless Agent Bottleneck in Technical Support
Support interactions do not occur in vacuum-sealed prompt windows; incidents span multiple days, diagnostic attempts, and environment constraints. Treating memory as ephemeral frontend state or simple session history causes three failure modes:
- Context Amputation: Once a conversation closes, technical details—such as Kubernetes cluster versions or ingress configurations—remain buried in closed ticket rows that specialists rarely inspect in real time.
- Repetition of Known Dead Ends: If a specialist previously confirmed that a TLS 1.2 downgrade failed, another specialist working on the same account might spend hours re-testing that same failed workaround.
- Information Fatigue: Customers are forced to act as the vendor's human memory, repeatedly detailing their clusters, operating systems, and network topologies.
Addressing these issues requires a dedicated memory layer capable of semantic extraction, customer-scoped querying, and continuous retention across ticket lifecycles.
RecallDesk: High-Level System Architecture
RecallDesk pairs a modern frontend workspace with an asynchronous FastAPI backend and an external persistent memory layer powered by Hindsight:
- Frontend (React 19, Vite, Tailwind CSS): A three-pane workspace featuring a ticket queue, an active conversation thread, and a Memory Intelligence panel separating operational ticket state from retrieved memory context.
- Backend (FastAPI, Pydantic, Python 3.10+): Exposes REST endpoints for customers, conversations, and health checks, managing validation, input sanitization, and asynchronous Hindsight communication.
-
Data Storage Separation: Operational customer profiles and tickets are maintained in an in-memory data store (
backend/app/services/mock_store.py), while long-term cognitive context is managed externally in Hindsight. - Memory Subsystem (Hindsight): Serves as our persistent memory engine, indexing dialogue turns, troubleshooting attempts, and ticket outcomes into dedicated memory banks for customer-scoped semantic recall.
+-------------------+ +--------------------+ +-----------------------+
| React 19 Client | ---> | FastAPI Backend | ---> | Hindsight Memory Bank |
| - Ticket Queue | REST | - conversations | Async | - arecall() |
| - Memory Panel | <--- | - customers | Client | - aretain() |
| - Chat Thread | | - hindsight_serv | <--- | - persistent bank |
+-------------------+ +--------------------+ +-----------------------+
Decoupling Memory: Why Hindsight Over Frontend State
In agent workflows, it is tempting to treat memory as an extension of application state—caching messages in React context, saving transcripts in relational tables, or dumping raw JSON into an ad-hoc database table.
Support memory is not simple chat history; it is an evolving corpus of facts, environment constraints, verified solutions, and failed hypotheses. Treating memory as frontend state breaks down when customers open tickets from different sessions or new specialists take over.
Instead, we adopted Vectorize agent memory via Hindsight. Hindsight provides a persistent memory bank accessible via its official Python SDK (hindsight_client), allowing our backend to store structured text interactions and execute semantic queries over previous customer history.
Our integration lives in backend/app/services/hindsight_service.py. Below is our client initialization:
from typing import Optional
from hindsight_client import Hindsight
from app.core.config import settings
class HindsightMemoryService:
def __init__(self):
self.base_url = settings.HINDSIGHT_BASE_URL
self.api_key = settings.HINDSIGHT_API_KEY
self.bank_id = settings.HINDSIGHT_BANK_ID
self._client: Optional[Hindsight] = None
self._init_client()
def _init_client(self):
try:
self._client = Hindsight(
base_url=self.base_url,
api_key=self.api_key if self.api_key else None,
timeout=15.0,
user_agent="RecallDesk-Support/0.1.0"
)
except Exception:
self._client = None
We configure a dedicated bank ID (recalldesk-support) with an explicit mission: "Persistent support memory and troubleshooting history for enterprise customer incidents."
The Memory Lifecycle: Retain, Sanitize, and Recall
The interaction lifecycle between our FastAPI routes and Hindsight follows a disciplined sequence: sanitize, recall, append, and retain.
1. Pre-Ingestion Sanitization
Before support dialogue touches the memory layer, it passes through a regex redaction pipeline in sanitize_content(). Support tickets frequently contain bearer tokens, temporary passwords, API credentials, and private keys. Our service strips these patterns prior to invoking Hindsight retention APIs.
2. Scoped Semantic Recall
When a specialist inspects a conversation, RecallDesk queries Hindsight before the specialist drafts a reply. We scope the recall using customer metadata tags:
async def recall_customer_memories(
self, customer_id: str, query: str, max_tokens: int = 2048, budget: str = "mid"
) -> Dict[str, Any]:
if not self._client:
return {"success": False, "memories": [], "bank_id": self.bank_id}
sanitized_query = sanitize_content(query)
recall_res = await asyncio.wait_for(
self._client.arecall(
bank_id=self.bank_id,
query=sanitized_query,
tags=[f"customer:{customer_id}"],
tags_match="any",
max_tokens=max_tokens,
budget=budget
),
timeout=8.0
)
return {"success": True, "count": len(recall_res.results), "memories": recall_res.results}
Filtering with tags=[f"customer:{customer_id}"] helps scope recall to memories associated with the active customer. Because the current call uses tags_match="any", the supplied customer tag is treated as an OR-style tag filter.
3. Contextual Retention
When a message is sent or an issue is marked as resolved, RecallDesk compiles a structured record containing account context, technical environment details, dialogue, and resolution outcomes.
We construct deterministic document IDs (cust_{customer_id}_conv_{conversation_id}) and issue an asynchronous retain call:
document_id = f"cust_{customer_id}_conv_{conversation_id}"
tags = [
f"customer:{customer_id}",
f"conv:{conversation_id}",
f"status:{status or 'open'}",
f"tier:{tier.lower()}"
]
retain_res = await asyncio.wait_for(
self._client.aretain(
bank_id=self.bank_id,
content=content,
document_id=document_id,
tags=tags,
metadata=metadata,
context=f"Support ticket interaction for {customer_name} at {company}"
),
timeout=8.0
)
4. Route Orchestration in FastAPI
In backend/app/api/routes/conversations.py, the messaging route links operational ticket storage with the Hindsight memory lifecycle:
@router.post("/{conversation_id}/messages", response_model=Message)
async def send_message(conversation_id: str, payload: SendMessageRequest):
conversation = data_store.get_conversation(conversation_id)
if not conversation:
raise HTTPException(status_code=404, detail="Conversation not found.")
# 1. Recall prior customer context to assist the specialist
query = f"{conversation.subject} - {payload.text}"
recall_res = await hindsight_service.recall_customer_memories(
customer_id=conversation.customer_id, query=query
)
# 2. Append support reply to operational thread
message = data_store.add_message(
conversation_id=conversation_id, text=payload.text, sender_type="agent"
)
# 3. Retain updated dialogue in Hindsight memory bank
await hindsight_service.retain_customer_conversation(
customer_id=conversation.customer_id,
conversation_id=conversation.id,
messages=[m.model_dump() for m in conversation.messages],
subject=conversation.subject,
status=conversation.status,
customer_context=customer_context
)
return message
The support specialist reviews the recalled context in the workspace UI to guide their recommendations, rather than relying on an autonomous generation step.
Concrete Behavior: What the Agent Recalls
To test and demonstrate this flow during development, we seeded representative technical support cases in mock_store.py (such as Elena Rostova at Acme Cloud Infrastructure troubleshooting an Envoy mTLS issue).
When an issue regarding TLS handshake failures is selected, RecallDesk issues an arecall() query to Hindsight using the ticket subject and customer messages. In the React frontend, the CustomerContextPanel displays these recollections in the Memory Intelligence tab:
-
What Worked: Seeded records indicating that Envoy v1.28+ requires
fullchain.pemin secret mounts rather than leaf-onlycert.pem. - What Failed: Earlier troubleshooting records demonstrating that a TLS 1.2 protocol downgrade failed to bypass intermediate CA validation.
- Customer Environment: Environmental metadata confirming that the account runs AWS EKS v1.29 with Envoy ingress edge proxies.
This provides the human specialist with immediate situational awareness before drafting a reply.
Implementation Experience: Dual Environments and Liveness
During development, managing service availability between local environments and cloud infrastructure was an instructive challenge.
Our configuration supports two operational targets:
- A self-hosted local Hindsight instance (
http://localhost:8888), common in local container workflows. - The managed cloud endpoint (
https://api.hindsight.vectorize.io).
Because external services can experience transient network latency, RecallDesk avoids fabricating synthetic memory items when the service is unreachable. Instead, we implemented active health checks using aget_version(), bounded by an 8.0-second timeout.
When Hindsight is offline, the backend reports its status as standby or unavailable, and the React UI indicates standby mode. When live, real semantic recollections populate immediately. This separation ensured debugging reflected actual client-server communication.
What Works in the Current Implementation
The following capabilities are currently implemented and demonstrated:
-
Asynchronous Hindsight Connectivity: Real connection verification against the Hindsight server via
aget_version(), surfaced through/healthendpoints to the frontend. - Scoped Retention: Support conversations and resolution updates serialize into Hindsight memory banks using customer, conversation, and status tags.
- Semantic Recall: Natural language queries execute against the Hindsight bank, returning recalled memory facts and document IDs.
- Frontend Memory Display: The React 19 interface triggers FastAPI endpoints, rendering ticket threads alongside recalled memory panels.
- Credential Redaction: Regex-based sanitization runs prior to payload retention to protect sensitive credentials.
Practical Engineering Lessons
Building this persistent memory architecture provided four practical takeaways:
-
Scope Queries With Entity Tags: Combining semantic search with explicit metadata tags (
customer:{id}) bounds retrieval to the appropriate customer context, avoiding multi-tenant cross-talk. - Index Failures Alongside Fixes: Retaining diagnostic dead ends prevents specialists from recommending workarounds that have already proven ineffective.
- Redact Prior to Ingestion: Scrubbing credentials before invoking retention APIs prevents memory banks from storing unintended secrets.
- Decouple Memory From Operational State: Tickets have rigid lifecycles (open, pending, resolved), while memories are continuous. Treating Hindsight as an independent memory service rather than an operational database kept both layers cleanly isolated.
Current Limitations
While the persistent memory pipeline with Hindsight is functional, our repository currently stores operational ticket and customer entities in an in-memory data store (mock_store.py) rather than an external relational database.
Furthermore, autonomous LLM reply generation is currently in standby mode within the UI; the current implementation uses recalled context to assist the human specialist rather than generating unreviewed automated replies.
Importantly, the Hindsight memory layer persists separately from local server restarts when connected to an external Hindsight instance. Memories stored in the Hindsight bank remain queryable across backend reboots, demonstrating the structural separation between operational tickets and persistent memory.
Conclusion
Building persistent memory into support workspaces changes how engineering teams handle recurring customer incidents. By architecting RecallDesk around Hindsight, we decoupled operational support tickets from long-term memory, gave specialists visibility into past resolutions, and reduced the need to ask customers repetitive questions during critical outages.
To explore Hindsight or integrate persistent memory into your own software, check out the Hindsight documentation, explore the Hindsight GitHub repository, and read more about Vectorize agent memory.

Top comments (0)