DEV Community

Cover image for Building, Debugging, and Testing RecallDesk: What We Learned While Connecting Persistent AI Memory to a Support Workflow
T. Mani Vardhan
T. Mani Vardhan

Posted on

Building, Debugging, and Testing RecallDesk: What We Learned While Connecting Persistent AI Memory to a Support Workflow

Building, Debugging, and Testing RecallDesk: What We Learned While Connecting Persistent AI Memory to a Support Workflow

By T ManiVardhan — RecallDesk Engineering Team


When building AI-assisted developer tools, the hardest engineering challenges rarely lie in prompt crafting. They emerge in the seams between systems: where a reactive frontend expects instant responsiveness, an asynchronous backend coordinates multi-step state transitions, and an external persistent memory engine indexes technical dialogue.

In previous articles, our team explored agent memory concepts and the FastAPI backend architecture. In this article, I share the practical engineering log: the friction points, debugging cycles, and testing patterns we encountered while building, integrating, debugging, and testing RecallDesk.

We will walk through connecting our React frontend and FastAPI backend to Hindsight—an open-source persistent memory system for AI agents—and what it takes to make persistent memory reliable in an enterprise support workflow.

Build ───► Integrate ───► Debug ───► Test ───► Improve
Enter fullscreen mode Exit fullscreen mode

1. The Starting Point: Why Persistent Memory Is Not Chat History

RecallDesk began with a functional support workspace: a React 19 frontend (Vite, Tailwind CSS), an asynchronous FastAPI backend (Uvicorn), and an in-memory mock store with realistic enterprise incident threads.

The UI was organized into a three-pane incident triage workspace: CustomerList (intake queue and status filters), ConversationView (active thread and message composer), and CustomerContextPanel (environment specs and memory intelligence).

RecallDesk dashboard

Visual Suggestion 1: RecallDesk Dashboard Workspace

Placement: Insert full dashboard screenshot here, showing the three-pane workspace with active ticket #conv_101 and the right-hand context panel.

Our initial instinct was simple: why not query past closed tickets from the local store and display them?

Displaying raw chat logs quickly failed the support specialist:

  • Vocabulary Drift: A customer's symptom description rarely matches the keywords documented in the resolution (e.g., SSL_ERROR_UNKNOWN_CA_ALERT versus patching Vault ConfigMaps to fullchain.pem).
  • Cognitive Overload: Dumping transcripts forces engineers under SLA pressure to parse dead ends and chatter.
  • Lack of Structure: Chat history records dialogue; it does not isolate what worked, what failed, or environment constraints.

Persistent memory required an engine that semantically indexes facts, separates solutions from failures, and scopes retrieval to the account: Hindsight.


2. Connecting the Hindsight Memory Layer

To integrate Hindsight into our FastAPI backend, we used the official Python SDK: hindsight-client (>=0.10.1).

Hindsight recall implementation

We encapsulated memory interactions within HindsightMemoryService in backend/app/services/hindsight_service.py. The service initializes the asynchronous Hindsight client using settings from backend/app/core/config.py:

# Initializing the official Hindsight client in hindsight_service.py
self._client = Hindsight(
    base_url=self.base_url,
    api_key=self.api_key if self.api_key else None,
    timeout=15.0,
    user_agent="RecallDesk-Support/0.1.0"
)
Enter fullscreen mode Exit fullscreen mode

Key configuration parameters include HINDSIGHT_BASE_URL (self-hosted or Hindsight Cloud), optional HINDSIGHT_API_KEY, target HINDSIGHT_BANK_ID (recalldesk-support), and a 15-second client timeout paired with 8.0-second operation timeouts via asyncio.wait_for().

During backend startup, FastAPI's lifespan handler invokes _ensure_bank_exists(), calling acreate_bank() to ensure the bank exists before handling requests.


3. Real Debugging: Transparency Over Silent Degradation

During development, we hit an integration hurdle: when testing against a local Hindsight service that was temporarily offline, backend logs recorded connection errors, and the memory subsystem reported as unavailable:

{
  "status": "healthy",
  "service": "RecallDesk API",
  "memory_subsystem": "unavailable (Cannot connect to host localhost:8888)"
}
Enter fullscreen mode Exit fullscreen mode

Earlier in development, our API caught this exception and quietly returned an empty memory list. While preventing crashes, it introduced silent degradation: developers could not tell whether a ticket had zero relevant memories or whether Hindsight was unreachable.

We resolved this with four improvements:

  1. Active Health Diagnostics: In get_connection_status(), we execute await self._client.aget_version() with an 8-second timeout, confirming real connectivity and returning the exact API version.
  2. Defensive Error Handling: When Hindsight is unreachable, the backend captures the exception and returns structured diagnostic metadata instead of crashing.
  3. Surfacing State via FastAPI: Our /health endpoint exposes memory_subsystem and memory_details, giving frontend clients full visibility into connection health.
  4. Decoupled Operation: Core ticketing functions (viewing tickets, sending replies, changing statuses) continue operating even when the memory service is offline.

4. Recall and Retention: The Dual Memory Lifecycle

Persistent memory involves two distinct operations: Recall (retrieving previous context) and Retention (storing resolved interaction records). Keeping the two operations conceptually separate helps avoid retaining unverified hypotheses as future support context.

Recall: Querying Past Experience

When an engineer opens or creates a ticket, the backend queries Hindsight using arecall():

# Recalling customer memories via arecall() in hindsight_service.py
recall_res: RecallResponse = await asyncio.wait_for(
    self._client.arecall(
        bank_id=self.bank_id,
        query=sanitized_query,
        tags=[f"customer:{customer_id}"],
        tags_match="any",
        max_tokens=max_tokens,
        budget=budget
    ),
    timeout=8.0
)
Enter fullscreen mode Exit fullscreen mode

Visual Suggestion 2: Hindsight Recall Implementation

Placement: Position beside the code snippet above, showing the query formation and tag parameter mapping.

The query combines the ticket subject and latest customer message, scoped by customer:{customer_id} with tags_match="any". Customer tags serve as an organizational indexing mechanism within Hindsight; they are not a cryptographic multi-tenant isolation boundary.

Retention: Retaining Resolved Interactions

Retention occurs after a support interaction is resolved and retained—specifically when a specialist clicks Resolve & Retain or triggers /retain:

# Retaining interaction knowledge via aretain() in hindsight_service.py
retain_res: RetainResponse = await asyncio.wait_for(
    self._client.aretain(
        bank_id=self.bank_id,
        content=content,
        document_id=document_id,
        tags=tags,
        metadata=metadata,
        context=f"Support ticket interaction for {customer_name} at {company}"
    ),
    timeout=8.0
)
Enter fullscreen mode Exit fullscreen mode

The application stores what was documented during the interaction; it does not independently verify the technical correctness of the solution.

Deterministic document IDs help prevent duplicate records and allow the same conversation document to be updated when the ticket status changes (cust_{clean_cust_id}_conv_{clean_conv_id}).


5. Debugging the Frontend Boundary: Banishing Silent Mock Fallbacks

Another debugging effort addressed client-side error masking. Early on, frontend/src/services/api.js returned static mock JSON whenever network requests failed.

This created a deceptive testing trap: the FastAPI server could be stopped completely, yet the UI still appeared functional. Backend serialization bugs and connectivity drops went unnoticed.

We eliminated all silent client fallbacks:

  • All methods in api.js now execute genuine fetch requests with AbortSignal.timeout().
  • When the backend is offline, App.jsx renders an explicit amber alert banner with a "Retry Connection" action.
  • In Header.jsx, a dynamic status pill displays Hindsight Live (green), Standby (amber), or Offline (slate).

Surfacing real backend state in the UI eliminated hours of ambiguous debugging.


6. Making Memory Useful: Turning Recalled Memories into Support Context

Support engineers need useful recalled context rather than low-level retrieval output or similarity scores. In CustomerContextPanel.jsx, RecallDesk categorizes recalled memories into actionable sections:

  1. What Worked (Emerald): Solutions reported as effective during previous troubleshooting (e.g., Vault agent config pointing to fullchain.pem).
  2. What Failed (Rose): Known dead ends (e.g., TLS 1.2 protocol downgrade attempts).
  3. Relevant Memories (Indigo): Environment details and general ticket context.

Visual Suggestion 3: RecallDesk Memory Hub

Placement: Place here to illustrate how the right panel categorizes recalled memories into "What Worked" and "What Failed".

The Human-in-the-Loop Workflow

RecallDesk does not dispatch automated responses directly to customers. The current workflow places human review between recalled suggestions and the outgoing customer response:

Customer Problem ──► Recall Previous Experience ──► Suggested Solution ──► Human Reviews & Edits ──► Sends Response ──► Resolve & Retain
Enter fullscreen mode Exit fullscreen mode

When an engineer clicks Use recalled solution, the frontend pre-fills the message composer (composerPrefill in ConversationView.jsx). The specialist reviews, edits, and checks the suggested solution before dispatching it.


7. Security Engineering: Content Sanitization

Support tickets frequently contain credential leaks: tokens, API keys, or private keys. If stored unredacted, these secrets persist across future ticket recalls.

In backend/app/services/hindsight_service.py, sanitize_content() runs regular expression scrubbers before content is passed to aretain():

# Sensitive pattern redactions in hindsight_service.py
(re.compile(r'(?i)(?:password|passwd|pwd|secret)\s*[:=]\s*([^\s\'";,]+)'), r'password=[REDACTED_SECRET]'),
(re.compile(r'(?i)\bbearer\s+[a-zA-Z0-9_\-\.]{20,}\b'), r'[REDACTED_BEARER_TOKEN]'),
(re.compile(r'(?i)(?:api[_-]?key|client[_-]?secret)\s*[:=]\s*([a-zA-Z0-9_\-]{16,})'), r'api_key=[REDACTED_API_KEY]'),
(re.compile(r'\b(?:\d{4}[ -]?){3}\d{4}\b'), r'[REDACTED_CARD_NUMBER]'),
(re.compile(r'(?i)\b(?:otp|one[- ]time code|pin|verification code)\s*[:=]?\s*\d{4,8}\b'), r'[REDACTED_OTP]'),
(re.compile(r'-----BEGIN [A-Z ]+ PRIVATE KEY-----[\s\S]*?-----END [A-Z ]+ PRIVATE KEY-----'), r'[REDACTED_PRIVATE_KEY]')
Enter fullscreen mode Exit fullscreen mode

While regex sanitization provides a baseline mitigation, it is not a complete security guarantee and does not replace enterprise data loss prevention systems.


8. Testing the Real Workflow: End-to-End Verification

To verify persistent memory across separate support sessions, we performed an end-to-end development verification using our seeded customer, Elena Rostova (cust_001):

  1. Open Ticket: Selected ticket #conv_101 ("mTLS handshake failure on ingress gateway during cert rotation").
  2. Inspect Context: Recalled context showing Envoy v1.28+ required fullchain.pem rather than cert.pem.
  3. Review & Send: Reviewed the suggested draft and dispatched the response.
  4. Resolve & Retain: Clicked Resolve & Retain, committing document cust_001_conv_101 to Hindsight bank recalldesk-support.
  5. Create Second Ticket: Opened a new ticket for Elena Rostova: "Ingress gateway rejecting TLS handshakes after cert renewal".
  6. Observe Semantic Recall: Hindsight recalled the fullchain.pem resolution note from #conv_101.
  7. Reuse Solution: Clicked Use recalled solution, reviewed the prefilled draft, and resolved the incident.

Visual Suggestion 4: Second Ticket Recalling Previous Experience

Placement: Insert screenshot showing the newly created ticket displaying the recalled fullchain.pem solution from ticket #conv_101.

This development verification confirmed that Hindsight retains unstructured resolution knowledge from one ticket and surfaces it during a subsequent ticket for the same customer.


9. Verification and Testing Checks

During development, we conducted practical verification checks:

  • Backend Imports: Confirmed clean module imports (hindsight_client, fastapi, pydantic).
  • FastAPI Health Endpoint: Queried /health and /api/v1/health to confirm valid JSON output and Hindsight connection reporting.
  • Hindsight Diagnostics: Tested timeout handling and error formatting when Hindsight was stopped.
  • Frontend Build & Lint: Executed npm run build with Vite 8 and ran npm run lint (oxlint).
  • Browser Verification: Inspected network tabs to ensure status changes and message sends triggered real API requests.

10. What Changed Through Debugging

Dimension Before Debugging After Debugging & Hardening
Frontend Network Silently fell back to mock data on network errors. Real HTTP fetch with timeouts; surfaces explicit offline alerts.
Hindsight Health Connection failures swallowed; memory appeared empty. Health check probes real service version (aget_version()); surfaced in /health.
System Status UI Static status text. Dynamic pill in Header: Hindsight Live, Standby, or Offline.
Document IDs Ad-hoc identifiers risked duplicate retention records. Deterministic IDs (cust_{id}_conv_{id}) help prevent duplicate records upon status updates.
Memory UX Recalled memory was less structured in the UI. Categorized into What Worked, What Failed, and Relevant Memories.
Response Flow Risk of unreviewed automated responses. Suggested solution pre-fills composer; the workflow relies on human review before sending.
Credential Hygiene Raw message strings retained directly. Multi-pattern regex sanitizer redacts secrets prior to aretain().

11. Engineering Lessons Learned

  1. Integrations Must Fail Transparently: Swallowing network exceptions destroys observability. If an external memory service is down, report it loudly while letting core ticketing continue.
  2. Mock Fallbacks Mask Real Bugs: Client-side mock fallbacks hide broken endpoints. Remove them before validating real workflows.
  3. Persistent Memory Needs Dual Lifecycles: Querying memory (arecall()) and persisting memory (aretain()) must have decoupled triggers to avoid indexing unverified hypotheses.
  4. Structured Memory Outperforms Raw Retrieval: Structuring recalled memory into What Worked and What Failed makes context immediately usable for triage.
  5. Human Review Remains Important: Recommending an incorrect configuration in infrastructure support can drop production traffic. AI memory is designed to inform the engineer, not replace them.

12. Honest Limitations

  • In-Memory Store: RecallDesk's operational ticket store (mock_store.py) runs in-memory and resets upon server restart. Production requires an external relational database.
  • Service Availability: Memory capabilities depend on the configured Hindsight server. If Hindsight is offline, RecallDesk operates as a standard ticketing system without memory.
  • Tag Scoping Is Not a Tenant Boundary: Scoping memories using customer:{customer_id} tags indexes search, but does not provide cryptographically isolated multi-tenant data partitioning.
  • Prototype Status: RecallDesk is a development workspace designed to validate persistent AI memory patterns, not an enterprise-certified production deployment.

Conclusion

Building RecallDesk showed that persistent memory transforms complex technical triage. By capturing solutions documented at resolution and recalling them during similar incidents, support teams avoid re-diagnosing solved issues.

Reliability comes down to disciplined engineering: transparent health checks, deterministic document keys, honest UI state, pre-retention sanitization, and human-in-the-loop review.


Resources

Top comments (0)