RecallDesk: Turning Persistent AI Memory into a Practical Support Workspace
When critical infrastructure fails, the engineers diagnosing the outage rarely suffer from a lack of information. Instead, they suffer from fragmented information.
During an active incident, a support specialist typically juggles customer CRM profiles, chat threads, closed ticket archives, and runbooks, all trying to answer one question: Have we seen this failure with this customer before, and what actually fixed it?
Most AI support prototypes attempt to solve this with autonomous chatbots. But in enterprise environments—where misconfiguring a proxy or certificate secret can take down production—autonomous agents introduce unacceptable risk.
In building RecallDesk, our team took a different approach. RecallDesk is an AI-assisted support workspace built around Hindsight, an open-source persistent memory system for AI agents. Rather than hiding memory behind a chatbot, we built a React frontend that surfaces long-term memory directly into the specialist’s natural triage workflow.
This article details how we built the React frontend of RecallDesk: how the three-pane interface organizes context, how live memory recall communicates with our FastAPI backend, and how human-in-the-loop design turns recalled memories into actionable responses.
[IMAGE 1: Full RecallDesk three-pane dashboard]
Designing the Three-Pane Workspace
Enterprise support engineers need context density without visual clutter. RecallDesk organizes the mental model into three coordinated visual panes:
-
Pane 1 (Left: CustomerList): The intake queue. Specialists filter tickets by status (
All,Open,Pending,Resolved), search error codes, track SLA timers, and trigger new ticket creation. -
Pane 2 (Center: ConversationView): The active resolution canvas. Renders chronological message bubbles, resolution controls (
Resolve & Retain,Mark Pending), a suggested solution notification bar, and the message composer. - Pane 3 (Right: CustomerContextPanel): The memory intelligence hub. Displays recalled facts from Hindsight (What Worked and What Failed), customer technical environment specs (cluster versions, cloud regions, tech stacks), and a dynamic activity stream.
Keeping these three dimensions visible simultaneously eliminates context switching. When an incident is selected, the conversation, the customer's infrastructure profile, and their historical troubleshooting memories appear in one unified view.
The Frontend Architecture
RecallDesk’s frontend is built with React 19, Vite, and Tailwind CSS, using Lucide React for icons.
The architecture follows a centralized state coordinator pattern:
React UI (Panes 1, 2, 3)
│
▼
frontend/src/services/api.js (Live Fetch with AbortSignal)
│
▼ HTTP JSON
FastAPI Backend (:8000)
│
▼ Python SDK (arecall / aretain)
Hindsight Cloud Memory Bank (recalldesk-support)
│
▼ Semantic Facts & Document IDs
FastAPI Response (recalled_memories payload)
│
▼ React State Coordination (App.jsx)
CustomerContextPanel & Composer Pre-fill
App.jsx coordinates global state: active conversations, the selected incident ID, health monitoring, and cross-pane state hand-offs.
All network interactions pass through frontend/src/services/api.js, with explicit AbortSignal.timeout controls:
// frontend/src/services/api.js
const BASE_URL = import.meta.env.VITE_API_BASE_URL || 'http://localhost:8000/api/v1';
export const api = {
async getConversation(id) {
const response = await fetch(`${BASE_URL}/conversations/${id}`, {
signal: AbortSignal.timeout(6000)
});
if (!response.ok) {
throw new Error(`Failed to load conversation ${id} (HTTP ${response.status})`);
}
return await response.json();
}
};
Making the UI Honest About Backend State
In early development, our frontend fell back to hardcoded mock data whenever an API request failed.
While mock fallbacks keep a prototype clicking, they are dangerous in an AI support system. An engineer might believe they are viewing genuine customer history when the memory subsystem is completely disconnected.
We eliminated all silent mock fallbacks:
- If the backend or Hindsight API fails,
api.jsthrows an explicit error. - A health probe checks
/healthon startup and every 15 seconds. If the backend is unreachable, an amber warning banner displays across the top:
{!backendConnected && (
<div className="bg-amber-500/15 border-b border-amber-500/30 px-4 py-2 flex items-center justify-between text-xs text-amber-200">
<div className="flex items-center gap-2">
<AlertTriangle className="w-4 h-4 text-amber-400 shrink-0" />
<span>
<strong>FastAPI backend is offline:</strong> Cannot reach{' '}
<code className="bg-amber-950/60 px-1 py-0.5 rounded font-mono">http://localhost:8000/api/v1</code>.
Live memory retrieval will activate once the backend is running.
</span>
</div>
<button onClick={loadData} className="px-2.5 py-1 rounded bg-amber-500/20 text-amber-100 text-[11px]">
Retry Connection
</button>
</div>
)}
Dynamic badges in the header and sidebar display Hindsight Live, Standby, or Offline. If memory cannot be retrieved, the specialist knows immediately.
Turning Memory into UI: "What Worked" vs "What Failed"
When an incident is selected, App.jsx issues a GET /conversations/{id} request. FastAPI calls hindsight_service.arecall(), scoped to the active customer tag (customer:cust_001).
Hindsight returns an array of semantic memory facts extracted from past retained tickets. In CustomerContextPanel.jsx, the frontend consumes this array and organizes the facts for rapid triage.
[IMAGE 2: Memory Intelligence panel showing recalled Hindsight memories]
Hindsight returns raw semantic facts; it does not automatically classify items into "What Worked" or "What Failed". RecallDesk applies deterministic client-side keyword heuristics to sort recalled facts into operational buckets:
// frontend/src/components/CustomerContextPanel.jsx
const recalledMemories = conversation?.recalled_memories || [];
// Categorize verified resolutions vs known dead-ends
const whatWorkedItems = recalledMemories.filter(m => {
const t = (m.text || '').toLowerCase();
return (
t.includes('resolved') ||
t.includes('what worked') ||
t.includes('solution') ||
t.includes('fixed') ||
t.includes('patch') ||
t.includes('fullchain.pem')
);
});
const whatFailedItems = recalledMemories.filter(m => {
const t = (m.text || '').toLowerCase();
return (
t.includes('failed') ||
t.includes('what failed') ||
t.includes('dead-end') ||
t.includes('error') ||
t.includes('reject') ||
t.includes('instead of')
);
});
This classification is simple keyword pattern matching, not autonomous AI cognition. However, for a support engineer under pressure, this structure is practical:
-
Relevant Memories: Lists all recalled facts with source document IDs (e.g.,
#cust_001_conv_101). - What Worked: Highlights verified fixes with emerald badges so the specialist can identify the root cause fix immediately.
- What Failed: Highlights dead-ends with rose badges so engineers avoid repeating failed diagnostics.
- Empty State: If Hindsight returns no recollections, the panel displays "No previous memory found for this customer." rather than inventing synthetic data.
From Memory to Action: "Use Recalled Solution"
Surfacing memory is helpful, but the test of a support workspace is whether it accelerates incident resolution.
Consider customer Elena Rostova at Acme Cloud Infrastructure:
- In Ticket #101, Elena reported that their Envoy ingress gateway was rejecting mutual TLS handshakes with
SSL_ERROR_UNKNOWN_CA_ALERT. - Troubleshooting revealed that Vault was writing only the leaf
cert.pemto the secret mount. Pointing Envoy tofullchain.pemresolved the error. - The conversation was resolved and retained into Hindsight.
When Elena opens a subsequent ticket regarding certificate errors, Hindsight recalls the fullchain.pem resolution.
[IMAGE 3: Suggested solution / specialist composer]
RecallDesk provides a "Use recalled solution" action button. Clicking it formats a suggested response and routes it into the composer via composerPrefill:
// frontend/src/components/ConversationView.jsx
export function ConversationView({ composerPrefill = '', onClearComposerPrefill, ...props }) {
const [inputText, setInputText] = useState('');
// Synchronize composer with recalled solution hand-off
useEffect(() => {
if (composerPrefill) {
setInputText(composerPrefill);
if (onClearComposerPrefill) onClearComposerPrefill();
}
}, [composerPrefill, onClearComposerPrefill]);
// ...
}
Crucially, the frontend does not automatically send the message.
The draft appears inside the composer accompanied by an indicator: ✨ Specialist Composer (Review and edit before sending). The specialist can review the guidance, adjust details specific to the new ticket, and manually click Send.
This human-in-the-loop pattern provides the speed of automated memory recall while keeping the engineer accountable for production recommendations.
Creating the Second Ticket
To test cross-ticket persistence, our team implemented a complete + New Ticket workflow in CustomerList.jsx.
Specialists click + New Ticket to open a modal where they select an existing customer, set priority, enter a subject ("Production ingress rejecting mutual TLS after certificate rotation"), and supply the customer's message.
[IMAGE 4: New Ticket modal or retention confirmation]
This demonstrates the difference between stateless ticketing and memory-augmented workspaces:
- Without Persistent Memory: The new ticket opens empty. The specialist sees only the fresh complaint and must repeat preliminary diagnostic questions from scratch.
-
With RecallDesk and Hindsight: The moment the ticket is created, the backend executes semantic recall. In our tested flow, the ticket opened with 12 recalled facts from Elena's previous incident already populated in the right pane, and the
fullchain.pemsolution ready to insert into the composer.
Showing That Memory Was Actually Saved
Persistent memory systems often suffer from an invisibility problem: users perform actions, but cannot tell whether the system actually retained anything.
When an incident is resolved in RecallDesk, the frontend issues a PATCH /conversations/{id}/status request. The backend triggers hindsight_service.aretain(), preserving the sanitized conversation and troubleshooting findings into the customer's permanent bank.
Once confirmed, ConversationView.jsx mounts a visible confirmation banner:
{retentionNotice && (
<div className="px-6 py-2.5 bg-emerald-950/40 border-b border-emerald-500/30 flex items-center justify-between text-xs animate-in fade-in">
<div className="flex items-center gap-2 text-emerald-300">
<CheckCircle2 className="w-4 h-4 text-emerald-400 shrink-0" />
<span>
<strong>Memory saved to Hindsight:</strong> Customer memory updated in bank{' '}
<code className="bg-emerald-900/60 px-1 py-0.5 rounded font-mono text-emerald-200">
{retentionNotice.bankId}
</code>{' '}
(Doc: <code className="font-mono text-[10px] text-emerald-300">{retentionNotice.docId}</code>).
</span>
</div>
<button onClick={() => setRetentionNotice(null)} className="text-[10px] text-emerald-400 underline font-mono">
Dismiss
</button>
</div>
)}
Simultaneously, the live activity stream in the right panel logs an audit event: "Ticket #conv_101 marked resolved & memory retained in Hindsight". This provides clear confirmation that the fix is now permanently indexed.
Honest Frontend Limitations
While the current interface provides a functional prototype, several engineering limitations should be noted:
- Client-Side Keyword Heuristics: Our "What Worked" and "What Failed" buckets rely on keyword matching. If Hindsight returns a fact phrased unconventionally, it displays safely in "Relevant Memories" but may miss the specific "What Worked" bucket.
-
Ephemeral Ticket Store: While Hindsight stores vector memories permanently in the cloud, our backend ticket store (
mock_store.py) is held in-memory and resets upon server restart. - Desktop-First Layout: The three-pane layout is optimized for widescreen monitors (> 1280px). On smaller screens, panels must be collapsed.
- Request/Response Model: The application uses fetch requests and polling rather than WebSockets or Server-Sent Events (SSE). Real-time typing indicators are not yet supported.
- Single-Agent Session: The interface models a single active specialist session ("Alex Vance") without multi-user concurrent collision handling.
Conclusion
The core challenge of AI in customer support is not generating text; it is managing context.
If an AI tool operates without long-term memory, customers must repeat themselves, and engineers re-solve problems that were already fixed weeks earlier. Conversely, if an AI agent is given autonomous authority without review, it risks deploying invalid recommendations to production infrastructure.
RecallDesk demonstrates a practical middle ground:
Past Customer Experience -> Persistent Hindsight Memory -> Actionable Workspace UI -> Human Review
By anchoring our React frontend around transparent system status, dynamic memory intelligence, and a human-in-the-loop composer hand-off, persistent AI memory becomes a practical tool that augments engineering teams rather than replacing them.
Resources and Further Reading
- Explore the Hindsight GitHub repository to inspect the persistent memory client and engine.
- Read the official Hindsight documentation for guides on configuring banks, document retention, and semantic recall.
- Learn more about long-term context retention from Vectorize agent memory.

Top comments (0)