RecallDesk: Building the FastAPI Backend Behind an AI Support Memory System
In customer support applications, the frontend is where engineers interact with tickets, but the backend is where architectural integrity is enforced. When adding a persistent cognitive memory layer to an enterprise helpdesk, the backend must coordinate three distinct responsibilities: managing operational ticket state, sanitizing dialogue before it leaves the internal boundary, and interfacing with an external persistent memory engine.
In RecallDesk, our team built the backend using Python and FastAPI. The backend acts as the central coordinator between a React 19 workspace, an operational customer and ticket repository, and an external persistent memory subsystem powered by Hindsight.
In this article, we walk through the backend engineering of RecallDesk: why we chose FastAPI, how we structured routes and Pydantic models, how we decoupled operational state from persistent memory, and how we handle asynchronous integration and failure modes in practice.
Why RecallDesk Needed a Backend
When designing AI-assisted support tools, it is tempting to connect the client directly to third-party APIs from the browser. For an enterprise support application, however, a dedicated backend service is essential for several reasons:
- Centralizing Secrets and Redaction on the Server: Support tickets frequently contain sensitive values like API keys, tokens, or credentials. Keeping memory service API keys on the server and routing all interactions through the backend ensures dialogue can pass through a redaction pipeline before leaving the application environment.
- Coordination of Divergent Data Lifecycles: Support tickets have rigid operational lifecycles (open, pending, resolved), while customer memory is continuous. A single action—such as sending a message—must write to the immediate ticket thread while also persisting updated dialogue to the memory bank via an asynchronous client call.
- Resilient Fallback Handling: If the external memory service experiences transient latency or downtime, the support desk must continue operating. The backend acts as a defensive buffer, reporting memory status cleanly without disrupting ticket workflows.
FastAPI as the API Layer
We selected FastAPI (fastapi>=0.115.0) running on Uvicorn (uvicorn[standard]>=0.32.0) as our backend framework. FastAPI provides several practical advantages for our service requirements:
-
Native Asynchronous I/O: Calls to external memory services are I/O-bound network requests. FastAPI's native support for
asyncandawaitallows our backend to handle concurrent requests without blocking worker threads. - Schema Validation via Pydantic v2: Automatic request parsing, type enforcement, and serialization reduce boilerplate and ensure invalid inputs are rejected before reaching business logic.
-
Interactive Documentation: Automatic OpenAPI generation (
/docsand/redoc) simplifies testing and endpoint inspection during development.
Centralized Configuration
Backend settings are centralized in backend/app/core/config.py using Pydantic's BaseModel and dotenv:
# backend/app/core/config.py
import os
from typing import List
from pydantic import BaseModel
from dotenv import load_dotenv
load_dotenv()
class Settings(BaseModel):
PROJECT_NAME: str = os.getenv("PROJECT_NAME", "RecallDesk API")
API_V1_PREFIX: str = os.getenv("API_V1_PREFIX", "/api/v1")
HOST: str = os.getenv("HOST", "127.0.0.1")
PORT: int = int(os.getenv("PORT", "8001"))
ENVIRONMENT: str = os.getenv("ENVIRONMENT", "development")
DEBUG: bool = os.getenv("DEBUG", "true").lower() in ("true", "1", "yes")
CORS_ORIGINS: List[str] = [
origin.strip()
for origin in os.getenv(
"CORS_ORIGINS",
"http://localhost:5173,http://127.0.0.1:5173,http://localhost:3000,http://localhost:5174,http://127.0.0.1:5174"
).split(",")
if origin.strip()
]
HINDSIGHT_BASE_URL: str = os.getenv("HINDSIGHT_BASE_URL", "https://api.hindsight.vectorize.io")
HINDSIGHT_API_KEY: str | None = os.getenv("HINDSIGHT_API_KEY", None)
HINDSIGHT_BANK_ID: str = os.getenv("HINDSIGHT_BANK_ID", "recalldesk-support")
settings = Settings()
CORS middleware in backend/app/main.py explicitly allows requests from the Vite development server origins defined in CORS_ORIGINS.
Organizing Backend Routes

To keep endpoint handlers focused, routes are partitioned by domain under backend/app/api/routes/ and combined in backend/app/api/__init__.py:
# backend/app/api/__init__.py
from fastapi import APIRouter
from app.api.routes import health_router, customers_router, conversations_router
api_router = APIRouter()
api_router.include_router(health_router)
api_router.include_router(customers_router)
api_router.include_router(conversations_router)
In main.py, this master router is mounted under the versioned prefix settings.API_V1_PREFIX (/api/v1), keeping endpoints organized and versioned:
-
/healthand/api/v1/health: Service liveness and Hindsight connection verification. -
/api/v1/customers: Customer listing and detailed profile retrieval. -
/api/v1/customers/{customer_id}/context: Account metadata and memory connection status. -
/api/v1/conversations: Filterable ticket queries bycustomer_idandstatus. -
/api/v1/conversations/{conversation_id}: Ticket transcript and message history, which automatically performs semantic memory recall for the active issue via Hindsight. -
/api/v1/conversations/{conversation_id}/messages: Message creation, prior context recall, thread update, and async Hindsight retention. -
/api/v1/conversations/{conversation_id}/status: Ticket status updates (open,pending,resolved) with outcome retention. -
/api/v1/conversations/{conversation_id}/memories: Explicit query-based semantic recall endpoint. -
/api/v1/conversations/{conversation_id}/retain: Manual retention trigger.
Pydantic Models and Request/Response Validation
Data entering and leaving the backend is strictly modeled with Pydantic v2. This guarantees consistent API contracts between the FastAPI backend and the React frontend.
In backend/app/models/conversation.py, we define schemas for both operational ticket structures and recalled memory items:
# backend/app/models/conversation.py
from typing import Any, Dict, List, Optional
from pydantic import BaseModel, Field
class RecalledMemoryItem(BaseModel):
id: Optional[str] = None
text: str
type: str = "fact"
document_id: Optional[str] = None
tags: List[str] = Field(default_factory=list)
metadata: Dict[str, Any] = Field(default_factory=dict)
scores: Optional[Any] = None
class Message(BaseModel):
id: str
conversation_id: str
sender_type: str = Field(..., description="'customer' or 'agent'")
sender_name: str
text: str
timestamp: str
is_automated: bool = False
recalled_context: Optional[List[RecalledMemoryItem]] = None
class SendMessageRequest(BaseModel):
text: str = Field(..., min_length=1, description="Message text content")
sender_name: Optional[str] = "Support Specialist"
class UpdateConversationStatusRequest(BaseModel):
status: str = Field(..., description="Target status: 'open', 'pending', 'resolved'")
When a specialist posts a reply, FastAPI validates incoming JSON against SendMessageRequest. If the payload is empty or missing required fields, FastAPI's validation exception handler returns a structured 422 Unprocessable Entity JSON response before any service logic executes.
Connecting the Backend to Hindsight
Communication with Hindsight is encapsulated in a dedicated singleton service class: HindsightMemoryService in backend/app/services/hindsight_service.py.
Client Initialization
The service initializes the official hindsight-client using configured settings:
# backend/app/services/hindsight_service.py
from typing import Optional
from hindsight_client import Hindsight
from app.core.config import settings
class HindsightMemoryService:
def __init__(self):
self.base_url = settings.HINDSIGHT_BASE_URL
self.api_key = settings.HINDSIGHT_API_KEY
self.bank_id = settings.HINDSIGHT_BANK_ID
self._client: Optional[Hindsight] = None
self._init_client()
def _init_client(self):
try:
self._client = Hindsight(
base_url=self.base_url,
api_key=self.api_key if self.api_key else None,
timeout=15.0,
user_agent="RecallDesk-Support/0.1.0"
)
except Exception:
self._client = None
Pre-Ingestion Regex Sanitization
Support interactions frequently include secrets such as private keys, API credentials, or bearer tokens. To protect customer privacy before storing data in an external memory bank, sanitize_content() passes text through six compiled regular expressions:
# backend/app/services/hindsight_service.py
SENSITIVE_PATTERNS = [
(re.compile(r'(?i)(?:password|passwd|pwd|secret)\s*[:=]\s*([^\s\'";,]+)'), r'password=[REDACTED_SECRET]'),
(re.compile(r'(?i)\bbearer\s+[a-zA-Z0-9_\-\.]{20,}\b'), r'[REDACTED_BEARER_TOKEN]'),
(re.compile(r'(?i)(?:api[_-]?key|client[_-]?secret)\s*[:=]\s*([a-zA-Z0-9_\-]{16,})'), r'api_key=[REDACTED_API_KEY]'),
(re.compile(r'\b(?:\d{4}[ -]?){3}\d{4}\b'), r'[REDACTED_CARD_NUMBER]'),
(re.compile(r'(?i)\b(?:otp|one[- ]time code|pin|verification code)\s*[:=]?\s*\d{4,8}\b'), r'[REDACTED_OTP]'),
(re.compile(r'-----BEGIN [A-Z ]+ PRIVATE KEY-----[\s\S]*?-----END [A-Z ]+ PRIVATE KEY-----'), r'[REDACTED_PRIVATE_KEY]')
]
def sanitize_content(text: str) -> str:
if not text:
return ""
sanitized = text
for pattern, replacement in SENSITIVE_PATTERNS:
sanitized = pattern.sub(replacement, sanitized)
return sanitized
Operational Support Data vs. Persistent Memory
A key design principle in RecallDesk is the structural separation between operational support data and persistent cognitive memory:
-
Operational Data (
backend/app/services/mock_store.py): Represents immediate ticketing state—customer profiles, open tickets, assigned categories, and message sequences. In our current implementation, this is stored in memory indata_store. -
Persistent Memory (
backend/app/services/hindsight_service.py): Represents long-term cognitive context—verified solutions, disproven hypotheses, and customer environmental constraints. This is indexed externally in the Hindsight memory bank (recalldesk-support).
This decoupling keeps operational ticket queries fast and isolated from external network calls, while allowing memory to persist across backend restarts when connected to an external Hindsight instance.
Async Processing and Failure Handling
Because Hindsight is an external service, network latency or server downtime must not crash our API routes. We implemented two defensive patterns:
1. Multi-Tiered Timeouts and Asynchronous Safeguards
Client initialization in HindsightMemoryService sets a 15.0-second HTTP request timeout on the client instance (timeout=15.0). At the application layer, our service further bounds individual calls to arecall(), aretain(), and aget_version() with asyncio.wait_for(..., timeout=8.0):
# backend/app/services/hindsight_service.py
recall_res = await asyncio.wait_for(
self._client.arecall(
bank_id=self.bank_id,
query=sanitized_query,
tags=[f"customer:{customer_id}"],
tags_match="any",
max_tokens=max_tokens,
budget=budget
),
timeout=8.0
)
If the call exceeds 8.0 seconds, the service catches asyncio.TimeoutError and logs a warning rather than letting the worker hang indefinitely.
2. Live Connection Status Verification
Rather than hardcoding status flags, our health check endpoint calls aget_version() to verify true end-to-end connectivity:
# backend/app/services/hindsight_service.py
async def get_connection_status(self) -> Dict[str, Any]:
if not self._client:
return {"connected": False, "status": "uninitialized", ...}
try:
version_response = await asyncio.wait_for(self._client.aget_version(), timeout=8.0)
api_version = getattr(version_response, "api_version", "unknown")
return {
"connected": True,
"status": f"connected (Hindsight API v{api_version})",
"version": api_version,
"bank_id": self.bank_id,
"base_url": self.base_url
}
except asyncio.TimeoutError:
return {"connected": False, "status": "unavailable (connection timed out)", ...}
except Exception as exc:
return {"connected": False, "status": f"unavailable ({type(exc).__name__})", ...}
In backend/app/api/routes/health.py, this information is surfaced directly to /health, allowing the frontend to reflect real backend-to-Hindsight connectivity.
Example Request Flow: Opening a Customer Conversation
When a support specialist selects a ticket in the React workspace, the frontend initiates a request to load the conversation:
-
Frontend Call: The client triggers
GET /api/v1/conversations/conv_101. -
Operational Lookup: The endpoint queries
data_store.get_conversation("conv_101"). If not found, it raises an HTTP 404 error. - Formulating the Recall Query: The route inspects the conversation. If customer messages exist, it combines the ticket subject with the latest customer message into a semantic search query:
query = f"{conversation.subject} - {customer_msgs[-1].text}"
-
Scoped Recall Execution: The route invokes
hindsight_service.recall_customer_memories()withtags=[f"customer:{conversation.customer_id}"]. This tag helps scope recall toward memories associated with the active customer (usingtags_match="any"). -
Response Enrichment: The recalled memory items are validated against
RecalledMemoryItemand assigned toconversation.recalled_memories. -
Payload Delivery: The enriched
Conversationmodel is returned to the client, delivering the ticket messages and historical memory in a single HTTP response.
Example Request Flow: Recalling Customer Context and Retaining Responses
When a support engineer replies to an active incident, the message route coordinates both operational appending and cognitive retention:
# backend/app/api/routes/conversations.py
@router.post("/{conversation_id}/messages", response_model=Message, status_code=status.HTTP_201_CREATED)
async def send_message(conversation_id: str, payload: SendMessageRequest):
conversation = data_store.get_conversation(conversation_id)
if not conversation:
raise HTTPException(
status_code=status.HTTP_404_NOT_FOUND,
detail=f"Conversation with ID '{conversation_id}' was not found."
)
# 1. Recall prior customer context to assist the specialist
query = payload.text
if conversation.messages:
customer_msgs = [m for m in conversation.messages if m.sender_type == "customer"]
if customer_msgs:
query = f"{conversation.subject} - {customer_msgs[-1].text}"
recall_res = await hindsight_service.recall_customer_memories(
customer_id=conversation.customer_id,
query=query,
conversation_id=conversation.id
)
# 2. Append support reply to operational thread
message = data_store.add_message(
conversation_id=conversation_id,
text=payload.text,
sender_name=payload.sender_name or "Support Specialist",
sender_type="agent"
)
# 3. Retain updated dialogue in Hindsight memory bank
customer = data_store.get_customer(conversation.customer_id)
customer_context = {
"name": customer.name if customer else conversation.customer_name,
"company": customer.company if customer else conversation.customer_company,
"tier": customer.tier if customer else "Standard",
"environment": customer.environment if customer else "Production"
}
await hindsight_service.retain_customer_conversation(
customer_id=conversation.customer_id,
conversation_id=conversation.id,
messages=[m.model_dump() for m in conversation.messages],
subject=conversation.subject,
category=conversation.category,
status=conversation.status,
priority=conversation.priority,
customer_context=customer_context
)
return message
This ensures that the operational ticket record updates immediately, while the conversation is synthesized, sanitized, and retained in the Hindsight memory bank via an awaited asynchronous call before returning the response.
Current Backend Limitations
To accurately reflect the repository's state, several engineering limitations should be noted:
-
In-Memory Operational Data Store: Operational customer records and ticket threads are maintained in memory via
mock_store.py. While Hindsight memories persist across restarts when connected to an external service, local ticket changes reset upon backend restart. -
Tag-Based Scoping Semantics:
tags=[f"customer:{customer_id}"]helps scope recall toward memories associated with the active customer. The current implementation usestags_match="any", so this tag filter should not be treated as a strict isolation or security boundary. - No Autonomous Backend LLM Loop: The backend performs memory retrieval and retention to inform human specialists; it does not currently execute an autonomous generative loop to dispatch automated replies without human review.
- No Authentication or Authorization Middleware: The API currently runs without JWT, session cookies, OAuth, or RBAC authorization. It is structured as an internal development service.
Conclusion
Building persistent memory into enterprise support requires more than just calling vector search from a frontend script. By structuring RecallDesk around an asynchronous FastAPI backend, we decoupled operational ticketing state from persistent cognitive memory, enforced pre-ingestion regex sanitization, and established resilient failure handling.
This architecture ensures that customer support workflows remain fast and reliable, while giving human specialists immediate access to historical context when resolving critical customer incidents.
Top comments (0)