DEV Community

Cover image for RecallDesk: Building the FastAPI Backend Behind an AI Support Memory System
T Deekshitha
T Deekshitha

Posted on

RecallDesk: Building the FastAPI Backend Behind an AI Support Memory System

RecallDesk: Building the FastAPI Backend Behind an AI Support Memory System

In customer support applications, the frontend is where engineers interact with tickets, but the backend is where architectural integrity is enforced. When adding a persistent cognitive memory layer to an enterprise helpdesk, the backend must coordinate three distinct responsibilities: managing operational ticket state, sanitizing dialogue before it leaves the internal boundary, and interfacing with an external persistent memory engine.

In RecallDesk, our team built the backend using Python and FastAPI. The backend acts as the central coordinator between a React 19 workspace, an operational customer and ticket repository, and an external persistent memory subsystem powered by Hindsight.

In this article, we walk through the backend engineering of RecallDesk: why we chose FastAPI, how we structured routes and Pydantic models, how we decoupled operational state from persistent memory, and how we handle asynchronous integration and failure modes in practice.


Why RecallDesk Needed a Backend

When designing AI-assisted support tools, it is tempting to connect the client directly to third-party APIs from the browser. For an enterprise support application, however, a dedicated backend service is essential for several reasons:

  • Centralizing Secrets and Redaction on the Server: Support tickets frequently contain sensitive values like API keys, tokens, or credentials. Keeping memory service API keys on the server and routing all interactions through the backend ensures dialogue can pass through a redaction pipeline before leaving the application environment.
  • Coordination of Divergent Data Lifecycles: Support tickets have rigid operational lifecycles (open, pending, resolved), while customer memory is continuous. A single action—such as sending a message—must write to the immediate ticket thread while also persisting updated dialogue to the memory bank via an asynchronous client call.
  • Resilient Fallback Handling: If the external memory service experiences transient latency or downtime, the support desk must continue operating. The backend acts as a defensive buffer, reporting memory status cleanly without disrupting ticket workflows.

FastAPI as the API Layer

We selected FastAPI (fastapi>=0.115.0) running on Uvicorn (uvicorn[standard]>=0.32.0) as our backend framework. FastAPI provides several practical advantages for our service requirements:

  1. Native Asynchronous I/O: Calls to external memory services are I/O-bound network requests. FastAPI's native support for async and await allows our backend to handle concurrent requests without blocking worker threads.
  2. Schema Validation via Pydantic v2: Automatic request parsing, type enforcement, and serialization reduce boilerplate and ensure invalid inputs are rejected before reaching business logic.
  3. Interactive Documentation: Automatic OpenAPI generation (/docs and /redoc) simplifies testing and endpoint inspection during development.

Centralized Configuration

Backend settings are centralized in backend/app/core/config.py using Pydantic's BaseModel and dotenv:

# backend/app/core/config.py
import os
from typing import List
from pydantic import BaseModel
from dotenv import load_dotenv

load_dotenv()

class Settings(BaseModel):
    PROJECT_NAME: str = os.getenv("PROJECT_NAME", "RecallDesk API")
    API_V1_PREFIX: str = os.getenv("API_V1_PREFIX", "/api/v1")
    HOST: str = os.getenv("HOST", "127.0.0.1")
    PORT: int = int(os.getenv("PORT", "8001"))
    ENVIRONMENT: str = os.getenv("ENVIRONMENT", "development")
    DEBUG: bool = os.getenv("DEBUG", "true").lower() in ("true", "1", "yes")

    CORS_ORIGINS: List[str] = [
        origin.strip()
        for origin in os.getenv(
            "CORS_ORIGINS",
            "http://localhost:5173,http://127.0.0.1:5173,http://localhost:3000,http://localhost:5174,http://127.0.0.1:5174"
        ).split(",")
        if origin.strip()
    ]

    HINDSIGHT_BASE_URL: str = os.getenv("HINDSIGHT_BASE_URL", "https://api.hindsight.vectorize.io")
    HINDSIGHT_API_KEY: str | None = os.getenv("HINDSIGHT_API_KEY", None)
    HINDSIGHT_BANK_ID: str = os.getenv("HINDSIGHT_BANK_ID", "recalldesk-support")

settings = Settings()
Enter fullscreen mode Exit fullscreen mode

CORS middleware in backend/app/main.py explicitly allows requests from the Vite development server origins defined in CORS_ORIGINS.


Organizing Backend Routes

RecallDesk API and FastAPI backend
To keep endpoint handlers focused, routes are partitioned by domain under backend/app/api/routes/ and combined in backend/app/api/__init__.py:

# backend/app/api/__init__.py
from fastapi import APIRouter
from app.api.routes import health_router, customers_router, conversations_router

api_router = APIRouter()
api_router.include_router(health_router)
api_router.include_router(customers_router)
api_router.include_router(conversations_router)
Enter fullscreen mode Exit fullscreen mode

In main.py, this master router is mounted under the versioned prefix settings.API_V1_PREFIX (/api/v1), keeping endpoints organized and versioned:

  • /health and /api/v1/health: Service liveness and Hindsight connection verification.
  • /api/v1/customers: Customer listing and detailed profile retrieval.
  • /api/v1/customers/{customer_id}/context: Account metadata and memory connection status.
  • /api/v1/conversations: Filterable ticket queries by customer_id and status.
  • /api/v1/conversations/{conversation_id}: Ticket transcript and message history, which automatically performs semantic memory recall for the active issue via Hindsight.
  • /api/v1/conversations/{conversation_id}/messages: Message creation, prior context recall, thread update, and async Hindsight retention.
  • /api/v1/conversations/{conversation_id}/status: Ticket status updates (open, pending, resolved) with outcome retention.
  • /api/v1/conversations/{conversation_id}/memories: Explicit query-based semantic recall endpoint.
  • /api/v1/conversations/{conversation_id}/retain: Manual retention trigger.

Pydantic Models and Request/Response Validation

Data entering and leaving the backend is strictly modeled with Pydantic v2. This guarantees consistent API contracts between the FastAPI backend and the React frontend.

In backend/app/models/conversation.py, we define schemas for both operational ticket structures and recalled memory items:

# backend/app/models/conversation.py
from typing import Any, Dict, List, Optional
from pydantic import BaseModel, Field

class RecalledMemoryItem(BaseModel):
    id: Optional[str] = None
    text: str
    type: str = "fact"
    document_id: Optional[str] = None
    tags: List[str] = Field(default_factory=list)
    metadata: Dict[str, Any] = Field(default_factory=dict)
    scores: Optional[Any] = None

class Message(BaseModel):
    id: str
    conversation_id: str
    sender_type: str = Field(..., description="'customer' or 'agent'")
    sender_name: str
    text: str
    timestamp: str
    is_automated: bool = False
    recalled_context: Optional[List[RecalledMemoryItem]] = None

class SendMessageRequest(BaseModel):
    text: str = Field(..., min_length=1, description="Message text content")
    sender_name: Optional[str] = "Support Specialist"

class UpdateConversationStatusRequest(BaseModel):
    status: str = Field(..., description="Target status: 'open', 'pending', 'resolved'")
Enter fullscreen mode Exit fullscreen mode

When a specialist posts a reply, FastAPI validates incoming JSON against SendMessageRequest. If the payload is empty or missing required fields, FastAPI's validation exception handler returns a structured 422 Unprocessable Entity JSON response before any service logic executes.


Connecting the Backend to Hindsight

Communication with Hindsight is encapsulated in a dedicated singleton service class: HindsightMemoryService in backend/app/services/hindsight_service.py.

Client Initialization

The service initializes the official hindsight-client using configured settings:

# backend/app/services/hindsight_service.py
from typing import Optional
from hindsight_client import Hindsight
from app.core.config import settings

class HindsightMemoryService:
    def __init__(self):
        self.base_url = settings.HINDSIGHT_BASE_URL
        self.api_key = settings.HINDSIGHT_API_KEY
        self.bank_id = settings.HINDSIGHT_BANK_ID
        self._client: Optional[Hindsight] = None
        self._init_client()

    def _init_client(self):
        try:
            self._client = Hindsight(
                base_url=self.base_url,
                api_key=self.api_key if self.api_key else None,
                timeout=15.0,
                user_agent="RecallDesk-Support/0.1.0"
            )
        except Exception:
            self._client = None
Enter fullscreen mode Exit fullscreen mode

Pre-Ingestion Regex Sanitization

Support interactions frequently include secrets such as private keys, API credentials, or bearer tokens. To protect customer privacy before storing data in an external memory bank, sanitize_content() passes text through six compiled regular expressions:

# backend/app/services/hindsight_service.py
SENSITIVE_PATTERNS = [
    (re.compile(r'(?i)(?:password|passwd|pwd|secret)\s*[:=]\s*([^\s\'";,]+)'), r'password=[REDACTED_SECRET]'),
    (re.compile(r'(?i)\bbearer\s+[a-zA-Z0-9_\-\.]{20,}\b'), r'[REDACTED_BEARER_TOKEN]'),
    (re.compile(r'(?i)(?:api[_-]?key|client[_-]?secret)\s*[:=]\s*([a-zA-Z0-9_\-]{16,})'), r'api_key=[REDACTED_API_KEY]'),
    (re.compile(r'\b(?:\d{4}[ -]?){3}\d{4}\b'), r'[REDACTED_CARD_NUMBER]'),
    (re.compile(r'(?i)\b(?:otp|one[- ]time code|pin|verification code)\s*[:=]?\s*\d{4,8}\b'), r'[REDACTED_OTP]'),
    (re.compile(r'-----BEGIN [A-Z ]+ PRIVATE KEY-----[\s\S]*?-----END [A-Z ]+ PRIVATE KEY-----'), r'[REDACTED_PRIVATE_KEY]')
]

def sanitize_content(text: str) -> str:
    if not text:
        return ""
    sanitized = text
    for pattern, replacement in SENSITIVE_PATTERNS:
        sanitized = pattern.sub(replacement, sanitized)
    return sanitized
Enter fullscreen mode Exit fullscreen mode

Operational Support Data vs. Persistent Memory

A key design principle in RecallDesk is the structural separation between operational support data and persistent cognitive memory:

  • Operational Data (backend/app/services/mock_store.py): Represents immediate ticketing state—customer profiles, open tickets, assigned categories, and message sequences. In our current implementation, this is stored in memory in data_store.
  • Persistent Memory (backend/app/services/hindsight_service.py): Represents long-term cognitive context—verified solutions, disproven hypotheses, and customer environmental constraints. This is indexed externally in the Hindsight memory bank (recalldesk-support).

This decoupling keeps operational ticket queries fast and isolated from external network calls, while allowing memory to persist across backend restarts when connected to an external Hindsight instance.


Async Processing and Failure Handling

Because Hindsight is an external service, network latency or server downtime must not crash our API routes. We implemented two defensive patterns:

1. Multi-Tiered Timeouts and Asynchronous Safeguards

Client initialization in HindsightMemoryService sets a 15.0-second HTTP request timeout on the client instance (timeout=15.0). At the application layer, our service further bounds individual calls to arecall(), aretain(), and aget_version() with asyncio.wait_for(..., timeout=8.0):

# backend/app/services/hindsight_service.py
recall_res = await asyncio.wait_for(
    self._client.arecall(
        bank_id=self.bank_id,
        query=sanitized_query,
        tags=[f"customer:{customer_id}"],
        tags_match="any",
        max_tokens=max_tokens,
        budget=budget
    ),
    timeout=8.0
)
Enter fullscreen mode Exit fullscreen mode

If the call exceeds 8.0 seconds, the service catches asyncio.TimeoutError and logs a warning rather than letting the worker hang indefinitely.

2. Live Connection Status Verification

Rather than hardcoding status flags, our health check endpoint calls aget_version() to verify true end-to-end connectivity:

# backend/app/services/hindsight_service.py
async def get_connection_status(self) -> Dict[str, Any]:
    if not self._client:
        return {"connected": False, "status": "uninitialized", ...}

    try:
        version_response = await asyncio.wait_for(self._client.aget_version(), timeout=8.0)
        api_version = getattr(version_response, "api_version", "unknown")
        return {
            "connected": True,
            "status": f"connected (Hindsight API v{api_version})",
            "version": api_version,
            "bank_id": self.bank_id,
            "base_url": self.base_url
        }
    except asyncio.TimeoutError:
        return {"connected": False, "status": "unavailable (connection timed out)", ...}
    except Exception as exc:
        return {"connected": False, "status": f"unavailable ({type(exc).__name__})", ...}
Enter fullscreen mode Exit fullscreen mode

In backend/app/api/routes/health.py, this information is surfaced directly to /health, allowing the frontend to reflect real backend-to-Hindsight connectivity.


Example Request Flow: Opening a Customer Conversation

When a support specialist selects a ticket in the React workspace, the frontend initiates a request to load the conversation:

  1. Frontend Call: The client triggers GET /api/v1/conversations/conv_101.
  2. Operational Lookup: The endpoint queries data_store.get_conversation("conv_101"). If not found, it raises an HTTP 404 error.
  3. Formulating the Recall Query: The route inspects the conversation. If customer messages exist, it combines the ticket subject with the latest customer message into a semantic search query:
   query = f"{conversation.subject} - {customer_msgs[-1].text}"
Enter fullscreen mode Exit fullscreen mode
  1. Scoped Recall Execution: The route invokes hindsight_service.recall_customer_memories() with tags=[f"customer:{conversation.customer_id}"]. This tag helps scope recall toward memories associated with the active customer (using tags_match="any").
  2. Response Enrichment: The recalled memory items are validated against RecalledMemoryItem and assigned to conversation.recalled_memories.
  3. Payload Delivery: The enriched Conversation model is returned to the client, delivering the ticket messages and historical memory in a single HTTP response.

Example Request Flow: Recalling Customer Context and Retaining Responses

When a support engineer replies to an active incident, the message route coordinates both operational appending and cognitive retention:

# backend/app/api/routes/conversations.py
@router.post("/{conversation_id}/messages", response_model=Message, status_code=status.HTTP_201_CREATED)
async def send_message(conversation_id: str, payload: SendMessageRequest):
    conversation = data_store.get_conversation(conversation_id)
    if not conversation:
        raise HTTPException(
            status_code=status.HTTP_404_NOT_FOUND,
            detail=f"Conversation with ID '{conversation_id}' was not found."
        )

    # 1. Recall prior customer context to assist the specialist
    query = payload.text
    if conversation.messages:
        customer_msgs = [m for m in conversation.messages if m.sender_type == "customer"]
        if customer_msgs:
            query = f"{conversation.subject} - {customer_msgs[-1].text}"

    recall_res = await hindsight_service.recall_customer_memories(
        customer_id=conversation.customer_id,
        query=query,
        conversation_id=conversation.id
    )

    # 2. Append support reply to operational thread
    message = data_store.add_message(
        conversation_id=conversation_id,
        text=payload.text,
        sender_name=payload.sender_name or "Support Specialist",
        sender_type="agent"
    )

    # 3. Retain updated dialogue in Hindsight memory bank
    customer = data_store.get_customer(conversation.customer_id)
    customer_context = {
        "name": customer.name if customer else conversation.customer_name,
        "company": customer.company if customer else conversation.customer_company,
        "tier": customer.tier if customer else "Standard",
        "environment": customer.environment if customer else "Production"
    }

    await hindsight_service.retain_customer_conversation(
        customer_id=conversation.customer_id,
        conversation_id=conversation.id,
        messages=[m.model_dump() for m in conversation.messages],
        subject=conversation.subject,
        category=conversation.category,
        status=conversation.status,
        priority=conversation.priority,
        customer_context=customer_context
    )

    return message
Enter fullscreen mode Exit fullscreen mode

This ensures that the operational ticket record updates immediately, while the conversation is synthesized, sanitized, and retained in the Hindsight memory bank via an awaited asynchronous call before returning the response.


Current Backend Limitations

To accurately reflect the repository's state, several engineering limitations should be noted:

  • In-Memory Operational Data Store: Operational customer records and ticket threads are maintained in memory via mock_store.py. While Hindsight memories persist across restarts when connected to an external service, local ticket changes reset upon backend restart.
  • Tag-Based Scoping Semantics: tags=[f"customer:{customer_id}"] helps scope recall toward memories associated with the active customer. The current implementation uses tags_match="any", so this tag filter should not be treated as a strict isolation or security boundary.
  • No Autonomous Backend LLM Loop: The backend performs memory retrieval and retention to inform human specialists; it does not currently execute an autonomous generative loop to dispatch automated replies without human review.
  • No Authentication or Authorization Middleware: The API currently runs without JWT, session cookies, OAuth, or RBAC authorization. It is structured as an internal development service.

Conclusion

Building persistent memory into enterprise support requires more than just calling vector search from a frontend script. By structuring RecallDesk around an asynchronous FastAPI backend, we decoupled operational ticketing state from persistent cognitive memory, enforced pre-ingestion regex sanitization, and established resilient failure handling.

This architecture ensures that customer support workflows remain fast and reliable, while giving human specialists immediate access to historical context when resolving critical customer incidents.

Top comments (0)