DEV Community

Manish Goud
Manish Goud

Posted on

"Sub-Second Clinical Briefings: FastAPI, Async Recall and Fast Inference "

A therapist has a few minutes between sessions. A briefing that takes eight seconds might as well not exist.

When you pair a memory layer with fast LLM inference, most of the remaining latency is plumbing. Here is how to keep the request path short.

Keep the request path thin

Ingestion (writing notes) and synthesis (answering a clinician) have different needs. Heavy work like pattern reflection runs in the background, so the query endpoint only does recall plus one generation call.

Go fully async

A naive endpoint blocks on each network call. An async client lets the server handle other requests while it waits:

import httpx
from fastapi import FastAPI

app = FastAPI()
client = httpx.AsyncClient(timeout=5.0)

@app.post("/api/copilot-ask")
async def copilot_ask(item: CopilotQuery):
    recall = await client.post(
        recall_url(item.child_id),
        json={"query": item.query},
        headers=HEADERS,
    )
    context = recall.text if recall.is_success else ""
    return await generate_briefing(item, context)
Enter fullscreen mode Exit fullscreen mode

Trim the prompt

Latency and cost scale with tokens. Recall only what's relevant, cap the number of memories you inject, and ask for a short output. One or two bullets is plenty for a pre-session briefing.

Plan for failure

Set timeouts, and decide what the UI shows if recall or inference is slow. A clear "briefing unavailable, showing last known protocol" beats a spinner.

Measure, don't assume

Log recall time and generation time separately, and look at p95, not just the average. Use your own numbers as the target.

Takeaways

  1. Separate the write path from the read path.
  2. Async I/O and small prompts do most of the work.
  3. Timeouts and fallbacks are part of the product, not an afterthought.

Top comments (0)