A therapist has a few minutes between sessions. A briefing that takes eight seconds might as well not exist.
When you pair a memory layer with fast LLM inference, most of the remaining latency is plumbing. Here is how to keep the request path short.
Keep the request path thin
Ingestion (writing notes) and synthesis (answering a clinician) have different needs. Heavy work like pattern reflection runs in the background, so the query endpoint only does recall plus one generation call.
Go fully async
A naive endpoint blocks on each network call. An async client lets the server handle other requests while it waits:
import httpx
from fastapi import FastAPI
app = FastAPI()
client = httpx.AsyncClient(timeout=5.0)
@app.post("/api/copilot-ask")
async def copilot_ask(item: CopilotQuery):
recall = await client.post(
recall_url(item.child_id),
json={"query": item.query},
headers=HEADERS,
)
context = recall.text if recall.is_success else ""
return await generate_briefing(item, context)
Trim the prompt
Latency and cost scale with tokens. Recall only what's relevant, cap the number of memories you inject, and ask for a short output. One or two bullets is plenty for a pre-session briefing.
Plan for failure
Set timeouts, and decide what the UI shows if recall or inference is slow. A clear "briefing unavailable, showing last known protocol" beats a spinner.
Measure, don't assume
Log recall time and generation time separately, and look at p95, not just the average. Use your own numbers as the target.
Takeaways
- Separate the write path from the read path.
- Async I/O and small prompts do most of the work.
- Timeouts and fallbacks are part of the product, not an afterthought.
Top comments (0)