Every LLM backend I've built started the same way.
FastAPI for the endpoints, a basic /chat route, and then a week of solving the same problems I'd already solved on the last project: how do you stream responses without blocking the event loop? What happens when an agent run takes 90 seconds and your reverse proxy kills it? How do you maintain conversation history for a WebSocket session without it leaking between users?
These aren't hard problems. They're just tedious — the kind of thing you solve once, copy from project to project, and then spend time adapting because the last solution was half-baked.
I packaged the full solution as the FastAPI + AI Agent Backend Scaffold.
What it is
A complete FastAPI application scaffold for teams building AI agent backends. Not a tutorial. Working code, with tests, Docker setup, and documentation for every config option.
app/
main.py — FastAPI factory, lifespan, middleware stack
config.py — Pydantic BaseSettings (env vars, .env file)
dependencies.py — Auth, LLM service, shared Depends()
routers/
chat.py — POST /chat/completions (streaming + non-streaming)
agents.py — POST /agents/run + GET /agents/{id}/status
websocket.py — WS /ws/chat (multi-turn sessions)
health.py — /health + /readiness probes
middleware/
rate_limit.py — Per-user token bucket (Redis-backed)
auth.py — Bearer token validation
logging.py — Structured JSON logging
services/
llm.py — OpenAI + Anthropic wrapper (your agent logic goes here)
task_queue.py — Asyncio background queue with concurrency limits
tests/ — Full test suite (mock LLM, no real API needed)
docker/ — Multi-stage Dockerfile + docker-compose
.env.example — Every config option documented
The four problems it solves
1. Streaming that doesn't block
The naive approach works until you try to stream to multiple clients simultaneously:
# Naive: blocks during generation
@app.post("/chat")
async def chat(request: ChatRequest):
response = await openai_client.chat.completions.create(...)
return {"content": response.choices[0].message.content}
The scaffold's streaming endpoint uses StreamingResponse with a proper async generator — no blocking, backpressure handled, Cache-Control and X-Accel-Buffering headers set correctly for nginx:
@router.post("/completions")
async def chat_completions(request: ChatRequest, llm: LLMService = Depends(get_llm)):
if request.stream:
return StreamingResponse(
llm.stream_chat(request.messages, model=request.model),
media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
content = await llm.chat(request.messages, model=request.model)
return ChatResponse.from_content(content, model=request.model or "default")
Works with any frontend that reads SSE — Vercel AI SDK, EventSource in the browser, httpx in Python.
2. Long agent runs that don't timeout
A 90-second agent run will get killed by most reverse proxy timeout defaults. The common fix (increasing the timeout globally) creates a new problem: you can no longer distinguish slow agent runs from hung processes.
The scaffold's background task pattern returns immediately with a task_id. The client polls for the result:
# POST /agents/run — returns in < 100ms
@router.post("/run", status_code=202)
async def run_agent(request: AgentRunRequest, req: Request, llm: LLMService = Depends(get_llm)):
queue: TaskQueue = req.app.state.task_queue
task_id = await queue.submit(_run_agent, request.prompt, llm, request.system, request.model)
return {"task_id": task_id, "status": "queued"}
# GET /agents/{task_id}/status — poll until completed or failed
@router.get("/{task_id}/status")
async def agent_status(task_id: str, req: Request):
record = req.app.state.task_queue.get_status(task_id)
if record is None:
raise HTTPException(status_code=404, detail=f"Task {task_id} not found")
return {"task_id": record.task_id, "status": record.status, "result": record.result}
No external broker needed. The TaskQueue uses asyncio with a semaphore-based concurrency limit (MAX_BACKGROUND_TASKS env var, default 10). Swap to Celery or ARQ later if you need distributed workers — the interface stays the same.
3. WebSocket multi-turn sessions
The tricky part of WebSocket agent sessions isn't the connection — it's session state. Conversation history should persist across reconnects, be isolated per session_id, and clean up if a session goes idle.
@router.websocket("/chat")
async def websocket_chat(
websocket: WebSocket,
session_id: str = Query(...),
system: str | None = Query(default=None),
):
await websocket.accept()
history = _sessions.setdefault(session_id, [])
llm = get_llm()
try:
while True:
user_text = await websocket.receive_text()
history.append(ChatMessage(role="user", content=user_text))
# Stream response back token by token
full_response = ""
async for chunk in llm.stream_chat(history, system=system):
if chunk.startswith("data: ") and "[DONE]" not in chunk:
text = chunk[6:].strip()
full_response += text
await websocket.send_text(text)
await websocket.send_text("[DONE]")
history.append(ChatMessage(role="assistant", content=full_response))
except WebSocketDisconnect:
pass
The in-memory session store (_sessions: dict[str, list[ChatMessage]]) works for single-instance deployments. For multi-instance: swap for a Redis-backed store. The interface doesn't change.
4. Per-user rate limiting that survives restarts
A fixed-rate limiter that resets on restart isn't useful in production. The scaffold uses a Redis-backed token bucket:
# 60 requests/minute, burst of 10 — configurable via RATE_LIMIT_RPM and RATE_LIMIT_BURST
# Falls back to in-memory if Redis is unavailable (single-instance safe)
class RateLimitMiddleware(BaseHTTPMiddleware):
async def dispatch(self, request: Request, call_next: Callable) -> Response:
if request.url.path in {"/health", "/readiness"}:
return await call_next(request)
key = _get_rate_limit_key(request) # keyed on Authorization header hash or client IP
allowed = await self._bucket.check(key)
if not allowed:
return JSONResponse(status_code=429, content={"error": "rate_limit_exceeded"})
return await call_next(request)
Rate limit keys are derived from the Authorization header (hashed, so raw tokens never appear in Redis). Falls back to client IP for unauthenticated requests.
Structured logging that's actually queryable
Every request gets a request_id. Every response logs status_code and latency_ms. All output is JSON by default:
{"timestamp": "2026-09-15T14:00:01", "level": "info", "request_id": "a1b2c3d4", "method": "POST", "path": "/chat/completions"}
{"timestamp": "2026-09-15T14:00:02", "level": "info", "request_id": "a1b2c3d4", "status_code": 200, "latency_ms": 1840}
Import directly into Datadog, CloudWatch, or Grafana Loki without a custom parser.
Switch to human-readable for local dev: LOG_JSON=false.
Adding your agent logic
The scaffold handles infrastructure. Your agent logic goes in services/llm.py:
class LLMService:
async def chat(self, messages: list[ChatMessage], model=None, system=None, **kwargs) -> str:
# Add your system prompt, tools, multi-step logic here
# The streaming and non-streaming paths call this method
formatted = self._format_messages(messages, system)
response = await self._client.chat.completions.create(
model=model or self._settings.llm_model,
messages=formatted,
**kwargs,
)
return response.choices[0].message.content or ""
Add tools, chain calls, implement retry logic — the endpoints stay the same. Your agent complexity lives in one place.
Testing setup
The included test fixtures let you test every endpoint without a real LLM API or Redis:
def test_non_streaming_chat(client, mock_llm):
mock_llm.queue_response("Hello! How can I help?")
response = client.post(
"/chat/completions",
json={"messages": [{"role": "user", "content": "Hi"}], "stream": False},
)
assert response.status_code == 200
assert response.json()["content"] == "Hello! How can I help?"
def test_agent_task_completes(client, mock_llm):
mock_llm.queue_response("Completed work.")
submit = client.post("/agents/run", json={"prompt": "Do some work"})
task_id = submit.json()["task_id"]
# Poll to completion
import time
deadline = time.time() + 5.0
while time.time() < deadline:
status = client.get(f"/agents/{task_id}/status")
if status.json()["status"] in ("completed", "failed"):
break
time.sleep(0.1)
assert status.json()["status"] == "completed"
assert status.json()["result"] == "Completed work."
The MockLLMService returns queued responses in order — no mocking of internal methods, no monkeypatching of the OpenAI SDK.
Quick start
cp .env.example .env
# Set OPENAI_API_KEY or ANTHROPIC_API_KEY in .env
docker-compose up
curl -X POST http://localhost:8000/chat/completions \
-H "Authorization: Bearer dev-key" \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Hello"}], "stream": false}'
The scaffold supports OpenAI and Anthropic out of the box. Switch providers with LLM_PROVIDER=anthropic in .env.
Getting the scaffold
The FastAPI + AI Agent Backend Scaffold is available on Gumroad for $69 — one-time purchase, GitHub repo delivery, MIT license.
Includes the full application (app/, tests/, docker/), requirements.txt, pyproject.toml, and .env.example.
If you're building agent tests alongside this, the Pytest for AI Agents Starter Kit uses the same MockLLMService pattern — the fixtures compose directly with this scaffold's test setup.
Top comments (0)