DEV Community

Peyton Green
Peyton Green

Posted on

The FastAPI Backend Boilerplate I Wish I Had When I Started Building LLM Apps

Every LLM backend I've built started the same way.

FastAPI for the endpoints, a basic /chat route, and then a week of solving the same problems I'd already solved on the last project: how do you stream responses without blocking the event loop? What happens when an agent run takes 90 seconds and your reverse proxy kills it? How do you maintain conversation history for a WebSocket session without it leaking between users?

These aren't hard problems. They're just tedious — the kind of thing you solve once, copy from project to project, and then spend time adapting because the last solution was half-baked.

I packaged the full solution as the FastAPI + AI Agent Backend Scaffold.


What it is

A complete FastAPI application scaffold for teams building AI agent backends. Not a tutorial. Working code, with tests, Docker setup, and documentation for every config option.

app/
  main.py               — FastAPI factory, lifespan, middleware stack
  config.py             — Pydantic BaseSettings (env vars, .env file)
  dependencies.py       — Auth, LLM service, shared Depends()

  routers/
    chat.py             — POST /chat/completions (streaming + non-streaming)
    agents.py           — POST /agents/run + GET /agents/{id}/status
    websocket.py        — WS /ws/chat (multi-turn sessions)
    health.py           — /health + /readiness probes

  middleware/
    rate_limit.py       — Per-user token bucket (Redis-backed)
    auth.py             — Bearer token validation
    logging.py          — Structured JSON logging

  services/
    llm.py              — OpenAI + Anthropic wrapper (your agent logic goes here)
    task_queue.py       — Asyncio background queue with concurrency limits

tests/                  — Full test suite (mock LLM, no real API needed)
docker/                 — Multi-stage Dockerfile + docker-compose
.env.example            — Every config option documented
Enter fullscreen mode Exit fullscreen mode

The four problems it solves

1. Streaming that doesn't block

The naive approach works until you try to stream to multiple clients simultaneously:

# Naive: blocks during generation
@app.post("/chat")
async def chat(request: ChatRequest):
    response = await openai_client.chat.completions.create(...)
    return {"content": response.choices[0].message.content}
Enter fullscreen mode Exit fullscreen mode

The scaffold's streaming endpoint uses StreamingResponse with a proper async generator — no blocking, backpressure handled, Cache-Control and X-Accel-Buffering headers set correctly for nginx:

@router.post("/completions")
async def chat_completions(request: ChatRequest, llm: LLMService = Depends(get_llm)):
    if request.stream:
        return StreamingResponse(
            llm.stream_chat(request.messages, model=request.model),
            media_type="text/event-stream",
            headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
        )
    content = await llm.chat(request.messages, model=request.model)
    return ChatResponse.from_content(content, model=request.model or "default")
Enter fullscreen mode Exit fullscreen mode

Works with any frontend that reads SSE — Vercel AI SDK, EventSource in the browser, httpx in Python.


2. Long agent runs that don't timeout

A 90-second agent run will get killed by most reverse proxy timeout defaults. The common fix (increasing the timeout globally) creates a new problem: you can no longer distinguish slow agent runs from hung processes.

The scaffold's background task pattern returns immediately with a task_id. The client polls for the result:

# POST /agents/run — returns in < 100ms
@router.post("/run", status_code=202)
async def run_agent(request: AgentRunRequest, req: Request, llm: LLMService = Depends(get_llm)):
    queue: TaskQueue = req.app.state.task_queue
    task_id = await queue.submit(_run_agent, request.prompt, llm, request.system, request.model)
    return {"task_id": task_id, "status": "queued"}

# GET /agents/{task_id}/status — poll until completed or failed
@router.get("/{task_id}/status")
async def agent_status(task_id: str, req: Request):
    record = req.app.state.task_queue.get_status(task_id)
    if record is None:
        raise HTTPException(status_code=404, detail=f"Task {task_id} not found")
    return {"task_id": record.task_id, "status": record.status, "result": record.result}
Enter fullscreen mode Exit fullscreen mode

No external broker needed. The TaskQueue uses asyncio with a semaphore-based concurrency limit (MAX_BACKGROUND_TASKS env var, default 10). Swap to Celery or ARQ later if you need distributed workers — the interface stays the same.


3. WebSocket multi-turn sessions

The tricky part of WebSocket agent sessions isn't the connection — it's session state. Conversation history should persist across reconnects, be isolated per session_id, and clean up if a session goes idle.

@router.websocket("/chat")
async def websocket_chat(
    websocket: WebSocket,
    session_id: str = Query(...),
    system: str | None = Query(default=None),
):
    await websocket.accept()
    history = _sessions.setdefault(session_id, [])
    llm = get_llm()

    try:
        while True:
            user_text = await websocket.receive_text()
            history.append(ChatMessage(role="user", content=user_text))

            # Stream response back token by token
            full_response = ""
            async for chunk in llm.stream_chat(history, system=system):
                if chunk.startswith("data: ") and "[DONE]" not in chunk:
                    text = chunk[6:].strip()
                    full_response += text
                    await websocket.send_text(text)

            await websocket.send_text("[DONE]")
            history.append(ChatMessage(role="assistant", content=full_response))
    except WebSocketDisconnect:
        pass
Enter fullscreen mode Exit fullscreen mode

The in-memory session store (_sessions: dict[str, list[ChatMessage]]) works for single-instance deployments. For multi-instance: swap for a Redis-backed store. The interface doesn't change.


4. Per-user rate limiting that survives restarts

A fixed-rate limiter that resets on restart isn't useful in production. The scaffold uses a Redis-backed token bucket:

# 60 requests/minute, burst of 10 — configurable via RATE_LIMIT_RPM and RATE_LIMIT_BURST
# Falls back to in-memory if Redis is unavailable (single-instance safe)
class RateLimitMiddleware(BaseHTTPMiddleware):
    async def dispatch(self, request: Request, call_next: Callable) -> Response:
        if request.url.path in {"/health", "/readiness"}:
            return await call_next(request)

        key = _get_rate_limit_key(request)  # keyed on Authorization header hash or client IP
        allowed = await self._bucket.check(key)

        if not allowed:
            return JSONResponse(status_code=429, content={"error": "rate_limit_exceeded"})

        return await call_next(request)
Enter fullscreen mode Exit fullscreen mode

Rate limit keys are derived from the Authorization header (hashed, so raw tokens never appear in Redis). Falls back to client IP for unauthenticated requests.


Structured logging that's actually queryable

Every request gets a request_id. Every response logs status_code and latency_ms. All output is JSON by default:

{"timestamp": "2026-09-15T14:00:01", "level": "info", "request_id": "a1b2c3d4", "method": "POST", "path": "/chat/completions"}
{"timestamp": "2026-09-15T14:00:02", "level": "info", "request_id": "a1b2c3d4", "status_code": 200, "latency_ms": 1840}
Enter fullscreen mode Exit fullscreen mode

Import directly into Datadog, CloudWatch, or Grafana Loki without a custom parser.

Switch to human-readable for local dev: LOG_JSON=false.


Adding your agent logic

The scaffold handles infrastructure. Your agent logic goes in services/llm.py:

class LLMService:
    async def chat(self, messages: list[ChatMessage], model=None, system=None, **kwargs) -> str:
        # Add your system prompt, tools, multi-step logic here
        # The streaming and non-streaming paths call this method
        formatted = self._format_messages(messages, system)
        response = await self._client.chat.completions.create(
            model=model or self._settings.llm_model,
            messages=formatted,
            **kwargs,
        )
        return response.choices[0].message.content or ""
Enter fullscreen mode Exit fullscreen mode

Add tools, chain calls, implement retry logic — the endpoints stay the same. Your agent complexity lives in one place.


Testing setup

The included test fixtures let you test every endpoint without a real LLM API or Redis:

def test_non_streaming_chat(client, mock_llm):
    mock_llm.queue_response("Hello! How can I help?")

    response = client.post(
        "/chat/completions",
        json={"messages": [{"role": "user", "content": "Hi"}], "stream": False},
    )

    assert response.status_code == 200
    assert response.json()["content"] == "Hello! How can I help?"

def test_agent_task_completes(client, mock_llm):
    mock_llm.queue_response("Completed work.")

    submit = client.post("/agents/run", json={"prompt": "Do some work"})
    task_id = submit.json()["task_id"]

    # Poll to completion
    import time
    deadline = time.time() + 5.0
    while time.time() < deadline:
        status = client.get(f"/agents/{task_id}/status")
        if status.json()["status"] in ("completed", "failed"):
            break
        time.sleep(0.1)

    assert status.json()["status"] == "completed"
    assert status.json()["result"] == "Completed work."
Enter fullscreen mode Exit fullscreen mode

The MockLLMService returns queued responses in order — no mocking of internal methods, no monkeypatching of the OpenAI SDK.


Quick start

cp .env.example .env
# Set OPENAI_API_KEY or ANTHROPIC_API_KEY in .env

docker-compose up

curl -X POST http://localhost:8000/chat/completions \
  -H "Authorization: Bearer dev-key" \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Hello"}], "stream": false}'
Enter fullscreen mode Exit fullscreen mode

The scaffold supports OpenAI and Anthropic out of the box. Switch providers with LLM_PROVIDER=anthropic in .env.


Getting the scaffold

The FastAPI + AI Agent Backend Scaffold is available on Gumroad for $69 — one-time purchase, GitHub repo delivery, MIT license.

Includes the full application (app/, tests/, docker/), requirements.txt, pyproject.toml, and .env.example.

If you're building agent tests alongside this, the Pytest for AI Agents Starter Kit uses the same MockLLMService pattern — the fixtures compose directly with this scaffold's test setup.

Get the FastAPI + AI Agent Backend Scaffold →

Top comments (0)