DEV Community

Armin Burger
Armin Burger

Posted on

Streaming an AI chat from Python FastAPI through a Next.js proxy: the four annoying details

Streaming chat looks trivial in tutorials: for chunk in response: print(chunk.text). It stops
looking trivial the moment you add the things a real product needs — a logged-in user, a
provider that can be swapped at runtime, a cost record per request, and an error that happens
after you already sent 200 OK.

Here's the path a message takes in the ChimerAI stack, with the parts that cost me time.

The shape

ChatWindow (React)
  → POST /api/v1/chat/stream        Next.js API route, TypeScript
      auth + credits + persist user message
  → POST /api/chat/stream           FastAPI, Python + LiteLLM
      async generator, yields SSE chunks
  ← data: {...}\n\n  …  data: [DONE]\n\n
Enter fullscreen mode Exit fullscreen mode

Two layers, because the LLM ecosystem is Python and the UI is TypeScript. The Next.js route is
not a pass-through — it's where auth, persistence, and metering happen, so the Python service
stays a stateless completion endpoint you could put behind any other frontend.

1. The frontend is a component, not a page

import { ChatWindow } from '@chimerai/chat-ui';

export default function ChatPage() {
  return (
    <ChatWindow
      apiEndpoint="/api/v1/chat/stream"
      placeholder="Ask me anything..."
      showModelSelector={true}
      models={['gpt-4', 'claude-3', 'ollama/llama3']}
    />
  );
}
Enter fullscreen mode Exit fullscreen mode

The component handles markdown rendering, code blocks with a copy button, auto-scroll, and the
streaming reader. If you're embedding into someone else's site rather than shipping your own
app, there's a separate chat-widget — a self-contained Web Component with Shadow DOM, mounted
per API key:

<script src="https://your-app.com/widget/chat.js"></script>
<div id="chat"></div>
<script>ChimerAI.mount('#chat', { apiKey: 'sk_live_...', theme: 'auto' });</script>
Enter fullscreen mode Exit fullscreen mode

2. Check what you can't verify: estimate tokens before the call

You don't know the completion length before the model generates it, so a credit check can't be
exact. The route estimates from the prompt and treats it as a precondition, then reconciles
afterwards:

function estimateTokens(messages: any[]): number {
  if (!messages || !Array.isArray(messages)) return 500;
  const totalChars = messages.reduce((sum, msg) => sum + (msg.content?.length || 0), 0);
  return Math.ceil((totalChars / 4) * 1.2);
}

const estimatedTokens = estimateTokens(payload.messages);
const creditResult = await requireCredits(request, estimatedTokens);
if (!creditResult.authorized) {
  return new Response(JSON.stringify(createErrorResponse(creditResult)), { status: 402 });
}
Enter fullscreen mode Exit fullscreen mode

chars / 4 is the usual rough tokenizer rule; the 1.2 is headroom. It's a heuristic and it
fails in both directions — a short prompt with a long completion overshoots the balance, a
user with exactly enough credits can get cut off mid-stream. That's acceptable for a soft
limit; it is not acceptable as billing. Actual usage is recorded from the provider's own
token counts at the end of the stream, and that's the number that goes into ApiUsage.

Note the ordering: permission check → credit check → then create the conversation row and
persist the user message. Doing the DB writes first means a rejected request still costs you a
conversation with an orphan message in it.

3. Errors after 200 OK must travel inside the stream

This is the one that bites everyone. StreamingResponse sends headers immediately, so a Python
try/except around the generator body is useless — by the time the provider call fails, you've
already committed to a 200:

@router.post("/chat/stream")
async def create_chat_completion_stream(request: ChatCompletionRequest):
    # No try/except here — StreamingResponse returns 200 immediately,
    # the generator runs lazily. Errors are handled inside the generator
    # as SSE error events.
    return StreamingResponse(
        chat_service.create_streaming_completion(request, provider_id=..., user_id=...),
        media_type="text/event-stream",
    )
Enter fullscreen mode Exit fullscreen mode

So the error handling lives inside the generator and speaks the protocol:

        except Exception as e:
            logger.error("streaming_completion_error", error=str(e))
            # Yield error as SSE event so the client receives it
            # (raising here would be swallowed — 200 OK was already sent)
            error_payload = json.dumps({"error": {"message": str(e), "type": "server_error"}})
            yield f"data: {error_payload}\n\n"
            yield "data: [DONE]\n\n"
Enter fullscreen mode Exit fullscreen mode

The client therefore has to handle an error key in the middle of an otherwise healthy stream,
and still terminate on [DONE]. If your reader assumes "chunks only contain content", a
provider outage renders as a silently truncated answer instead of an error message — which is
the worst possible failure mode, because it looks like the model just stopped talking.

4. Normalize chunks into your own schema

The service doesn't forward provider chunks verbatim; it rebuilds each one:

            async for chunk in response:
                if not chunk.choices:
                    continue
                choice = chunk.choices[0]

                if hasattr(chunk, "usage") and chunk.usage:
                    total_prompt_tokens = getattr(chunk.usage, "prompt_tokens", 0) or 0
                    total_completion_tokens = getattr(chunk.usage, "completion_tokens", 0) or 0

                chunk_response = ChatCompletionChunk(
                    id=chunk.id or f"chatcmpl-{uuid.uuid4().hex[:8]}",
                    object="chat.completion.chunk",
                    created=int(time.time()),
                    model=chunk.model,
                    choices=[ChatCompletionChunkChoice(
                        index=choice.index,
                        delta=DeltaMessage(
                            role=getattr(choice.delta, "role", None),
                            content=getattr(choice.delta, "content", None),
                        ),
                        finish_reason=choice.finish_reason,
                    )],
                )
                yield f"data: {chunk_response.model_dump_json()}\n\n"
Enter fullscreen mode Exit fullscreen mode

Two reasons this indirection is worth the boilerplate:

  • if not chunk.choices: continue — providers emit usage-only or role-only frames. Anthropic and Ollama don't frame deltas the way OpenAI does; LiteLLM gets you most of the way, and these edge cases are the rest.
  • id=chunk.id or f"chatcmpl-{uuid4().hex[:8]}" — some providers omit the id on chunks. The client uses it to correlate, so a missing id becomes a dropped message.

Usage is accumulated from whichever chunks carry it and reported once, after the stream closes:

            if provider_id and user_id and (total_prompt_tokens or total_completion_tokens):
                await provider_client.report_usage(
                    provider_id=provider_id, user_id=user_id, model=request.model,
                    prompt_tokens=total_prompt_tokens,
                    completion_tokens=total_completion_tokens,
                    endpoint="/api/chat/stream",
                )
Enter fullscreen mode Exit fullscreen mode

Swapping providers without touching the chat code

The model isn't hardcoded in the route. The request may omit it, in which case it falls back to
the active provider's config.defaultModel:

let model = payload.model; // Optional — resolved after provider loading
// Model validation is deferred until after provider loading,
// so we can fall back to provider.config.defaultModel
Enter fullscreen mode Exit fullscreen mode

That's why the permission check has a deferred branch (requireModelPermission(request,
'__deferred__')
) — you can't validate a model you haven't resolved yet. Providers are DB rows
with AES-256-encrypted keys, so adding Claude or a local Ollama instance is an admin-UI action,
not a deploy:

await fetch('/api/providers', {
  method: 'POST',
  body: JSON.stringify({ name: 'Local Llama', provider: 'ollama',
                         baseUrl: 'http://localhost:11434', isActive: true }),
});
Enter fullscreen mode Exit fullscreen mode

What I'd do differently

The token estimator is the weakest link — it's a guess in the request path. Cleaner would be a
queue/worker split: accept the request, enqueue it, stream from the worker, and enforce budget
per user per window instead of per request. I haven't done that because at current traffic the
heuristic's failure mode (a request rejected slightly early) is cheaper than the complexity.

Also: 1.2 is a magic number I chose. If you have real prompt/completion distributions,
replace it with a p95 multiplier from your own logs.

Repo: github.com/armbur19-collab/chimerai-kickstart

Top comments (0)