DEV Community

Ubaid Ullah
Ubaid Ullah

Posted on Originally published at djangix.com

Streaming OpenAI API Responses: How It Works, in Python, Django & FastAPI

Originally published on the Djangix blog. This is a condensed version — the full article, with complete code, is linked at the end.

A non-streamed chat answer leaves the user staring at a spinner for the whole generation. Streaming changes the feel, not the total time: the first tokens arrive in under a second and the answer appears to type itself.

What is actually sent

The transport is Server-Sent Events — one HTTP response that stays open and delivers small data: messages. Each message carries a JSON chunk, and the useful text lives in a tiny delta field. Most chunks are fragments, not whole words, so the client simply concatenates them until a final done marker arrives.

Going through your own backend

Browsers should not call the provider directly, because that would expose your API key. So your server receives the chunks and re-emits them — and that middle step is where many implementations break.

  • Wrap each token as JSON. Raw token text can contain newlines, which break the event framing. Encoding the token inside a JSON object avoids mangled multi-line answers.
  • Read with fetch, not EventSource, for POST chats. The built-in event API cannot send a JSON body or an authorization header, so read the response stream manually and keep a buffer for events split across network chunks.
  • Ask for usage explicitly. In the example in the full article, a stream option is needed if you want a final usage chunk for billing or budgeting.

Production pitfalls worth planning for

  • Buffering proxies: Reverse proxies and CDNs may hold the response and deliver it all at once, silently defeating streaming. Disable buffering for that route and test through the real production path.
  • Errors after the stream starts: Once the first chunk is sent, you cannot change the status code. Send an error as an event instead and render it in the UI.
  • Disconnected users and saving: Stop consuming when the client goes away, and save the assembled answer once at the end rather than writing to the database per token.
  • Idle timeouts: Periodic keepalive comments can keep a long generation from being cut off.

When not to stream

Short answers gain little, machine-consumed structured output usually needs to be complete before it can be parsed, and background jobs have no one watching. Streaming is a user-experience feature — use it where a person is waiting.


Read the full article on Djangix: Streaming OpenAI API Responses: How It Works, in Python, Django & FastAPI — with the complete Python, Django StreamingHttpResponse and FastAPI StreamingResponse code.

Top comments (0)