DEV Community

Cover image for "LLM streaming works in the demo. These 4 hops break it in prod"
Ajay Vishwakarma
Ajay Vishwakarma

Posted on

"LLM streaming works in the demo. These 4 hops break it in prod"

Words appear one by one on localhost, and LLM streaming looks done. In production the answer lands in one lump, the model keeps generating after the user leaves, and half a sentence shows up as the full answer.

None of these are LLM problems. They live in the plumbing between the model and the browser.

The rule: your server is not a pipe. It sits between two connections, and either one can misbehave.

Scope: text chat over fetch, examples in Node/Express. The ideas carry over to other stacks.

Where does a streamed answer actually break?

Here's the flow:

Browser  ──POST /chat──►  Your API  ──stream:true──►  LLM provider
   ▲                          │                            │
   └──── tokens as events ◄───┴──────── tokens ◄───────────┘
Enter fullscreen mode Exit fullscreen mode

The 30-line demo teaches the pipe model: tokens go in one side and come out the other. That took me a while to unlearn.

Your server sits in the middle of two connections. Something in front of it can buffer. The client behind it can leave. The provider underneath it can fail. And the browser can misread what arrives.

![Four hops where a streamed LLM answer can break: browser, proxy, API, provider]

Most streaming bugs are one of these four hops going wrong. The sections below take them in order.

I use fetch with a streamed response for all of this. A chat request needs a message history and an auth token, and fetch gives me a POST body, headers, and abort support. The alternatives are near the end.

Why does the answer arrive all at once in production?

Reverse proxies like Nginx collect the upstream response and send it in bigger chunks by default. That's great for normal pages and terrible for a stream. Your tokens sit in a buffer, and the user sees nothing until it fills or the response ends.

I've seen this called out as the most frequent production surprise with streaming, and it matches what the infrastructure docs say.

In your app, send headers that tell proxies to back off:

res.writeHead(200, {
  "Content-Type": "text/event-stream",
  "Cache-Control": "no-cache, no-transform",
  "Connection": "keep-alive",
  "X-Accel-Buffering": "no", // Nginx-specific: don't buffer this response
});
Enter fullscreen mode Exit fullscreen mode

In Nginx, for the streaming route:

location /api/chat {
    proxy_pass http://app;
    proxy_http_version 1.1;
    proxy_set_header Connection "";
    proxy_buffering off;      # the important one
    gzip off;                 # compression can re-buffer the stream
    proxy_read_timeout 300s;  # longer than your slowest answer
}
Enter fullscreen mode Exit fullscreen mode

A CDN, load balancer, or API gateway in front of you may have its own buffering and idle timeout settings. I can't tell you the right config for each one, so check your provider's docs. What works everywhere is a test:

curl -N -X POST https://your-domain.com/api/chat \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"count slowly from 1 to 20"}]}'
Enter fullscreen mode Exit fullscreen mode

-N turns off curl's own buffering. If lines appear gradually, the whole path streams. If they arrive in one lump, something is buffering. Run it against the real production URL, and again after any infrastructure change.

The same hop kills quiet streams. When the model is thinking or a tool call takes several seconds, some proxies and load balancers see silence and decide the connection is dead. Raise timeouts above your slowest realistic answer on every layer, and send a heartbeat. In SSE, a line starting with : is a comment that clients ignore:

const heartbeat = setInterval(() => res.write(": ping\n\n"), 15000);
res.on("close", () => clearInterval(heartbeat));
Enter fullscreen mode Exit fullscreen mode

If your answer involves tools, also stream status events like {"status": "searching documents"}. Ten seconds of silence feels broken. Ten seconds with "Looking that up…" feels normal.

Deploys count too. A rolling restart can kill streams mid-answer, so give your servers a graceful shutdown period at least as long as your longest answer.

What happens when the user closes the tab?

This is the hop I think is most underrated.

A user asks a question, sees it start answering, and closes the tab. In a lot of apps the server does nothing. The loop reading from the LLM keeps going, the provider keeps generating, and you pay for output nobody sees. A Stop button that only hides text in the UI is the same problem with a friendlier face.

A real cancel has to travel the whole chain:

![Without cancellation the model keeps generating after the user leaves; with cancellation one abort signal stops it end to end]

Browser:

const controller = new AbortController();

const res = await fetch("/api/chat", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ messages }),
  signal: controller.signal,
});

// Stop button:
stopButton.onclick = () => controller.abort();
Enter fullscreen mode Exit fullscreen mode

Server:

app.post("/api/chat", async (req, res) => {
  const upstream = new AbortController();

  // If the client leaves before we finish, cancel the LLM call.
  res.on("close", () => {
    if (!res.writableEnded) upstream.abort();
  });

  res.writeHead(200, { "Content-Type": "text/event-stream" /* + headers above */ });

  try {
    const stream = await llm.chat.completions.create(
      { model, messages: req.body.messages, stream: true },
      { signal: upstream.signal }
    );

    for await (const chunk of stream) {
      const text = chunk.choices[0]?.delta?.content;
      if (text) res.write(`data: ${JSON.stringify({ text })}\n\n`);
    }

    res.write(`data: ${JSON.stringify({ done: true })}\n\n`);
  } catch (err) {
    if (!upstream.signal.aborted) {
      res.write(`data: ${JSON.stringify({ error: "generation_failed" })}\n\n`);
    }
  } finally {
    res.end();
  }
});
Enter fullscreen mode Exit fullscreen mode

Choices I'd make again:

  • Listen on the response's close event, and check writableEnded. That separates "the client left early" from "we finished normally." In recent Node versions, I wouldn't rely on the request's close event for this.
  • Pass the same signal to everything the request started. If the question also triggered a vector search, a reranker, or tool calls, they should stop too. Otherwise you cancel the visible part and keep paying for the invisible part.
  • Don't assume the provider stops billing immediately. Aborting the connection usually stops generation, but behavior varies by provider. I'd test it with a short script and check the usage numbers, instead of trusting the docs alone.
  • Decide what to do with partial answers. If a user stops mid-answer, do you save what was generated? Choose on purpose.

How do you report an error after you've already sent 200 OK?

Normally errors are status codes: 400, 429, 500. In a stream, you've already sent 200 OK and some tokens before the problem happens. The provider times out, you hit a rate limit, or a tool call throws. The status code is already gone.

So errors have to travel inside the stream as events. Two rules I follow, both in the server code above:

  1. Send an explicit error event, so the UI can say something useful.
  2. Always send a final done event on success.

The second one is easy to skip and it matters. Without done, the client can't tell "complete" from "the connection dropped after the third sentence." Users will read the truncated answer as the real answer.

In the UI I treat these as three separate end states:

  • Complete: got done
  • Failed: got error
  • Interrupted: the stream ended without either (show "response was cut off, try again")

I'd also avoid silently auto-retrying a half-finished generation. A retry costs a second generation, and it might give a different answer from the one the user already half-read.

Why does the browser show broken or half-parsed text?

The client code looks simple. Two bugs got me.

Bug 1: network chunks don't line up with your events. One read can contain half an event, or three and a half. If you parse each chunk directly, you'll eventually hit a JSON.parse error on a partial line. Keep a buffer, split on the blank line that ends each event, and carry the leftover:

const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";

while (true) {
  const { value, done } = await reader.read();
  if (done) break;

  // stream: true keeps multi-byte characters intact across chunks
  buffer += decoder.decode(value, { stream: true });

  const events = buffer.split("\n\n");
  buffer = events.pop(); // the last piece may be incomplete, keep it

  for (const evt of events) {
    if (!evt.startsWith("data: ")) continue; // skips heartbeat comments
    handle(JSON.parse(evt.slice(6)));
  }
}
Enter fullscreen mode Exit fullscreen mode

Bug 2: broken characters. An emoji or a non-English character can be several bytes, and a chunk can end in the middle of one. Decoding each chunk separately gives you �. The { stream: true } option on TextDecoder fixes it. This one is nasty because it passes every English-only test.

Rendering has the same trap. If you re-render the whole message on every token, especially if you re-parse Markdown each time, the page starts to stutter. What helps:

  • Batch updates. Update the screen once per animation frame, not once per token.
  • Treat the streaming message as append-only. Avoid rebuilding earlier content.
  • Plan for half-finished Markdown. A code fence that hasn't closed yet can make a renderer flicker.
  • Respect the user's scroll. If they scrolled up to read, don't pull them back down on every token.

And measure the right thing. Users will wait through a long answer if the first token arrives quickly. Time to first token is worth logging more than total duration.

When is this the wrong approach?

When the transport doesn't fit. You have three realistic options:

Option Good for Catch
SSE via EventSource Simple server-to-browser streams GET only, no custom headers, auto-reconnects by itself
Streamed response via fetch Chat: POST body, auth headers, abort support You parse the stream yourself
WebSocket True two-way, low latency (voice, live collaboration) More infrastructure, harder to scale and debug

EventSource reconnecting on its own sounds helpful. For LLMs it can mean a second paid generation for the same question that nobody planned for. If you use it, make retries something you decide.

I'd reach for WebSockets when audio or real two-way traffic is involved. I've used one for a realtime voice assistant, and it's the right tool there. "The user interrupted" is just cancellation with a microphone.

When you shouldn't stream at all.

  • Structured output (JSON, function arguments) is useless half-finished. Wait and validate the whole thing.
  • Server-to-server calls gain nothing from streaming.
  • Answers you need to check before showing (safety filters, "does this match the tool results?") can't be checked until they're done. Streaming and validation pull in opposite directions, and you have to pick which matters more.

Where my own experience runs out. I haven't run streaming at massive scale. Three things I'm still figuring out:

  • Resumable streams. If a user refreshes mid-answer, can they pick up where they left off? I've read about approaches that write chunks to something like Redis Streams so a reconnect can resume, but I haven't built this myself yet.
  • Scaling beyond one server. Long-lived connections behave differently under load balancing and autoscaling, and I'm still learning where the limits are.
  • Backpressure. If the client reads slower than you write, buffers grow on your side. I know the concept, but I haven't pushed it hard enough in practice to have strong opinions.

What should you check before you ship?

One check per hop, plus one metric:

  1. Proxy: run the curl -N command against your real production URL. Do lines arrive gradually, or in one lump?
  2. Your API: start a long answer, close the tab, and watch your server logs. Does the LLM call stop? Do the tools and searches stop with it?
  3. Provider: force an error mid-answer. Does the UI say "failed" or "cut off", or does it show half an answer as if it were complete?
  4. Browser: ask for an answer full of emoji and non-English text. Any �?
  5. Metric: log time to first token, not just total duration.

Next post: keeping LLM output trustworthy when you can't validate it before it streams, covering structured outputs, retries, and what to do with almost-valid JSON.

Which of these four hops broke first for you in production, and what was sitting in front of your server when it did?

Four things to do before you publish:

Images: upload where-streams-break.png and cancellation-chain.png in the DEV editor and replace the two REPLACE-WITH-DEV-URL placeholders. Upload the cover separately at 1000×420; your current cover carries the old title.
Front matter: the block between the --- lines sets title, description and tags in DEV's Markdown editor. If your editor shows separate Title and Tags boxes instead, delete that block and type those values into the boxes.
Draft flag: it is set to published: false so you can preview first. Flip it to true when you publish.
AI disclosure: set the field to AI-Assisted.

The same text is in dev-llm-streaming-full.md. This is the single-post route; the shorter cancellation-only post from earlier is still the alternative if you decide to run it as a series.

Dev llm streaming full
Document·MD 

Lastly,
Convert this whole text into a single copy-paste format so that I can just copy and paste it onto Dev Community and follow the Dev Community copy-paste algorithm. The formatting should be structured, according to Dev Community.

Here is the whole post as one block in DEV's Markdown format: front matter on top (title, description, four tags), then the body starting at H2 with language-tagged code blocks. I removed the two image lines so nothing in it needs editing before you paste.

markdown

title: "LLM streaming works in the demo. These 4 hops break it in prod"
published: false
description: "Your server is not a pipe. It sits between two connections, and four hops can fail: proxy buffering, cancellation, mid-stream errors, client parsing."

tags: ai, node, backend, architecture

Words appear one by one on localhost, and LLM streaming looks done. In production the answer lands in one lump, the model keeps generating after the user leaves, and half a sentence shows up as the full answer.

None of these are LLM problems. They live in the plumbing between the model and the browser.

The rule: your server is not a pipe. It sits between two connections, and either one can misbehave.

Scope: text chat over fetch, examples in Node/Express. The ideas carry over to other stacks.

Where does a streamed answer actually break?

Here's the flow:

Browser  ──POST /chat──►  Your API  ──stream:true──►  LLM provider
   ▲                          │                            │
   └──── tokens as events ◄───┴──────── tokens ◄───────────┘
Enter fullscreen mode Exit fullscreen mode

The 30-line demo teaches the pipe model: tokens go in one side and come out the other. That took me a while to unlearn.

Your server sits in the middle of two connections. Something in front of it can buffer. The client behind it can leave. The provider underneath it can fail. And the browser can misread what arrives.

Most streaming bugs are one of these four hops going wrong. The sections below take them in order.

I use fetch with a streamed response for all of this. A chat request needs a message history and an auth token, and fetch gives me a POST body, headers, and abort support. The alternatives are near the end.

Why does the answer arrive all at once in production?

Reverse proxies like Nginx collect the upstream response and send it in bigger chunks by default. That's great for normal pages and terrible for a stream. Your tokens sit in a buffer, and the user sees nothing until it fills or the response ends.

I've seen this called out as the most frequent production surprise with streaming, and it matches what the infrastructure docs say.

In your app, send headers that tell proxies to back off:

res.writeHead(200, {
  "Content-Type": "text/event-stream",
  "Cache-Control": "no-cache, no-transform",
  "Connection": "keep-alive",
  "X-Accel-Buffering": "no", // Nginx-specific: don't buffer this response
});
Enter fullscreen mode Exit fullscreen mode

In Nginx, for the streaming route:

location /api/chat {
    proxy_pass http://app;
    proxy_http_version 1.1;
    proxy_set_header Connection "";
    proxy_buffering off;      # the important one
    gzip off;                 # compression can re-buffer the stream
    proxy_read_timeout 300s;  # longer than your slowest answer
}
Enter fullscreen mode Exit fullscreen mode

A CDN, load balancer, or API gateway in front of you may have its own buffering and idle timeout settings. I can't tell you the right config for each one, so check your provider's docs. What works everywhere is a test:

curl -N -X POST https://your-domain.com/api/chat \
  -H "Content-Type: application/json" \
  -d '{"messages":[{"role":"user","content":"count slowly from 1 to 20"}]}'
Enter fullscreen mode Exit fullscreen mode

-N turns off curl's own buffering. If lines appear gradually, the whole path streams. If they arrive in one lump, something is buffering. Run it against the real production URL, and again after any infrastructure change.

The same hop kills quiet streams. When the model is thinking or a tool call takes several seconds, some proxies and load balancers see silence and decide the connection is dead. Raise timeouts above your slowest realistic answer on every layer, and send a heartbeat. In SSE, a line starting with : is a comment that clients ignore:

const heartbeat = setInterval(() => res.write(": ping\n\n"), 15000);
res.on("close", () => clearInterval(heartbeat));
Enter fullscreen mode Exit fullscreen mode

If your answer involves tools, also stream status events like {"status": "searching documents"}. Ten seconds of silence feels broken. Ten seconds with "Looking that up…" feels normal.

Deploys count too. A rolling restart can kill streams mid-answer, so give your servers a graceful shutdown period at least as long as your longest answer.

What happens when the user closes the tab?

This is the hop I think is most underrated.

A user asks a question, sees it start answering, and closes the tab. In a lot of apps the server does nothing. The loop reading from the LLM keeps going, the provider keeps generating, and you pay for output nobody sees. A Stop button that only hides text in the UI is the same problem with a friendlier face.

A real cancel has to travel the whole chain:

Browser:

const controller = new AbortController();

const res = await fetch("/api/chat", {
  method: "POST",
  headers: { "Content-Type": "application/json" },
  body: JSON.stringify({ messages }),
  signal: controller.signal,
});

// Stop button:
stopButton.onclick = () => controller.abort();
Enter fullscreen mode Exit fullscreen mode

Server:

app.post("/api/chat", async (req, res) => {
  const upstream = new AbortController();

  // If the client leaves before we finish, cancel the LLM call.
  res.on("close", () => {
    if (!res.writableEnded) upstream.abort();
  });

  res.writeHead(200, { "Content-Type": "text/event-stream" /* + headers above */ });

  try {
    const stream = await llm.chat.completions.create(
      { model, messages: req.body.messages, stream: true },
      { signal: upstream.signal }
    );

    for await (const chunk of stream) {
      const text = chunk.choices[0]?.delta?.content;
      if (text) res.write(`data: ${JSON.stringify({ text })}\n\n`);
    }

    res.write(`data: ${JSON.stringify({ done: true })}\n\n`);
  } catch (err) {
    if (!upstream.signal.aborted) {
      res.write(`data: ${JSON.stringify({ error: "generation_failed" })}\n\n`);
    }
  } finally {
    res.end();
  }
});
Enter fullscreen mode Exit fullscreen mode

Choices I'd make again:

  • Listen on the response's close event, and check writableEnded. That separates "the client left early" from "we finished normally." In recent Node versions, I wouldn't rely on the request's close event for this.
  • Pass the same signal to everything the request started. If the question also triggered a vector search, a reranker, or tool calls, they should stop too. Otherwise you cancel the visible part and keep paying for the invisible part.
  • Don't assume the provider stops billing immediately. Aborting the connection usually stops generation, but behavior varies by provider. I'd test it with a short script and check the usage numbers, instead of trusting the docs alone.
  • Decide what to do with partial answers. If a user stops mid-answer, do you save what was generated? Choose on purpose.

How do you report an error after you've already sent 200 OK?

Normally errors are status codes: 400, 429, 500. In a stream, you've already sent 200 OK and some tokens before the problem happens. The provider times out, you hit a rate limit, or a tool call throws. The status code is already gone.

So errors have to travel inside the stream as events. Two rules I follow, both in the server code above:

  1. Send an explicit error event, so the UI can say something useful.
  2. Always send a final done event on success.

The second one is easy to skip and it matters. Without done, the client can't tell "complete" from "the connection dropped after the third sentence." Users will read the truncated answer as the real answer.

In the UI I treat these as three separate end states:

  • Complete: got done
  • Failed: got error
  • Interrupted: the stream ended without either (show "response was cut off, try again")

I'd also avoid silently auto-retrying a half-finished generation. A retry costs a second generation, and it might give a different answer from the one the user already half-read.

Why does the browser show broken or half-parsed text?

The client code looks simple. Two bugs got me.

Bug 1: network chunks don't line up with your events. One read can contain half an event, or three and a half. If you parse each chunk directly, you'll eventually hit a JSON.parse error on a partial line. Keep a buffer, split on the blank line that ends each event, and carry the leftover:

const reader = res.body.getReader();
const decoder = new TextDecoder();
let buffer = "";

while (true) {
  const { value, done } = await reader.read();
  if (done) break;

  // stream: true keeps multi-byte characters intact across chunks
  buffer += decoder.decode(value, { stream: true });

  const events = buffer.split("\n\n");
  buffer = events.pop(); // the last piece may be incomplete, keep it

  for (const evt of events) {
    if (!evt.startsWith("data: ")) continue; // skips heartbeat comments
    handle(JSON.parse(evt.slice(6)));
  }
}
Enter fullscreen mode Exit fullscreen mode

Bug 2: broken characters. An emoji or a non-English character can be several bytes, and a chunk can end in the middle of one. Decoding each chunk separately gives you �. The { stream: true } option on TextDecoder fixes it. This one is nasty because it passes every English-only test.

Rendering has the same trap. If you re-render the whole message on every token, especially if you re-parse Markdown each time, the page starts to stutter. What helps:

  • Batch updates. Update the screen once per animation frame, not once per token.
  • Treat the streaming message as append-only. Avoid rebuilding earlier content.
  • Plan for half-finished Markdown. A code fence that hasn't closed yet can make a renderer flicker.
  • Respect the user's scroll. If they scrolled up to read, don't pull them back down on every token.

And measure the right thing. Users will wait through a long answer if the first token arrives quickly. Time to first token is worth logging more than total duration.

When is this the wrong approach?

When the transport doesn't fit. You have three realistic options:

Option Good for Catch
SSE via EventSource Simple server-to-browser streams GET only, no custom headers, auto-reconnects by itself
Streamed response via fetch Chat: POST body, auth headers, abort support You parse the stream yourself
WebSocket True two-way, low latency (voice, live collaboration) More infrastructure, harder to scale and debug

EventSource reconnecting on its own sounds helpful. For LLMs it can mean a second paid generation for the same question that nobody planned for. If you use it, make retries something you decide.

I'd reach for WebSockets when audio or real two-way traffic is involved. I've used one for a realtime voice assistant, and it's the right tool there. "The user interrupted" is just cancellation with a microphone.

When you shouldn't stream at all.

  • Structured output (JSON, function arguments) is useless half-finished. Wait and validate the whole thing.
  • Server-to-server calls gain nothing from streaming.
  • Answers you need to check before showing (safety filters, "does this match the tool results?") can't be checked until they're done. Streaming and validation pull in opposite directions, and you have to pick which matters more.

Where my own experience runs out. I haven't run streaming at massive scale. Three things I'm still figuring out:

  • Resumable streams. If a user refreshes mid-answer, can they pick up where they left off? I've read about approaches that write chunks to something like Redis Streams so a reconnect can resume, but I haven't built this myself yet.
  • Scaling beyond one server. Long-lived connections behave differently under load balancing and autoscaling, and I'm still learning where the limits are.
  • Backpressure. If the client reads slower than you write, buffers grow on your side. I know the concept, but I haven't pushed it hard enough in practice to have strong opinions.

What should you check before you ship?

One check per hop, plus one metric:

  1. Proxy: run the curl -N command against your real production URL. Do lines arrive gradually, or in one lump?
  2. Your API: start a long answer, close the tab, and watch your server logs. Does the LLM call stop? Do the tools and searches stop with it?
  3. Provider: force an error mid-answer. Does the UI say "failed" or "cut off", or does it show half an answer as if it were complete?
  4. Browser: ask for an answer full of emoji and non-English text. Any �?
  5. Metric: log time to first token, not just total duration.

Next post: keeping LLM output trustworthy when you can't validate it before it streams, covering structured outputs, retries, and what to do with almost-valid JSON.

Which of these four hops broke first for you in production, and what was sitting in front of your server when it did?

Top comments (0)