DEV Community

jidonglab
jidonglab

Posted on

Nginx proxy_buffering Broke My LLM SSE Stream: 0.4s Became 3.4s

My LLM was streaming its first token in 410 milliseconds. My users were seeing it after 3.4 seconds. Same server, same model, same prompt. The only thing between them was Nginx, and Nginx proxy_buffering was quietly holding every SSE stream until it had something worth sending.

Locally it was perfect. In production, the "streaming" UI just sat there, then dumped the whole reply at once like a fax machine. This is the autopsy, with the numbers from my own setup and the one header that fixed it.

TL;DR

  • Nginx buffers proxied responses by default (proxy_buffering on). Your SSE events go into a buffer instead of straight to the client.
  • Default buffers are one memory page (4KB or 8KB). An LLM reply that serializes to less than that never fills a buffer, so it arrives all at once when the stream closes.
  • In my setup, 195 of 214 turns (91%) arrived as a single lump. Median time to first event at the browser: 3.4s. At the origin: 0.41s.
  • Fix: send X-Accel-Buffering: no from the streaming endpoint, or set proxy_buffering off; on that one location. Median dropped to 0.46s.
  • Measure time to first event at the client, not time to first byte at the server. Server-side logs will lie to you.

What was I building?

The system is a voice interviewer. A candidate answers out loud, speech-to-text produces a transcript, Claude writes the interviewer's follow-up question, and text-to-speech reads it back. The follow-up streams token by token, and I cut it at sentence boundaries so TTS can start speaking the first sentence while the model is still writing the second.

That pipeline lives inside Preterview, an interview practice tool that runs voice interviews with three interviewer styles and writes a report afterwards (full disclosure: I built it, at preterview.com). In a voice product, latency is the whole experience. Two seconds of silence after you finish talking feels like the other side is judging you.

Preterview — an interview session in progress

My budget from "user stops talking" to "interviewer starts talking" was about 1.2 seconds. Locally I hit it. In production, median time to first audio was 4.1 seconds.

Why does Nginx buffer SSE streams?

Nginx buffers proxied responses by default to protect your backend from slow clients. With proxy_buffering on, Nginx reads the upstream response into memory buffers as fast as the backend can produce it, frees the backend, and then feeds the data to the client at whatever speed the client can take.

For a normal JSON API, that is a good trade. For Server-Sent Events it is a disaster, because the entire point of SSE is that each event reaches the client now.

The defaults that matter:

proxy_buffering    on;
proxy_buffer_size  4k;    # or 8k, one memory page
proxy_buffers      8 4k;  # or 8 8k
Enter fullscreen mode Exit fullscreen mode

In practice, small SSE events sit in a buffer until it fills or the upstream response ends. Then they go out together.

Why did short LLM replies arrive all at once?

Short LLM replies arrived in one lump because they never filled a single buffer. The stream ended before the buffer did, so the first flush to the client was also the last one.

Here is the math from my own stream. I re-emit model tokens as compact events:

data: {"t":" Can"}

data: {"t":" you"}

data: {"t":" walk"}
Enter fullscreen mode Exit fullscreen mode

That is roughly 20 to 25 bytes per event. A typical interviewer follow-up in my logs is 40 to 80 tokens. Call it 60 tokens times 25 bytes: about 1.5KB. The buffer is 4KB.

So for 91% of turns, the user got nothing, then everything. The other 9% were long multi-part questions, and those arrived in 4KB chunks, which looked almost like streaming.

That is what made this so annoying to debug. Every time I tested a long prompt in production, it seemed to stream "kind of." Every time I tested locally, it streamed perfectly, because there was no Nginx on my laptop. I spent most of an evening blaming the TTS provider.

How do you measure SSE latency correctly?

Measure time to first event at the client, through the same proxy your users hit. My server logs said "first token sent at 410ms" the whole time, and they were right. The server did send it. Nginx just didn't pass it on.

This is the script I ended up using. It hits the real public URL, so the measurement includes every proxy in the path:

import time, httpx, statistics

def time_to_first_event(url, payload):
    t0 = time.perf_counter()
    first = None
    lumps = 0
    last_chunk_at = None
    with httpx.stream("POST", url, json=payload, timeout=30) as r:
        for chunk in r.iter_raw():
            now = time.perf_counter()
            if first is None:
                first = now - t0
            # chunks arriving >50ms apart count as separate deliveries
            if last_chunk_at is None or now - last_chunk_at > 0.05:
                lumps += 1
            last_chunk_at = now
    return first, lumps, time.perf_counter() - t0

results = [time_to_first_event(URL, p) for p in replayed_turns]
print("median first event:", statistics.median(r[0] for r in results))
print("single-lump turns:", sum(1 for r in results if r[1] == 1))
Enter fullscreen mode Exit fullscreen mode

The lumps counter is the important part. A healthy stream of 60 tokens shows dozens of separate deliveries. A buffered one shows exactly one. I replayed 214 turns from test sessions over a weekend:

Metric Origin (direct) Through Nginx (before) Through Nginx (after)
Median first event 0.41s 3.4s 0.46s
p95 first event 0.9s 6.2s 1.0s
Single-lump turns 0 / 214 195 / 214 0 / 214
Median first audio 1.1s 4.1s 1.2s

The 50ms gap between origin and fixed-Nginx is just the extra network hop. I can live with that.

How do you disable Nginx proxy_buffering for SSE?

Disable it per response with the X-Accel-Buffering: no header, or per route with proxy_buffering off;. Nginx honors the header unless someone has configured proxy_ignore_headers X-Accel-Buffering.

I prefer the header, because the app knows which responses are streams and Nginx doesn't. In FastAPI:

from fastapi.responses import StreamingResponse

@app.post("/api/turn/stream")
async def stream_turn(req: TurnRequest):
    return StreamingResponse(
        generate_events(req),
        media_type="text/event-stream",
        headers={
            "X-Accel-Buffering": "no",
            "Cache-Control": "no-cache",
        },
    )
Enter fullscreen mode Exit fullscreen mode

I also scoped it in Nginx as a belt-and-braces backup, only for the streaming path:

location /api/turn/stream {
    proxy_pass http://app;
    proxy_http_version 1.1;
    proxy_set_header Connection "";
    proxy_buffering off;
    proxy_cache off;
    proxy_read_timeout 300s;
}
Enter fullscreen mode Exit fullscreen mode

Do not turn buffering off globally. Every other endpoint still benefits from it, and a slow mobile client on a normal JSON route will happily tie up a backend worker if you do.

What else breaks LLM streaming behind a proxy?

Three more things bit me or nearly did while I was in there:

  1. gzip on the event stream. Nginx's default gzip_types only covers text/html, so SSE is safe out of the box. But if someone "optimized bandwidth" by adding text/event-stream to gzip_types, compression will batch your events again. I checked mine. It was clean, but I added an explicit gzip off; to the stream location anyway.
  2. proxy_read_timeout during long thinking. The default is 60 seconds between reads. If your model thinks for a long time before the first token, or a tool call stalls, Nginx closes the connection. I send an SSE comment line (: ping) every 15 seconds while waiting. Clients ignore comment lines.
  3. Your own framework buffering. Some app servers and middleware collect the response before sending. If the lumps count is 1 even when hitting the origin directly, the problem is upstream of Nginx, not in it.

Was it worth an evening?

It was the cheapest latency win I've ever shipped. No model change, no prompt change, no new infrastructure. One response header took median time to first audio from 4.1s to 1.2s, which is the difference between a conversation and a walkie-talkie.

The lesson I keep relearning: when streaming works locally and not in production, stop looking at the model. Look at every box between the model and the user, and measure from the user's side.

So, why is your LLM SSE stream slow behind Nginx? Because proxy_buffering is on by default, and Nginx holds events in 4KB or 8KB buffers until a buffer fills or the response ends. Short LLM replies never fill a buffer, so the client receives the whole answer at once at the end. Add X-Accel-Buffering: no to your streaming response (or proxy_buffering off; on that one location), keep buffering on everywhere else, and verify by counting how many separate chunks reach a real client.


Written by the developer behind Preterview, an interview prep platform.

Top comments (0)