The advice is everywhere: if you are streaming Server-Sent Events through
nginx, turn off proxy_buffering or your tokens will arrive in one lump.
...
For further actions, you may consider blocking this person and/or reporting abuse
@remdore : good data, and it holds up. Thanks for sharing it! Disclosure: I work at HAProxy, so weigh this however you like.
I dug into where the 206ms comes from. It isn't an application buffer. By default, HAProxy tags in-flight HTTP body writes with the kernel's
MSG_MOREflag, so the kernel waits to fill packets. That's a smart default for bulk HTTP, and the manual documents the ~200ms behavior and the way to adjust it:option http-no-delay.Your "what I would still test" list already had the right question: does
option http-no-delayremove it? Running that cell, using 60-byte SSE events every 50 ms. Default: 212 ms first event, 5 frames per read. Withoption http-no-delay: 1 ms, 1 frame per read. Might be fun to try on your rig.One note: it's best enabled only on the backends that carry SSE or token streams, so everything else keeps the default's throughput benefits.
Full write-up with the mechanism, sources, and reproduction: dev.to/rlnorthcutt/haproxys-200-ms...
Ran it.
option http-no-delaydoes exactly what you said, and there is a second condition worth adding.HAProxy 2.9.15 in Docker, 60-byte SSE events every 50ms, raw socket client recording arrival time and events per read:
So the fix lands where you said it would, and the MSG_MORE explanation fits the shape of it.
The extra condition: the hold only appears when the upstream sends
Transfer-Encoding: chunked. Same proxy, same default config, same client, switching nothing but the upstream framing:I did not expect that and I do not have a mechanism for it, so I am stating it as measured rather than explaining it. It may be part of why this bites some people and not others, since a close-delimited stream never got held on my rig and plenty of small SSE endpoints are close-delimited.
Point taken on scoping the option to the backends that actually carry streams rather than turning it on globally.
UPDATE: Willy has implemented the check for text/event-stream to automatically apply http-no-delay, we're now at +0.04ms (40 microseconds) per token and no longer 200ms :-)
github.com/haproxy/haproxy/commit/...
He also added a simple latency measurement tool for SSE through a gateway (in the previous commit). You may be interested in running your own tests with it on various tools.
This is a great reminder that sharing your accomplishments doesn't make you egotistical. It's okay to be proud of your work and celebrate your progress.
I also share my coding journey and achievements on Codecan.net, and I believe there's nothing wrong with that. We all have different paths, and someone's success doesn't take away from our own.
Keep sharing your wins. You deserve to feel good about what you've achieved!
The frames-per-read ratio is the part I keep thinking about. Most buffering debugging I've seen anchors on absolute latency deltas, which is exactly the number a noisy shared host destroys — and then you start blaming a proxy for a 40ms scheduler hiccup. We serve everything behind one reverse proxy and our streams are all small token-sized frames on a steady interval. Never once thought to test whether the proxy was coalescing them; we assumed nginx. What's the emitter you used — is the rig published somewhere? I'd like to point it at our config before I conclude we're fine.
Not published, so the honest answer is that you cannot point it at your config today.
The emitter is about 180 lines of stdlib and the shape is more useful than the code: a plain http.server that sends SSE frames against a monotonic deadline computed from a fixed start, never sleep(interval) in a loop, and which records its own send timestamps and exposes them on an endpoint so you can audit whether it actually paced before believing anything it produced.
The client matters more than the emitter and is the part I got wrong first. Do not use urllib or requests: they buffer, and the buffering is the thing you are measuring. Raw socket, write the request by hand, TCP_NODELAY, Connection: close so the loop ends on EOF, and timestamp every recv() the moment it returns. Then count arrivals per read, not per frame. Two frames in one recv is one arrival carrying two events, and getting that backwards produces numbers that look entirely plausible and mean nothing.
Given your setup, a fleet of agents behind one proxy with small token-sized frames on a steady interval, you are in the exact regime where this shows up. Frame size is the first thing I would check, before touching any config. At roughly 60-byte frames HAProxy held first token 206ms; at roughly 1.1KB the same stock config gave 53ms. If you are already batching several tokens per frame you may have nothing to fix.
Correcting myself: it is published now, mostly because you asked twice and the second time I had no good reason not to.
github.com/DimitrovK/sse-proxy-buf...
Everything that produced the numbers is in there, including the image digests and every proxy config as run, so the configs in the post cannot have drifted from the ones measured. Raw per-run rows too, not just the aggregates, so you can check the exclusions rather than take my word for them.
One thing before you point it at your own config: on macOS the pacing test fails and the self-test refuses to certify. That is the rig working, not a broken checkout, and I left the tolerance alone rather than relaxing it so a laptop passes. Develop wherever, measure on Linux. A busy host is how I got a 42ms finding that did not exist.
For your fleet I would still check frame size before touching any proxy config.
This is why "turn off proxy_buffering" might be the most repeated non-fix in streaming: the header is an nginx convention, and nginx was never the one holding your tokens. The frames-per-read metric is the smart part - a ratio that survives a noisy host and needs no baseline, which is exactly what you want when the claim is "this proxy batches." Bursts of five with zero gap, 206ms late, is a beautifully damning row. I have debugged a "streaming is slow" complaint that was really "someone put a buffering proxy in front of it," and the folk advice sent everyone looking at the wrong layer.
Mis-aimed rather than a non-fix, I think, and the distinction took me an extra cell to find.
Every small-frame condition says the directive buys nothing: nginx and nginx with proxy_buffering off are identical at 1.02, both indistinguishable from no proxy. But that is a client reading promptly and a stream of about 2.4KB total, against nginx's default proxy_buffers of 8 by 4k or 8k. The buffers are never close to full, so the directive has nothing to do.
Push 328KB through a client that sleeps 200ms between reads and it separates: 53ms to first token with buffering on, 3ms with it off. Consistent too, median 53 and max 54 over ten runs, never once fast. So the advice is right in a regime nobody mentions and irrelevant in the one most people are actually in.
What it never did, in any of the nine conditions, is the thing it is famous for. Frames per read tracked the direct baseline everywhere. It cost latency; it did not turn the stream into one lump.
the hold time difference is wild. i'd also compare sticky-session behavior and maxconn under the same load so the 206ms isn't just buffering hiding a backend stall.
Neither applies here, and the reason is the topology rather than the tuning. The backend is one server, there is no maxconn anywhere in the config, and nothing sticky to be sticky about, because the client opens one connection at a time and sends Connection: close. No session to pin and no queue to saturate.
The backend-stall version of your question is the real one though, and it is the thing the rig is built to rule out. The upstream is a synthetic emitter that paces frames against a monotonic deadline and records its own send timestamps. After every run the harness asks the emitter whether it actually paced as instructed, and a run that drifted past tolerance is discarded rather than attributed to a proxy. So an upstream that stalled shows up as emitter drift and voids the cell. It cannot arrive as a 206ms proxy row.
The direct, unproxied baseline is measured in the same cell as the proxies: 2ms to first token, frames 50ms apart. Same emitter, same client, same host, same moment. A backend stall would have to be selective about which of six endpoints it stalled for.