DEV Community

Cover image for 5 Mistakes That Make LLM Streaming Break on a Flaky Network
MICHAEL-MAURICE
MICHAEL-MAURICE

Posted on

5 Mistakes That Make LLM Streaming Break on a Flaky Network

Your chat UI looks fine on the office Wi-Fi. Then someone joins from a train, the SSE connection drops mid-sentence, and either the answer restarts from scratch or the model keeps burning tokens while nobody is listening. Same feature, different network.

I spent a stretch of time turning a "works in demos" stream into something that survives reconnects. The full working project (ASP.NET Core 10, TypedResults.ServerSentEvents, tests) lives in Tech Skill Builder. This post is the short checklist: five mistakes I kept seeing, and the shape that fixes them.

Mistake 1: Tying generation to the HTTP request

The five-minute version wires GetStreamingResponseAsync to HttpContext.RequestAborted. Close the tab, and the model call stops. That feels thrifty until you realize a flaky proxy is doing the same thing: every blink cancels work you already paid for, and a reconnect starts a new prompt.

Fix: split start from subscribe.

  1. POST /api/chat/streams starts generation in a registry and returns 202 with streamId, eventsUrl, cancelUrl.
  2. GET /api/chat/streams/{id}/events is the SSE subscription.
  3. Generation keeps running even if the subscriber disconnects (until a sweeper decides nobody cares anymore).

Skeleton:

// POST /api/chat/streams
var stream = registry.TryStart(messages, owner);
return TypedResults.Accepted(
    $"/api/chat/streams/{stream.Id}/events",
    new StreamStarted(stream.Id, eventsUrl, cancelUrl));

// GET /api/chat/streams/{id}/events
// Last-Event-ID header (or ?lastEventId=) → replay missed events
return TypedResults.ServerSentEvents(
    StreamSubscriber.ReadAsync(stream, after, options, time, http.RequestAborted));
// full implementation in the complete project
Enter fullscreen mode Exit fullscreen mode

EventSource only speaks GET. That alone pushes you toward this split: you cannot POST a body on reconnect.

Mistake 2: No sequence ids, so reconnect means "start over"

Without an SSE id: on each event, the browser has nothing to put in Last-Event-ID. Your client either duplicates tokens or asks the server to regenerate.

Fix: keep an append-only StreamLog per answer. Every event gets a sequence number. On subscribe, parse Last-Event-ID, then replay only what was missed. Bound the buffer; a very late client gets one snapshot with the text so far instead of every erased delta.

public long Append(ChatStreamEvent e)
{
    // assign next sequence, store event, wake waiters
    // if over capacity → trim oldest; keep aggregated text for snapshots
    // full implementation in the complete project
}
Enter fullscreen mode Exit fullscreen mode

Caught-up client on a finished stream? Return 204 No Content. That is the SSE-spec way to stop EventSource from reconnecting forever.

Mistake 3: Naming a server event error

Browsers fire EventSource's own error for connection trouble. If your payload also uses event: error, both land in the same mental bucket (and often the same handler). You will debug "is the model broken or is the Wi-Fi broken?" for longer than you want.

Fix: use a small, boring vocabulary:

Event Meaning
delta text fragment
tool name + status only (not arguments/results)
done finished (stop, cancelled, …)
failed in-band model/server failure after HTTP 200
heartbeat keep proxies from killing an idle stream
snapshot replace local text after a truncated buffer

Call the failure event failed, not error. Put a user-safe message and an errorId in the payload. Keep partial text on screen.

Mistake 4: Silent streams and no Stop button

Proxies and load balancers love idle connections. While the model thinks or a tool runs, your SSE can sit quiet long enough to get cut. Separately, users hit Stop and expect the bill to stop too. If cancel only closes the browser side, the server may keep generating.

Fix:

  • Emit heartbeat on a timer (for example every 15 seconds) while work is in progress.
  • Expose POST /api/chat/streams/{id}/cancel that cancels the model call and ends with done / finishReason: cancelled.
  • Send X-Accel-Buffering: no so nginx does not buffer the stream into one giant blob.
// Cancel endpoint shape
if (registry.Find(id) is not { } stream || stream.Owner != owner)
    return TypedResults.NotFound();
return stream.RequestCancel()
    ? TypedResults.Accepted((string?)null)
    : TypedResults.Conflict(); // already finished
// full implementation in the complete project
Enter fullscreen mode Exit fullscreen mode

Mistake 5: Shipping the "quick" stream as production

TypedResults.ServerSentEvents over a raw token loop is great for a spike. It is not a product API. You get no resume, no ownership check, no rate limit on active streams, and no story for multi-instance hosts.

Checklist before you call it done:

  • [ ] POST start + GET subscribe (EventSource-friendly)
  • [ ] Sequence ids + Last-Event-ID replay
  • [ ] Generation outlives the request; sweeper cleans orphans
  • [ ] Event names avoid colliding with EventSource (failed not error)
  • [ ] Heartbeats + cancel endpoint
  • [ ] Owner binding (cookie auth or signed URL; EventSource cannot set Authorization)
  • [ ] Sticky sessions or a shared log (Redis Streams, etc.) if you scale out
  • [ ] Prefer HTTP/2 or HTTP/3 (browsers cap ~6 SSE connections per domain on HTTP/1.1)

What I leave out here on purpose

The complete project has the coalescer (batch tiny token fragments), the sweeper timings, the offline streaming model for demos without a key, and the full test suite. This article is the map, not the hike.

If you want the working source and the deeper write-up package, grab it from Tech Skill Builder. Membership includes the full resumable streaming solution ready to run with dotnet test and dotnet run. Limited-time member pricing is on the product page; if you are building anything that streams tokens to a browser, this is the piece you want finished before the next flaky demo.

Flaky networks are not an edge case. Treat resume as part of the feature, not a patch after the first support ticket.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to