TL;DR: When building an AI chat assistant with LLM token streaming, most tutorials recommend WebSockets or short polling. In production, WebSockets introduce stateful connection complexity, load balancer friction, and proxy issues, while polling creates unnecessary server load. Here is why we migrated IntelliDesk AI from WebSockets to Server-Sent Events (SSE), how it brought our Time-to-First-Token under 180ms, and the four concrete engineering challenges we solved in our code.
The Initial Architectural Decision and Its Limits
In our initial version of IntelliDesk AI, we used WebSockets (via Flask-SocketIO) to stream responses from Groq's Llama 3.3 70B model to our React frontend.
On paper, WebSockets seemed like the default choice for real-time applications. However, as we tested multi-user workloads and prepared our deployment architecture for NGINX and AWS Application Load Balancers, significant architectural drawbacks emerged:
- Stateful Connection Overhead: WebSockets require a persistent, bi-directional TCP connection. In distributed or containerized environments behind load balancers, scaling requires sticky sessions or a Redis Pub/Sub adapter to manage client-server mapping.
- Context and Threading Friction: In Python WSGI/ASGI environments, Socket.IO often relies on greenlet context switching. Managing request application contexts, database sessions, and auth lifecycles across greenlets introduces subtle concurrency edge cases.
-
Proxy and Firewall Friction: Reverse proxies and enterprise security appliances often buffer or terminate HTTP protocol upgrade requests (
Upgrade: websocket) unless explicitly configured with specialized timeouts.
When analyzing the actual communication pattern of our conversational assistant, the reality became clear:
The user sends a single prompt, and the server streams hundreds of tokens back.
This is an inherently unidirectional communication pattern. Using a full-duplex, stateful WebSocket connection for a one-way stream was unnecessary technical debt.
Evaluating the Three Transport Options
1. SHORT POLLING
Client: Sends request every 500ms -> Server: Returns partial/empty status
Result: High network overhead, SSL handshakes, and delayed perceived response.
2. WEBSOCKETS
Client <============== Full-Duplex Stateful TCP Connection ==============> Server
Result: Bi-directional, stateful, requires sticky sessions and persistent connection state.
3. SERVER-SENT EVENTS (SSE)
Client: Sends 1 HTTP POST with query -> Server: Streams text chunks over keep-alive HTTP
Result: Unidirectional, stateless, works natively with standard HTTP infrastructure and load balancers.
Why Short Polling Failed
Short polling introduces severe resource waste:
- A typical 6-to-8 second AI response requires 12 to 16 separate HTTP requests per user query.
- Each request incurs round-trip latency, authentication checks, and database queries.
- Perceived latency is choppy because tokens appear in staggered batches rather than as a smooth reading stream.
Why Server-Sent Events (SSE) Was the Superior Fit
Server-Sent Events operates on standard HTTP:
- The client sends an HTTP request.
- The server responds with
Content-Type: text/event-streamandTransfer-Encoding: chunked. - The server pushes data chunks line-by-line (
data: ...\n\n) over the open HTTP connection. - When generation ends, the connection closes cleanly.
SSE requires no protocol switching, is completely stateless at the transport layer, and works seamlessly through standard reverse proxies, CDNs, and load balancers.
How We Implemented SSE in IntelliDesk AI
Here is the architecture of our current streaming pipeline:
[React 18 SPA]
│
│ POST /api/v1/ai/chat/stream
│ Authorization: Bearer <JWT>
▼
[NGINX Reverse Proxy] (proxy_buffering off, X-Accel-Buffering: no)
│
▼
[Flask Backend Controller] (stream_with_context, text/event-stream)
│
│ Semantic Search (RAG) + Groq Token Stream
▼
[Groq LPU Engine] (Llama 3.3 70B)
The Backend: Flask Stream Generator
In backend/app/controllers/ai_controller.py, we created an authenticated, rate-limited POST endpoint:
@ai_bp.route("/chat/stream", methods=["POST"])
@jwt_required()
@limiter.limit("30/minute")
def chat_stream_sse():
body = request.get_json(silent=True) or {}
query = (body.get("query") or "").strip()
if not query:
return {"error": {"message": "query is required"}}, 400
user_id = get_current_user_id()
session_uuid = body.get("session_uuid")
generator = AIChatService.generate_chat_sse(
user_id=user_id,
query=query,
session_uuid=session_uuid,
)
return Response(
stream_with_context(generator),
mimetype="text/event-stream",
headers={
"Cache-Control": "no-cache",
"X-Accel-Buffering": "no", # Disables NGINX response buffering
"Connection": "keep-alive",
},
)
In backend/app/services/ai_service.py, the generator yields JSON payloads formatted according to the SSE wire standard:
def _sse(payload: dict) -> bytes:
return f"data: {json.dumps(payload)}\n\n".encode("utf-8")
# 1. Start event
yield _sse({"type": "start", "session_uuid": session.session_uuid})
# 2. Token chunks from Groq API
for chunk in LLMService.chat_completion_stream(messages=messages):
if "chunk" in chunk:
yield _sse({"type": "chunk", "content": chunk["chunk"]})
# 3. Final event with metadata and RAG sources
yield _sse({
"type": "done",
"sources": rag_sources,
"tokens_used": total_tokens,
"latency_ms": latency_ms,
})
Four Concrete Engineering Challenges We Resolved
Implementing SSE in production required addressing four specific technical hurdles:
1. Native EventSource Authentication and Request Method Limitations
The Problem:
The browser's native new EventSource(url) API only supports GET requests and does not allow custom HTTP request headers.
- Passing JWTs in URL query parameters (
/stream?token=xyz) exposes access tokens in access logs, CDN request history, and browser caches. - Lengthy prompts and system context can exceed URL length limitations (HTTP 414).
The Solution:
Instead of using new EventSource(), we used the standard fetch() API combined with ReadableStream (response.body.getReader()):
const response = await fetch(`${API_BASE}/ai/chat/stream`, {
method: "POST",
headers: {
"Content-Type": "application/json",
"Authorization": `Bearer ${token}`,
},
body: JSON.stringify({
query: userMessageContent,
session_uuid: activeSessionUuid,
}),
signal: abortCtrl.signal,
});
This allows us to maintain a POST interface, pass large JSON payloads cleanly, and include standard Authorization headers.
2. TCP Chunk Fragmentation and UTF-8 Boundaries
The Problem:
Network chunks returned by reader.read() do not strictly align with SSE message delimiters (\n\n). A single network read can cut an SSE JSON frame in half:
Chunk 1: data: {"type": "chunk", "content": "The syst
Chunk 2: em is operational"}\n\n
Calling JSON.parse() immediately on incoming bytes causes syntax parse errors. Additionally, multi-byte UTF-8 sequences (such as symbols or accented letters) can be split across packet boundaries.
The Solution:
We implemented a stateful buffer accumulator using TextDecoder({ stream: true }) and delimiter splitting:
const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";
while (true) {
const { done, value } = await reader.read();
if (done) break;
buffer += decoder.decode(value, { stream: true });
const parts = buffer.split("\n\n");
buffer = parts.pop() ?? ""; // Retain incomplete frame in buffer
for (const frame of parts) {
for (const line of frame.split("\n")) {
if (!line.startsWith("data: ")) continue;
try {
const event = JSON.parse(line.slice(6));
if (event.type === "chunk") {
appendContent(event.content);
}
} catch {
// Ignore unparseable or partial lines safely
}
}
}
}
3. Reverse Proxy Buffering (NGINX and Load Balancers)
The Problem:
By default, reverse proxies such as NGINX buffer upstream responses to optimize packet sizes before transmitting data to the client. This behavior prevents tokens from rendering in real-time, resulting in a delay followed by a sudden burst of text.
The Solution:
We addressed this at two levels:
-
Application Layer: Added the
X-Accel-Buffering: noheader to the Flask Response to instruct NGINX to disable proxy buffering for the stream. - Proxy Layer: Configured NGINX with unbuffered settings for the streaming location:
location /api/v1/ai/chat/stream {
proxy_pass http://flask_backend;
proxy_buffering off;
proxy_cache off;
proxy_set_header Connection '';
proxy_http_version 1.1;
chunked_transfer_encoding on;
}
4. Client Disconnections and Component Unmounting
The Problem:
If a user navigated away from the assistant or switched chat sessions while the model was generating, the stream continued running in the background. This resulted in React unmounted component warnings and consumed unnecessary LLM inference quota.
The Solution:
We integrated AbortController into the component lifecycle:
const abortCtrl = new AbortController();
abortCtrlRef.current = abortCtrl;
try {
const response = await fetch(endpoint, {
method: "POST",
signal: abortCtrl.signal,
// ...
});
// ...
} catch (err: any) {
if (err.name === "AbortError") return; // Handled gracefully
}
// Cleanup on unmount or session switch
useEffect(() => {
return () => {
abortCtrlRef.current?.abort();
};
}, [activeSessionUuid]);
Aborting the fetch request drops the HTTP stream immediately, allowing the backend generator to stop execution.
Comparison: WebSockets vs. SSE in Production
| Dimension | WebSockets | Server-Sent Events (SSE) |
|---|---|---|
| Protocol |
ws:// / wss:// (Protocol Upgrade) |
Standard http:// / https://
|
| Connection State | Stateful | Stateless at transport layer |
| Load Balancing | Requires sticky sessions or Redis Pub/Sub | Works natively behind standard ALBs |
| Time-to-First-Token | ~400ms (Handshake + Socket connection) | < 180ms (Direct HTTP POST) |
| Authentication | Custom query/handshake logic | Standard Authorization: Bearer <JWT>
|
| Client Implementation | Socket.IO library dependency | Native fetch() + ReadableStream
|
Conclusion
Choosing the right communication protocol depends on the directional requirements of your data:
- If your application requires bidirectional, high-frequency exchanges (such as collaborative editing or gaming), WebSockets remains the correct choice.
- If your application streams unidirectional data from server to client (such as LLM tokens or real-time metric feeds), Server-Sent Events (SSE) offers lower operational overhead, native load balancer compatibility, and simpler client-side state handling.
Source Code
The implementation discussed here is part of the open-source IntelliDesk AI project:
- GitHub Repository: github.com/Pruthviraj-333/intellidesk-ai
- Architecture Walkthrough: YouTube Demo
Top comments (0)