DEV Community

Cover image for Stop Using WebSockets for AI Chat Apps: Why We Switched to Server-Sent Events (SSE) in Production
Pruthviraj Janwade
Pruthviraj Janwade

Posted on Originally published at pruthvirajjanwade.hashnode.dev

Stop Using WebSockets for AI Chat Apps: Why We Switched to Server-Sent Events (SSE) in Production

TL;DR: When building an AI chat assistant with LLM token streaming, most tutorials recommend WebSockets or short polling. In production, WebSockets introduce stateful connection complexity, load balancer friction, and proxy issues, while polling creates unnecessary server load. Here is why we migrated IntelliDesk AI from WebSockets to Server-Sent Events (SSE), how it brought our Time-to-First-Token under 180ms, and the four concrete engineering challenges we solved in our code.


The Initial Architectural Decision and Its Limits

In our initial version of IntelliDesk AI, we used WebSockets (via Flask-SocketIO) to stream responses from Groq's Llama 3.3 70B model to our React frontend.

On paper, WebSockets seemed like the default choice for real-time applications. However, as we tested multi-user workloads and prepared our deployment architecture for NGINX and AWS Application Load Balancers, significant architectural drawbacks emerged:

  1. Stateful Connection Overhead: WebSockets require a persistent, bi-directional TCP connection. In distributed or containerized environments behind load balancers, scaling requires sticky sessions or a Redis Pub/Sub adapter to manage client-server mapping.
  2. Context and Threading Friction: In Python WSGI/ASGI environments, Socket.IO often relies on greenlet context switching. Managing request application contexts, database sessions, and auth lifecycles across greenlets introduces subtle concurrency edge cases.
  3. Proxy and Firewall Friction: Reverse proxies and enterprise security appliances often buffer or terminate HTTP protocol upgrade requests (Upgrade: websocket) unless explicitly configured with specialized timeouts.

When analyzing the actual communication pattern of our conversational assistant, the reality became clear:

The user sends a single prompt, and the server streams hundreds of tokens back.

This is an inherently unidirectional communication pattern. Using a full-duplex, stateful WebSocket connection for a one-way stream was unnecessary technical debt.


Evaluating the Three Transport Options

1. SHORT POLLING
Client: Sends request every 500ms -> Server: Returns partial/empty status
Result: High network overhead, SSL handshakes, and delayed perceived response.

2. WEBSOCKETS
Client <============== Full-Duplex Stateful TCP Connection ==============> Server
Result: Bi-directional, stateful, requires sticky sessions and persistent connection state.

3. SERVER-SENT EVENTS (SSE)
Client: Sends 1 HTTP POST with query -> Server: Streams text chunks over keep-alive HTTP
Result: Unidirectional, stateless, works natively with standard HTTP infrastructure and load balancers.
Enter fullscreen mode Exit fullscreen mode

Why Short Polling Failed

Short polling introduces severe resource waste:

  • A typical 6-to-8 second AI response requires 12 to 16 separate HTTP requests per user query.
  • Each request incurs round-trip latency, authentication checks, and database queries.
  • Perceived latency is choppy because tokens appear in staggered batches rather than as a smooth reading stream.

Why Server-Sent Events (SSE) Was the Superior Fit

Server-Sent Events operates on standard HTTP:

  • The client sends an HTTP request.
  • The server responds with Content-Type: text/event-stream and Transfer-Encoding: chunked.
  • The server pushes data chunks line-by-line (data: ...\n\n) over the open HTTP connection.
  • When generation ends, the connection closes cleanly.

SSE requires no protocol switching, is completely stateless at the transport layer, and works seamlessly through standard reverse proxies, CDNs, and load balancers.


How We Implemented SSE in IntelliDesk AI

Here is the architecture of our current streaming pipeline:

[React 18 SPA]
     │
     │  POST /api/v1/ai/chat/stream
     │  Authorization: Bearer <JWT>
     ▼
[NGINX Reverse Proxy] (proxy_buffering off, X-Accel-Buffering: no)
     │
     ▼
[Flask Backend Controller] (stream_with_context, text/event-stream)
     │
     │  Semantic Search (RAG) + Groq Token Stream
     ▼
[Groq LPU Engine] (Llama 3.3 70B)
Enter fullscreen mode Exit fullscreen mode

The Backend: Flask Stream Generator

In backend/app/controllers/ai_controller.py, we created an authenticated, rate-limited POST endpoint:

@ai_bp.route("/chat/stream", methods=["POST"])
@jwt_required()
@limiter.limit("30/minute")
def chat_stream_sse():
    body = request.get_json(silent=True) or {}
    query = (body.get("query") or "").strip()
    if not query:
        return {"error": {"message": "query is required"}}, 400

    user_id = get_current_user_id()
    session_uuid = body.get("session_uuid")

    generator = AIChatService.generate_chat_sse(
        user_id=user_id,
        query=query,
        session_uuid=session_uuid,
    )

    return Response(
        stream_with_context(generator),
        mimetype="text/event-stream",
        headers={
            "Cache-Control":    "no-cache",
            "X-Accel-Buffering": "no",       # Disables NGINX response buffering
            "Connection":       "keep-alive",
        },
    )
Enter fullscreen mode Exit fullscreen mode

In backend/app/services/ai_service.py, the generator yields JSON payloads formatted according to the SSE wire standard:

def _sse(payload: dict) -> bytes:
    return f"data: {json.dumps(payload)}\n\n".encode("utf-8")

# 1. Start event
yield _sse({"type": "start", "session_uuid": session.session_uuid})

# 2. Token chunks from Groq API
for chunk in LLMService.chat_completion_stream(messages=messages):
    if "chunk" in chunk:
        yield _sse({"type": "chunk", "content": chunk["chunk"]})

# 3. Final event with metadata and RAG sources
yield _sse({
    "type": "done",
    "sources": rag_sources,
    "tokens_used": total_tokens,
    "latency_ms": latency_ms,
})
Enter fullscreen mode Exit fullscreen mode

Four Concrete Engineering Challenges We Resolved

Implementing SSE in production required addressing four specific technical hurdles:


1. Native EventSource Authentication and Request Method Limitations

The Problem:
The browser's native new EventSource(url) API only supports GET requests and does not allow custom HTTP request headers.

  • Passing JWTs in URL query parameters (/stream?token=xyz) exposes access tokens in access logs, CDN request history, and browser caches.
  • Lengthy prompts and system context can exceed URL length limitations (HTTP 414).

The Solution:
Instead of using new EventSource(), we used the standard fetch() API combined with ReadableStream (response.body.getReader()):

const response = await fetch(`${API_BASE}/ai/chat/stream`, {
  method: "POST",
  headers: {
    "Content-Type": "application/json",
    "Authorization": `Bearer ${token}`,
  },
  body: JSON.stringify({
    query: userMessageContent,
    session_uuid: activeSessionUuid,
  }),
  signal: abortCtrl.signal,
});
Enter fullscreen mode Exit fullscreen mode

This allows us to maintain a POST interface, pass large JSON payloads cleanly, and include standard Authorization headers.


2. TCP Chunk Fragmentation and UTF-8 Boundaries

The Problem:
Network chunks returned by reader.read() do not strictly align with SSE message delimiters (\n\n). A single network read can cut an SSE JSON frame in half:

Chunk 1: data: {"type": "chunk", "content": "The syst
Chunk 2: em is operational"}\n\n
Enter fullscreen mode Exit fullscreen mode

Calling JSON.parse() immediately on incoming bytes causes syntax parse errors. Additionally, multi-byte UTF-8 sequences (such as symbols or accented letters) can be split across packet boundaries.

The Solution:
We implemented a stateful buffer accumulator using TextDecoder({ stream: true }) and delimiter splitting:

const reader = response.body.getReader();
const decoder = new TextDecoder();
let buffer = "";

while (true) {
  const { done, value } = await reader.read();
  if (done) break;

  buffer += decoder.decode(value, { stream: true });

  const parts = buffer.split("\n\n");
  buffer = parts.pop() ?? ""; // Retain incomplete frame in buffer

  for (const frame of parts) {
    for (const line of frame.split("\n")) {
      if (!line.startsWith("data: ")) continue;
      try {
        const event = JSON.parse(line.slice(6));
        if (event.type === "chunk") {
          appendContent(event.content);
        }
      } catch {
        // Ignore unparseable or partial lines safely
      }
    }
  }
}
Enter fullscreen mode Exit fullscreen mode

3. Reverse Proxy Buffering (NGINX and Load Balancers)

The Problem:
By default, reverse proxies such as NGINX buffer upstream responses to optimize packet sizes before transmitting data to the client. This behavior prevents tokens from rendering in real-time, resulting in a delay followed by a sudden burst of text.

The Solution:
We addressed this at two levels:

  1. Application Layer: Added the X-Accel-Buffering: no header to the Flask Response to instruct NGINX to disable proxy buffering for the stream.
  2. Proxy Layer: Configured NGINX with unbuffered settings for the streaming location:
location /api/v1/ai/chat/stream {
    proxy_pass http://flask_backend;
    proxy_buffering off;
    proxy_cache off;
    proxy_set_header Connection '';
    proxy_http_version 1.1;
    chunked_transfer_encoding on;
}
Enter fullscreen mode Exit fullscreen mode

4. Client Disconnections and Component Unmounting

The Problem:
If a user navigated away from the assistant or switched chat sessions while the model was generating, the stream continued running in the background. This resulted in React unmounted component warnings and consumed unnecessary LLM inference quota.

The Solution:
We integrated AbortController into the component lifecycle:

const abortCtrl = new AbortController();
abortCtrlRef.current = abortCtrl;

try {
  const response = await fetch(endpoint, {
    method: "POST",
    signal: abortCtrl.signal,
    // ...
  });
  // ...
} catch (err: any) {
  if (err.name === "AbortError") return; // Handled gracefully
}

// Cleanup on unmount or session switch
useEffect(() => {
  return () => {
    abortCtrlRef.current?.abort();
  };
}, [activeSessionUuid]);
Enter fullscreen mode Exit fullscreen mode

Aborting the fetch request drops the HTTP stream immediately, allowing the backend generator to stop execution.


Comparison: WebSockets vs. SSE in Production

Dimension WebSockets Server-Sent Events (SSE)
Protocol ws:// / wss:// (Protocol Upgrade) Standard http:// / https://
Connection State Stateful Stateless at transport layer
Load Balancing Requires sticky sessions or Redis Pub/Sub Works natively behind standard ALBs
Time-to-First-Token ~400ms (Handshake + Socket connection) < 180ms (Direct HTTP POST)
Authentication Custom query/handshake logic Standard Authorization: Bearer <JWT>
Client Implementation Socket.IO library dependency Native fetch() + ReadableStream

Conclusion

Choosing the right communication protocol depends on the directional requirements of your data:

  • If your application requires bidirectional, high-frequency exchanges (such as collaborative editing or gaming), WebSockets remains the correct choice.
  • If your application streams unidirectional data from server to client (such as LLM tokens or real-time metric feeds), Server-Sent Events (SSE) offers lower operational overhead, native load balancer compatibility, and simpler client-side state handling.

Source Code

The implementation discussed here is part of the open-source IntelliDesk AI project:

Top comments (0)