DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Browsers answer WebSocket pings and never tell your JavaScript, so we send two heartbeats

A WebSocket that has died does not tell you it has died. Both ends go on believing in a connection that no longer exists: the server keeps writing into a socket nobody is reading, and the browser keeps showing whatever arrived last. For a dashboard that is a stale number. For a quiz, it is a player staring at question four while the room has moved to question six, and no error anywhere.

Our app runs in pubs, which is roughly the worst connectivity environment there is that still counts as indoors. Captive portals. Routers nobody has rebooted since 2019. People wandering outside for a cigarette and handing off to patchy 4G. Phones that lock, freeze the tab, and come back forty minutes later. Every one of those produces a socket that is open according to readyState and dead according to reality.

Here is the whole ladder we ended up with, and the one API fact that forces the shape of it.

The fact: ping and pong are invisible to the page

The WebSocket protocol has ping and pong control frames built in, exactly for liveness. A browser implements them and answers a ping automatically. It does not expose either to JavaScript. There is no event, no callback, no onping.

That asymmetry decides everything. The server can use the protocol. The client cannot.

So the server pings, and counts:

const HEARTBEAT_INTERVAL_MS = 20_000;
/** Missed pongs before a socket is considered dead. Two, to spare slow phones. */
const MAX_MISSED_PONGS = 2;

const heartbeat = setInterval(() => {
  for (const ws of wss.clients) {
    const missed = missedPongs.get(ws) ?? 0;
    if (missed >= MAX_MISSED_PONGS) {
      ws.terminate();
      continue;
    }
    missedPongs.set(ws, missed + 1);
    ws.ping();
    if (ws.readyState === ws.OPEN) ws.send(heartbeatMessage);
  }
}, HEARTBEAT_INTERVAL_MS);
Enter fullscreen mode Exit fullscreen mode

The counter is incremented when the ping is sent and reset when the pong arrives, so a socket that has stopped answering accumulates misses and gets terminated on the third pass. Two misses rather than one because a phone on a congested access point can genuinely take twenty seconds to answer, and dropping it for being slow is a worse outcome than keeping a doubtful socket for another twenty.

And in the same loop, the part that exists purely because the page cannot see the protocol:

const heartbeatMessage = JSON.stringify({ type: 'heartbeat' });
Enter fullscreen mode Exit fullscreen mode

An application-level message with no content, whose only job is to make noise the client can observe. The client's handler for it is one line:

case 'heartbeat':
  return
Enter fullscreen mode Exit fullscreen mode

It does nothing, and that is the point. Everything it needed to accomplish happened before the switch statement:

ws.onmessage = (event: MessageEvent) => {
  lastMessageAtRef.current = Date.now()
Enter fullscreen mode Exit fullscreen mode

The watchdog: silence is the signal

Once every message stamps a timestamp, "is this socket alive" becomes "when did we last hear anything":

/** The server sends a heartbeat every 20 s; this much silence means the socket is dead. */
const STALE_SOCKET_MS = 50_000
const WATCHDOG_INTERVAL_MS = 5_000
Enter fullscreen mode Exit fullscreen mode

Fifty seconds against a twenty-second heartbeat allows two missed beats plus slack. Checking every five seconds means detection lands within fifty-five seconds of the socket going quiet.

if (Date.now() - lastMessageAtRef.current > STALE_SOCKET_MS) {
  console.warn('[useGameWs] No message from the server in 50s, reconnecting')
  reconnectNow()
}
Enter fullscreen mode Exit fullscreen mode

Note what it does not do: inspect readyState. A half-open socket reports OPEN and will go on reporting OPEN indefinitely. The only honest evidence of life is traffic, so the watchdog replaces the socket rather than asking it how it feels.

Reconnection: back off, but not when you already gave up

export function computeBackoff(attempt: number): number {
    return Math.min(1000 * Math.pow(2, attempt), 30000)
}
Enter fullscreen mode Exit fullscreen mode

One second, two, four, up to a thirty-second ceiling. Standard, and worth saying out loud why it is standard: the failure mode a heartbeat introduces is a server restart handing you every client at once, each of them retrying. Back-off is what stops your own liveness checking from being a thundering herd.

There is a second branch, and it is the one most implementations miss. Once the client has already fallen back to polling, exponential back-off is wrong, because the attempt counter is high and the delay has reached its cap for reasons that are no longer relevant:

const degraded = isPollingRef.current || statusRef.current === 'error'
const delay = degraded ? WS_RETRY_WHILE_DEGRADED_MS : computeBackoff(attempt)
Enter fullscreen mode Exit fullscreen mode

A flat thirty seconds. The polling fallback is working, so there is no urgency, but there is also no reason to let the retry interval drift towards hours. The socket gets a probe every half minute for as long as the session lasts, and the moment one succeeds the client switches back and stops polling.

Waking up: three events, one handler

A phone in a pocket is not a phone with a slow connection. The OS freezes the tab. Timers do not fire. The watchdog does not run. The socket may or may not have survived, and nothing in the page knows which.

document.addEventListener('visibilitychange', wake)
window.addEventListener('online', wake)
window.addEventListener('pageshow', wake)
Enter fullscreen mode Exit fullscreen mode

Unlocked, network returned, restored from the back-forward cache. The handler makes one decision:

const fresh = Date.now() - lastMessageAtRef.current < WAKE_FRESH_MS
if (ws && ws.readyState === WebSocket.OPEN && fresh && statusRef.current === 'connected') {
  // Alive, but events may have been missed while the page was frozen.
  ws.send(JSON.stringify({ type: 'sync' }))
  return
}
reconnectNow()
Enter fullscreen mode Exit fullscreen mode

If the socket has heard from the server within the last twenty-five seconds, it survived the freeze and only the state is suspect, so ask for a fresh snapshot over the connection you already have. Otherwise throw it away and reconnect. The distinction matters at scale: a room of sixty phones unlocking as a question lands should produce sixty small sync requests, not sixty new WebSocket handshakes.

The fallback, and the jitter that makes it safe

After twenty seconds of failed reconnection the client stops waiting and starts polling for state over plain HTTP:

const POLL_INTERVAL_MS = 5_000
const POLL_JITTER_MS = 1_500
Enter fullscreen mode Exit fullscreen mode
pollTimerRef.current = setTimeout(poll, POLL_INTERVAL_MS + Math.random() * POLL_JITTER_MS)
Enter fullscreen mode Exit fullscreen mode

The jitter is not politeness. When a venue's access point drops and recovers, every phone in the building transitions to polling inside the same second, and without jitter they stay in lockstep for the rest of the night: a flat background load with a spike every five seconds, from one IP address, which is exactly the shape that trips a rate limiter. Spreading them over 1.5 seconds turns a spike into a floor.

That the polling fallback is rate limited per player rather than per IP is its own story, and the reason is the same one: under NAT, a whole pub is one address.

The ladder, bottom to top

Layer Detects Within
Server ping and pong count Dead client sockets ~60 s
App-level heartbeat message Nothing by itself; it is the traffic the client measures 20 s cadence
Client watchdog Silence on an apparently open socket ~55 s
Wake handlers A tab that was frozen Immediately on unlock
Back-off reconnect Transient drops 1 s to 30 s
Polling fallback A WebSocket that will not come back 20 s after first failure

Six mechanisms for one concern, which sounds like over-engineering until you notice that each catches a failure the others structurally cannot see. The server's ping cannot reach a frozen tab. The watchdog cannot run inside one. The wake handler cannot help a phone with no signal. And none of them can do anything about a captive portal that is cheerfully returning a login page for your WebSocket upgrade, which is what the polling fallback is for.

The product side of this

  • pub-trivia.app/faq answers the bad-WiFi question in customer language. It is this post with the arithmetic removed, and it is the question venues actually ask first.
  • pub-trivia.app/features/live-leaderboard is what all of the above protects: a leaderboard that moves when a question closes is only a feature if it keeps moving on a bad night.
  • pub-trivia.app/solutions/pubs is the environment these numbers were chosen for, written for the person who owns the room rather than the socket.
  • To watch the ladder work, start a free session, join on a phone, and put the phone in flight mode for a minute. Reconnect, snapshot, and the quiz carries on from where the room is, not from where the phone left off.

Top comments (0)