A WebSocket that has died does not tell you it has died. Both ends go on believing in a connection that no longer exists: the server keeps writing into a socket nobody is reading, and the browser keeps showing whatever arrived last. For a dashboard that is a stale number. For a quiz, it is a player staring at question four while the room has moved to question six, and no error anywhere.
Our app runs in pubs, which is roughly the worst connectivity environment there is that still counts as indoors. Captive portals. Routers nobody has rebooted since 2019. People wandering outside for a cigarette and handing off to patchy 4G. Phones that lock, freeze the tab, and come back forty minutes later. Every one of those produces a socket that is open according to readyState and dead according to reality.
Here is the whole ladder we ended up with, and the one API fact that forces the shape of it.
The fact: ping and pong are invisible to the page
The WebSocket protocol has ping and pong control frames built in, exactly for liveness. A browser implements them and answers a ping automatically. It does not expose either to JavaScript. There is no event, no callback, no onping.
That asymmetry decides everything. The server can use the protocol. The client cannot.
So the server pings, and counts:
const HEARTBEAT_INTERVAL_MS = 20_000;
/** Missed pongs before a socket is considered dead. Two, to spare slow phones. */
const MAX_MISSED_PONGS = 2;
const heartbeat = setInterval(() => {
for (const ws of wss.clients) {
const missed = missedPongs.get(ws) ?? 0;
if (missed >= MAX_MISSED_PONGS) {
ws.terminate();
continue;
}
missedPongs.set(ws, missed + 1);
ws.ping();
if (ws.readyState === ws.OPEN) ws.send(heartbeatMessage);
}
}, HEARTBEAT_INTERVAL_MS);
The counter is incremented when the ping is sent and reset when the pong arrives, so a socket that has stopped answering accumulates misses and gets terminated on the third pass. Two misses rather than one because a phone on a congested access point can genuinely take twenty seconds to answer, and dropping it for being slow is a worse outcome than keeping a doubtful socket for another twenty.
And in the same loop, the part that exists purely because the page cannot see the protocol:
const heartbeatMessage = JSON.stringify({ type: 'heartbeat' });
An application-level message with no content, whose only job is to make noise the client can observe. The client's handler for it is one line:
case 'heartbeat':
return
It does nothing, and that is the point. Everything it needed to accomplish happened before the switch statement:
ws.onmessage = (event: MessageEvent) => {
lastMessageAtRef.current = Date.now()
The watchdog: silence is the signal
Once every message stamps a timestamp, "is this socket alive" becomes "when did we last hear anything":
/** The server sends a heartbeat every 20 s; this much silence means the socket is dead. */
const STALE_SOCKET_MS = 50_000
const WATCHDOG_INTERVAL_MS = 5_000
Fifty seconds against a twenty-second heartbeat allows two missed beats plus slack. Checking every five seconds means detection lands within fifty-five seconds of the socket going quiet.
if (Date.now() - lastMessageAtRef.current > STALE_SOCKET_MS) {
console.warn('[useGameWs] No message from the server in 50s, reconnecting')
reconnectNow()
}
Note what it does not do: inspect readyState. A half-open socket reports OPEN and will go on reporting OPEN indefinitely. The only honest evidence of life is traffic, so the watchdog replaces the socket rather than asking it how it feels.
Reconnection: back off, but not when you already gave up
export function computeBackoff(attempt: number): number {
return Math.min(1000 * Math.pow(2, attempt), 30000)
}
One second, two, four, up to a thirty-second ceiling. Standard, and worth saying out loud why it is standard: the failure mode a heartbeat introduces is a server restart handing you every client at once, each of them retrying. Back-off is what stops your own liveness checking from being a thundering herd.
There is a second branch, and it is the one most implementations miss. Once the client has already fallen back to polling, exponential back-off is wrong, because the attempt counter is high and the delay has reached its cap for reasons that are no longer relevant:
const degraded = isPollingRef.current || statusRef.current === 'error'
const delay = degraded ? WS_RETRY_WHILE_DEGRADED_MS : computeBackoff(attempt)
A flat thirty seconds. The polling fallback is working, so there is no urgency, but there is also no reason to let the retry interval drift towards hours. The socket gets a probe every half minute for as long as the session lasts, and the moment one succeeds the client switches back and stops polling.
Waking up: three events, one handler
A phone in a pocket is not a phone with a slow connection. The OS freezes the tab. Timers do not fire. The watchdog does not run. The socket may or may not have survived, and nothing in the page knows which.
document.addEventListener('visibilitychange', wake)
window.addEventListener('online', wake)
window.addEventListener('pageshow', wake)
Unlocked, network returned, restored from the back-forward cache. The handler makes one decision:
const fresh = Date.now() - lastMessageAtRef.current < WAKE_FRESH_MS
if (ws && ws.readyState === WebSocket.OPEN && fresh && statusRef.current === 'connected') {
// Alive, but events may have been missed while the page was frozen.
ws.send(JSON.stringify({ type: 'sync' }))
return
}
reconnectNow()
If the socket has heard from the server within the last twenty-five seconds, it survived the freeze and only the state is suspect, so ask for a fresh snapshot over the connection you already have. Otherwise throw it away and reconnect. The distinction matters at scale: a room of sixty phones unlocking as a question lands should produce sixty small sync requests, not sixty new WebSocket handshakes.
The fallback, and the jitter that makes it safe
After twenty seconds of failed reconnection the client stops waiting and starts polling for state over plain HTTP:
const POLL_INTERVAL_MS = 5_000
const POLL_JITTER_MS = 1_500
pollTimerRef.current = setTimeout(poll, POLL_INTERVAL_MS + Math.random() * POLL_JITTER_MS)
The jitter is not politeness. When a venue's access point drops and recovers, every phone in the building transitions to polling inside the same second, and without jitter they stay in lockstep for the rest of the night: a flat background load with a spike every five seconds, from one IP address, which is exactly the shape that trips a rate limiter. Spreading them over 1.5 seconds turns a spike into a floor.
That the polling fallback is rate limited per player rather than per IP is its own story, and the reason is the same one: under NAT, a whole pub is one address.
The ladder, bottom to top
| Layer | Detects | Within |
|---|---|---|
| Server ping and pong count | Dead client sockets | ~60 s |
| App-level heartbeat message | Nothing by itself; it is the traffic the client measures | 20 s cadence |
| Client watchdog | Silence on an apparently open socket | ~55 s |
| Wake handlers | A tab that was frozen | Immediately on unlock |
| Back-off reconnect | Transient drops | 1 s to 30 s |
| Polling fallback | A WebSocket that will not come back | 20 s after first failure |
Six mechanisms for one concern, which sounds like over-engineering until you notice that each catches a failure the others structurally cannot see. The server's ping cannot reach a frozen tab. The watchdog cannot run inside one. The wake handler cannot help a phone with no signal. And none of them can do anything about a captive portal that is cheerfully returning a login page for your WebSocket upgrade, which is what the polling fallback is for.
The product side of this
- pub-trivia.app/faq answers the bad-WiFi question in customer language. It is this post with the arithmetic removed, and it is the question venues actually ask first.
- pub-trivia.app/features/live-leaderboard is what all of the above protects: a leaderboard that moves when a question closes is only a feature if it keeps moving on a bad night.
- pub-trivia.app/solutions/pubs is the environment these numbers were chosen for, written for the person who owns the room rather than the socket.
- To watch the ladder work, start a free session, join on a phone, and put the phone in flight mode for a minute. Reconnect, snapshot, and the quiz carries on from where the room is, not from where the phone left off.
Top comments (0)