DEV Community

Daniel Pertu
Daniel Pertu

Posted on

Sixty phones reconnect in the same second, and they all share one snapshot build

The hard part of a real-time connection is not the first one. It is the two hundredth, when the venue's access point has just come back and every phone in the building is handshaking at once, in a quiz that is currently mid-question and cannot wait for them.

Two decisions carry that load. One is about what you send a client that reconnects. The other is about what it costs you when they all reconnect together.

Send state, not history

The tempting model for a reconnecting client is a replay: here is what you missed. It is tempting because it is what the client appears to be asking for, and because a message log feels like the honest representation of a real-time system.

It is wrong for anything with a clock. A phone that missed questions four and five does not want them. The room is on question six and nothing about four and five is actionable: the windows are closed, the answers are in, the standings already reflect them. A replay would walk the player forward through two dead questions while the live one ticks down.

So the handshake ends with a snapshot:

const meta: ClientMeta = { sessionId, role, participantId: player };
addClient(ws, meta);
subscribeSession(sessionId);
sendToClient(ws, { type: 'auth_ok' });
sendToClient(ws, { type: 'state_snapshot', payload: snapshot });
Enter fullscreen mode Exit fullscreen mode

The whole of now: session status, the current question and when it launched, whether it has closed, the standings, and, for an identified player, their own answer to the question in progress. The client renders that and is caught up. There is no catch-up state, no buffer to drain, no ordering question about whether a replayed event should overwrite a live one that arrived first.

It is also the thing that makes everything else in the system allowed to be lossy. A dropped broadcast, a terminated slow socket, a deploy that closes every connection: all of them are recoverable because reconnection is not a resumption, it is a refresh. You get to be cheerful about losing messages once losing a message costs a client nothing but a round trip.

The same shape is what the polling fallback asks for over HTTP when the WebSocket cannot be re-established at all. One snapshot format, two transports.

The reconnect storm

A snapshot is not free. It is a session row, the current question, the answer progress, the standings with participant names, and on top of that the connecting player's own answer. Call it six queries.

One client, six queries, nobody cares. Sixty clients in the same second is three hundred and sixty queries in a burst, against a database that is at that moment also serving the live question. And the sixty results would be identical, because they describe the same session at the same instant.

The fix is a memo on the promise rather than on the result:

const SNAPSHOT_TTL_MS = 750;
const sharedSnapshots = new Map<string, { at: number; promise: Promise<SharedSnapshot | null> }>();

function sharedSnapshot(sessionId: string): Promise<SharedSnapshot | null> {
  const now = Date.now();
  const cached = sharedSnapshots.get(sessionId);
  if (cached && now - cached.at < SNAPSHOT_TTL_MS) return cached.promise;
  const promise = buildSharedSnapshot(sessionId);
  sharedSnapshots.set(sessionId, { at: now, promise });
  promise.catch(() => sharedSnapshots.delete(sessionId));
  setTimeout(() => {
    if (sharedSnapshots.get(sessionId)?.promise === promise) sharedSnapshots.delete(sessionId);
  }, SNAPSHOT_TTL_MS).unref();
  return promise;
}
Enter fullscreen mode Exit fullscreen mode

Caching the promise rather than the value is the whole trick, and it is worth being precise about why. If you cache results, the sixty clients that arrive while the first build is in flight all miss the cache and start their own build: you have cached nothing at the only moment it mattered. Caching the in-flight promise means the second through sixtieth arrivals await the first one's work. Three hundred and sixty queries become six.

Three details that are not decoration:

promise.catch(() => delete) so a failed build is not served to everyone for 750 ms. Without it, one transient database error becomes sixty failed connections.

The setTimeout eviction, because otherwise the map accumulates an entry for every session the process has ever seen, and a long-lived server with a thousand quizzes a week is a slow leak with a timestamp attached. unref() so that timer never keeps the process alive during shutdown.

And the TTL itself. 750 ms is chosen against the thing the snapshot describes: a question lasts seconds, so three quarters of a second of staleness is tolerable, and it comfortably covers the burst of a venue's WiFi recovering. It would be the wrong number for a slower domain in either direction.

It is a floor, not a ceiling. Anything that changes the game invalidates it immediately:

/** Forget the memo, so the next snapshot reflects a change that just happened. */
export function invalidateSnapshot(sessionId: string): void {
  sharedSnapshots.delete(sessionId);
}
Enter fullscreen mode Exit fullscreen mode

Called on every launch, close and join, before the broadcast goes out. A client connecting one millisecond after a question opened gets the new question, not the previous one with 749 ms left on the clock.

The per-client part stays per-client

export async function buildSnapshot(sessionId, participantId) {
  const shared = await sharedSnapshot(sessionId);
  if (!shared) return null;
  const { questionId, ...snapshot } = shared;

  let you: PlayerQuestionState | null = null;
  if (participantId && questionId) {
    const answer = await fetchPlayerAnswer(sessionId, questionId, participantId);
    you = toPlayerState(snapshot.currentQuestionIndex, answer, snapshot.questionClosed);
  }
  return { ...snapshot, you };
}
Enter fullscreen mode Exit fullscreen mode

Shared work memoised, personal work not. Sixty reconnects cost one shared build and sixty single-row lookups, rather than sixty of everything or a cache keyed per participant that would never be hit twice.

Note also that questionId is destructured off before the snapshot is returned. It is needed to look up the player's answer and is not part of the client's view of the world, so it leaves at the boundary.

The handshake rules that make the above safe

The snapshot is built during authentication, which means the handshake is doing real work and has to be strict about it.

/** Clients that don't auth within this window are disconnected. */
const AUTH_TIMEOUT_MS = 5_000;
Enter fullscreen mode Exit fullscreen mode

A socket that connects and says nothing is dropped after five seconds. Without a timeout, anything that can open a TCP connection can hold a slot indefinitely for free.

Session and participant ids are checked against a UUID pattern before they reach the database:

// The client may send a placeholder while no real session exists yet.
// Reject it here instead of letting Postgres throw on an invalid uuid.
if (typeof sessionId !== 'string' || !UUID_RE.test(sessionId)) { ... }
Enter fullscreen mode Exit fullscreen mode

The frames themselves are tiny, and the server says so:

// Clients only ever send the auth message and the occasional sync request. The
// ws default accepts frames up to 100 MB.
const wss = new WebSocketServer({ noServer: true, maxPayload: 4 * 1024 });
Enter fullscreen mode Exit fullscreen mode

Four kilobytes. Our clients send an auth object and the word sync; a hundred-megabyte default is a hundred megabytes of attack surface for no feature.

And the upgrade is refused before any of it if the origin is not ours:

if (!allowedOrigins.includes(origin)) {
  socket.write('HTTP/1.1 403 Forbidden\r\n\r\n');
  socket.destroy();
  return;
}
Enter fullscreen mode Exit fullscreen mode

One more case the snapshot handles that an error would not. A phone that reconnects just after the last question, in the pause before the final leaderboard, is not told the session is gone:

// Hand over the final standings rather than an error. A phone that
// reconnects in the pause before the final leaderboard (or just after
// it) would otherwise be left with an empty results screen.
sendToClient(ws, { type: 'session_ended', payload: { leaderboard: snapshot.leaderboard } });
Enter fullscreen mode Exit fullscreen mode

The session is over, so the socket closes, but the last thing down it is the result. The player sees who won. This is the sort of case a replay model cannot serve at all, because there is nothing left to replay.

Where to look

  • pub-trivia.app/faq has the plain-English version: a phone that drops out is sent the current state of the game rather than being made to start again. That sentence is this entire post.
  • pub-trivia.app/features/qr-code-quiz-joining is the join flow whose first and worst burst is a whole room scanning at once, which is the same shape of load as the reconnect storm.
  • pub-trivia.app/features/live-leaderboard is the heaviest thing in the snapshot, and the reason six queries per client was worth memoising.
  • Start a free session and reload a player's phone mid-question. The question reappears with the right time remaining, and the answer already submitted is still there.

Top comments (0)