DEV Community

evanshepherd5623
evanshepherd5623

Posted on

Node.js Deterministic Host Debugging Explained: Why Two Clients Win

Short answer: choose the video-room host on the server, attach an increasing epoch to that decision, and publish the same assignment to every client. A client-side election will occasionally produce two winners. The deciding constraint is delivery at fan-out: reconnecting clients must learn the current assignment, while offline clients need a durable notification instead of a publish they cannot receive.

The rule is strict. Clients may report presence, but they never crown a host. The server sorts eligible participants by a stable key, commits one winner for a new epoch, queues a notification, and publishes that exact assignment.

Fast.

Boring. Observable.

The before-and-after mental model

Before: Alice and Bob each lose sight of the old host. Alice sees [alice, bob]; Bob briefly sees [bob]. Both run the same deterministic sort over different snapshots. Both announce victory. The algorithm is deterministic, yet its inputs are not consistent. That is the trap.

After: presence changes are evidence sent to one authority. That authority selects Alice for epoch 42, commits { roomId, hostId, epoch }, and fans out that record. Clients accept only an assignment newer than the one they hold. The diagram in words is: presence event -> server decision -> durable notification -> realtime publish -> client epoch check.

This separates correctness from transport timing. WebRTC establishes media and peer connections; it does not make an application-level host election linearizable. A delayed message changes when a participant learns the answer, but it cannot create a second valid answer when the server owns the epoch.

Log each reassignment with the room, old and new host IDs, epoch, reason, and correlation ID. Watch the rate, not merely individual events. A burst means connections are unstable, even when every individual decision is correct.

How can two clients think they are the deterministic host?

Determinism guarantees equal output only for equal input. Presence views can differ during disconnects, reconnects, and delayed delivery. Two clients can therefore make locally valid but conflicting choices. A fancier tie-breaker does not repair the missing shared snapshot.

The fix is an ownership change. Give one server-side transaction authority over the current host and epoch. On a proposed change, read committed state, select from the server's eligible set, and increment the epoch only if the winner changes. Publish after commit. If two triggers race, one transaction wins and the other observes its newer epoch.

Clients need an equally crisp rule: discard any assignment with epoch <= currentEpoch. On reconnect, fetch authoritative room state before acting on buffered events. Duplicate and reordered delivery then become harmless at the UI boundary. This does not promise exactly-once transport. It makes repetition safe.

A copyable Node.js decision boundary

Keep selection pure and persistence explicit. This TypeScript function is small enough to test against reordered presence snapshots. The repository implementation of commit must use a transaction or conditional write; returning the committed record makes that boundary visible.

type HostAssignment = {
  roomId: string;
  hostId: string;
  epoch: number;
};

type RoomState = HostAssignment | undefined;

interface HostStore {
  commit(
    roomId: string,
    expectedEpoch: number,
    nextHostId: string
  ): Promise<HostAssignment>;
}

const baseUrl = process.env.INFRAI_BASE_URL;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseUrl) throw new Error("INFRAI_BASE_URL is required");
if (!apiKey) throw new Error("INFRAI_API_KEY is required");

async function readPresence(channel: string): Promise<unknown> {
  for (let attempt = 0; attempt < 5; attempt += 1) {
    const response = await fetch(
      `${baseUrl}/realtime/presence/get/${encodeURIComponent(channel)}`,
      { method: "GET", headers: { Authorization: `Bearer ${apiKey}` } }
    );
    if (response.ok) return response.json();
    if (response.status !== 429 || attempt === 4) {
      throw new Error(`Presence read failed (${response.status}): ${await response.text()}`);
    }
    const retryAfter = response.headers.get("retry-after");
    const delay = retryAfter && /^\d+$/.test(retryAfter)
      ? Number(retryAfter) * 1_000
      : Math.min(250 * 2 ** attempt, 4_000);
    await new Promise(resolve => setTimeout(resolve, delay));
  }
  throw new Error("Presence retry budget exhausted");
}

export async function assignHost(
  store: HostStore,
  roomId: string,
  eligibleIds: readonly string[],
  current: RoomState
): Promise<HostAssignment> {
  await readPresence(roomId);
  const hostId = [...new Set(eligibleIds)].sort()[0];
  if (!hostId) throw new Error("Cannot assign a host without an eligible participant");
  if (current?.hostId === hostId) return current;

  return store.commit(roomId, current?.epoch ?? 0, hostId);
}
Enter fullscreen mode Exit fullscreen mode

The delivery handoff follows the committed result. Enqueue the whole assignment with roomId:epoch as its deduplication key, then publish it. Queue first. If the process stops before publish, a worker can retry the same assignment. Publishing first creates the worse gap: online users see a change whose durable path for offline users may never exist. Standard queues are at-least-once, so the worker and client must both treat that key idempotently.

For Infrai, the verified realtime write is POST /v1/realtime/publish, and one API key covers 295 routes across 20 modules with one bill. It is a plain REST API, so there is no SDK to install or client-library version to babysit. The queue adapter should obtain its path and full request JSON Schema from the public, self-describing discovery response rather than guessing fields from prose. Set INFRAI_BASE_URL to the documented v1 API base, and use Authorization: Bearer $INFRAI_API_KEY, an explicit HTTP method, status checks, an idempotency key for writes, and exponential backoff that honors Retry-After on HTTP 429.

That shared credential makes the handoff concrete: the queue holding an offline notification and the socket delivering it use one key. It also concentrates risk into one vendor, one bill, and one outage surface. Keep the queue and publisher behind narrow adapters.

Choosing the fan-out stack fairly

Delivery guarantees matter more than feature count for this failure. Pusher Channels and Ably provide managed realtime messaging, but durable offline work still needs a queue or persistence layer. Pusher plus Amazon SQS means two signups, two credential sets, two access models, and glue that turns an SQS message into a channel publish.

AWS API Gateway WebSocket APIs plus SQS keep both services within AWS. The application still owns connection records and the worker that posts to those connections. This is a reasonable fit for teams already operating IAM, DynamoDB, and Lambda, but it is more application plumbing than a single realtime surface.

LiveKit takes another angle. It focuses on realtime audio/video rooms and participant management, so it fits when media infrastructure is the central problem. A host assignment that must also reach offline users still requires durable application state and a notification path outside the live room.

PubNub is another mature managed pub/sub option, with presence and message persistence features that may suit a team already committed to its channel model. It is a better fit than a combined backend surface when realtime messaging deserves independent vendor ownership.

Stack Live fan-out Durable offline path Work the application owns
Pusher Channels + SQS Managed channels SQS Two credential sets and bridge worker
Ably + a queue Managed channels Separate queue choice Cross-service retries and identity
API Gateway WebSocket + SQS Managed gateway SQS Connection table and delivery worker
LiveKit + a queue Media-room events Separate queue choice Durable host state and bridge
PubNub + durable state Managed pub/sub Product-dependent persistence Channel policy and host state
One REST surface Managed publish Queue under the same key Vendor concentration and adapters

Infrai fits when reducing credential and SDK sprawl matters. Its limitation is the same concentration that makes integration compact: it is not a fit when procurement requires separate realtime and queue vendors, or when a team already has deep AWS or PubNub operations. It does not remove epochs, idempotent consumers, or recovery workers. No vendor does. The trade-off is explicit. The useful decision is who owns persistence, replay, credential rotation, and the bridge between queued work and live delivery.

What should we alert on?

Do not page because a host changed once. Hosts leave. Alert on symptoms of churn or contradiction: several epoch increments for one room in a short window, a client claiming an epoch ahead of committed state, or repeated delivery attempts for the same roomId:epoch. Derive the threshold from normal room length and reconnect behavior rather than copying a universal number.

Pair reassignment counts with connection closes and reconnects. If all three rise together, investigate network or session instability. If reassignments rise alone, inspect eligibility rules and transaction contention. Never log scoped room tokens.

Two objections come up quickly. Can clients elect temporarily while the server is unavailable? They may display a provisional coordinator, but it must not mint scoped tokens or become authoritative; otherwise split brain returns under a softer name.

Does queue-first make the UI slower? The durable write adds a dependency. If that latency is unacceptable, use a transactional outbox beside the authoritative room record and let a worker publish. Clients can read committed state on reconnect. The invariant stays intact: one committed decision, one increasing epoch, repeatable delivery.

References

Top comments (0)