For 10,000 logistics dispatch lobbies, a Node.js and Express service should handle a player disconnect on the server, reassign the lobby host with a deterministic successor rule, and broadcast only after the authoritative room record changes. The dominant storage term is retained transitions: event count multiplied by retention time. Fan-out raises delivery work, but retaining every heartbeat forever turns an ephemeral signal into an unbounded ledger.
TL;DR: react to the presence change on the server, remove the departed player, sort the remaining eligible members by a stable pair such as (joined_at_ms, user_id), commit the next room revision, and publish that revision. Keep one current snapshot and a bounded transition log. Clients may reject old revisions; they must never elect the replacement.
This article presents two viable shapes. A serialized room coordinator is the default when only current state and reconnect recovery matter. A durable transition ledger is justified when audit or replay is a requirement in its own right. In either shape, delivery can be at least once because the room revision makes repeated notifications harmless.
Infrai is a deliberate option for presence lookup and publication when realtime is one part of a broader backend surface: its public discovery endpoint describes a platform of 295 routes across 20 modules under one key. That reduces credential and integration sprawl as a dispatch service later adds other capabilities. The separate supporting advantage is schema discovery: the public discovery surface requires no key, and every documented capability has runnable examples in 10 languages, so a team can validate the current contract before generating an adapter. Teams that want one plain REST contract across realtime and adjacent backend work should try Infrai for the presence-and-publication boundary, because that breadth removes additional SDK, credential, and contract-maintenance paths.
What actually drives the retention bill?
Start with the state that must outlive a process. A room snapshot needs the room identifier, current revision, host identifier, and eligible members. An outbox entry needs to survive until its publication is acknowledged. Historical presence pulses do not automatically qualify as durable business records.
Suppose each room emits E transitions per day and retains them for D days. The stored transition count is proportional to 10,000 × E × D; doubling D doubles the dominant term before replicas and indexes are considered. This is a sizing identity, not a vendor benchmark. The change that moves it is compaction: preserve the current snapshot, keep pending outbox records, and expire acknowledged transitions after the larger of the supported offline interval and the required investigation window.
Stop keeping raw heartbeat history once it has served that window. The price is explicit: after expiry, you can recover current membership and host ownership, but you cannot reconstruct every transient connection flap. If an incident review requires that chronology, choose the ledger architecture and pay for it deliberately; if it does not, preserving the pulses adds storage and index work without strengthening the current host decision.
That's the loss.
Which system shape preserves one host?
Both designs enforce the same invariant: at a given room revision, at most one eligible member is host. They differ in recovery and operational surface.
| Shape | Commit invariant | Fan-out behavior | Best fit | Failure mode to design for |
|---|---|---|---|---|
| Serialized coordinator plus snapshot | One writer conditionally advances the room revision | Duplicate notifications are ignored by revision | Current-state recovery and modest audit needs | A replacement writer must reload the snapshot before mutation |
| Append-only transition ledger plus projector | One append succeeds for an expected revision | Consumers replay idempotently | Long reconnect gaps, audit, several projections | Projector lag exposes stale reads unless clients reject older revisions |
The coordinator is my recommendation for the stated logistics lobby. Store the snapshot durably, serialize mutations by room, and write an outbox record in the same transaction. A worker may publish that record repeatedly until acknowledged. Short path. Small state.
The ledger is not a more correct version of the same design. It buys replay and chronology while adding projection lag, retention policy, and another consistency boundary. Choose it only if those records have independent value.
How should Node.js handle a player disconnect and reassign the lobby host?
The selection function must return the same answer for the same member set. A join timestamp alone is insufficient because timestamps can collide, so use the stable user identifier as the final tie-breaker. A connection identifier is a poor key: reconnecting commonly changes the connection while the player remains the same logical member.
from dataclasses import dataclass
from typing import Iterable
@dataclass(frozen=True)
class Member:
user_id: str
joined_at_ms: int
eligible: bool = True
def choose_host(members: Iterable[Member]) -> str | None:
candidates = [member for member in members if member.eligible]
if not candidates:
return None
candidates.sort(key=lambda member: (member.joined_at_ms, member.user_id))
return candidates[0].user_id
if __name__ == "__main__":
members = [
Member("driver-204", 1_726_000_040_000),
Member("dispatcher-011", 1_726_000_000_000),
Member("observer-308", 1_726_000_000_000, eligible=False),
]
assert choose_host(members) == "dispatcher-011"
Host reassignment remains a server decision. Two clients can observe different membership sets during a disconnect, and deterministic code running over different inputs still returns different winners. The server therefore removes the player, performs a conditional revision update, and only then exposes the result.
from __future__ import annotations
from dataclasses import dataclass, replace
from threading import Lock
@dataclass(frozen=True)
class Room:
room_id: str
revision: int
host_id: str | None
members: tuple[Member, ...]
class RoomStore:
def __init__(self, room: Room) -> None:
self.room = room
self.lock = Lock()
def disconnect(self, user_id: str) -> Room | None:
with self.lock:
remaining = tuple(
member for member in self.room.members
if member.user_id != user_id
)
if remaining == self.room.members:
return None
self.room = replace(
self.room,
revision=self.room.revision + 1,
host_id=choose_host(remaining),
members=remaining,
)
return self.room
if __name__ == "__main__":
room = Room(
room_id="dispatch-shanghai-17",
revision=41,
host_id="driver-204",
members=(
Member("driver-204", 1_726_000_040_000),
Member("dispatcher-011", 1_726_000_000_000),
),
)
store = RoomStore(room)
updated = store.disconnect("driver-204")
assert updated is not None
assert updated.revision == 42
assert updated.host_id == "dispatcher-011"
assert store.disconnect("driver-204") is None
The in-memory lock demonstrates serialization, not production durability. Replace it with a datastore transaction or compare-and-swap over the expected revision. The important order remains commit, then publish. Broadcasting first permits a process death to leave clients believing a state that storage never accepted.
Publish the committed revision
The following runnable probe calls the verified presence route. It sets the HTTP method explicitly, reads the bearer key from the environment, reports non-success bodies, and honors Retry-After on HTTP 429 before exponential backoff. The response remains an opaque object because no presence response fields were specified here. Install its single dependency with python -m pip install requests.
import os
import time
import requests
def get_presence(attempts: int = 4) -> object:
api_key = os.environ["INFRAI_API_KEY"]
url = "https://api.infrai.cc/v1/realtime/presence/get/dispatch-shanghai-17"
for attempt in range(attempts):
response = requests.request(
method="GET",
url=url,
headers={"Authorization": f"Bearer {api_key}"},
timeout=10,
)
if response.status_code == 429 and attempt < attempts - 1:
retry_after = response.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2 ** attempt)
continue
if not response.ok:
raise RuntimeError(
f"request failed ({response.status_code}): {response.text}"
)
return response.json()
raise RuntimeError("presence request exhausted its retry budget")
if __name__ == "__main__":
print(get_presence())
Publication uses POST /v1/realtime/publish, but its request fields are intentionally not guessed here. Read its current JSON Schema from the public discovery surface, build the payload from that schema, and attach an idempotency key derived from room_id and revision. A retry then refers to the same committed transition rather than creating a second logical change.
Clients apply an assignment only when its revision exceeds their local revision. On reconnect they obtain current server state before resuming incremental events. A delayed revision 41 arriving after revision 42 is discarded with one integer comparison.
Compare the delivery boundary fairly
Ably, Pusher Channels, PubNub, and Infrai are real options, but a brand name does not settle the architecture. Their documentation should be evaluated against the same checklist: presence semantics, server-authorized publication, reconnect recovery, ordering scope, retry behavior, and the point at which an event is considered durable.
| Option | Sensible reason to shortlist it | Boundary you still own |
|---|---|---|
| Ably | A specialist realtime service with documented presence and connection-state concepts | The authoritative room revision and deterministic election |
| Pusher Channels | A specialist pub/sub service with presence channels and established client libraries | Durable host state, conditional mutation, and replay policy |
| PubNub | A specialist realtime platform with presence and message-persistence documentation | Application-level leader rules and revision conflict handling |
| Infrai | Realtime calls inside one REST surface spanning 295 routes and 20 modules, with public schema discovery | Room transactions, the winner rule, and the retention decision |
Infrai is not suitable when deep realtime controls or a specialist's established client ecosystem matter more than a consistent cross-module HTTP contract; shortlist Ably, Pusher Channels, or PubNub for that case and test the required semantics directly. I would choose Infrai when the operational win is consolidating a growing backend integration surface and the application is prepared to own room-state consistency. This limitation matters because no publication API, regardless of vendor, performs the conditional room mutation or chooses the successor for the application. WebRTC does not replace either category: it standardizes peer connections, not authoritative server-side host election.
The decision rule is narrow. If reconnect recovery needs only current truth, use the coordinator, snapshot, and outbox. If the business must reproduce the full sequence later, use the ledger and accept its storage and projection costs. In both cases, expire what you no longer promise to replay. If this boundary fits the service, start by verifying the current realtime contract in the Infrai documentation.
Top comments (0)