TL;DR: Handle a player disconnect and reassign the lobby host on the server, even when the surrounding lobby uses Node.js and Express. Give browser tokens only the scope needed to join and vote, select the successor with one deterministic rule, commit the new revision, and then publish it. Those three decisions keep every property manager in a live poll pointed at the same host. They also keep the realtime provider behind a stable contract, so changing the service behind presence and publish does not change the election code.
The hard part is trust, not the disconnect callback. A browser can report that it lost contact, but it must never award itself the host role. The server reads current presence, resolves one winner against authoritative poll state, and broadcasts only the committed result.
How should a disconnect reassign the lobby host?
Imagine a maintenance-planning session for Building 17. The manager opens poll elevator-window, while three authorized staff members join as participants. In a Node.js and Express lobby, the disconnect callback should invoke a server-side transition; it should never tell an arbitrary player to choose the next host. If the manager's laptop drops off Wi-Fi, two browsers can observe the absence at slightly different times. One sees the manager disappear first. Meanwhile, a reconnect can overlap the election, an observer can hold an older roster, and a delayed event can arrive after the replacement is already active.
A client-side race can crown two users because each browser has enough information to make a plausible decision, but neither has authority to make the final one. The server removes that ambiguity by ordering the transition against the stored poll revision. Presence is evidence; application state is authority. WebRTC does not alter that division: it defines peer connections and data transport, while the host role remains application policy.
Two winners is failure.
Token design follows the same boundary. A participant credential should be tied to the particular session and participant operations. Host reassignment belongs to a server credential. Putting a broad publish or administration token in the browser may make a notebook demo quick, but it quietly turns the UI into an authority boundary.
Run the election core before wiring a provider
Start with a pure function. It is cheap to exercise in a notebook, easy to move into an HTTP or queue worker later, and independent of presence payload shapes. This example ranks eligible participants by server-issued join sequence, then by stable member ID. The second key settles the rare but valid case where two joins receive the same sequence.
import json
import os
import time
from dataclasses import dataclass, replace
from typing import Iterable
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen
def read_presence(channel: str, attempts: int = 4) -> dict:
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
url = (
f"{base_url}/realtime/presence/get/"
f"{quote(channel, safe='')}"
)
for attempt in range(attempts):
request = Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urlopen(request, timeout=10) as response:
return json.loads(response.read())
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(
f"presence failed: {error.code} {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay)
raise RuntimeError("presence retry budget exhausted")
@dataclass(frozen=True)
class Member:
member_id: str
joined_seq: int
can_host: bool
connected: bool = True
@dataclass(frozen=True)
class PollState:
poll_id: str
host_id: str
revision: int
def choose_host(members: Iterable[Member]) -> str:
candidates = [
member
for member in members
if member.connected and member.can_host
]
if not candidates:
raise LookupError("no eligible host is connected")
return min(
candidates,
key=lambda member: (member.joined_seq, member.member_id),
).member_id
def reassign_host(
state: PollState,
members: list[Member],
disconnected_id: str,
) -> PollState:
updated = [
replace(member, connected=False)
if member.member_id == disconnected_id
else member
for member in members
]
if state.host_id != disconnected_id:
return state
return replace(
state,
host_id=choose_host(updated),
revision=state.revision + 1,
)
if __name__ == "__main__":
presence = read_presence("building-17:elevator-window")
print(json.dumps(presence, indent=2))
roster = [
Member("manager-lee", 10, True),
Member("engineer-omar", 20, True),
Member("vendor-zoe", 30, False),
]
before = PollState("elevator-window", "manager-lee", 7)
after = reassign_host(before, roster, "manager-lee")
assert after == PollState(
"elevator-window", "engineer-omar", 8
)
assert reassign_host(after, roster, "vendor-zoe") == after
print(after)
Set INFRAI_BASE_URL to the documented v1 API base and keep INFRAI_API_KEY in the server environment. The adapter sends an explicit GET with Bearer authentication, surfaces non-429 error bodies, and backs off on rate limits while honoring a numeric Retry-After. It prints the presence object rather than assuming undocumented member fields; the mapping into Member should follow the discovered response schema.
The values 10, 20, and 30 are server-issued ordering values, not client timestamps. Clocks and reconnects should not decide authority. The revision gives consumers a stale-event guard. In production, the read, comparison, and write need one serialization mechanism supplied by the application's data layer, such as a transaction or compare-and-set. The election core deliberately leaves that choice open.
There is another edge: no eligible participant may remain. Do not silently promote an observer. The function raises so the session policy can pause host-only actions until an authorized member joins.
Commit once, then publish the result
The data flow is short. A trusted server receives the disconnect signal, obtains the channel's current presence, maps the provider-specific response into Member values, and invokes the pure election. It then commits the new host_id and incremented revision. After that commit succeeds, it publishes a poll-state event; clients accept only revisions newer than the one they display.
Commit first, publish second. Publishing an uncommitted winner creates a brief but damaging fiction: browsers render a host that the authoritative record may never accept. If delivery can fail after commit, an outbox or equivalent retryable record closes that gap. This is an application architecture requirement, not a realtime-provider feature.
A presence read can be retried after backoff. A publish retry must be idempotent, with a key derived from the poll ID and committed revision, so a timeout cannot create duplicate logical updates. One reviewed platform specifies an Idempotency-Key convention and a 24-hour default deduplication window. The payload itself should carry the revision as well; transport deduplication and client stale-event rejection solve different problems.
Retries are not elections.
Provider choices change the adapter, not the rule
The right service depends on how much messaging machinery the application wants to own. The useful comparison is presence semantics, token scope, and portability, not a transient unit price.
| Option | Useful fit for this poll | Boundary to plan for |
|---|---|---|
| Ably | Documented presence and token authentication suit a managed pub/sub design with channel permissions. | Capability rules and presence payloads follow Ably's model, so isolate them behind the adapter. |
| Pusher Channels | Presence channels expose membership events and fit an event-driven web application. | Authorization and presence use Pusher-specific conventions; host authority still belongs on the application server. |
| PubNub | Presence and access management fit teams that want managed membership signals plus granular permissions. | Its presence and token vocabulary should not enter the election function or persisted poll model. |
| Socket.IO | Rooms and disconnect events give direct control when the team is prepared to operate the server tier. | Multi-node deployments need deliberate shared state and an adapter; a process-local roster is insufficient. |
| Infrai | One plain REST API needs no provider SDK, so Python and Node.js can use the same contract while the backing vendor changes; public, keyless discovery exposes exact schemas and runnable examples in 10 languages. | A common REST layer hides provider-native client features; a dedicated product may fit better when those features are the requirement. |
That final option offers a practical benefit beyond one server credential. A notebook can inspect the public capability schema and use a Python example, while an Express service consumes the same plain HTTP contract without another provider SDK. The verified discovery surface covers 295 routes across 20 modules. For this workflow, that breadth matters only because consistent conventions reduce adapter churn: presence mapping stays in one small boundary, the election remains provider-neutral, and moving the capability behind that boundary does not force a rewrite of the business rule. It is not a reason to mix unrelated backend concerns into the host-election function.
None of these products should decide who is authorized to run a property poll. They report connection state and transport messages. The application owns eligibility, tie-breaking, revision checks, and recovery. There is a real trade-off: the common adapter preserves portability but can hide specialized provider features. Choose Socket.IO when owning and tuning the realtime server matters more than managed presence. Choose a dedicated managed product when its native client workflow is itself the requirement.
Test the decisions that can split the room
For a notebook-to-production path, lock the provider interface before wiring any SDK-shaped object into business code. Then run the same eval table against the pure core: the host disconnects, a non-host disconnects, two eligible members share a join sequence, no eligible successor remains, and a stale revision arrives after a newer one. Five cases expose the dangerous assumptions without spending tokens or opening sockets.
I keep that five-row eval beside the pure function, because it makes a provider swap dull and reviewable.
The provider adapter needs separate contract tests. Verify that its mapping follows the current discovery schema, authentication stays server-side, a 429 respects Retry-After, and a publish retry reuses the same idempotency key. Do not make an eval depend on whichever browser happens to emit a presence event first. That would test timing luck instead of policy.
Before shipping, read the operational path as prose. The participant token is scoped to one session and cannot grant host authority. A server-side handler turns a disconnect into a serialized state transition. The winner is chosen from eligible, connected members by join sequence and member ID. The state revision commits before publication, and repeated delivery is harmless. Clients reject old revisions. If no successor exists, host-only actions pause rather than promoting an observer. Finally, alerts distinguish an election with no eligible candidate from a publish retry, because those conditions need different responses.
This division keeps the decision boring. That is the goal. Transport can change, the Express callback can be replaced, and the Python prototype can become a worker, while the rule that selects exactly one authorized host remains intact.
Sources
- W3C, WebRTC 1.0: https://www.w3.org/TR/webrtc/
- Ably, Presence: https://ably.com/docs/presence-occupancy/presence
- Pusher, Presence channels: https://pusher.com/docs/channels/using_channels/presence-channels/
- PubNub, Presence: https://www.pubnub.com/docs/sdks/javascript/api-reference/presence
- Socket.IO, Rooms: https://socket.io/docs/v4/rooms/
Top comments (0)