Use a pushed presence stream for immediate room changes, but make every reconnect query an authoritative snapshot and make the live metrics dashboard query aggregated metrics on its own cadence. The deciding constraint is accuracy after a broken connection: a channel can tell a client what changed, but it cannot prove that the client observed every change while it was away.
Short answer: separate the two jobs. Presence is state reconstruction with low-latency hints; dashboard metrics are sampled observations. Trying to carry both over one channel couples correctness to connection continuity, the least dependable part of a browser session.
This architecture decision record uses four invariants: one active session lease per player connection, monotonic room revisions, snapshot-before-delta recovery, and metrics that never participate in presence correctness. The examples use Python and Postgres, but the boundaries are generic.
What must remain true after a reconnect?
A gaming chat room can tolerate a counter being a few seconds old. It cannot casually claim that a player is present because a disconnect event vanished between a load balancer and a process, or omit a player because the browser missed an arrival while its radio changed networks. Presence accuracy starts with a less glamorous definition: presence is a leased, versioned server-side fact, not a side effect of an open socket.
Be strict here.
A session has an expiry time and must be renewed. Every committed membership change advances a room revision. A reconnecting client installs a snapshot and its revision before accepting later deltas. Telemetry reads may lag or fail without changing membership state.
WebSocket supplies a two-way browser/server channel, while WebRTC data channels can carry peer-to-peer arbitrary data; neither transport specification turns transient delivery into an authoritative durable membership record. The system establishes that record above the transport layer. The W3C WebRTC specification exposes connection state transitions rather than promising that application state survived a disconnected interval.
The concrete failure boundaries are missed deltas, duplicate or late retries, a server disappearing before it emits a leave event, a lease expiring while an old connection still looks healthy, and a dashboard displaying a stale sample. The last failure is different. An active_sessions metric is derived from leases at a point in time; it is not the lease table itself.
Record the decision before choosing the transport
The useful comparison is not push versus pull in the abstract. It is which path owns truth, how gaps are repaired, and what load shape the system accepts.
| Option | Freshness model | Reconnect behavior | Primary failure mode | Valid use |
|---|---|---|---|---|
| Push deltas only | Immediate on a healthy channel | Completeness needs retained history and a cursor | A missed delta leaves silent drift | Typing indicators |
| Query snapshots only | Bounded by polling interval | Next query replaces local state | Aligned clients create read bursts | Small, delay-tolerant rooms |
| Push plus snapshot recovery | Immediate hints, authoritative repair | Snapshot establishes revision | A bad handoff creates a gap | Accurate player presence |
| Query aggregated metrics | Deliberately sampled | Failed reads retry independently | Stale display or excess reads | Live dashboards |
Choose the third row for room presence and the fourth for the dashboard. This is a split decision. Each path gets semantics appropriate to its job.
It costs something.
The hybrid design's main limitation is operational complexity: it requires a lease reaper, revision storage, snapshot reads, channel infrastructure, and client recovery logic, while query-only presence can run behind one ordinary API and is easier to inspect with familiar request logs. A small turn-based game whose roster may be five seconds stale should probably choose query-only snapshots and avoid that machinery. At the other extreme, a room with a replayable, durably retained event log may let reconnecting clients resume from a cursor instead of fetching a full snapshot, provided retention exceeds the supported offline window and the service can prove a cursor has not fallen behind. The trade-off is concrete: snapshot recovery bounds client reasoning and database reads by room size, whereas replay bounds them by missed-event volume and retention. Neither model wins without those workload limits.
The design needs explicit limits: delta retention, snapshot size, lease duration, reconnect buffer size, and dashboard freshness. Reject a delta whose revision does not follow the client revision. Those limits determine behavior under pressure; a vague promise of real time does not.
Postgres can serialize a room revision update and membership mutation in one transaction. It also gives a snapshot query a defined database view at the selected transaction isolation level. Do not infer more: replication topology, failover policy, and durability settings still define when a committed write is visible elsewhere and what loss envelope operators accept.
Implement the critical path in Python
The following boundary avoids a framework-specific server. store is an asynchronous repository whose operations use database transactions; channel is only the notification path. Thirty seconds is an example policy value, not a universal limit.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from typing import Protocol
LEASE_TTL = timedelta(seconds=30)
@dataclass(frozen=True)
class PresenceDelta:
room_id: str
revision: int
player_id: str
state: str
class PresenceStore(Protocol):
async def renew_session(
self, room_id: str, player_id: str, session_id: str,
expires_at: datetime
) -> PresenceDelta | None: ...
class EventChannel(Protocol):
async def publish(self, delta: PresenceDelta) -> None: ...
A renewal advances the revision only when visible membership changes. Repeated heartbeats extend the lease without manufacturing room churn. Publication happens after commit; if publication fails, a later snapshot repairs the client because the database, not the channel, owns truth.
async def heartbeat(store, channel, room_id, player_id, session_id):
delta = await store.renew_session(
room_id=room_id,
player_id=player_id,
session_id=session_id,
expires_at=datetime.now(timezone.utc) + LEASE_TTL,
)
if delta is not None:
await channel.publish(delta)
There is an uncomfortable edge in that function: commit and publication are two operations. If the process stops between them, connected clients miss a hint. Snapshot recovery preserves eventual correctness, but clients remain temporarily stale. A system needing dependable delta delivery can write an outbox record in the membership transaction and publish it asynchronously; consumers still need idempotence because retries can duplicate delivery.
Reconnect order matters more than the channel API. Subscribe into a bounded buffer, read and install the snapshot, then apply buffered deltas newer than the snapshot revision. On any revision gap, discard the partial sequence and fetch another snapshot.
async def recover_room(room_id, store, channel, view):
subscription = await channel.subscribe(room_id, buffer_limit=256)
snapshot = await store.read_room_snapshot(room_id)
view.replace(snapshot.player_ids, snapshot.revision)
async for delta in subscription:
if delta.revision <= view.revision:
continue
if delta.revision != view.revision + 1:
snapshot = await store.read_room_snapshot(room_id)
view.replace(snapshot.player_ids, snapshot.revision)
continue
view.apply(delta)
The 256 limit is visible for a reason. A full buffer is a correctness signal, not permission to drop the oldest event and continue. Production code also needs cancellation, authorization on subscription and snapshot reads, and a maximum snapshot response size appropriate to room limits.
Let the dashboard query its own metrics API
Presence events look like convenient dashboard input. They are the wrong contract. A tab opened after ten minutes has no starting state; a suspended tab can miss traffic; replaying membership events merely to render a chart forces an observer to reconstruct data the server can aggregate consistently.
Give metrics a query interface with an observation timestamp and freshness budget. Poll with jitter so clients do not align on one boundary, cancel obsolete requests, and retain the last successful value with a stale marker when a request fails.
from dataclasses import dataclass
from datetime import datetime
@dataclass(frozen=True)
class RoomMetrics:
room_id: str
active_sessions: int
observed_at: datetime
source_revision: int
async def read_dashboard_metrics(store, room_id):
row = await store.aggregate_active_sessions(room_id)
return RoomMetrics(room_id, row.count, row.observed_at, row.revision)
source_revision relates a sample to presence state without pretending telemetry is authoritative. Cache duration and query interval should follow the dashboard freshness objective and measured database capacity. There is no defensible universal polling number.
Observe both paths. For presence, track lease expiry, revision-gap resnapshots, outbox age, snapshot latency, and snapshot size. For metrics, track response age, query latency, failures, and concurrency. Keep labels bounded; raw player or room identifiers in time-series labels can create unbounded cardinality.
During a schema change, old and new application versions must agree on revision and lease semantics. Add fields compatibly and test rolling restarts while clients reconnect. A clean steady-state test proves very little.
Test the failure boundaries, not the happy animation
Pause publication after database commit. Duplicate a delta, reorder two deliveries, overflow the buffer, expire a lease without a disconnect signal, and restart a server during reconnect. Every test has one durable assertion: after recovery, the client player set and revision equal an authoritative snapshot.
Then stop metric reads while presence renewals continue; membership must remain correct. Slow the aggregate query; channel delivery must not wait for it. Resume reads and verify that observed_at advances and the stale marker clears.
Use a controllable clock for lease tests. Waiting 30 real seconds makes tests slow and still leaves boundary behavior nondeterministic. The expiry comparison, transaction ordering, and database clock policy must be explicit.
Why reject one channel for everything?
A single pushed stream for presence and metrics assigns durable reconstruction, transient UI updates, and sampled telemetry the same retention and delivery contract. The result is too weak for presence or unnecessarily expensive for metrics. It also lets a metrics consumer pressure the correctness path during a burst.
The rejected option has a valid use case. If a displayed value is transient, needs no missed-history reconstruction, and can truthfully disappear on reconnect, pushing it over an existing channel is reasonable. Typing indicators and short-lived animation cues fit that boundary. Accurate player presence does not.
The decision is narrow and testable: persist leased membership with monotonic revisions, push committed changes as hints, repair every uncertain sequence from a snapshot, and query aggregated dashboard metrics independently. Revisit it when freshness requirements, room-size bounds, retention, or database capacity changes; transport fashion is not a reason.
Top comments (0)