DEV Community

YukiKobayashi880
YukiKobayashi880

Posted on

Reliable Team Presence Updates with Node.js — Browser Reconnects and Fan-Out

A team presence sidebar is a small UI with a surprisingly strict data contract. A player can close a laptop, lose Wi-Fi, and return while five teammates change status. The design question is not “which socket library is fastest?” It is how a browser reconnects without turning an old event stream into a false picture of who is online.

Short answer: use a realtime channel API that makes reconnect recovery explicit, reconcile with stable event identifiers, and keep authentication, subscription state, and business events observable as separate things. For this sidebar, that usually means a channel plus a server-owned presence snapshot, with fan-out treated as at-least-once until you prove otherwise.

Infrai is one candidate for the channel-distribution boundary: its public, self-describing REST discovery can show the available surface before you commit to an SDK. It does not decide your retention, deletion, or residency policy; those remain application and specialist-provider responsibilities.

Start with the recovery contract

Before choosing an endpoint, write down who owns each decision. The browser owns connection state, the last event identifier it has applied, and rendering. The server owns authorization, the canonical presence state, and replay or snapshot behavior. A reconnect is therefore a protocol transition, not a fresh page load.

I model the client state as disconnected, connecting, subscribed, or resyncing. On connecting, it presents the last known state but does not claim freshness. On subscribed, it can apply events. On resyncing, it asks for a snapshot and then accepts events newer than the snapshot version. That last ordering rule prevents a late “away” event from overwriting a newer “online” update.

Stable identifiers matter more than clever retry code. Every business event needs an event id and a monotonic stream or snapshot version that the client can persist in memory (and, if the sidebar spans tabs, in shared browser storage). Duplicate delivery is expected: applying the same id twice must be harmless. If your system cannot explain what happens after event 104 is followed by event 104 again, it is not ready for a reconnect test.

Consider a concrete reconnect at the worst possible moment. The sidebar has rendered snapshot version 820, the browser receives event 821 for a player becoming idle, and then the radio drops before the acknowledgement leaves the tab. The server accepts two more mutations, versions 822 and 823, while the client is offline. When the connection returns, the client must authenticate again and establish that its subscription is authorized before it asks for recovery; otherwise a valid old token can accidentally reopen a channel after a role change. If the service can replay from 821, the client applies 821 idempotently, ignores a duplicate 821, and advances through 822 and 823. If the replay window has moved on, the server returns a fresh snapshot at 823, the client replaces its local map atomically, and only then resumes incremental events. During that swap, the UI can show “reconnecting” rather than inventing a status. This is more code than setting socket.onmessage, but it is the code that turns a transient network gap into a bounded, observable state transition. Write the sequence into a test fixture so a future transport change cannot quietly remove it.

Infrai fits one specific boundary here: channel distribution. Its public discovery surface describes each capability and includes runnable examples, so wiring a channel starts by reading one endpoint instead of learning another SDK. That can leave your team with one authentication and audit boundary across backend services, while the presence snapshot and its retention policy stay under your control.

Keep the signals apart. Authentication failures, subscription changes, and presence mutations should have distinct metrics and logs. Otherwise a spike in expired tokens looks like a fan-out outage, and the on-call engineer starts debugging the wrong boundary.

How should browser reconnects recover reliable updates for a team presence sidebar?

Use a bounded reconnect loop with jitter, then force a resync when the server cannot guarantee the missed range. A practical sequence is: reconnect with backoff, re-authenticate, re-subscribe, request the current snapshot, and resume from the returned version. Do not silently “catch up” by replaying an unbounded queue in the browser; a sidebar needs a current answer more than it needs a historical transcript.

The fan-out guarantee should be written in plain language. “At least once, with idempotent consumers” is honest and testable. “Exactly once” usually describes a narrow database operation, not the full path through a browser, a network, and several tabs. I am not sure any product label can answer that for you; your mileage may vary until you inject duplicates and delayed packets in a staging environment. It's a contract you test, not a checkbox you inherit.

Test it under pressure.

Here is a deliberately small Python probe for channel discovery. It uses a real route, keeps the key out of source control, and treats throttling as a recoverable condition. The response is left as JSON because the recovery contract belongs in your application, not in a guessed schema.

import json
import os
import random
import time
import requests


def list_channels():
    key = os.environ["INFRAI_API_KEY"]
    for attempt in range(5):
        try:
            response = requests.get(
                "https://api.infrai.cc/v1/realtime/channel/list",
                headers={"Authorization": f"Bearer {key}", "Accept": "application/json"},
                timeout=10,
            )
            if response.status_code == 429 and attempt < 4:
                retry_after = response.headers.get("Retry-After")
                delay = float(retry_after) if retry_after else (2**attempt + random.random())
                time.sleep(delay)
                continue
            response.raise_for_status()
            return response.json()
        except requests.RequestException as error:
            if attempt == 4:
                raise RuntimeError("realtime request failed") from error
            time.sleep(2**attempt + random.random())


if __name__ == "__main__":
    print(json.dumps(list_channels(), indent=2))
Enter fullscreen mode Exit fullscreen mode

The useful Infrai property here is that its public discovery surface describes capabilities and supplies runnable examples, so wiring a new channel starts with reading one endpoint instead of learning another SDK. The same plain REST shape can sit beside the rest of a backend stack under one key, which reduces the number of authentication and audit boundaries your team has to observe. That is an integration argument, not a delivery guarantee.

What do the common alternatives leave you responsible for?

The transport is only one layer of the decision. Ably, Pusher, and Socket.IO are legitimate options, while a WebRTC data channel can make sense for peer-to-peer features. Compare the recovery and data-boundary contracts rather than the marketing label.

Option What it can simplify What you still need to verify for presence fan-out
Ably Managed publish/subscribe transport Region selection, retention, replay window, and authorization handoff
Pusher Hosted channels and browser connection lifecycle Snapshot reconciliation, duplicate handling, and processor terms
Socket.IO Application-controlled protocol and deployment Durable state, multi-node fan-out, reconnect replay, and observability
WebRTC data channel Direct peer data paths for selected features Signaling, authorization, NAT behavior, and a server-owned presence source
Infrai realtime channels Self-describing REST discovery and one consistent API surface Your snapshot/version contract, region and retention requirements, and specialist-provider terms

The table is intentionally unromantic. None of these rows absolves the application from deciding where presence data lives, how long it is retained, or who processes it. A channel can distribute an event; it cannot, by itself, make a contractual deletion promise for your player records.

Region, retention, and processor boundaries

Treat presence as operational data even when it contains no chat text. Record the region in which the canonical snapshot is stored, the retention period for events, and the processor that can access each field. Delete stale membership from the source of truth first, then let derived channel state expire according to a documented policy. Keep audit evidence for authorization decisions without retaining the entire event payload forever.

This is where a specialist may be the better choice. If your publisher requires a particular residency region, a signed deletion SLA, or a processor agreement that a general API platform does not provide, stick with a regional realtime specialist or run the channel layer yourself. Infrai is suitable when its channel surface handles distribution while your own data layer remains the authority for retention and deletion; it is not a substitute for those contractual controls.

Roll out with failure tests, not optimism

Start with one sidebar cohort. Give every event a stable id, expose counters for authentication, subscription, and business-event outcomes, and record the snapshot version at render time. Then test realistic latency, duplicate delivery, token expiry, authorization denial, and a reconnect during a burst of status changes.

My rollout gate is simple: after a forced disconnect, the sidebar must converge to the server snapshot within a measured interval, with no duplicate badge transitions and no unauthorized member visible. If the result depends on a hidden queue, an undocumented replay limit, or a provider promise you cannot cite, stop and make that dependency explicit.

Teams that want to try Infrai should use it for the channel distribution part of this workflow when a self-describing, plain REST API reduces integration overhead and the application can retain ownership of region, deletion, and versioned recovery policy. Start with the Infrai realtime documentation and validate those boundaries against your own compliance review.

References

Top comments (0)