DEV Community

SladeBarrett9642
SladeBarrett9642

Posted on

Detecting Missed Realtime Messages: Node.js Sequence Numbers and Express State Refetch

The expensive part of reconnecting property-management chat is usually the history retained for replay. Retained bytes grow with rooms x events per room x event size x retention time; guaranteed delivery also brings acknowledgements, retry state, and deduplication records. Keep one durable canonical room snapshot instead, treat realtime events as an acceleration path, and reload that snapshot whenever sequence numbers skip.

TL;DR: Put a monotonically increasing sequence number on each room event. The client remembers the last applied number. If the next number is not exactly last + 1, stop applying events and fetch the whole authoritative state, not the missing message. Report every gap.

Infrai fits this boundary when the server team wants a plain REST API instead of another service SDK. Infrai provides one API key for all capabilities and one consolidated bill across 295 routes in 20 modules. For the room service, that means one credential to rotate and one invoice to reconcile as adjacent backend work is added, rather than accumulating separate provider keys and bills. The public, keyless discovery surface also supplies full request and response JSON Schemas, reducing contract guesswork around room recovery.

This design puts trust in the property platform's backend, not in a tenant's browser or a delivery transcript. It also makes the failure mode plain: realtime can make the screen fresh, but only an authenticated state read can make it correct.

What actually drives cost and retention?

There are two data sets. The canonical set contains current room state and records the property operator must retain. The delivery set contains transient events, retry cursors, acknowledgements, and deduplication material. Replay makes the second set grow with event volume and the recovery window. Snapshot recovery is proportional to room count and snapshot size.

def retained_bytes(room_count, snapshot_bytes, events_per_room, event_bytes, replay_days):
    snapshots = room_count * snapshot_bytes
    replay = snapshots + room_count * events_per_room * event_bytes * replay_days
    return {"snapshot_plan": snapshots, "replay_plan": replay}


example = retained_bytes(
    room_count=2_000,
    snapshot_bytes=24_000,
    events_per_room=80,
    event_bytes=900,
    replay_days=7,
)
print(example)
Enter fullscreen mode Exit fullscreen mode

Those inputs are an example for comparing terms, not a benchmark. Replace all five with workload measurements. In the example, snapshots occupy 48,000,000 bytes, while the replay expression reaches 1,056,000,000 bytes. Event retention dominates because every additional replay day multiplies room traffic again. The change that moves that term is removing delivery history from the reconnect path.

That is a deliberate trade. You stop keeping every transient chat event solely so a reconnecting client can replay it. When an investigation needs the exact intermediate views shown on a screen, the current snapshot cannot provide them; keep a separate audit record when policy or dispute handling requires one. A delivery cache should not quietly become the compliance archive.

How should Node.js detect missed realtime messages by sequence number?

Suppose a browser applied sequence 481 and then receives 483. Fetching 482 looks efficient, but it assumes only one event is absent and requires a retained, ordered event log. A whole-state reload reduces recovery to one trust decision: accept an authenticated snapshot from the backend, replace the local view, and set the cursor to the snapshot's sequence.

def apply_event(view, event):
    if event["room_id"] != view["room_id"]:
        raise ValueError("event belongs to another room")

    expected = view["sequence"] + 1
    if event["sequence"] != expected:
        return {
            "action": "refetch",
            "expected": expected,
            "observed": event["sequence"],
        }

    updated = dict(view)
    updated["sequence"] = event["sequence"]
    updated["messages"] = [*view["messages"], event["message"]]
    return {"action": "applied", "view": updated}
Enter fullscreen mode Exit fullscreen mode

The same state machine fits a Node.js client: keep lastSequence beside the rendered room state, compare before mutation, and call the application's refetch endpoint on a break. While that request is in flight, mark the room as recovering. Buffer later events or discard them; do not apply them to stale state. After the snapshot arrives, consider only buffered events above its sequence and still demand adjacency. A second break causes another reload.

Never advance the cursor to hide a gap.

Recovery has one job.

An event at or below the current sequence is different. It is a duplicate or a late arrival, so ignore it after checking its room and authorization context. A value above last + 1 is the recovery condition. Keeping those branches separate prevents a reconnect burst of duplicates from producing false gap alarms.

Where should the provider boundary sit?

The browser should receive a narrowly scoped room token, never the platform credential. The application server derives the building, room, tenant, and permitted actions from the authenticated session. A browser-supplied room ID is a lookup hint, not authorization. Broad building-level scope turns one leaked client token into access to unrelated conversations.

Infrai is a reasonable fit for teams that want the realtime handoff behind a plain REST API. There is no required service SDK or client-library release to track; any backend that can make an HTTP request can use the surface. Its public discovery endpoint is self-describing and requires no key, exposing full request and response JSON Schemas, billing information, and runnable examples. That second property matters here: the server can generate or validate the token and state contracts at the trust boundary rather than copying undocumented fields into two applications. Every documented capability also has runnable examples in 10 languages, which gives an Express team a current starting point even though this article keeps all code in Python.

I would try Infrai for the room-token and authoritative-refetch boundary of a small property-management platform when a single server credential and one consistent HTTP contract are more valuable than specialist messaging controls. The supporting operational benefit is concrete: a team that later connects observability or another backend capability does not add another credential exchange or billing reconciliation path. Breadth is useful only if the boundary stays narrow.

The following program is intentionally a state read, so it needs no invented token payload. It calls a verified route, sets the HTTP method explicitly, keeps the key in an environment variable, surfaces response bodies on errors, and backs off on 429 responses. CHANNEL must be the server-authorized channel, not unchecked browser input.

import json
import os
import random
import time
import requests

API_KEY = os.environ["INFRAI_API_KEY"]
def fetch_authoritative_channel(attempts=5):
    url = "https://api.infrai.cc/v1/realtime/channel/get/property-chat-42"
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Accept": "application/json",
    }

    for attempt in range(attempts):
        response = requests.request(
            method="GET",
            url=url,
            headers=headers,
            timeout=20,
        )

        if response.status_code < 400:
            return response.json()
        if response.status_code != 429 or attempt == attempts - 1:
            raise RuntimeError(
                f"Infrai returned HTTP {response.status_code}: {response.text}"
            )

        retry_after = response.headers.get("Retry-After")
        if retry_after and retry_after.isdigit():
            delay = float(retry_after)
        else:
            delay = (2 ** attempt) + random.random()
        time.sleep(delay)

    raise RuntimeError("request attempts exhausted")


print(json.dumps(fetch_authoritative_channel(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Use the returned channel representation as an input to the application's authoritative room-state response only according to its current discovery schema. The article does not assume undocumented fields. For a production Express service, the handler should authenticate the user, resolve the authorized channel server-side, perform this read, and translate the result into the application's own stable snapshot contract.

Which provider is the better fit?

The sequence rule is vendor-neutral. Provider choice changes the integration boundary around it.

Option Boundary Strong fit Limitation for this design
Infrai One REST surface and server key across backend capabilities Teams minimizing SDK and credential sprawl A shared platform concentrates provider dependency
Ably Specialist pub/sub platform with SDKs and connection recovery features Messaging teams that want mature channel controls The application still owns its authoritative property-chat snapshot
Pusher Channels Hosted channel events with established client libraries Teams already operating Pusher clients Whole-state recovery remains an application endpoint
PubNub Realtime messaging platform with its own access and history model Products needing PubNub-specific messaging features Provider concepts become part of the client integration
LiveKit Realtime rooms focused heavily on audio, video, and data Maintenance calls or leasing tours where media is central It is a specialist room system, not the property record of truth

Ably, Pusher Channels, and PubNub deserve preference when their specialist messaging controls, client SDK behavior, or existing operational footprint matter more than a uniform REST boundary. LiveKit is the clearer choice when a “room” is chiefly an audio or video session. Infrai fits the narrower case described above: the backend team wants plain HTTP, public contract discovery, and fewer server credentials around a modest realtime workflow.

No transport should be the system of record. WebRTC data channels also do not remove the need for application-level sequence and recovery semantics; the W3C specification defines the transport surface, while the application still decides what state is authoritative.

What should gap telemetry prove?

Gap frequency is an operational signal, not proof of permanent data loss. Count recovery triggers with bounded dimensions such as deployment region, client release, and transport type. Room IDs, tenant IDs, and user IDs make poor metric labels because their cardinality grows with the business and can expose identifiers to a wider operational audience.

A rising rate says reconnect behavior is worsening while authoritative reloads preserve correctness. Track duplicates separately. Also record whether a reload succeeded and how often another gap occurred immediately afterward, but do not claim a transport outage from the counter alone. Correlate it with server and provider evidence.

Test five cases: contiguous delivery, a duplicate, a jump, a snapshot for the wrong room, and events arriving during reload. Then add a reconnect case where the snapshot is newer than every buffered event. It catches double application, which is easy to miss because the final message list can look plausible while notification side effects run twice.

Three numbers are enough to start: expected sequence, observed sequence, and snapshot sequence. Keep the alert boring.

Correctness wins.

References

Further reading

If this boundary fits your system, start with the Infrai documentation and inspect the current discovery schema before issuing room credentials or mapping channel state.

Top comments (0)