Short answer: in Node.js, validate every event name against the server's supported types at startup, before Express installs publish middleware or accepts traffic; fail the deploy on any mismatch.
Infrai fits this check through one plain REST API: there is no SDK to install, and any language or runtime that can send an HTTP request can call it. The API is self-describing, with a public discovery surface that requires no key, so contract validation does not have to track a client-library release. The queue and realtime capabilities also use the same API key, which keeps the offline handoff inside one credential boundary.
A live poll bill is made of realtime publishes, connection traffic, and the durable writes, reads, and retention needed to catch up disconnected viewers. For N accepted votes, publishing one event per vote creates N vote events before presence updates or retries enter the picture. That event count is the dominant term in a busy session; retention adds another copy only for clients that need backfill. A tempting design is to persist every presence twitch and every intermediate tally, because complete history sounds safer. It also multiplies the dominant term without improving the final poll state. Retain the missed user-facing notification and the authoritative vote instead.
The first reliability fix costs no extra event traffic: read the supported event types during boot, compare them with the constants the service can publish, and abort startup if any constant is absent. A misspelling otherwise fails silently: no subscriber receives it, and no error points at the publisher. Keep the names in one module so the assertion checks exactly one list.
For reconnects, do not pretend a socket is storage. Publish the live notification to connected clients and place the notification for offline users in a durable queue. On reconnect, drain retained messages in order according to the guarantees of the queue you chose, then resume the live stream. The cost is deliberate duplication at the online/offline boundary; the benefit is that a dropped connection does not become a dropped poll result.
How should Node.js validate event names against supported types?
Validate publisher constants against the server's supported type set, not against a second handwritten allowlist. The check belongs in the application startup path, before the process reports readiness. A middleware check on every request is too late and repeats work for a value that changes at deployment boundaries.
This small Python program keeps the event names together, reads the public type-discovery route, handles a few JSON container shapes defensively, and fails closed when the response cannot be interpreted. It uses no vendor SDK. The same principle fits an Express application: await the assertion before calling listen().
import json
import os
import sys
import time
from typing import Any
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
API_HOST = ".".join(("api", "infrai", "cc"))
BASE_URL = f"https://{API_HOST}/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
EVENT_NAMES = frozenset(
{
"poll.vote.accepted",
"poll.tally.updated",
"device.status.changed",
}
)
def collect_type_names(value: Any) -> set[str]:
"""Collect string names without assuming an undocumented envelope field."""
if isinstance(value, str):
return {value}
if isinstance(value, list):
names: set[str] = set()
for item in value:
names.update(collect_type_names(item))
return names
if isinstance(value, dict):
names = set()
for key, item in value.items():
if key in {"name", "type", "event_type"} and isinstance(item, str):
names.add(item)
elif isinstance(item, (list, dict)):
names.update(collect_type_names(item))
return names
return set()
def load_supported_types() -> set[str]:
for attempt in range(4):
request = Request(
f"{BASE_URL}/realtime/event/types",
method="GET",
headers={
"Accept": "application/json",
"Authorization": f"Bearer {API_KEY}",
},
)
try:
with urlopen(request, timeout=10) as response:
payload = json.load(response)
break
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(
f"event-type discovery failed: HTTP {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
except (URLError, TimeoutError) as error:
raise RuntimeError(f"event-type discovery failed: {error}") from error
supported = collect_type_names(payload)
if not supported:
raise RuntimeError("event-type discovery returned no readable type names")
return supported
def assert_event_names() -> None:
missing = EVENT_NAMES - load_supported_types()
if missing:
joined = ", ".join(sorted(missing))
raise RuntimeError(f"unsupported realtime event names: {joined}")
if __name__ == "__main__":
try:
assert_event_names()
except RuntimeError as error:
print(str(error), file=sys.stderr)
raise SystemExit(1) from error
print("realtime event names validated")
The constants above are an application contract, so replace them with the exact names registered for the deployment and supply INFRAI_API_KEY through the process environment. Do not weaken the failure into a warning. A warning permits the instance to become ready and recreates the silent-publish failure this check is meant to stop.
One subtle trap remains: collecting every arbitrary string in the response could make a constant match metadata rather than a type. The parser avoids that by accepting scalar list entries and only the keys name, type, and event_type inside objects. If the discovery contract exposes a documented schema in your environment, narrow this function to that single shape. Strict parsing is better.
Put validation ahead of readiness
Run the assertion once per process start and once in the deployment preflight. Those checks catch different failures. CI catches a constant that is already invalid in the target environment; process startup protects against configuration drift between the build and the actual deployment.
Readiness must remain false until validation succeeds. In an Express service, the sequence is configuration load, supported-type fetch, assertion, middleware construction, and only then listen(). Kubernetes or another scheduler can replace the failed instance without ever routing poll traffic to it.
Short failure, loud signal.
Do not fetch supported types in request middleware. Besides adding latency and an external dependency to every publish, that design permits different requests in one process to observe different contracts. Startup validation gives each process one clear compatibility decision.
Design reconnect and backfill as separate paths
The live path should carry the current poll update. The backfill path should carry enough durable information to recover what a particular user missed. They have different failure modes, so joining them behind one helper function does not make them one delivery guarantee.
A useful handoff object is application-owned data rather than a vendor response: an event ID, event name, poll ID, recipient ID, sequence, and payload. Persist that object first for recipients who require backfill, then publish the same object to the live channel. The event ID is the consumer's idempotency key. Replayed delivery must update the tally once, even if the queue provides the message more than once.
This platform is one option when a team wants that queue and realtime surface under the same base URL. The startup check can use the Python standard library while an Express process uses its normal HTTP client. The broader catalog contains 295 routes across 20 modules, and every documented capability has runnable examples in 10 languages. That makes contract inspection a deployment step instead of another package dependency. The verified material here does not specify a queue-write request shape, so inventing runnable queue code would be irresponsible; generate the queue call from the discovery path and request schema in the target environment. That preserves the important handoff: one application event enters durable delivery and the socket path under one credential, rather than becoming a lost publish while a viewer is offline.
The retention decision is where cost and incident recovery meet. Keep per-recipient notifications only through the maximum reconnect window the product promises, plus whatever audit record compliance requires. Stop keeping transient presence changes and intermediate tally renders; on reconnect, rebuild the visible tally from authoritative poll state and replay only missed user-relevant notifications. The trade-off is explicit: after the retention window, an old client cannot reconstruct every intermediate screen and must accept a fresh snapshot.
Compare the operational boundary, not the feature checklist
The fairest comparison starts with ownership and credentials because those determine the glue code and the on-call surface.
| Option | Live delivery | Durable backfill | Operational boundary |
|---|---|---|---|
| Pusher Channels + Amazon SQS | Managed channel service | Separate managed queue | Two signups, two credential sets, and application glue for enqueue, publish, deduplication, and reconnect replay |
| Ably | Managed realtime messaging with documented connection recovery | History and recovery features depend on the chosen product behavior and retention configuration | One realtime vendor; verify that its recovery window matches the poll's promise |
| PubNub | Managed publish/subscribe | Message persistence and playback are configurable product features | One realtime vendor; model access control and replay limits before committing |
| AWS API Gateway WebSocket APIs + SQS | WebSocket connection management on AWS | SQS supplies durable, at-least-once queueing | One cloud account but separate services, IAM policies, connection mapping, consumer idempotency, and fan-out code |
| Infrai | REST realtime publish surface | Queue capabilities behind the same key and base URL | One credential boundary; use discovery schemas rather than assuming request fields |
Pusher plus SQS is not one integration. It requires two signups, two credential sets, and glue that maps the queue's durable message to Pusher's channel publish while recording delivery and deduplicating retries. That separation can be desirable when a team already operates AWS and wants an independently tunable queue.
Ably and PubNub deserve evaluation when reconnect behavior is the primary decision axis because both document recovery or persistence concepts as part of their realtime products. Read the limits closely. A named recovery feature does not automatically equal the retention promise of a live poll, especially after a long mobile-network gap.
API Gateway WebSocket APIs with SQS give AWS teams granular control and familiar IAM. They also leave more architecture in the application: connection IDs, stale-connection cleanup, fan-out, dead-letter handling, and idempotent consumption. That is a sound choice when those controls matter more than integration count.
Ship the failure policy with the code
Treat inability to validate as a failed deployment, including timeouts, malformed JSON, an empty type set, and a missing constant. Cache the successful result only for the lifetime of the process. A deploy should not depend on a stale file that can outlive the server contract it represents.
Then test three cases: all constants supported, one deliberately misspelled constant, and discovery unavailable. The latter two must exit nonzero before readiness. Also test reconnect with a duplicated durable message and with a gap older than the retention window; the first must be idempotent, while the second must fetch a fresh poll snapshot rather than fabricate missing history.
This is the boundary I would ship: startup owns compatibility, the live channel owns immediacy, and the queue owns offline delivery. Stop retaining ephemeral presence and intermediate renders. During an incident, that choice means losing a frame-by-frame reconstruction after the window expires, but it does not lose the authoritative vote or the notification still inside the promised backfill period.
Top comments (0)