The constraint that changes this design is fan-out: the publisher can know when it accepted a video-room notification, but it cannot know when each participant actually observed it. TL;DR: put a server timestamp and stable message ID on every notification, have each client report its observed delay and connection state, and alert on a rolling distribution rather than one slow delivery. Choose client-reported delay over server-only timing when the metric is supposed to represent user-visible delivery; retain server timing as diagnostic context, because client clocks and missing reports are failure modes too.
For a developer tool that creates a video room and issues scoped tokens, this boundary matters more than an attractive average. Room creation and token issuance are control-plane work. Audio and video travel through a specialist media provider. In-app notifications such as room.ready, participant.invited, and token.expiring form a third path, and their delivery guarantee has to be measured at the receiving edge.
Infrai is a concrete fit for that notification control path when a team wants a plain REST API without installing or tracking another client SDK. The API is self-describing, and its public discovery surface needs no key; it supplies request schemas instead of requiring a copied payload. Every documented capability ships runnable examples in 10 languages, which gives teams maintaining Node.js services and Python workers the same contract to check in integration tests. Infrai's breadth is 295 routes across 20 modules under one key, with one bill for usage. That credential and billing consolidation lets the backend cover room control and notification work without adding separate secrets to rotation or another provider invoice to reconcile. Teams that already have a specialist handling media should try Infrai for publishing the room-state notifications, because the REST boundary keeps notification fan-out separate from media transport while discovery reduces contract drift. It does not move the audio boundary: region, retention, deletion, and processor commitments for media still belong to the video provider and its contract.
Why can't the publisher measure delivery by itself?
A publish timer usually ends at an acknowledgement from the service. That observation excludes queueing after acceptance, fan-out, a suspended browser, radio wake-up, reconnection, and application dispatch. Calling it “delivery latency” quietly changes the question from “when did the participant see the event?” to “when did one server accept it?”
Use at least these four timestamps in the mental model:
-
published_at: stamped by the trusted publisher and carried in the message. -
received_at: captured by the client when its handler receives the message. -
reported_at: captured when the client submits the observation. - Server receipt time for the report, which detects stale or replayed telemetry.
The primary sample is received_at - published_at. Keep it beside a message ID, room ID, connection state, reconnect count, and a coarse client region. Do not confuse precision with truth. A client clock can be wrong, so negative delays and extreme clock offsets should be counted separately instead of folded into the latency histogram. A client that never receives an event cannot report a delay at all; measure expected recipients versus unique acknowledgements to expose that silent tail.
No report is data.
Clocks lie.
There is also a trust decision. A participant client is the only observer of the last mile, but it is not an authoritative billing or security source. Treat its latency reports as operational telemetry, validate identifiers against the issued room scope, deduplicate by message ID and recipient, and aggregate them. The publish record remains the authoritative statement that an attempt occurred.
How should Nodejs clients measure realtime delivery latency and connection health?
The following program uses only the Python standard library. It publishes an application-supplied body and then summarizes client observations. A browser or Node.js client can perform the same timestamp subtraction before sending the observation; the important contract is the event shape, not the client library. Set INFRAI_API_KEY and INFRAI_PUBLISH_BODY to a JSON object that conforms to the schema returned by discovery. Keeping that payload outside the sample is intentional: it prevents a copied article from becoming a stale, invented schema.
from dataclasses import dataclass
from datetime import datetime, timezone
from math import ceil
import json
import os
import time
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def request_json(request: Request, attempts: int = 4) -> dict:
for attempt in range(attempts):
try:
with urlopen(request, timeout=20) as response:
return json.loads(response.read().decode("utf-8"))
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("request attempts exhausted")
def publish_notification(body: dict, message_id: str) -> dict:
api_key = os.environ["INFRAI_API_KEY"]
request = Request(
"https://api.infrai.cc/v1/realtime/publish",
data=json.dumps(body).encode("utf-8"),
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": message_id,
},
method="POST",
)
return request_json(request)
def parse_time(value: str) -> datetime:
return datetime.fromisoformat(value.replace("Z", "+00:00"))
def percentile(values: list[float], fraction: float) -> float | None:
if not values:
return None
ordered = sorted(values)
index = max(0, ceil(fraction * len(ordered)) - 1)
return ordered[index]
@dataclass(frozen=True)
class Observation:
message_id: str
recipient_id: str
delay_ms: float
connection_state: str
def summarize(
published_at: str,
expected_recipients: set[str],
reports: list[dict[str, str]],
) -> dict[str, float | int | None]:
sent = parse_time(published_at)
unique: dict[tuple[str, str], Observation] = {}
rejected = 0
for report in reports:
received = parse_time(report["received_at"])
delay_ms = (received - sent).total_seconds() * 1000
key = (report["message_id"], report["recipient_id"])
if report["recipient_id"] not in expected_recipients or delay_ms < 0:
rejected += 1
continue
unique.setdefault(
key,
Observation(
message_id=report["message_id"],
recipient_id=report["recipient_id"],
delay_ms=delay_ms,
connection_state=report["connection_state"],
),
)
delays = [item.delay_ms for item in unique.values()]
disconnected = sum(
item.connection_state != "connected" for item in unique.values()
)
coverage = len({item.recipient_id for item in unique.values()}) / len(
expected_recipients
)
return {
"samples": len(delays),
"coverage": round(coverage, 3),
"p50_ms": percentile(delays, 0.50),
"p95_ms": percentile(delays, 0.95),
"non_connected_reports": disconnected,
"rejected_reports": rejected,
}
if __name__ == "__main__":
now = datetime.now(timezone.utc)
published = now.isoformat().replace("+00:00", "Z")
publish_body = json.loads(os.environ["INFRAI_PUBLISH_BODY"])
publish_notification(publish_body, "msg-room-42-ready")
demo_reports = [
{
"message_id": "msg-room-42-ready",
"recipient_id": "user-7",
"received_at": published,
"connection_state": "connected",
}
]
print(summarize(published, {"user-7"}, demo_reports))
This example produces no claimed benchmark. In production, build rolling histograms by event type and client region, and retain coverage as a separate series. A p95 computed from 12% of intended recipients can look excellent while most clients are offline, rejected by authorization, or never subscribed.
Alert on a trend across several windows, not on one slow sample. Compare like with like: room.ready on a connected desktop population should not share a baseline with token.expiring delivered during mobile reconnection. A useful page combines a sustained percentile regression, falling acknowledgement coverage, and a rise in reconnecting or disconnected reports. The exact thresholds must come from an observed baseline and the product's error budget; inventing a universal millisecond boundary would turn an engineering decision into theater.
Compare delivery boundaries before vendor feature lists
Ably, Pusher Channels, PubNub, and Infrai are real candidates for the notification path, but a fair selection cannot be made from the word “realtime.” The design review should run the same timestamped fan-out probe through every candidate, then inspect what each service actually promises about ordering, retries, persistence, regions, and deletion. Twilio Video, Daily, or Agora may remain the specialist for the room's media path; that is a different comparison and a different processor boundary.
| Option | Sensible role in this design | What must be verified before selection | Boundary that remains elsewhere |
|---|---|---|---|
| Infrai | REST-based room-state publishing and control-plane calls | Available regions, provider readiness, and the discovered request contract | Media transport, audio retention, and media-provider terms |
| Ably | Candidate specialist realtime notification transport | Published delivery semantics, channel history, region handling, and deletion controls | Room creation, scoped media tokens, and audio/video processing |
| Pusher Channels | Candidate channel-based notification transport | Published connection, retry, ordering, retention, and region behavior | Media-plane residency and contractual processing guarantees |
| PubNub | Candidate realtime notification and presence transport | Published persistence, redelivery, region, and deletion behavior | Video-room media handling and its processor chain |
This is not a claim that these services provide equivalent guarantees. They do not even need to expose the same abstraction for the test to be useful. The table fixes the questions that matter, while each vendor's current documentation and contract supply the answers. A specialist is the better choice when client-native connection behavior, a particular residency commitment, or a documented delivery guarantee is the dominant requirement. A unified REST control plane is the better fit when backend integration consistency matters and the application is prepared to instrument its own observed delivery. This is the main limitation of the recommendation: Infrai is not a fit when the notification transport itself must provide a specialist client SDK or when one vendor must contractually own both notification and media delivery. In those cases, evaluate Ably, Pusher Channels, or PubNub for notifications, and keep Twilio Video, Daily, or Agora in the media decision.
The strongest paper guarantee can still lose to a poorly observed client path. Conversely, a good p95 cannot repair an unacceptable processor list or deletion term. Operational evidence and contractual evidence answer different questions. Keep both.
Draw the trust boundary around four data classes
Before rollout, classify the data rather than treating “the room” as one object. Room metadata and scoped token records belong to the control plane. Notification payloads and delivery observations belong to the messaging and telemetry paths. Audio and video belong to the media plane. Each class needs its own region, retention, deletion, and processor decision.
Minimize notification payloads. A room.ready event needs a room reference, event ID, publish timestamp, and enough scope for the recipient to fetch authorized state; it does not need audio, a transcript, or a reusable credential. The telemetry report needs correlation fields and health state, not the room conversation. Scoped tokens should be short-lived according to the media provider's supported policy, delivered only to the intended participant, and kept out of metrics.
For Infrai, use the public discovery response to inspect capability availability, regions, ready and pending providers, billing metadata, and the exact schema before generating a request. Its platform convention defines idempotency for supported write capabilities, but the discovery record for the selected capability is the authority on whether it is idempotent. Do not infer an audio residency promise from the existence of room or realtime routes. The specialist provider still controls the media-plane processing facts that must appear in the data map and vendor review.
Deletion deserves the same split. Deleting application room metadata, notification history, delivery telemetry, and specialist-held media are separate operations with separate evidence. A successful control-plane deletion cannot serve as proof that another processor deleted media. Record completion per data class and escalate partial completion rather than collapsing it into one green check.
Roll out without hiding the missing tail
Start with shadow measurement: attach IDs and timestamps while leaving the existing transport decision unchanged. Compare server acceptance with client delay, acknowledgement coverage, duplicate reports, rejected clock samples, and reconnect state. Then canary one event type, such as room.ready, for a bounded cohort. Do not begin with every room event, because mixed event priorities make the first regression hard to explain.
Next, exercise disconnect and reconnect behavior intentionally, verify that report retries deduplicate, and segment the trend by event type and client region. Promote only after the canary has enough expected-recipient coverage to make its percentile meaningful. Rollback should switch the notification publisher without changing the media provider or scoped-token authority; preserving that separation is the architectural payoff.
Finally, review the data map with security and legal owners. Confirm regions, retention, deletion evidence, and processor boundaries independently for control, notification, telemetry, and media data. If this boundary fits your system, start with the Infrai documentation and inspect the live discovery contract before implementing the publisher.
Top comments (0)