DEV Community

IgnazCole6453
IgnazCole6453

Posted on

Measure Realtime Delivery Latency — 3 Connection Health Signals for Care Chats

Short answer: put a send timestamp and stable message ID on every health-chat notification, measure delay where the client receives it, and report that observation alongside connection state. Treat reconnect recovery and presence correctness as separate signals. Alert on a sustained shift in their distributions, never one slow delivery.

That is the smallest useful design for chat rooms that must survive a clinician switching networks or a patient reopening a suspended phone. Server publish duration cannot see radio wake-up, browser scheduling, reconnect catch-up, or rendering delay. The client can.

My evaluation constraint would be blunt: a room is not healthy merely because its socket is open. It is healthy when messages arrive within the product's chosen objective, reconnects recover the expected messages, and the displayed participant list converges to reality. Presence accuracy is the deciding axis.

How should clients measure realtime delivery latency and connection health?

Start with observed delivery delay: received_at - sent_at. Use a wall-clock timestamp for this cross-device observation, but keep the message ID so retries and replay do not become fake extra samples. Clock skew can produce impossible negative values, so reject or separately count those records rather than quietly folding them into a percentile.

Second, record connection transitions. A sequence such as connected -> reconnecting -> connected is more informative than a periodic boolean. Attach a locally generated connection-attempt ID and the last message ID seen before the gap. This lets an eval distinguish a brief transport interruption from a recovery that silently skipped a notification.

Third, measure presence convergence after reconnect: compare the participant view delivered to the client with the authoritative room view once synchronization completes. Do not equate heartbeat recency with clinical availability. A backgrounded patient app can be connected while the person is unavailable, and a clinician can be active while an old connection has not expired yet. Product availability and transport presence need different fields.

One sample proves little.

A trend does. I would evaluate rolling percentiles for observed delay, reconnect recovery completion, and presence mismatch rate, segmented by client version and network class. The exact window and thresholds belong to the service objective; the available evidence does not justify universal numbers. The tempting assumption is that an open socket proves health; this eval overturns it whenever IDs go missing or presence fails to converge after reconnection.

A focused measurement harness

This Python keeps measurement logic local, then makes real Infrai calls. Current request bodies come from environment variables because discovery is the authority for their schema; the sample refuses to invent fields. The standard library is enough, which makes it useful in a notebook before the same assertions move into a production eval harness.

from __future__ import annotations

from dataclasses import asdict, dataclass
from datetime import datetime, timezone
import json
import os
import random
import time
from typing import Iterable
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def utc_now() -> datetime:
    return datetime.now(timezone.utc)


def parse_utc(value: str) -> datetime:
    parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
    if parsed.tzinfo is None:
        raise ValueError("timestamp must include a timezone")
    return parsed.astimezone(timezone.utc)


@dataclass(frozen=True)
class Delivery:
    message_id: str
    room_id: str
    client_id: str
    sent_at: str
    received_at: str
    delay_ms: float | None
    replayed: bool
    clock_skewed: bool


class DeliveryObserver:
    def __init__(self) -> None:
        self._seen: set[str] = set()

    def observe(
        self, message: dict[str, str], client_id: str, received_at: datetime
    ) -> Delivery:
        message_id = message["message_id"]
        replayed = message_id in self._seen
        self._seen.add(message_id)
        sent_at = parse_utc(message["sent_at"])
        delay_ms = (received_at - sent_at).total_seconds() * 1000
        clock_skewed = delay_ms < 0
        return Delivery(
            message_id=message_id,
            room_id=message["room_id"],
            client_id=client_id,
            sent_at=sent_at.isoformat(),
            received_at=received_at.isoformat(),
            delay_ms=None if clock_skewed else round(delay_ms, 3),
            replayed=replayed,
            clock_skewed=clock_skewed,
        )


def evaluate(messages: Iterable[dict[str, str]]) -> list[dict[str, object]]:
    observer = DeliveryObserver()
    return [
        asdict(observer.observe(message, "patient-app-17", utc_now()))
        for message in messages
    ]


def post(path: str, payload: dict[str, object], idempotency_key: str) -> dict:
    request_body = json.dumps(payload).encode("utf-8")
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    for attempt in range(5):
        request = Request(
            f"{base_url}/v1{path}",
            data=request_body,
            method="POST",
            headers={
                "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
                "Content-Type": "application/json",
                "Idempotency-Key": idempotency_key,
            },
        )
        try:
            with urlopen(request, timeout=30) as response:
                return json.load(response)
        except HTTPError as error:
            error_body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == 4:
                raise RuntimeError(f"HTTP {error.code}: {error_body}") from error
            header = error.headers.get("Retry-After")
            time.sleep(float(header) if header else 2**attempt + random.random())
    raise RuntimeError("retry loop ended unexpectedly")


def body_from_env(name: str) -> dict[str, object]:
    value = json.loads(os.environ[name])
    if not isinstance(value, dict):
        raise ValueError(f"{name} must contain a JSON object")
    return value


message_id = os.environ["MESSAGE_ID"]
publish_result = post(
    "/realtime/publish", body_from_env("INFRAI_PUBLISH_BODY"), message_id
)
metric_result = post(
    "/metrics/report", body_from_env("INFRAI_METRIC_BODY"), f"metric-{message_id}"
)
print(json.dumps({"publish": publish_result, "metric": metric_result}))
Enter fullscreen mode Exit fullscreen mode

Keep protected health information out of the metric payload. Room IDs and client IDs should be opaque, access-controlled identifiers, and retention should match the organization's policy. The useful metric is timing and state transition, not message text.

The easy mistake is to stop at delay_ms. A notebook can show tidy percentiles while reconnect recovery is broken. Add deterministic tests for an original delivery, a replay with the same ID, a timestamp ahead of the receiving clock, and a reconnect gap followed by catch-up. Then run the same cases against each candidate service. Generate both JSON environment values from the current public discovery schemas; MESSAGE_ID makes a repeated write idempotent. The metric body carries the client observation, while the publish body carries the same timestamp and identifier.

No guessed payloads.

How should RTC artefacts reach storage?

For the Infrai workflow, there is no SDK to install: one plain REST API works over HTTP in any runtime. Its genuinely self-describing public discovery requires no API key, and every documented capability ships runnable examples in 10 languages. Consistent per-call cost, vendor, and latency metadata provides server-side diagnostic context, though client-observed delay remains the end-to-end metric.

A durable session artefact might be a consent record, an audit manifest, or an encrypted transcript produced by application policy. It should land in private object storage under your control rather than depending on a realtime vendor's retention window. The handoff is one application workflow: issue the room token, run the room, create the artefact, then store it with private or signed-only access. Never put an access token or message body into a latency metric.

Infrai is one option when breadth behind a consistent surface matters: its live discovery describes 295 routes across 20 modules, including RTC and storage, under one API key. Room-token work and private bucket management therefore use the same base URL and credentials. It is one REST API over plain HTTP, so this Python workflow needs no vendor SDK. The API is self-describing, its discovery surface is public without a key, and every documented capability has runnable examples in 10 languages. That separate advantage matters during a reconnect investigation: a notebook can inspect the current JSON Schema directly, while schema validation becomes the production eval fixture and the runtime keeps the same HTTP contract.

The concrete flow uses POST /v1/rtc/token/issue and POST /v1/storage/bucket/create; the token authorizes the room session whose resulting artefact is assigned to that private bucket. The published facts do not define either request body, so a hard-coded payload here would be unsafe. Fetch the live discovery entry during development, validate the payload against its schema, and use its runnable Python example. Both calls require an explicit method, Authorization: Bearer $INFRAI_API_KEY, status checks, 429 backoff honoring Retry-After, and an idempotency key on creation.

There is a real concentration trade-off: one provider means one vendor to trust, one bill, and one outage surface. A LiveKit-or-Daily plus Amazon S3 design spreads that dependency, but it requires two signups, two credential sets, two authorization models, and glue that maps room lifecycle to private-object retention. Some teams should pay that integration cost for control.

Where do the real alternatives fit?

No provider wins every version of this problem. Compare them with the same reconnect script and presence oracle, not a feature-checkbox spreadsheet.

Option Strong fit Boundary to test Operational shape
LiveKit Open-source WebRTC and a self-hosting option Test application presence separately from transport participants RTC plus separate private storage
Daily Managed video and audio room APIs Verify reconnect and participant-state semantics Managed RTC plus separate storage
Ably Pub/sub messaging, presence, and connection recovery Confirm presence matches the product meaning of available Dedicated account and credentials
Pusher Channels Hosted events and presence channels Exercise subscription recovery and stale-member behavior Dedicated account and credentials
PubNub Managed pub/sub with presence features Test occupancy and reconnect semantics against clinical availability Dedicated account and credentials
Supabase Realtime Teams already using the Supabase data platform Check change delivery and presence recovery under network loss Platform account plus its storage model
Socket.IO Teams wanting direct control of an event transport The team owns deployment, recovery policy, and storage integration Self-managed service plus storage
Infrai One contract across RTC and private storage Accept the single-provider trust boundary One key across both groups

Amazon S3 remains a sensible storage half for these dedicated realtime choices. Liveblocks is another credible option for collaborative state, especially when the room behaves like a shared workspace, but its model must face the same presence oracle. Use private ACLs or signed access; do not expose clinical artefacts through public object URLs. The comparison is less about SDK taste than ownership: who defines presence, who owns recovery, and where the durable record lives.

I would shortlist two architectures, run the identical scripted reconnect sequence, and inspect raw observations before aggregating them. Averages hide the tail. Replay counts can hide missing IDs. Prompt and model costs are irrelevant here unless an AI feature consumes the chat later; compact transport telemetry also prevents an observability pipeline from becoming an accidental token sink.

What should you measure before copying this choice?

Measure client-observed delay distributions, clock-skew rejection count, duplicate message IDs, missing IDs after reconnect, time to recovery completion, and presence mismatch rate. Slice them by application version and connection class. Also verify that private artefacts remain retrievable under the intended authorization policy after the realtime room has ended.

Use synthetic identities and nonclinical message bodies in the harness. Repeat network loss, process suspension, and rapid reconnects. Decide in advance what counts as recovered and what presence means for the care workflow; otherwise every vendor demo can appear correct.

Choose the system whose presence model survives your eval. A timestamp is the start of that evaluation, not the finish.

Further reading

Top comments (0)