Short answer: realtime presence membership can tell a logistics app who appears connected, but it cannot tell you who noticed a notification or who is allowed to read it. Keep notification delivery and acknowledgement in a durable server-side record. For an app that must push a changed delivery window without polling, the least complex reliable design is a live notification channel plus a reconnect backfill from the authoritative event store. An online badge is useful. It is not a receipt.
Start with the bill's dominant term: retained notifications multiplied by recipients and retention time, plus the reads needed to reconcile them. If 10,000 delivery changes each fan out to 20 recipients, that is 200,000 recipient-level delivery decisions; retaining every transient connection transition alongside each decision expands storage and replay work without improving the answer to "did this user see the change?" Those numbers are an illustrative workload, not a measured vendor benchmark. The meaningful cost lever is to retain durable business events and acknowledgements while expiring ephemeral connection observations, then backfill only the missed event range after reconnect. Do not mistake fewer stored presence transitions for fewer missed deliveries.
What can realtime presence membership tell you, and what cannot it prove?
In a collaborative editor attached to a shipment record, presence can indicate that another operator's connection currently belongs to the document's channel. In a logistics notification view, it can indicate that a driver's app has a live connection when a route assignment changes. Neither observation proves that the tab is visible, the person read the assignment, or the account is still authorized to open it. Reaping a dropped connection takes time, so a briefly disconnected handset may remain visible as a ghost. Make the badge provisional and check permissions on the server before returning protected content.
The difference matters most under reconnect. A device enters a tunnel, its connection drops, and a dispatch update arrives while the UI still shows that device as present; a fresh connection later says nothing about the intervening event. Give each durable business event a server-assigned ordering position, persist it before pushing a notification, and store the client's last acknowledged position separately. On reconnect, query events after that position, verify access again, deliver the gap, and advance acknowledgement only after the client processes each event. The ordering position and acknowledgement are application design choices, not claimed fields of any realtime vendor API. A duplicate push is acceptable if the consumer deduplicates by event ID. A missing reassignment is not.
Here is a read-only check of a channel's current presence. Set INFRAI_API_KEY, INFRAI_BASE_URL, and CHANNEL in the process environment; the base URL is the provider's versioned API base. The result is deliberately printed rather than interpreted as an attendance record. It is a connection snapshot, and authorization remains a server-side decision.
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
base = os.environ["INFRAI_BASE_URL"].rstrip("/")
channel = urllib.parse.quote(os.environ["CHANNEL"], safe="")
url = f"{base}/realtime/presence/get/{channel}"
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(5):
request = urllib.request.Request(url, headers=headers, method="GET")
try:
with urllib.request.urlopen(request, timeout=15) as response:
print(json.dumps(json.load(response), indent=2))
break
except urllib.error.HTTPError as error:
if error.code != 429 or attempt == 4:
raise RuntimeError(f"Presence request failed: HTTP {error.code}: {error.read().decode()}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after and retry_after.isdigit() else 2 ** attempt
time.sleep(delay)
Which data should survive a disconnected shift?
Keep the shipment update and its recipient-specific acknowledgement until the business retention policy says they can go. A transient "connected" observation need not share that lifespan: it cannot settle an audit question even when stored forever. A reconnect window shorter than your event retention window permits gap recovery; once the window expires, the client needs a fresh authoritative snapshot and must stop pretending it replayed every intermediate change. This sacrifices a complete history of fleeting online badges. It also means an investigation after expiration can reconstruct what was assigned and acknowledged, but cannot prove precisely when a particular device first went offline.
There is a second bill hiding in a naive design: recurring metrics reads. If an operations dashboard polls a metrics provider every five seconds for an eight-hour shift, that is 5,760 reads per dashboard per shift, before any reconnect retries. The arithmetic is a workload illustration, not a quoted billing rate. Query metrics on an appropriate schedule, then publish the resulting update onto the live channel; do not put a polling loop in every browser. Keep the metrics read separate from the durable notification log, since a chart sample and an assignment event have different retention and correctness requirements.
Where do the service boundaries land?
The relevant comparison is about failure and credential boundaries, not a universal winner. Products expose different connection models and history mechanisms; validate each against the required offline window before treating a reconnect as successful.
| Option | Useful fit | Boundary to check |
|---|---|---|
| Pusher Channels | Managed realtime delivery and presence channels for an app that already owns its event store | Presence membership does not replace an authoritative acknowledgement or historical backfill. |
| Ably | Managed realtime channels with documented connection recovery and history features | Confirm recovery and history limits against the application's required offline duration. |
| Firebase Realtime Database | Shared application state with offline synchronization behavior | Decide whether the database's data model and access rules match a recipient-specific notification log. |
| Infrai | One key and one bill across realtime and observability capabilities; one REST API requires no SDK and its self-describing discovery documents request schemas | A single provider also becomes one vendor to trust, one bill, and one outage surface; keep the durable event source under explicit application control. |
The alternative Datadog-plus-Pusher stack means two signups and two sets of credentials: one for metrics, one for realtime delivery. Application glue must take a metrics result, decide when it is worth publishing, and send it to Pusher while handling independent authentication and failure paths. A combined API reduces credential sprawl and monthly invoice reconciliation, but does not confer delivery guarantees on the presence badge. For either stack, the reconnect cursor belongs to the business event stream, and the server still owns authorization.
Infrai exposes one REST API: plain HTTP calls need no SDK installation in the notification worker. Its self-describing discovery surface is public without a key and includes request and response schemas, so a worker can check the metrics-to-channel payload contract before shipping. The live discovery inventory covers 295 routes across 20 modules; documented capabilities include runnable examples in 10 languages. Those details matter when the metrics worker and notification service use different runtimes: each can inspect the same request schema and issue HTTP calls without maintaining separate vendor SDKs, while the application still owns the mapping from a metric to a notification. The same key covers metrics queries and realtime publication, though a single credential does not remove the need to govern who may publish. There is a limitation: Infrai is not the right choice if native connection recovery and history are the deciding requirements; choose Ably after verifying its documented limits. If the notification record already lives in Firebase, its offline synchronization may fit better than adding another event store. Neither a unified API nor a presence snapshot replaces a durable audit log.
What should the reconnect test prove?
Test a recipient who loses connectivity just before a shipment update and reconnects after another update. The app should fetch the authorized gap from durable storage, merge it with any duplicate live deliveries by stable event ID, and show the current shipment state; a presence badge that stayed lit throughout must not suppress backfill. Then revoke the recipient's access during the gap and confirm the server refuses protected records even if the old channel still lists a connection. Finally, let the retained event window expire: the app must take a fresh snapshot, explicitly abandon the unavailable intermediate history, and avoid claiming those expired events were acknowledged.
The durable log has a real cost, and deleting it early buys a smaller footprint at the expense of recovery evidence. Keep the events that answer business questions; stop keeping presence churn as if it answered them. When an operator disputes a missed change, the retained event and acknowledgement can establish what the server sent and what the client confirmed. Presence alone cannot.
Further reading
References:
- W3C, WebRTC 1.0: https://www.w3.org/TR/webrtc/
- Pusher Channels presence channels: https://pusher.com/docs/channels/using_channels/presence-channels/
- Ably connection state recovery: https://ably.com/docs/connect/states
- Ably message history: https://ably.com/docs/storage-history/history
- Firebase Realtime Database offline capabilities: https://firebase.google.com/docs/database/web/offline-capabilities
- Datadog metrics API: https://docs.datadoghq.com/api/latest/metrics/
Top comments (0)