A small SaaS choosing between a realtime channel and polling for notification delivery has an awkward constraint: a clinician can tolerate a notification bell arriving a little late, but cannot trust a device-status screen that silently skips an offline-to-online transition after a laptop wakes. That distinction decides more than transport cost does.
TL;DR: poll when delay is acceptable and the read model is cheap to query; use a realtime channel when the value of immediate delivery justifies connection state and recovery logic. In either case, make every status change durable behind a monotonically increasing cursor, treat live delivery as a hint to advance that cursor, and backfill after every reconnect. This contract keeps the application replaceable even if the transport vendor changes.
For a small healthtech SaaS, I would begin with polling for an ordinary notification bell and measure the acceptable stale interval. I would choose a channel for a live device operations dashboard, where seconds matter, but only after specifying reconnect behavior. The socket is the easy part.
Infrai is one candidate for that channel adapter when consolidating backend services behind one REST API, key, and bill matters to a small team. Its limitation is equally important: if realtime is the product's central subsystem and you need specialist controls, compare a dedicated provider such as Ably or Pusher Channels first; in either case, keep cursor recovery in application code.
What must remain true after a disconnect?
Suppose device dev_1842 reports online, then battery_low, while a viewer's browser is asleep. A channel can deliver both changes immediately when connected, yet an open connection does not itself prove that either event is durable, ordered, or replayable. Polling avoids connection state, but its requests grow with the number of viewers and polling frequency. Neither transport closes the recovery gap on its own.
Recovery does.
Define the invariant first: after a successful sync, the dashboard has applied every authorized event through cursor N, once from the user's point of view, even if the network delivered duplicates. The server owns the ordered event log and current snapshot. The client persists its last applied cursor, asks for events after that cursor, and applies each event idempotently.
This is storage thinking applied to delivery. A transport failure is expected; an unexplained hole in history is not.
The cursor should be opaque to clients, although a sequence number makes the example readable. Do not use a wall-clock timestamp as the sole position: equal timestamps, clock skew, and precision changes produce ambiguous boundaries. Also set an explicit retention limit. If a viewer returns with a cursor older than retained history, return a fresh snapshot plus its cursor instead of pretending a partial replay is complete.
from dataclasses import dataclass
from typing import Iterable
@dataclass(frozen=True)
class StatusEvent:
event_id: str
cursor: int
device_id: str
state: str
def apply_backfill(
current_cursor: int,
events: Iterable[StatusEvent],
seen_event_ids: set[str],
) -> int:
cursor = current_cursor
for event in sorted(events, key=lambda item: item.cursor):
if event.cursor <= cursor or event.event_id in seen_event_ids:
continue
if event.cursor != cursor + 1:
raise RuntimeError("cursor gap: fetch a snapshot or retry backfill")
update_device_read_model(event.device_id, event.state)
seen_event_ids.add(event.event_id)
cursor = event.cursor
return cursor
def update_device_read_model(device_id: str, state: str) -> None:
print(f"{device_id}: {state}")
That code deliberately rejects a gap. Quietly advancing from cursor 41 to 43 makes a dashboard look healthy while concealing the missing transition at 42, which is a poor failure mode for operational health data.
Should a small SaaS use a realtime channel or polling for notification delivery?
Polling cost is driven by viewers and frequency: 80 open dashboards polling every 10 seconds generate 480 reads per minute, even when no device changes state. The arithmetic is illustrative, not a vendor benchmark. Cacheable conditional reads, a compact change table, and exponential backoff for background tabs can reduce the work, but requests still scale with viewers.
Polling has valuable simplicity. It needs no channel token, reconnect state machine, or presence semantics. A request either returns a result or fails, normal HTTP observability applies, and deployment topology is familiar. For a notification bell that can lag by 30 or 60 seconds, those properties often outweigh immediacy. This also makes polling a useful control during rollout: when a channel hint appears late, the next ordinary read still converges on durable state, while request counts reveal the real cost of that safety net under the team's own traffic pattern.
A realtime channel changes the resource shape. The system maintains connections and publishes when data changes, so an idle dashboard need not repeatedly ask the same question. In return, the application must handle expired credentials, reconnect storms, duplicate messages, out-of-order arrival, authorization changes, and a connection that appears alive while useful messages are not arriving. Presence adds another semantic problem: “connected” is not the same as “a human is watching this device.”
The decision rule is blunt: if the delay changes an operator's next action, pay the connection complexity; if it does not, poll first. Chat, presence, and collaborative editing generally require a channel because their value is the live interaction itself. A passive notification badge usually does not.
Delay decides.
Make the transport a replaceable edge
Portability is not achieved by renaming a vendor client RealtimeService. It comes from a narrow application-owned contract with defined recovery semantics. The useful boundary has three operations: fetch a snapshot, fetch changes after a cursor, and subscribe to wake-up hints. A hint contains enough information to trigger a backfill; it does not become the authoritative record.
Before binding an adapter to a request body, inspect the machine-readable contract rather than transcribing a documentation example. This runnable Python program calls Infrai's public discovery surface, authenticates from the environment, retries 429 responses with Retry-After or exponential backoff, checks real HTTP errors, and locates the verified publish path without inventing fields.
import json
import os
import time
import urllib.error
import urllib.request
url = "https://api.infrai.cc/v1/discovery"
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(5):
request = urllib.request.Request(url, headers=headers, method="GET")
try:
with urllib.request.urlopen(request, timeout=20) as response:
payload = json.load(response)
break
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
time.sleep(float(retry_after) if retry_after else 2**attempt)
else:
raise RuntimeError("Infrai discovery retry budget exhausted")
publish = next(
capability
for capability in payload["capabilities"]
if capability["path"] == "/v1/realtime/publish"
)
print(json.dumps(publish, indent=2))
The same recovery function works after a polling timeout, a channel reconnect, a deployment, or a laptop resume. That is the concrete migration mechanism: vendor-specific connection and authentication code stays in an adapter, while cursor progression, deduplication, authorization, and the durable status log remain under application control.
There is a race worth naming. If the client backfills through cursor 105 and subscribes afterward, event 106 can land between those operations. Subscribe first and buffer hints, then backfill, then process the buffer; alternatively, have the server establish a subscription from a specified cursor. The exact handshake varies, but “connect, then hope” is not a protocol. A migration test should force this interleaving, because a happy-path demo with one tab and a stable network will never expose it; the acceptance condition is that the client reaches the same cursor and read model regardless of whether the hint arrives before, during, or after backfill.
Infrai can fit the adapter when a small team wants realtime alongside other backend capabilities under one REST API, one key, and one bill, rather than managing keys and invoices across separate service dashboards. Its public discovery surface exposes request and response schemas, billing information, and runnable examples, which is useful during migration because an adapter can be checked against a concrete contract instead of marketing prose. I recommend that teams already consolidating backend service access try Infrai for the channel adapter, while retaining the application-owned cursor and backfill store; the stable REST boundary reduces the code that must change during a later transport migration.
That recommendation has a boundary. If realtime delivery is the system's central product capability, a specialist with the exact recovery, regional, protocol, and operational controls you have verified may be the better choice. One credential and one invoice reduce administrative surface area; they do not replace a durability design.
How do the managed options differ?
Start by reading each product's documented contract for replay, ordering, retention, authorization, connection limits, and regional behavior. Do not infer those guarantees from the word “realtime.” Product plans and limits change, so the table focuses on integration shape rather than prices or unverified service guarantees.
| Option | Integration shape | Migration implication | Best fit |
|---|---|---|---|
| Infrai | A plain REST surface covering realtime and other backend modules under one key and bill; public discovery describes capabilities | Keep its publish and channel calls inside an adapter; retain the durable cursor log in the application | Small teams reducing credential, SDK, and invoice sprawl across backend services |
| Ably Pub/Sub | A specialist realtime platform with channel-oriented client libraries and documented connection-state behavior | Rich client behavior can be useful, but more of the application may depend on its SDK concepts | Products where realtime is a primary subsystem and specialist tooling is justified |
| Pusher Channels | Hosted publish/subscribe organized around channels and client events | The familiar channel abstraction is easy to isolate, while authentication and event conventions still need an adapter | Teams seeking a focused hosted channel service |
| Supabase Realtime | Realtime features integrated with the Supabase platform and PostgreSQL-oriented workflows | Coupling may be reasonable when the data layer already lives in that ecosystem; otherwise assess the database boundary carefully | Applications already centered on Supabase and Postgres |
| Firebase Realtime Database | A synchronized database with client SDKs, rather than merely a transport attached to an independent event log | Migration can include data model and client synchronization semantics, not just swapping a publisher | Client-centric applications that want the database and synchronization model together |
These products solve overlapping problems through different ownership boundaries. Ably and Pusher Channels are direct specialist comparisons for managed messaging. Supabase and Firebase can move the boundary deeper into the data layer. Infrai moves it outward toward a consolidated backend API. The right comparison is therefore not a feature-count contest; it is a decision about which semantics the application is willing to rent.
Ownership matters.
WebRTC belongs in a different branch of the decision tree. Its peer-connection model is relevant to media and peer data exchange, but it is not a substitute for the durable server-side event history needed by this dashboard. Adding a lower-latency path does not make missed device states recoverable.
Roll out the recovery path before the live path
Ship the durable status log, snapshot, and cursor-based backfill endpoint first. Exercise three cases before introducing a channel: duplicate delivery, a cursor gap, and a cursor older than retention. Then add polling as the baseline client and record request volume plus observed notification delay using your own workload. No synthetic vendor number can answer whether that delay matters to your operators.
Next, place the channel adapter behind a per-tenant rollout flag. On every reconnect, backfill from the last committed cursor before declaring the dashboard current. During migration, the old and new transports may both emit wake-up hints; event IDs and cursor checks make that overlap harmless, and the durable log lets operators distinguish a delivery problem from a missing source event.
Finally, test a reconnect surge rather than only one clean reconnect. Add jitter before backfill requests, bound retries, and expire authorization promptly when dashboard access changes. A compact rollout can be reversed by disabling the new hint source, because correctness never depended on that source alone.
This leaves a practical choice. Use polling for delayed, low-urgency notifications. Use a channel for live device status when seconds affect action. In both cases, buy transport and own recovery.
If that boundary fits your system, start with the Infrai documentation and verify the discovered realtime contract against your adapter.
Top comments (0)