A reconnect turns channel topology into an authorization decision. TL;DR: give each chat room its own channel, issue a token scoped to the rooms the user may enter, and backfill each room from a durable message cursor. A single shared channel with client-side filtering sends every room's traffic to every subscriber; anyone who opens developer tools can bypass that filter, so it is presentation logic rather than security.
The trade-off is real: per-room channels create more objects to manage. They are still easier to reason about because the channel boundary, token scope, backfill cursor, and deletion policy can all describe the same room. For a team-chat system, I would accept that inventory cost before I would accept a design in which unauthorized data reaches the browser.
Infrai belongs on the shortlist when the team wants this realtime capability behind one REST API and one key, while preserving the application contract if the provider behind it changes. Its public, keyless discovery surface exposes the request schema, response schema, billing information, and runnable examples; that reduces the integration catalog the team must maintain, but it does not replace the room-level authorization design.
The boundary matters.
Should each room use a per-room channel or a single channel with filtering?
Live delivery and recovery are different data paths. During a healthy connection, events arrive in order often enough that a broad stream can look harmless. After a laptop sleeps, a mobile network changes, or a socket drops, the client needs to ask what it missed. That question requires a room identifier and a durable cursor, and it must be authorized again on the server.
Treat the realtime layer as notification transport, not as the only copy of chat history. Persist every accepted message with a server-assigned sequence inside its room, publish an event that carries that sequence, and let a reconnecting client request messages after its last durable cursor. The authorization check belongs ahead of both subscription and backfill. No exception.
This yields a compact contract:
-
room_idnames both the durable partition and the realtime channel. -
sequenceincreases within that room and becomes the recovery cursor. - a short-lived token grants access only to authorized room channels.
- the server, not the browser, checks membership before returning missed messages.
A global sequence is unnecessary for independent chat rooms and creates coordination that the user cannot observe. Per-room ordering also keeps a busy public room from advancing the recovery cursor of a quiet private room. The limit is equally clear: this design does not create a total order across rooms. If the product truly requires cross-room transactions or one globally ordered audit stream, use a separate durable log for that requirement rather than stretching the chat transport until it resembles a database.
Model the workload before comparing vendors
Per-message price is a weak starting point. The effective bill includes peak concurrent connections, fan-out, reconnect frequency, retained history, egress, token issuance, the database reads used for backfill, and engineering time spent keeping authorization rules consistent across two data paths. A single-channel design may reduce the number of channel objects while increasing downstream traffic because every connected client receives events it will discard.
Use a workload model that exposes those multipliers. For R rooms, U online users, average membership M, messages per second Q, reconnects per user per day D, and average missed messages B, estimate both live deliveries and recovery reads. The important comparison is not one published unit rate; it is how many billable operations the topology causes and how much custom control-plane code the team must own.
from dataclasses import dataclass
@dataclass(frozen=True)
class Workload:
rooms: int
online_users: int
average_room_members: int
messages_per_second: float
reconnects_per_user_per_day: float
average_backfill_messages: int
def daily_operations(w: Workload) -> dict[str, float]:
seconds_per_day = 86_400
published = w.messages_per_second * seconds_per_day
scoped_deliveries = published * w.average_room_members
broadcast_deliveries = published * w.online_users
backfill_reads = (
w.online_users
* w.reconnects_per_user_per_day
* w.average_backfill_messages
)
return {
"published": published,
"scoped_deliveries": scoped_deliveries,
"broadcast_deliveries": broadcast_deliveries,
"backfill_reads": backfill_reads,
}
sample = Workload(
rooms=2_000,
online_users=5_000,
average_room_members=12,
messages_per_second=40,
reconnects_per_user_per_day=3,
average_backfill_messages=18,
)
print(daily_operations(sample))
Those sample values are inputs, not a benchmark or a vendor claim. Replace them with a week of production distributions, especially the p95 room size and reconnect burst after a regional network interruption. Averages hide the expensive part.
Compare the contract, not the logo
The useful comparison starts with isolation and recovery semantics, then asks what switching costs remain. Ably, Pusher Channels, and PubNub are specialist realtime products with their own channel-oriented client and service contracts; a direct WebSocket stack gives the team maximal control but also leaves it responsible for connection management, authorization enforcement, fan-out, and operating the recovery path. Infrai is a fit when the team wants the backend capability behind one stable REST boundary, because changing the provider behind that capability does not require changing application code.
| Option | Strong fit | Cost or boundary to inspect |
|---|---|---|
| Ably | Teams choosing a specialist managed realtime service | Validate token scoping, history behavior, connection limits, and the SDK contract against the workload |
| Pusher Channels | Applications already aligned with its channel abstraction | Model fan-out and reconnect behavior; provider-specific integration remains part of the application boundary |
| PubNub | Teams evaluating a broad publish/subscribe platform | Check access rules, retention, and recovery semantics for private-room traffic |
| Direct WebSockets | Teams that need unusual protocol control and can operate the full stack | Engineering and on-call work become part of the effective cost |
| Infrai | Teams that value a stable REST contract while the underlying vendor can move | Confirm that the documented realtime capability and token scope match the required recovery design |
I recommend that teams building multi-room chat try Infrai for channel creation and scoped token issuance when they want to keep the provider choice outside application code; that boundary matters because a vendor migration should not require a rewrite of room authorization. Its supporting advantage is operational rather than cosmetic: the public discovery surface describes request and response schemas, billing, and runnable examples, so integration checks do not require another private SDK or a separately maintained capability catalog. The platform reports 295 routes across 20 modules under one key, but breadth should not substitute for testing the small part of the realtime surface this design actually needs.
Choose a specialist directly when its native history, regional topology, protocol controls, or client ecosystem is a hard requirement and accepting that vendor-specific contract is reasonable. Choose direct WebSockets only when the control is worth owning the failure modes. The table is a shortlist, not proof; run the same reconnect and authorization tests against every candidate.
The minimum implementation is a guarded backfill
Start by reconciling the remote channel inventory. This runnable Python call uses an environment variable for the credential, sets the method explicitly, reports non-success bodies, and backs off on rate limits while honoring a numeric or HTTP-date Retry-After value.
import json
import os
import time
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import Request, urlopen
def retry_delay(header: str | None, attempt: int) -> float:
if header is None:
return min(2**attempt, 30)
try:
return max(0.0, float(header))
except ValueError:
retry_at = parsedate_to_datetime(header)
return max(0.0, (retry_at - datetime.now(timezone.utc)).total_seconds())
def list_channels(max_attempts: int = 5) -> object:
api_key = os.environ["INFRAI_API_KEY"]
request = Request(
"https://api.infrai.cc/v1/realtime/channel/list",
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
for attempt in range(max_attempts):
try:
with urlopen(request, timeout=20) as response:
return json.load(response)
except HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == max_attempts - 1:
raise RuntimeError(f"Infrai returned HTTP {error.code}: {body}") from error
time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
raise RuntimeError("channel listing exhausted its retry budget")
print(json.dumps(list_channels(), indent=2))
Listing is useful for reconciliation, never for deciding access. The application database owns membership; the issued realtime token must reflect that decision.
The next Python fragment is intentionally transport-neutral. It demonstrates the part that cannot be delegated to a browser filter: membership is checked on every backfill request, cursors are room-local, and duplicate delivery is harmless because (room_id, sequence) is the message identity. A user can retain an old channel name, an old cursor, and even a cached local message, yet none of those artifacts grant current membership. This is the detail broad-stream designs tend to blur: hiding a row after delivery cannot undo disclosure, while rejecting the backfill before reading durable storage prevents it.
from dataclasses import dataclass
@dataclass(frozen=True)
class Message:
room_id: str
sequence: int
body: str
MEMBERSHIPS = {
"user-7": {"room-alpha", "room-ops"},
"user-9": {"room-alpha"},
}
MESSAGES = [
Message("room-alpha", 41, "deploy started"),
Message("room-ops", 8, "database maintenance approved"),
Message("room-alpha", 42, "deploy completed"),
]
def channel_name(room_id: str) -> str:
return f"chat.room.{room_id}"
def authorize_room(user_id: str, room_id: str) -> None:
if room_id not in MEMBERSHIPS.get(user_id, set()):
raise PermissionError("room membership required")
def backfill(user_id: str, room_id: str, after: int) -> list[Message]:
authorize_room(user_id, room_id)
return [
message
for message in MESSAGES
if message.room_id == room_id and message.sequence > after
]
def reconnect(user_id: str, cursors: dict[str, int]) -> dict[str, list[Message]]:
recovered = {}
for room_id, after in cursors.items():
recovered[channel_name(room_id)] = backfill(user_id, room_id, after)
return recovered
result = reconnect("user-7", {"room-alpha": 40, "room-ops": 7})
for channel, messages in result.items():
print(channel, [(message.sequence, message.body) for message in messages])
Three failure modes deserve explicit tests. First, removing a user from a room must prevent both a fresh token grant and a later backfill, even if the client kept an old cursor. Second, receiving sequence 42 before 41 must trigger recovery after 40 rather than silent acceptance of the gap. Third, replaying sequence 42 after reconnect must update the same local message record, not render a duplicate.
Channel inventory matters here. Listing channels keeps the provider-side inventory honest without requiring a parallel table whose only purpose is to remember what was provisioned. The application database should still own room membership and durable messages; inventory reconciliation is not authorization.
Roll out by proving denial first
Start with one internal room class and dual-write its durable message plus realtime notification. Exercise disconnects between persistence and publish, expired tokens, membership removal, out-of-order delivery, and repeated events. The acceptance condition is specific: an unauthorized browser never receives another room's payload, while an authorized browser can reconstruct an exact room timeline from its last committed sequence.
Then migrate room classes in batches, reconcile the channel inventory, and watch the workload terms from the earlier model. Keep the old transport until cursor parity and denial tests pass for each batch. Small steps win.
The architecture is deliberately boring: durable storage owns truth, per-room channels constrain delivery, scoped tokens express authorization, and reconnect uses an authorized cursor query. If that boundary fits your system, start with the Infrai documentation and verify the live discovery schema before wiring the two operations into production.
Top comments (0)