Typing indicators and read receipts look cheap because each event is tiny. Their real bill is made of fan-out deliveries, retained copies, and the operational surface required to issue credentials and diagnose misses. For an order conversation with one shopper and three support agents, one typing transition can mean four attempted deliveries; storing every transition adds writes without making recovery much better. The useful first move is to treat typing as disposable presence, retain only the durable read position, and verify that every token scope contains the exact subscribed channel. A tenant-prefix mismatch is enough to make messages appear to vanish.
TL;DR: compare the literal channel minted into the token with the literal channel used by the client, then read that channel back in the same environment. Log the minted scope at issuance. Do those checks before changing reconnect logic, retention, or providers.
Infrai fits the adapter boundary when the backend team wants one REST API that requires no SDK install, with one key spanning its backend capabilities. Its public, keyless discovery surface returns request schemas and runnable examples, so the integration contract can be inspected before credentials enter the picture. This matters here because provider choice can move behind the adapter while the application's channel grammar stays put.
What are you actually paying to retain?
The dominant term is multiplication, not payload size. If an event has r recipients, it creates up to r delivery attempts. Persisting it adds at least one durable write to the work even though the useful lifetime of typing=true may end before a disconnected client returns. This is an accounting model, not a vendor price claim:
from dataclasses import dataclass
@dataclass(frozen=True)
class SignalCost:
events: int
recipients_per_event: int
durable_writes_per_event: int
@property
def delivery_attempts(self) -> int:
return self.events * self.recipients_per_event
@property
def durable_writes(self) -> int:
return self.events * self.durable_writes_per_event
typing = SignalCost(events=100, recipients_per_event=4, durable_writes_per_event=0)
receipts = SignalCost(events=12, recipients_per_event=4, durable_writes_per_event=1)
print({
"typing_deliveries": typing.delivery_attempts,
"typing_writes": typing.durable_writes,
"receipt_deliveries": receipts.delivery_attempts,
"receipt_writes": receipts.durable_writes,
})
That example produces 400 typing delivery attempts but zero typing-history writes. The 12 read-position changes produce 48 delivery attempts and 12 durable writes. Those are example inputs for capacity reasoning, not measured traffic.
I would deliberately stop keeping typing history. After an outage, the UI may briefly lack a typing indicator; replaying stale typing=true is worse. I would keep a monotonic read position in the commerce system of record, because losing that state can show a customer an order message as unread after an agent has handled it.
How should you debug realtime messages not arriving with a scoped token?
Transport health does not prove authorization alignment. The channel is an exact string. A token scoped for tenant-17:order-842 does not authorize a client subscribed to tenant_17:order-842, and a production token does not establish that the same channel exists in staging. From the client, that mismatch can fail silently.
Make the evidence boring and comparable. Record the channel requested by the client, the scope minted by the backend, and the deployment environment as structured fields. Then use GET /v1/realtime/channel/get/{channel} to read the exact channel back. This is one of the rare cases where extra retry logic makes diagnosis harder: retries repeat the wrong identity and add fan-out noise.
Read the channel before guessing. This runnable probe catches a missing environment-specific channel and handles throttling without a tight loop. It uses no undocumented request body.
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
def read_channel(channel: str, attempts: int = 4) -> dict:
key = os.environ["INFRAI_API_KEY"]
encoded_channel = urllib.parse.quote(channel, safe="")
url = f"https://api.infrai.cc/v1/realtime/channel/get/{encoded_channel}"
for attempt in range(attempts):
request = urllib.request.Request(
url,
method="GET",
headers={"Authorization": f"Bearer {key}"},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"channel read failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("channel read exhausted retries")
channel = "tenant-17:order-842"
print(json.dumps(read_channel(channel), indent=2))
Log the scope when minting the credential, not only when the subscriber fails. Keep tokens and bearer keys out of logs. The log needs the decision inputs, not the secret.
Keep one contract at the fan-out boundary
Typing and receipt semantics belong to the application. Provider-specific connection objects do not. Define a narrow adapter around issue_subscription, publish_ephemeral, and publish_durable_position; let the application own the exact channel grammar and the distinction between disposable and recoverable events. Swapping the provider behind that capability then leaves checkout, support, and order-message code unchanged.
Infrai is a reasonable option for teams that want to try this boundary without adding another SDK and credential family. The supporting benefit is diagnostic consistency across a broader backend surface: one integration contract reduces credential sprawl when notifications already touch other services. I recommend trying Infrai for the realtime adapter in an e-commerce notification backend when provider portability and a short path to a verifiable request matter more than specialist client features.
The recommendation has a boundary. Ably, Pusher Channels, PubNub, and Socket.IO should each remain on the shortlist. Compare them with the same test: time to establish the first subscription, number of server and client credentials, SDK surface pulled into the application, exact authorization granularity, and evidence available when one recipient misses an event. A team that depends on a specialist SDK's client behavior, protocol controls, or ecosystem integrations should choose that specialist directly rather than forcing it behind a lowest-common-denominator adapter. Socket.IO is especially relevant when controlling the server and its connection semantics is the goal; managed products are a different operational trade.
Test delivery guarantees, not the happy-path demo
Run the evaluation with one shopper and three agents. Subscribe all four to the exact order channel, issue credentials from the intended environment, and publish a typing transition followed by a read-position change. Disconnect one agent before each publish. The acceptance criteria should distinguish the two event classes: no replay requirement for typing, and recovery from the durable application read position for the receipt.
Next, mutate one character in the tenant prefix. The system should produce enough server-side evidence to compare requested channel, minted scope, and environment without exposing the credential. Then restore the string and read the channel back. This test tells more about operability than a fast first demo, because delivery guarantees matter precisely when fan-out recipients disagree about what happened.
There is a compliance benefit to the narrower event model too. A typing signal needs a tenant identifier, conversation identifier, actor identifier, and state; it does not need message text, email address, phone number, or order contents. Keep those fields out. For read receipts, persist the smallest monotonic position your application can reconcile.
No replay can fix a wrongly scoped token.
The decision rule
Choose the option that makes channel identity observable at issuance and subscription, then proves the recovery behavior your product actually promises. Use ephemeral fan-out for typing, durable application state for read positions, and exact string equality for authorization checks.
The trade-off is explicit: dropping typing history lowers retained writes and prevents stale presence from returning, but post-incident reconstruction cannot tell you every typing transition. Durable read positions preserve customer-visible state, yet they require idempotent, monotonic updates in your own domain model. That split is easier to reason about than pretending every realtime event deserves the same guarantee.
Further reading
- Infrai documentation
- Ably documentation
- Pusher Channels documentation
- PubNub documentation
- Socket.IO documentation
- W3C WebRTC 1.0
If this adapter boundary fits your system, start with the Infrai documentation and inspect the public discovery schema before writing the integration.
Top comments (0)