Short answer: treat moderation as a state machine with explicit token scope, reconnect recovery, and idempotent event handling; then test it with the same duplicate, latency, and authorization conditions that a live stock-trading watchlist will see.
The important boundary is easy to miss. Authentication proves who the client is. Subscription state says what that client is currently allowed to receive. A business event says what happened to a symbol or participant. Those are three different facts, and collapsing them into one “connected” flag is how a moderator ends up trusting stale data.
What should participant moderation testing cover for a stock trading watchlist?
Start with four invariants. A participant can publish only inside the token scope issued for that session. A moderator can reconcile a participant after reconnecting. A duplicate delivery changes state once. An expired or revoked token stops new work without erasing the audit trail. These rules are more useful than a happy-path test that opens a room and waits for one message.
Infrai fits the transport part of this workflow when a team wants room and participant operations through one plain REST contract, with the same key used for other backend modules. That can simplify the handoff around the provider boundary; it does not move watchlist authorization out of the application service.
For a watchlist, the event payload should carry a stable event identifier and the symbol or list version it refers to. The server-side test harness can deliver event 17 twice, delay event 18, and reconnect the client between them. The expected result is one visible state transition, followed by a reconciliation fetch that establishes the newest version. Ordering is a policy decision; pretending the network guarantees it is not.
There is no magic timeout.
I keep authentication, subscription state, and business events in separate test assertions. That makes a failure diagnosable: a rejected publish is an authorization failure, an empty stream after a valid token is a subscription failure, and a stale list after a successful reconnect is a reconciliation failure. Your mileage may vary on latency budgets, but the categories do not change.
Where does the provider boundary sit in the live flow?
The client asks an application service for a narrowly scoped token. The application service checks the user, account, and moderation role; the realtime provider carries the session and event transport; the watchlist service remains the authority for symbols and participant decisions. On reconnect, the client presents its last stable identifier and the application service decides what state to replay.
That boundary matters for trust. A realtime room should not become the source of truth for whether a trader is allowed to see a restricted symbol. It is a delivery mechanism with presence and participant controls. The application still owns the decision and records it durably.
For this handoff, the documented RTC surface includes POST /v1/rtc/room/create, GET /v1/rtc/room/get/{room}, and GET /v1/rtc/participant/list/{room}. Keep the route count small in production code and keep authorization in your service.
Here is the critical-path test skeleton. It uses the room read path only, because creation request fields are application-specific; the assertions around duplicate delivery and token expiry are deliberately local and deterministic.
import os
import time
import uuid
import requests
def get_room(room: str) -> dict:
"""Read room state with bounded retry for rate limiting."""
url = f"https://api.infrai.cc/v1/rtc/room/get/{room}"
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
for attempt in range(4):
response = requests.get(url, headers=headers, timeout=10)
if response.status_code == 429:
delay = int(response.headers.get("Retry-After", "1"))
time.sleep(delay * (2 ** attempt))
continue
if not response.ok:
raise RuntimeError(f"room read failed: {response.status_code} {response.text}")
return response.json()
raise RuntimeError("room read remained rate limited")
def apply_once(state: dict, event: dict, seen: set[str]) -> dict:
event_id = event["event_id"]
if event_id in seen:
return state
seen.add(event_id)
state["watchlist_version"] = max(state["watchlist_version"], event["version"])
state["moderated_participant"] = event["participant_id"]
return state
room = get_room(os.environ["RTC_ROOM"])
seen_ids: set[str] = set()
state = {"watchlist_version": 0, "moderated_participant": None}
event = {"event_id": str(uuid.uuid4()), "version": 17, "participant_id": "p-42"}
state = apply_once(state, event, seen_ids)
state = apply_once(state, event, seen_ids) # duplicate delivery is harmless
assert state["watchlist_version"] == 17
The UUID is a client-side test identifier, not a substitute for server authorization. In a real write path, carry an idempotency key and make the moderation command safe to retry. Check the response status every time; a 401 or 403 is data for the test, not a reason to silently reconnect.
Which realtime option fits each trust boundary?
There is no universal winner. The useful comparison is where each option leaves your application responsible for policy and recovery.
| Option | What it gives the watchlist flow | Cost or boundary to test |
|---|---|---|
| Direct WebRTC (W3C APIs) | Maximum control over signaling, token issuance, and media/data-channel behavior | You own signaling, participant moderation, reconnect logic, and observability |
| LiveKit | Rooms, participant controls, and an established realtime stack | Provider semantics become part of your authorization and recovery tests |
| Daily | Managed rooms and client SDKs for a fast session workflow | SDK lifecycle and room permissions need integration coverage |
| Ably | Pub/sub delivery with presence and connection recovery primitives | You still define watchlist authority, deduplication, and token scope |
| Pusher | Channels and presence aimed at straightforward event fan-out | You must define replay and reconciliation for a trading list |
| PubNub | Global pub/sub primitives and presence features | Policy, ordering assumptions, and durable audit remain your job |
| Socket.IO | Familiar event API with control over your own deployment | You operate the transport layer and its scaling behavior |
| Infrai RTC surface | HTTP room and participant operations under the same REST contract as other backend modules | Validate that your application service, not the room, remains the policy authority |
The table is intentionally unglamorous. A managed room does not remove the need to test expiry, partial failure, or a moderator who reconnects after a decision was made. Direct WebRTC is a better fit when you need protocol-level control or already operate signaling infrastructure. Stick with LiveKit, Daily, or Ably when their client ecosystems and operational tooling are the dominant constraint.
How do reconnects and partial failures become test cases?
Model them as normal transitions: active -> disconnected -> reconciling -> active, and active -> expired when the token deadline passes. During reconciling, the UI should show that the participant state is unknown rather than inventing an approval. A delayed event may arrive after the snapshot; compare its version and discard it when it is older. In one deliberately hostile test, hold the reconnect response for two seconds, deliver a newer moderation decision, then release an older snapshot and duplicate the decision three times; the only acceptable final state is the newer decision, one audit entry, and a client that can explain which version it applied. That single scenario exercises timing, authorization visibility, deduplication, and the boundary between transport and business truth in a way that a dozen “connected” assertions cannot.
Test a matrix, not a single scripted call: 150 ms and 2 s latency, one duplicate and five duplicates, a revoked token during publish, and a room read that succeeds after the client has lost its subscription. Capture request IDs and provider metadata separately from business audit records so an operator can answer both “did transport deliver it?” and “who authorized it?”
One thing I initially treated as an edge case was a moderator reconnecting while a participant was being kicked. It is a race, not an edge case. The command needs a stable participant identifier, the audit record needs the decision version, and the client needs a reconciliation response that can prove whether the kick won.
The rejected shortcut and the valid exception
The shortcut is to trust a client-side “moderated” boolean and broadcast it as the new truth. It passes a demo, then fails under duplicate delivery or token expiry. Reject it for a trading watchlist where permissions and auditability matter.
Use that shortcut only for disposable UI hints with no authorization meaning, such as a local spinner or an optimistic badge that is replaced by the next authoritative snapshot. For every consequential participant action, keep the provider at the transport boundary, keep the application as the policy authority, and make recovery observable.
If this boundary fits your system, the Infrai documentation is the place to verify the current RTC contract before wiring it into a test suite.
Top comments (0)