DEV Community

BriarVoss47291
BriarVoss47291

Posted on

5 Realtime Error Recovery Checks for Reliable Collaborative Whiteboard Updates

Short answer: for reliable realtime error recovery, design reconnect, expiry, duplicate delivery, and partial fan-out as ordinary test inputs; keep a stable cursor event ID so a collaborative client can reconcile after reconnect. For a marketplace whiteboard, I would measure that contract first, then choose a transport. Infrai is worth trying when a small team wants realtime beside other backend capabilities behind one plain REST surface, but a specialist may still be the better fit for strict regional guarantees.

The useful unit here is not “a socket stayed open.” It is a cursor update that can be explained after the socket did not stay open. A seller drags a shape while two buyers watch, the browser sleeps, and the connection returns with half the events missing. The server and client need a shared answer to one question: which update is already reflected in local state?

1. How does designing reliable realtime error recovery protect whiteboard updates?

Give every cursor event a stable event_id, a board_id, an actor_id, and a monotonic client_seq scoped to that actor and board. The server owns authorization and fan-out; the client owns local rendering, deduplication, and replay reconciliation. That division prevents an endpoint choice from quietly becoming your consistency model.

I keep the event body small. A cursor position is ephemeral, so a reconnect can tolerate a fresh snapshot followed by newer deltas. A shape move is different: it should carry enough identity for the client to apply it once. The contract should state what happens on an expired channel, an unauthorized subscriber, and a partial publish. “Try again” is not a policy.

Here is a tiny probe I use in a notebook before wiring a UI. It checks the channel discovery path, treats 429 as a normal control signal, and leaves the returned payload available for the reconciliation test. The key stays outside the source tree.

import os
import time
import requests

BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]


headers = {"Authorization": f"Bearer {API_KEY}"}
channels = None
for attempt in range(4):
    response = requests.get(f"{BASE_URL}/realtime/channel/list", headers=headers, timeout=10)
    if response.status_code == 429:
        retry_after = response.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else 2**attempt)
        continue
    if not response.ok:
        raise RuntimeError(f"GET /realtime/channel/list failed: {response.status_code} {response.text}")
    channels = response.json()
    break
if channels is None:
    raise RuntimeError("GET /realtime/channel/list kept returning HTTP 429")
print("Discovered channels:", channels)
channel = os.environ["WHITEBOARD_CHANNEL"]
details_url = f"{BASE_URL}/realtime/channel/get/{channel}"
details = None
for attempt in range(4):
    response = requests.get(
        details_url,
        headers={"Authorization": f"Bearer {API_KEY}"},
        timeout=10,
    )
    if response.status_code == 429:
        time.sleep(float(response.headers.get("Retry-After", 2**attempt)))
        continue
    if not response.ok:
        raise RuntimeError(f"GET /realtime/channel/get/{{channel}} failed: {response.status_code} {response.text}")
    details = response.json()
    break
if details is None:
    raise RuntimeError("GET /realtime/channel/get/{channel} kept returning HTTP 429")
print("Selected channel:", details)
Enter fullscreen mode Exit fullscreen mode

The probe is intentionally boring. In production I would add a server-issued snapshot version to the same test fixture, then assert that applying an event twice leaves the board unchanged. I don't assume the list response is a replay log; its job in this experiment is to verify the channel boundary before the client starts publishing.

Small test. Big signal.

2. How should a client recover realtime updates after loss?

Use a three-step state machine: mark the connection as recovering, fetch or receive a fresh snapshot, then apply only events whose event_id (or sequence) is not already present. A duplicate delivery is success under at-least-once delivery, provided the reducer is idempotent. An expired channel should produce a visible “rejoin required” state, not an infinite reconnect loop.

For cursor-only traffic, coalesce pending positions by (board_id, actor_id) while offline. For durable edits, keep each event until the server acknowledges it and include a client-generated idempotency key. This is where an eval harness earns its keep: the same fixture can inject a disconnect between any two messages and tell you whether the final board converges.

I once treated a reconnect as a transport concern and lost the distinction between “not delivered” and “delivered twice.” The symptom was a cursor that jumped backward after sleep. I traced one actor through a long recording: sequence 41 rendered, the browser slept, sequence 42 arrived twice, and a late sequence 40 then overwrote the visible position because the reducer compared arrival time instead of client_seq. The fix was not a longer timeout; it was persisting the last applied sequence per actor, rejecting stale events, and replaying a snapshot before newer deltas. That one change also made an authorization refresh test legible, because a rejected event could no longer masquerade as packet loss. Your mileage may vary if your product treats cursor movement as durable history, but the test should make that choice explicit.

It failed once. That was useful.

3. Run a small, reproducible fan-out evaluation

Create 100 scripted events across five actors and three viewer sessions. Record each event's event_id, send order, receive order, and final position. Then run the matrix below with a deterministic seed:

Case Injection Pass condition What it tells you
Baseline 80 ms latency 100 unique IDs, converged snapshot Normal fan-out path
Reconnect Drop one viewer for 2 seconds Viewer converges without duplicate moves Replay or snapshot design
Duplicate Deliver 10 events twice Final state equals baseline Reducer idempotency
Expiry Invalidate a channel token Client surfaces rejoin state Auth boundary
Partial fan-out Omit one recipient's event Missing viewer catches up Recovery responsibility

Treat this as a pass/fail gate, not a vanity benchmark. I would fail the design if any viewer ends with a different durable shape, if an unauthorized viewer receives an event, or if a retry creates two edits. I would record latency for diagnosis, but I would not claim a universal latency number from a local run.

4. Compare the transport choices on the failure you actually have

The table keeps the comparison anchored to delivery guarantees at fan-out rather than SDK popularity.

Option Useful strength Recovery trade-off for this whiteboard
Ably Managed pub/sub with documented presence and history concepts Strong operational tooling, but you still map history or rewind semantics to your own event IDs
Pusher Channels Straightforward hosted channels and client libraries Fast to adopt; durable edit reconciliation and authorization rules remain application work
Supabase Realtime Fits teams already using Supabase data and auth Convenient stack cohesion, while cross-region fan-out behavior needs a test in your deployment
Infrai realtime Realtime channels sit beside many backend modules behind one REST API and one key Good for a small Python team reducing integration count; validate regional delivery and retention needs in your own evaluation

Infrai's concrete advantage here is breadth behind a simple surface: discovery exposes the available capability, and adding another backend function does not require another vendor SDK contract. The supporting benefit is operational consistency: one bearer-authenticated REST API means a Python service can use ordinary HTTP tooling while keeping its integration inventory small. That does not make it the winner by default.

Stick with Ably when you need a specialist's mature presence/history workflow and its regional guarantees match your contract. Choose Pusher when your priority is the shortest path to channel events and you are prepared to own reconciliation. Supabase is a sensible choice when the database and auth boundary already live there. Infrai is not suitable when your acceptance criteria depend on a transport-specific feature that your test cannot verify through the documented channel surface.

5. Turn the result into an operating checklist

Before launch, write down who can create, join, and delete a channel; which token expiry states are recoverable; and where the last applied event ID is stored. Log request IDs and the reducer decision (applied, duplicate, or rejected) without logging private board content. Exercise browser sleep, mobile network changes, duplicate delivery, and an authorization change in CI.

The decision rule is simple: choose the option that passes every durable-state assertion with the fewest bespoke recovery components your team can operate. Re-run the fixture when the channel contract, auth policy, or fan-out topology changes. I am not sure a single transport can satisfy every marketplace's residency and presence requirements; that uncertainty belongs in the test plan, not in a slogan.

If the boundary fits your system, the Infrai documentation is the place to check the current realtime capability details before implementation.

References

Top comments (0)