Short answer: keep the board's event log authoritative, apply backpressure at the connection boundary, and test state convergence instead of hoping wall-clock sleeps line up. For a healthtech chat room or shared kanban board, presence accuracy is the decision axis: a delayed cursor is annoying, but a stale "online" clinician can misroute work.
The useful contract is small. Every mutation gets an ordered sequence, every reconnect asks for a cursor, and presence is explicitly soft state with a lease. A test should control delivery and acknowledgements with a virtual clock. Three seconds is a policy value, not a timing oracle.
What should realtime backpressure tests prove for a shared kanban board?
Start with invariants, because throughput numbers hide correctness failures:
- A card transition is applied once per sequence number, even after a reconnect.
- A client that falls behind is told to resume from a cursor or fetch a snapshot; it is never silently given a partial stream.
- Presence expires when its lease expires, not when an arbitrary test sleep happens to finish.
- A slow consumer cannot make the server retain unbounded events for every room.
The WebRTC recommendation describes connection and transport behavior, but it does not define your board's ordering or presence policy. Those belong in your application protocol. WebSocket or WebRTC can carry the bytes; neither one decides whether card-17 moved before card-22.
One sentence matters here.
Presence should be modeled as a lease: expires_at = last_heartbeat + lease, with the server clock as the authority. The UI may display "reconnecting" while a lease is still valid, but it must not manufacture an online member after expiry. I'm not sure a five-second lease is right for your network and clinical workflow; measure reconnect distributions, then choose it deliberately.
The architecture decision record
The critical path has four boundaries. The command handler validates a mutation and appends it to durable room storage. A fan-out worker reads the append-only sequence and offers it to each connection. The connection has a bounded queue. Finally, the client acknowledges the highest contiguous sequence it has rendered.
| Choice | Strength | Failure boundary | Use it when |
|---|---|---|---|
| Per-connection bounded queue | Makes memory pressure visible | A slow client must resync | You need predictable resource limits |
| Room-wide queue | Simple fan-out accounting | One slow member can delay everyone | Rooms are tiny and latency is less important |
| Snapshot plus cursor replay | Fast recovery after a gap | Snapshot and replay need one consistent cut | Reconnects are common |
| Drop-oldest presence updates | Keeps transient state fresh | Never safe for card mutations | The message is replaceable soft state |
The rejected option is an unbounded per-client buffer. It looks friendly during a demo, then turns a tablet with a suspended background tab into a memory leak. It is valid only for a bounded, offline export where the producer has a hard item limit and no live room is sharing that buffer.
Here is a deterministic harness. It does not sleep; it advances delivery and the lease clock as explicit events.
from dataclasses import dataclass
from collections import deque
@dataclass(frozen=True)
class Event:
sequence: int
kind: str
payload: dict
class Connection:
def __init__(self, capacity: int = 4):
self.queue = deque(maxlen=capacity)
self.ack = 0
self.needs_snapshot = False
def offer(self, event: Event) -> None:
if event.kind == "presence":
self.queue = deque(
(item for item in self.queue if item.kind != "presence"),
maxlen=self.queue.maxlen,
)
elif len(self.queue) == self.queue.maxlen:
self.needs_snapshot = True
return
self.queue.append(event)
def acknowledge(self, sequence: int) -> None:
if sequence < self.ack:
return
self.ack = sequence
def replay_after(events: list[Event], cursor: int) -> list[Event]:
return [event for event in events if event.sequence > cursor]
The important assertion is not that four items fit. It is that a dropped mutation flips needs_snapshot, while a replaceable presence notice can be coalesced. A test can enqueue five card moves, inspect the flag, load a snapshot at sequence 5, and replay from there with no dependence on scheduler luck. That catches the real bug class: confusing freshness with durability.
How do you remove flaky timing from presence and reconnect tests?
Give the system three clocks in the test: command time, delivery time, and lease time. They can share an integer tick, but each transition must be explicit. A heartbeat at tick 10 extends a lease to tick 13; advancing to tick 13 marks the member offline. No sleep(3) belongs in this test.
Test transitions, not milliseconds:
- Connect client A and issue a heartbeat.
- Pause delivery while appending card moves 1 through 5.
- Advance the lease clock past expiry and assert that presence is offline.
- Reconnect A with cursor 0; assert that the server requests a snapshot because its queue overflowed.
- Install the snapshot at sequence 5, replay later events, and assert one final board hash.
Use a fixed event fixture with two members, six cards, and one duplicate command. The duplicate should produce the same board hash and no second sequence. A 409 for a stale mutation is a useful protocol result; it is different from a transport retry and should be asserted separately. Keep the fixture small enough to read in a code review, then run the same state machine with randomized delivery permutations.
Your mileage may vary across browser transports. The test remains stable because it controls the schedule and checks convergence. Production telemetry should record queue depth, oldest unacknowledged sequence, snapshot frequency, reconnect count, and presence-expiry lag. Those measurements tell you whether the lease or the queue capacity needs changing.
Choosing a policy under clinical workflow constraints
Presence accuracy is not identical to message latency. A board can tolerate a 400 ms card update if the update is ordered and durable; it cannot tolerate an operator being shown as available after the lease has expired. Define separate service objectives and alert on each.
The catch is operational cost. Snapshotting every reconnect increases read load, while long leases improve apparent availability but make offline members linger. A room with ten collaborators may use a short lease and frequent heartbeats; a mobile workflow on an intermittent network may need a longer lease plus a prominent reconnect state. Stick with a simple ordered log when auditability dominates. Choose a CRDT only when concurrent offline edits are a real requirement and the team can operate its merge semantics.
Do not make a vendor choice the test oracle. Put transport adapters behind the same append, resume, snapshot, and heartbeat interface, then run the identical schedule against each adapter. The platform that cannot expose bounded queues, cursor replay, and request-level telemetry is not suitable for this workflow, regardless of its peak message rate.
Release gate
Before shipping, fail the build when any of these assertions fail: no duplicate sequence is applied; no mutation disappears without a snapshot marker; expired presence is not rendered online; reconnecting from a valid cursor converges to the same board hash; and queue depth stays below the configured bound. Run the suite with deterministic seeds and retain the seed on failure so a timing report is reproducible.
This is a boring gate. That's the point. Realtime correctness should be explainable at 02:00, when a reconnect storm is a storage and ordering problem rather than a mystery about which timer fired first.
Top comments (0)