DEV Community

YukiKobayashi880
YukiKobayashi880

Posted on

Realtime Testing for Shared Kanban Poll Recovery: Four Gates Beyond Timing

Realtime Testing for Shared Kanban Poll Recovery: Four Gates Beyond Timing

A customer-support team running a live poll during a session has a nasty constraint: an agent can lose a connection while the poll is changing, then reconnect after several updates have already happened. The test must prove the recovered board is correct without waiting for wall-clock timing.

Short answer: model connection drain as an explicit state transition, persist an ordered event cursor, and assert the post-reconnect snapshot plus backfill contents; use virtual time only for deadlines, never as evidence that a message arrived.

Start with the invariant, not the socket

For a support poll, the useful invariant is simple: every participant sees one valid projection of the same ordered poll history. A card can move from “open” to “pending,” and a vote can change from “yes” to “no,” but a reconnect must not create a second move or erase an earlier vote. The transport is an implementation detail below that invariant.

I write the test around four values: an event sequence, a durable cursor, a projection version, and a connection state. The sequence is monotonic per board. The cursor is the last sequence included in the client projection. The version identifies the schema used to build that projection. Connection state is deliberately boring: connected, draining, disconnected, or catching_up.

That naming matters. “Wait 200 ms and expect three messages” describes a scheduler, not a contract. A busy CI worker can deliver the messages in 20 ms or 2 seconds while the contract remains identical.

No clocks.

Here is a small in-memory model for a poll board. It has no sleeps, network calls, or test-only branches.

from dataclasses import dataclass

@dataclass(frozen=True)
class Event:
    seq: int
    kind: str
    card_id: str
    value: str

class BoardLog:
    def __init__(self):
        self.events = []

    def append(self, kind, card_id, value):
        seq = len(self.events) + 1
        event = Event(seq, kind, card_id, value)
        self.events.append(event)
        return event

    def after(self, cursor):
        return [event for event in self.events if event.seq > cursor]

class Projection:
    def __init__(self):
        self.cursor = 0
        self.cards = {}

    def apply(self, event):
        if event.seq != self.cursor + 1:
            raise ValueError("out_of_order_event")
        self.cards[event.card_id] = event.value
        self.cursor = event.seq
Enter fullscreen mode Exit fullscreen mode

The one-line failure is intentional. If a test feeds sequence 4 before sequence 3, it fails on the invariant instead of becoming a flaky race.

What should realtime connection drain testing assert?

A drain is not the same as a clean close. During deploy or browser handoff, the server may stop accepting new writes on a connection while already-buffered events still need acknowledgement. The client should finish or abandon that stream, record its cursor, and open a new stream that asks for events after that cursor.

The test therefore has observable checkpoints rather than timing guesses:

  1. The old stream enters draining and emits no new application command.
  2. The last acknowledged sequence is stored.
  3. A new stream authenticates and requests a backfill from that sequence.
  4. Backfill is applied exactly once before live delivery resumes.

For a live customer-support poll, seed events 1 through 5, drain the connection after sequence 3, then append votes 6 and 7 while the client is disconnected. The expected result is a projection at cursor 7, with each card's final value determined by the last event for that card. No sleep is needed; the harness controls delivery order directly.

def reconnect_and_backfill(log, projection, acknowledged):
    missing = log.after(acknowledged)
    for event in missing:
        projection.apply(event)
    return projection.cursor

log = BoardLog()
projection = Projection()
for value in ("open", "pending", "open", "yes", "no"):
    event = log.append("poll_update", "case-17", value)
    projection.apply(event)

acknowledged = 3
log.append("poll_update", "case-17", "yes")
log.append("poll_update", "case-22", "no")
assert reconnect_and_backfill(log, projection, acknowledged) == 7
assert projection.cards == {"case-17": "yes", "case-22": "no"}
Enter fullscreen mode Exit fullscreen mode

Notice that the example does not claim a particular WebRTC data-channel behavior. The W3C recommendation defines the browser API and its connection states; your application still owns event ordering, idempotency, and replay policy. That boundary is where many “realtime” tests quietly cheat.

Four failure modes that timing hides

The first failure is a lost cursor. If the client acknowledges a message before its local projection is committed, reconnecting from that cursor skips data. Ack after projection commit, and test a process crash between those operations.

The second is duplicate backfill. A reconnect can receive sequence 6 from replay and then sequence 6 again from a buffered live frame. Apply by sequence, or carry an idempotency key and make duplicate application a no-op.

The third is stale authorization. A support agent may still have a browser tab open after leaving a queue. Backfill must re-check board membership; a valid cursor does not grant access.

The fourth is an unbounded log. Keeping every poll mutation forever makes replay slower and retention harder to reason about. Compact snapshots only when you can prove that the snapshot cursor and subsequent events produce the same projection.

I once started with a timer-based test because it looked realistic. It passed locally, then failed with error code E_TIMEOUT on a loaded runner. The useful discovery was that the test had no assertion about the cursor at all. I replaced the timer with a delivery script and found a real duplicate event, then kept replaying the same seven-event support poll while changing only the order of acknowledgements, authorization revocation, and browser visibility changes; that longer exercise exposed that our snapshot writer could lag the event log even though every transport callback had returned successfully. The fix was to make the snapshot cursor part of the commit and to compare it with the acknowledged cursor after each drain. Your mileage may vary when browser scheduling is involved, so keep one small end-to-end timing test, but put the correctness suite on explicit queues.

Choosing a transport and test boundary

Transport choice should follow the failure you can observe. WebSockets are a common fit for ordered application frames, while WebRTC data channels can be useful when peers already negotiate a session; WebTransport offers HTTP/3 streams and datagrams but introduces different server and proxy requirements. None of these choices supplies durable backfill by itself.

Boundary Strength Cost or limit
In-process event bus Fast deterministic tests Does not exercise browser or proxy behavior
WebSocket integration Exercises framing and reconnect Needs a controllable server clock and network harness
WebRTC data channel Tests peer negotiation and channel state Signaling and NAT paths add moving parts
WebTransport session Tests HTTP/3 stream semantics Operational support varies across clients and gateways

The practical split is two suites. A large model-based suite drives event queues, drains, reorders, duplicates, and authorization changes. A smaller browser suite checks that the chosen API reaches the expected state transitions in a real user agent. The browser suite should consume a prepared event log so it is not also trying to prove your database.

Roll out a deterministic reconnect contract

Document the reconnect request as a contract: board identifier, authenticated principal, last applied sequence, projection version, and a maximum backfill window. Return a snapshot plus a bounded event range, or return a clear “resync required” result when compaction removed the requested cursor. That is a capability boundary, not a server failure.

Instrument the contract with counters for drain starts, reconnects, duplicate sequences, rejected cursors, and resyncs. Alert on a rising resync rate, not on a single slow message. During rollout, mirror production event shapes into the deterministic harness and compare projections by cursor.

The catch is that this design is not suitable when your product cannot retain an ordered change log or when updates are intentionally lossy, such as a transient typing indicator. For those signals, subscribe to the latest state and accept that reconnect means refresh. Stick with a simpler polling endpoint when the poll changes only a few times per session and the cost of a durable stream exceeds the user impact.

The decision is consequently about recovery semantics, not a fashionable protocol. If a support poll must survive a connection drain, make the cursor and replay path first-class, then let the transport earn its place through a small, honest browser test.

References

Top comments (0)