A live auction poll is a small write path with a surprisingly large failure surface. The useful load test is not a crowd of clients clicking at once; it is a crowd that disconnects, reconnects, and asks for the history it missed while bids and poll answers continue to arrive.
Short answer: make the fixture protocol explicit about event IDs, replay windows, and snapshot boundaries, then test reconnect and backfill as first-class API operations rather than hoping a WebSocket library hides them.
How should realtime load test fixtures model API boundaries for a live auction dashboard?
Start with a contract that can survive a broken connection. A client subscribes to a session, receives a monotonically increasing event ID, and records the highest contiguous ID it has applied. On reconnect it sends that cursor. The server either replays events after the cursor or returns a snapshot plus the first event ID that follows the snapshot. Those are different responses and the fixture must tell them apart.
For a media company running a poll during a live auction session, the event stream might contain poll_opened, answer_counted, and poll_closed. A bid feed can share the transport, but it should not share assumptions about ordering: a late answer may be valid after a bid, while a bid with an old sequence number is usually stale. Keep ordering guarantees scoped to a session and stream, and document what is merely best effort.
Here is a compact fixture model. It uses a generic HTTP boundary for backfill so the same assertions work behind WebSocket, Server-Sent Events, or WebRTC data channels.
from dataclasses import dataclass
@dataclass
class ReconnectCase:
session_id: str
last_event_id: int
expected_mode: str # "replay" or "snapshot"
def assert_backfill(response: dict, case: ReconnectCase) -> None:
mode = response["mode"]
assert mode == case.expected_mode
if mode == "replay":
ids = [event["id"] for event in response["events"]]
assert ids == sorted(ids)
assert all(event_id > case.last_event_id for event_id in ids)
else:
assert response["snapshot"]["session_id"] == case.session_id
assert response["next_event_id"] > case.last_event_id
The fixture should generate a gap deliberately. Disconnect after event 120, publish events 121 through 137, reconnect with cursor 120, and verify that applying the response produces the same state as a continuously connected client. Repeat with a cursor older than the retention window. A correct system will force a snapshot in that case; silently returning an empty replay creates a dashboard that looks healthy while displaying obsolete totals.
That last case is where many load tests lie. They measure connected throughput and never measure recovery correctness.
The cursor is the product.
Failure modes that matter more than peak messages per second
The first failure mode is duplicate delivery. At-least-once transports make duplicates normal, so event application needs an idempotency key such as (session_id, event_id). The second is a split brain between snapshot and replay: if the snapshot is read at time T but the replay cursor is chosen at T+delta, events can be skipped. Define a snapshot watermark and include it in the response.
The third is a retention mismatch. A 30-second replay cache does not satisfy a viewer who backgrounds a phone for five minutes. Measure the age of the oldest reconnect in the fixture, not just the number of connected sockets. Also test a session ending while a reconnect is in flight; the client should receive a terminal state and stop retrying instead of resurrecting a closed poll.
One particularly revealing trace starts with a clean connection at event 120. The browser loses radio coverage for 42 seconds while the auction host closes the poll and reopens a second one with the same session identifier. Events 121-146 now include both the terminal marker and the new poll's opening state. If the reconnect handler treats the cursor as a transport offset only, it can apply the second poll's answers to the first poll, or discard the terminal marker as a duplicate. The fixture must therefore assert stream identity and poll generation as well as event ID. I keep those fields in the test record even when the production payload does not expose them in the UI; they are the only practical explanation when a count differs after recovery. This is a longer test than a throughput loop, but it mirrors the failure that viewers actually report.
Small detail. Big consequence.
Use distinct metrics for transport and data correctness: reconnect latency, replay count, snapshot count, duplicate suppression, cursor gaps, and convergence time. A low error rate says little if convergence time is unbounded. I would rather see a controlled snapshot than a fast response that cannot prove which answers it contains.
Choosing a boundary: transport, broker, or storage
The transport is responsible for liveness and framing. A broker or append-only log is responsible for ordering and replay. Durable storage is responsible for the session snapshot and audit trail. Collapsing all three into one component makes a demo tidy and an incident opaque.
| Boundary choice | Strength | Cost or limit | Good fit for the poll |
|---|---|---|---|
| In-memory replay ring | Very low latency | Gaps after restart or eviction | Short reconnects during a single session |
| Durable append-only log | Clear cursor semantics and auditability | More I/O and retention work | Regulated results or long pauses |
| Snapshot plus ephemeral events | Bounded replay work | Snapshot watermark protocol is easy to get wrong | High fan-out dashboards |
| Peer data channel | Can reduce server fan-out for some media paths | Signaling, NAT, and delivery guarantees remain your problem | Small, cooperative audiences |
WebRTC is a transport standard, not a durable event-history specification. Its recommendation describes data channels, signaling, and connection behavior, but your application still has to define cursor exchange, authorization, and replay. That distinction is useful: it keeps a reconnect bug in the API contract where a load test can expose it, rather than attributing it to the wire protocol.
A rollout sequence that keeps the fixture honest
Begin with deterministic data: 1,000 sessions, a fixed poll vocabulary, and seeded disconnect points. Run a single-client golden test first. Then add concurrency while preserving the same event traces, so a change in result is attributable to scheduling or capacity instead of random input.
Inject one fault at a time: dropped connection, delayed replay, duplicate event, expired cursor, and session close. Capture the cursor and snapshot watermark in every failure artifact. Without those two fields, a retry log is mostly theater.
Finally, compare the recovered state with an offline fold of the event log. The dashboard is correct only when both folds agree. I'm not sure any team can infer that guarantee from a vendor's “realtime” label; the protocol and the fixture have to make it observable.
The catch is operational: durable replay and audit storage add retention policy, privacy review, and deletion jobs. For a disposable internal poll, an in-memory ring may be the right answer. Stick with a durable log when results affect payouts, compliance, or post-session editing. The boundary should follow the consequence of a missed event, not a benchmark headline.
Top comments (0)