A live auction has an unforgiving constraint: every bidder must converge on the same accepted high bid, even when connections pause, reconnect, or deliver updates at different times. The choice is Postgres-backed WebSockets for authoritative bid fan-out, with a monotonically increasing auction sequence and replay on reconnect. WebRTC data channels can serve latency-sensitive peer interaction, but they are the wrong authority boundary for this job because their delivery properties are configured per channel and a peer topology complicates one global order.
Short answer: when two clients show different high bids, stop comparing their wall clocks. Compare the highest contiguous server sequence each client applied, the durable sequence at the authority, and the oldest retained event. That separates an ordering defect from a missed-message defect in minutes, while latency histograms alone cannot.
Why can two clients show different winning bids?
The screen is a projection, not the ledger. A bid becomes authoritative only after one serialization point accepts it and assigns its position. Before that point, network arrival order says little: two bid requests can traverse different connections, a retry can arrive after its original request, and browser timestamps can disagree because wall clocks are neither synchronized nor monotonic across machines.
After acceptance, divergence usually enters through one of four paths. The fan-out service may publish before the transaction commits, so viewers briefly see a bid that the durable state cannot reproduce. A client may apply event 418 after 420 because callbacks from separate streams race. A reconnect may resume at 421 even though 419 and 420 were never applied. Or duplicate delivery may cause a non-idempotent reducer to update derived state twice. These are different failures; calling all of them "WebSocket lag" hides the evidence needed to repair them.
The invariant should be small enough to put on a dashboard: for one auction_id, accepted events have a unique contiguous auction_seq; a client applies event N only after N - 1, or deliberately requests a snapshot plus a later replay boundary. The high bid shown is the result of reducing that ordered prefix.
That last word matters. Prefix. A client at sequence 418 and a client at 420 can disagree temporarily without either inventing data, but the first client must know it is behind and must have a bounded route to convergence. Silent gaps are the real defect.
No gap, no guess.
Put order beside durable state
Use one database transaction to lock the auction row, validate the new amount against current state, advance its sequence, write the accepted bid, and append an outbox event. PostgreSQL documents row-level locking and transaction isolation behavior; the design still needs retry handling because concurrent transactions can abort under stricter isolation. The outbox publisher reads committed rows, publishes them, and records progress. Publishing can therefore be at least once, so consumers must deduplicate.
Here is the shape of the authoritative operation. It is deliberately generic Python, not an SDK recipe:
def accept_bid(db, auction_id, bidder_id, amount, request_id):
with db.transaction() as tx:
prior = tx.fetch_one(
"SELECT result FROM bid_requests WHERE request_id = %s",
(request_id,),
)
if prior is not None:
return prior["result"]
auction = tx.fetch_one(
"SELECT high_bid, auction_seq FROM auctions "
"WHERE auction_id = %s FOR UPDATE",
(auction_id,),
)
if amount <= auction["high_bid"]:
result = {"accepted": False, "reason": "bid_not_high_enough"}
else:
next_seq = auction["auction_seq"] + 1
tx.execute(
"UPDATE auctions SET high_bid = %s, auction_seq = %s "
"WHERE auction_id = %s",
(amount, next_seq, auction_id),
)
event = {
"type": "bid.accepted",
"auction_id": auction_id,
"auction_seq": next_seq,
"request_id": request_id,
"bidder_id": bidder_id,
"amount": str(amount),
}
tx.insert("auction_outbox", event)
result = {"accepted": True, "auction_seq": next_seq}
tx.insert("bid_requests", {"request_id": request_id, "result": result})
return result
The request_id handles submission retries; (auction_id, auction_seq) handles event deduplication and gap detection. They solve separate problems. A UUID alone does not express order, while a sequence alone cannot tell whether two repeated submissions are one intent.
Do not promise exactly-once delivery. The useful promise is narrower and testable: each accepted bid is committed once for a request id, the stream can redeliver it, every client reducer is idempotent, and any sequence gap triggers recovery. Durability belongs to the ledger and outbox, not to the lifetime of a socket buffer.
Delivery guarantees decide the transport
WebSocket, standardized by RFC 6455, provides a two-way connection over TCP. TCP preserves byte order within one connection, but that does not create global order across several publishers or across a disconnect. Application sequence numbers remain necessary.
WebRTC data channels use SCTP over DTLS and can be configured for ordered or unordered delivery, with reliable or partially reliable retransmission behavior. Those controls are valuable for media-adjacent state where an obsolete update may be discarded. They are dangerous defaults for accepted bids: partial reliability can intentionally abandon data, and per-channel order still is not an auction-wide serialization rule.
| Constraint | WebSocket to an authoritative fan-out | WebRTC data channels |
|---|---|---|
| One global bid order | Still requires server sequence; fits a central authority | Still requires an authority above peer or channel order |
| Reconnect recovery | Application resume token and replay are required | Application replay is also required |
| Selective loss tolerance | Usually handled with separate message classes or connections | Configurable reliability and ordering per data channel |
| Topology and signaling | Direct client-to-service connection | Requires signaling; connectivity commonly involves ICE candidates and may use relays |
| Best fit here | Accepted bids, close state, and resumable projections | Ephemeral reactions or hints that do not determine the winner |
The choice is not a claim that WebSocket is intrinsically more durable. It is not. Both transports need an application ledger and recovery protocol. The choice follows from ownership: a live auction already needs a single acceptance decision, so sending its ordered log from that authority creates fewer independently ordered paths.
There is a condition boundary. If the room only exchanges disposable cursor positions or audience reactions, an unordered or partially reliable data channel can avoid waiting behind stale messages. Keep that traffic off the accepted-bid stream. Mixing the two delivery classes invites a configuration change intended for reactions to weaken the auction path.
This approach has limitations: a central WebSocket fan-out is not suitable for disposable peer-to-peer updates, and its trade-off is the operational burden of a replay store, connection backpressure, and a highly available authority. Choose an ordered WebRTC data channel when direct peer exchange is the actual requirement and every participant can tolerate the absence of one server-wide order; choose partial reliability only when an older message becoming irrelevant is part of the data model. Neither choice removes the need for signaling, authentication, scoped room tokens, or authorization at bid acceptance.
That is the trade-off.
Metrics that distinguish lag from loss
Start with sequence gauges, not averages. For every active auction, record authority_committed_seq, fanout_published_seq, and each connection's client_applied_seq. The useful derived values are publication lag and client gap. Avoid bidder identifiers in metric labels; auction IDs can also create unbounded cardinality, so retain per-auction detail in sampled traces or logs and aggregate metrics by room shard or status.
A compact diagnostic payload makes support evidence comparable:
def connection_diagnostics(state, authority_seq, retained_from_seq):
return {
"connection_id": state.connection_id,
"auction_id": state.auction_id,
"client_applied_seq": state.applied_seq,
"authority_committed_seq": authority_seq,
"gap_size": max(0, authority_seq - state.applied_seq),
"oldest_replayable_seq": retained_from_seq,
"duplicate_events_total": state.duplicate_events,
"gap_events_total": state.gap_events,
"reconnects_total": state.reconnects,
}
Add a histogram from commit time to client acknowledgment, plus counters for replay requests, replay misses, duplicate events, reducer rejections, and snapshot fallbacks. Measure end-to-end latency by carrying a server commit timestamp, but use it for latency, not ordering. Client clock timestamps can help reconstruct local UI behavior only when their uncertainty is understood.
Then debug by finding the first boundary where sequences diverge. If committed is 420 and published is 418, inspect the outbox publisher, including its last durable cursor and any row it repeatedly rejects. If published is 420 but a client acknowledged 418, inspect connection backpressure, disconnect history, and the resume token actually received by the service rather than the value the UI claims it sent. If the client acknowledged 420 yet renders the amount from 418, the transport worked; capture the reducer inputs in sequence order and inspect view state. If a client requests 419 but retention begins at 425, replay cannot close the gap, so return a snapshot whose included sequence is explicit and then replay only events after that sequence. A latency chart can look healthy in all four cases because it samples messages that arrived; the missing sequence is the evidence it never observed.
Backpressure needs a policy before launch. Set a bounded per-connection queue; when it fills, close the connection with an application-recognizable reason and force replay or snapshot recovery. Dropping an accepted bid while keeping the socket apparently healthy converts overload into undetectable corruption.
Test gaps rather than happy paths
A useful test does more than send bids quickly. It pauses one subscriber after sequence 103, accepts 104 through 117, reconnects with last_applied_seq=103, duplicates 109, and verifies that the final projection matches the authoritative snapshot. Another test commits an outbox row, interrupts publishing before its delivery marker is stored, and verifies that redelivery changes no result.
Also test the uncomfortable boundary where replay retention has expired. The server should reject the stale resume point, send or direct the client to a snapshot tagged with sequence N, and continue at N + 1. During deployment, old and new clients must agree on event versioning and unknown-field behavior; otherwise a rolling release can look exactly like packet loss.
Three checks deserve release gates: no displayed winner comes from an uncommitted event; every injected gap is detected; and reconnect convergence completes within the service's declared objective under the tested load. The objective is a local engineering decision, so inventing a universal millisecond threshold would be dishonest. Capacity tests should discover it.
Roll out without changing the winner
Introduce sequence fields and client acknowledgments in observation mode first. Compare the client-reported sequence with the authority while the existing projection remains unchanged. Next, enable idempotent reduction and gap alarms, then enable replay for a small room cohort. Snapshot fallback comes before enforcing bounded queues because forced reconnects without recovery merely make missing updates more repeatable.
Finally, separate ephemeral room traffic from authoritative auction events and document their different delivery contracts. Choose the central WebSocket stream when an event can change the winner; choose a data channel only when losing or superseding that event cannot change the auction result. That rule is more durable than a transport benchmark, and it gives operators a concrete place to look when two screens disagree.
References
- https://www.rfc-editor.org/rfc/rfc6455
- https://www.rfc-editor.org/rfc/rfc8831
- https://www.w3.org/TR/webrtc/
- https://www.w3.org/TR/webrtc-stats/
- https://www.postgresql.org/docs/current/explicit-locking.html
- https://www.postgresql.org/docs/current/transaction-iso.html
Sources
The protocol, metrics, and database references above are the primary sources for the transport and transaction boundaries used in this design.
Top comments (1)
Some comments may only be visible to logged-in visitors. Sign in to view all comments.