Short answer: treat an optimistic update as a proposal, not evidence. In a logistics workspace where a dispatcher and a clinician join a video consultation room, presence is accurate only when the server reconciles client intent with WebRTC connection state, heartbeats, and an expiry policy. The fastest UI is the one that can admit it is provisional.
The real failure is a false green dot
A dispatcher clicks Join, the browser paints online, and the room sends an optimistic participant_added event. Then the camera permission prompt is denied, ICE negotiation stalls, or a laptop changes networks. If the green dot remains, another operator may wait for a person who is not reachable. That is a scheduling error, not a cosmetic glitch.
Not every disconnect is a departure.
I once assumed a WebSocket close was a clean signal. It was not. A mobile client can disappear without a close frame, while a healthy peer connection can outlive the tab that rendered it. The correction is to separate three facts: intent (the user pressed Join), transport liveness (the session still sends heartbeats), and media readiness (the WebRTC connection reaches a usable state). They need different clocks and different owners.
Consider a concrete handoff: at 10:00:00 the dispatcher presses Join, at 10:00:01 the UI shows joining, and at 10:00:03 the membership service accepts the request. At 10:00:04 the clinician browser reports connected, but the dispatcher is already on a train and the heartbeat is delayed until 10:00:27. If the UI promotes on the button click, the room appears healthy for 27 seconds without proof. If it promotes on media alone, a stale membership can still keep the badge green. The server needs a lease, a version, and an explicit confirmation event so those clocks meet in one place; that record also gives support staff an audit trail when operators disagree.
The client may render a pending badge immediately, but the authoritative presence record should carry a monotonically increasing version, an idempotency key, and an expiration timestamp. A late rejection must not overwrite a newer accepted state.
How should a video consultation room reconcile optimistic updates and presence accuracy?
Use a small state machine instead of toggling a boolean. For example, joining can become online only after the server observes a valid room membership and the media layer reports a usable connection. A failed attempt becomes offline with a reason safe to show to the operator; a silent client becomes stale after the lease expires.
Here is a framework-neutral sketch. The command endpoint is deliberately separate from the event stream, so replaying an event cannot create a second membership.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
LEASE = timedelta(seconds=20)
@dataclass
class Presence:
user_id: str
room_id: str
state: str
version: int
lease_until: datetime
def accept_join(current, user_id, room_id, request_id, now):
if current and current.state in {"joining", "online"} and current.lease_until > now:
return current
next_version = (current.version + 1) if current else 1
return Presence(user_id, room_id, "joining", next_version, now + LEASE)
def confirm_media(presence, connection_state, now):
if connection_state in {"connected", "completed"} and presence.lease_until > now:
presence.state = "online"
return presence
The important detail is not the syntax. It is the transition rule: a client acknowledgement can update the display, but only a server-side observation can confirm presence. On reconnect, send the last seen version and request ID; the server can answer with the current snapshot, then resume events from a cursor.
Keep media and membership separate. The W3C WebRTC model exposes connection states such as connecting, connected, disconnected, and failed, but those states do not decide your business lease. A dispatcher might still be reachable by chat while the video path is failed, so collapsing both into offline hides a useful operational distinction.
Failure boundaries worth testing
Test the timeline, not just the happy path. Inject a delayed join response, duplicate commands, out-of-order events, clock skew, browser sleep, a revoked room token, and an ICE failure after the participant was marked online. Advance a fake clock past the lease and verify that the sweeper emits one expiry event. Then reconnect with an old cursor and ensure the snapshot wins over stale local optimism.
A useful invariant is: no presence shown as online lacks a fresh lease and a server-observed membership. Another is: every visible transition has a correlation ID that appears in logs, metrics, and the operator audit trail.
Short tests catch long incidents.
For observability, count optimistic proposals, confirmations, expiries, reconciliation corrections, and media-state changes separately. Alert on the correction ratio and on leases expiring during an active consultation; a raw WebSocket disconnect count cannot tell those stories. Record reason codes without storing video content or unnecessary health information.
Choosing the contract and rollout shape
A single-process in-memory map is adequate for a prototype with one room worker. A multi-region service needs durable membership, a clock policy, and an event delivery contract that tolerates duplicates. At-least-once delivery plus idempotent consumers is usually easier to operate than pretending a network can provide exactly-once effects.
The catch is that lease-based presence intentionally shows stale for a short interval after a real disconnect. That is safer than a false online signal, but it may not suit dispatch decisions that require sub-second certainty; in that case, keep a dedicated signaling or paging channel and treat the video badge as advisory. Stick with a simpler server push model when rooms are tiny and the cost of a missed update is low.
Roll out by logging decisions before changing UI behavior. Compare the proposed state with the authoritative state, sample reconciliation traces, and only then enable optimistic badges for a cohort. Your mileage may vary with mobile sleep policies and regional network paths, so keep lease duration configurable and tie the final choice to observed correction rates rather than a fashionable latency target.
References
- W3C, WebRTC Recommendation: https://www.w3.org/TR/webrtc/
- MDN, RTCPeerConnection connectionState: https://developer.mozilla.org/en-US/docs/Web/API/RTCPeerConnection/connectionState
- RFC 8085, UDP usage guidelines: https://www.rfc-editor.org/rfc/rfc8085
Top comments (1)
Fascinating deep dive into AI workflows. Grounding outputs and evaluating latency budgets will be key as more agentic systems reach production.