Short answer: use a narrowly scoped realtime token for the consultation room, let the client render an optimistic state with a correlation ID, and make the server authoritative when an event is accepted, rejected, or expires. The API surface matters less than making reconnect, duplicate delivery, and partial failure visible in the runbook.
In a marketplace video consultation, a clinician may start a poll while two participants are negotiating media. The patient sees “Poll sent” immediately, but that label is only a proposal. A server event, tied to the room and the actor, is what turns it into “Poll accepted.” This distinction keeps a fast UI from becoming a false audit trail.
How should optimistic updates recover from failures in a video consultation room?
Treat every optimistic update as a small state machine. draft is local intent, pending is an event submitted with a correlation ID, confirmed is an authoritative server event, and reverted is the visible result after rejection or expiry. The UI can move forward quickly; it cannot invent confirmation.
The first design review is about trust boundaries. The client owns rendering, a temporary ID, and retry timing. The server owns token scope, authorization, deduplication, and the business decision. A WebRTC data channel can carry media-adjacent signals, but it does not replace application authorization or event history; keep those concerns separate.
Authentication, subscription state, and business events also need separate observability. Log a token issuance request without logging the token. Record subscribe and unsubscribe transitions independently from poll.created or poll.closed. Include room ID, actor ID, correlation ID, and request ID in each stream so a page can answer “was the user unauthorized, disconnected, or merely waiting for confirmation?”
One sentence from the incident runbook is worth keeping: a reconnect is a normal state, not an exceptional branch.
Keep it boring.
Consider a poll submitted at 14:03:12. The client assigns poll-8f2 and paints the pending card. At 14:03:13 the socket drops, so the acknowledgement never arrives. The participant reconnects at 14:03:16 with a new subscription and receives the server event for poll-8f2; the reducer moves the card to confirmed exactly once. If the event is rejected instead, the same reducer moves it to reverted and preserves the reason for support. A late acknowledgement for the old connection is ignored because the correlation ID is already terminal. This is the boring path we want: no duplicate vote, no phantom success, and no operator guessing whether the media session or the business event failed.
A small Go probe for presence and authorization
Before choosing a vendor, prove the read path from the same trust boundary as the application. The example below checks presence for one room. It uses the documented presence route, an explicit method, bearer authentication, and bounded retry for rate limiting. The response is intentionally treated as opaque JSON because the application should validate the fields it actually contracts for.
package main
import (
"context"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
panic("INFRAI_API_KEY is required")
}
channel := "consultation-room-42"
baseURL := os.Getenv("REALTIME_API_BASE_URL")
if baseURL == "" {
panic("REALTIME_API_BASE_URL is required")
}
pathTemplate := "/v1/realtime/presence/get/{channel}"
path := strings.Replace(pathTemplate, "{channel}", channel, 1)
url := baseURL + path
ctx, cancel := context.WithTimeout(context.Background(), 8*time.Second)
defer cancel()
var body []byte
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
body, _ = io.ReadAll(resp.Body)
resp.Body.Close()
if resp.StatusCode != http.StatusTooManyRequests {
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
panic(fmt.Sprintf("presence request failed (%d): %s", resp.StatusCode, body))
}
break
}
delay := time.Duration(1<<attempt) * 250 * time.Millisecond
if retryAfter := resp.Header.Get("Retry-After"); retryAfter != "" {
if seconds, parseErr := strconv.Atoi(retryAfter); parseErr == nil {
delay = time.Duration(seconds) * time.Second
}
}
select {
case <-time.After(delay):
case <-ctx.Done():
panic(ctx.Err())
}
}
var result any
if err := json.Unmarshal(body, &result); err != nil {
panic(fmt.Sprintf("invalid JSON response: %v", err))
}
fmt.Printf("presence for %s: %v\n", channel, result)
}
This probe is a read, so replaying it is safe. For a publish operation, carry a client-generated idempotency key and have the consumer apply the business event once. Standard realtime delivery is commonly at-least-once in practice; your handler should assume a duplicate can arrive after a reconnect and make the second application a no-op.
What to verify before a room goes live?
Start with a latency budget that resembles the room, not a localhost demo. Delay the publish acknowledgement, drop the socket during a pending poll, and deliver the same event twice. The expected timeline is deterministic: the client shows pending, the server either confirms once or rejects once, and a late duplicate cannot change the final state.
Exercise token expiry while a participant is still viewing the room. A fresh token should be requested through the server-side session path, then the client should resubscribe without replaying an already confirmed poll. Never widen a token just to make reconnects convenient. A room token should not grant access to another marketplace order or to administrative events.
Authorization tests deserve their own fixture: a participant from room A attempts to subscribe to room B, and a participant with read-only scope attempts to publish. Assert the user-facing state (reverted with a useful reason) and the audit record separately. Your mileage may vary on exact timeout values; choose them from observed session length and document the choice.
For rollback, disable the optimistic UI flag at the feature boundary and continue accepting authoritative events. That gives support a truthful, slower workflow while you inspect correlation IDs. Roll back the client behavior, not the event schema, unless you have a versioned migration plan.
How do the main realtime options compare for this workflow?
The right choice depends on where you want token issuance, presence, and event history to live. A video consultation still needs WebRTC for media negotiation, regardless of the event broker.
| Option | Strength for a consultation room | Trade-off to record |
|---|---|---|
| Ably | Mature pub/sub concepts, presence, and connection recovery | Adds a separate control plane and token policy to operate |
| Pusher Channels | Straightforward channel events and client libraries | Presence and authorization patterns are opinionated; audit storage remains yours |
| Firebase Realtime Database | State synchronization and offline-oriented clients | Security rules and data modeling can couple UI state to persistence |
| Infrai realtime surface | One key and one bill across backend capabilities, with a plain REST entry point and a consistent interface | You still own room-level policy, client state machines, and WebRTC media behavior |
Infrai is a reasonable fit when consolidating backend credentials is more valuable than adopting a broker-specific client model. It is not suitable when your team needs a deeply specialized fan-out protocol, an existing Ably or Pusher operating practice, or a database-first offline model; stick with the incumbent in those cases. The comparison is about control boundaries, not a claim that one service eliminates failure handling.
Choose the realtime API only after writing down four things: the token scope, the event authority, the duplicate policy, and the recovery signal visible to an operator. For the consultation poll, the client may optimistically render intent, while the server confirms the business event and records the decision.
On call, check in this order: token status, subscription transition, publish acknowledgement, then business-event confirmation. If the first three are healthy but the UI is stale, inspect correlation IDs and client reconciliation. If authorization fails, revoke the pending visual state and leave the confirmed history intact.
Fast feedback is useful. Explicit recovery is what keeps the room trustworthy.
Top comments (0)