Use one host as the authoritative playback clock, and treat host departure as a planned state transition. For a media workspace that shows who is online while editors review the same cut, peer consensus adds a distributed-systems problem without improving the part viewers actually judge: whether play, pause, and seek converge quickly enough to feel shared.
TL;DR: send intents to the host, publish monotonically numbered snapshots, and let clients correct toward the latest snapshot. Presence is evidence that a connection is currently visible, not durable membership and not permission to elect a leader. If the host leaves, pause the room, elect or appoint a replacement under an explicit policy, increment the epoch, and resume. Perfect synchronization is unavailable; the interface should say “syncing” when drift crosses the product's chosen tolerance instead of pretending every screen is on the same frame.
This is an architecture decision record, not a claim that a central host cannot fail. It can. The decision is that a small, named failure boundary is easier to test and explain than consensus spread across browsers, mobile networks, background tabs, and players with different buffering behavior.
1. Should a Watch Party Use an Authoritative Host or Peer Consensus?
The first invariant is single-writer authority within an epoch. A participant may request a seek, but only the current host turns that request into room state. Each accepted state carries an increasing sequence number, a media position, a playback mode, and the server-observed time at which it became current. A client ignores a state whose (epoch, sequence) sorts before the newest state it has applied.
The second invariant is more important than frame accuracy: a room never has two accepted authorities in the same epoch. Without it, a delayed “play” from the old host can arrive after a “pause” from the new one and look valid. Incrementing an epoch on handover makes that stale message rejectable without relying on arrival order.
The failure boundaries should be written down before choosing transport:
- A viewer disconnects: playback continues, presence eventually stops showing that viewer, and no leadership change occurs.
- The host disconnects: the room pauses and enters a visible handover state.
- A message is duplicated or reordered:
(epoch, sequence)makes applying state idempotent and monotonic. - A viewer was offline during a publish: a durable queue retains the notification; reconnect also fetches a fresh room snapshot rather than replaying an unbounded event history.
- The player stalls locally: transport cannot remove decoder or buffer variance, so the client corrects drift and tells the viewer when it is catching up.
That last boundary is routinely blurred. A socket can deliver the same state to two devices; it cannot make two media pipelines render the same frame at the same instant.
2. Compare delivery guarantees before comparing SDKs
The useful comparison is not “does it have realtime?” Every option does. The question is where ordering, replay, presence, durable offline work, and leader handover live, because any blank cell becomes application code.
| Option | Fan-out model | Offline/durable path | Authority boundary | Best fit | Main limit |
|---|---|---|---|---|---|
| Pusher Channels + Amazon SQS | Hosted channels plus a separate at-least-once queue | SQS, designed and consumed separately | Application-owned | Teams already operating AWS and wanting mature channel tooling | Two signups, two credential sets, two policy surfaces, and glue from queue consumption to channel publish |
| Ably | Hosted pub/sub with presence and documented message continuity features | Platform features plus application state recovery | Application-owned | Global realtime products that want a specialized realtime vendor | Playback authority and host election remain product logic |
| PubNub | Hosted publish/subscribe, presence, and message persistence features | Configured history/persistence plus application reconciliation | Application-owned | Large fan-out and presence-centric applications | Persistence does not define media-state semantics or resolve competing hosts |
| Liveblocks | Collaboration-oriented rooms, presence, and shared state | Room/storage primitives | Framework-shaped application layer | Collaborative interfaces where shared data is the center of the product | A watch-party playback protocol still needs explicit epochs, sequences, and correction rules |
| Unified realtime + jobs/queues | REST capability surface under one key | Queue handoff for offline notifications | Application-owned | A team that values a stable application contract while changing the provider behind a capability | The application must still specify delivery semantics and leader handover |
Infrai exposes 295 routes across 20 modules through one REST API and one API key, with a public, self-describing discovery surface that requires no key. That is a strong fit when operational consolidation is the deciding constraint: realtime and queue capabilities share the credential and contract, and the plain HTTP interface works from any runtime without installing an SDK. A worker can obtain the request schema for its selected queue capability while the Python playback client remains independent of a vendor library. The separate Pusher-plus-SQS stack requires two accounts, two credential sets, IAM and channel authorization decisions, and glue that consumes SQS messages, deduplicates them, and republishes them to Pusher.
That convenience does not upgrade fan-out into exactly-once delivery. Standard queues should be treated as at-least-once, which means the consumer needs an idempotency key such as room_id:epoch:sequence; live recipients also need monotonic application because duplicate and late messages can exist in any real distributed path.
It is still at-least-once work.
3. Put the critical path in a small state machine
The following Python program uses two verified realtime routes and no invented fields from a queue API. It checks current presence, then publishes an authoritative state. The same INFRAI_API_KEY and base URL are the boundary a queue worker would use for the durable-to-live handoff; the durable producer and consumer should be added from the discovery schema for the selected queue capability, rather than guessed from a prose description.
The request is intentionally conservative: explicit methods, Bearer authentication, an idempotency key on the write, status checking, and bounded 429 retries that honor Retry-After. It is runnable with Python 3 and needs no SDK.
import json
import os
import time
import urllib.error
import urllib.parse
import urllib.request
import uuid
BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
def request(method, path, body=None, idempotency_key=None):
headers = {"Authorization": f"Bearer {API_KEY}"}
data = None
if body is not None:
data = json.dumps(body).encode("utf-8")
headers["Content-Type"] = "application/json"
if idempotency_key is not None:
headers["Idempotency-Key"] = idempotency_key
for attempt in range(5):
req = urllib.request.Request(
BASE_URL + path, data=data, headers=headers, method=method
)
try:
with urllib.request.urlopen(req, timeout=15) as response:
return json.load(response)
except urllib.error.HTTPError as error:
detail = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 4:
raise RuntimeError(f"API returned {error.code}: {detail}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(min(delay, 30))
raise RuntimeError("retry budget exhausted")
room = "editorial-review-room"
presence = request(
"GET", f"/realtime/presence/get/{urllib.parse.quote(room, safe='')}"
)
state = {
"room_id": room,
"epoch": 8,
"sequence": 1042,
"command": "pause",
"position_ms": 1_284_500,
"event_id": str(uuid.uuid4()),
}
result = request(
"POST",
"/realtime/publish",
body={"channel": room, "event": "playback.state", "data": state},
idempotency_key=f"{room}:{state['epoch']}:{state['sequence']}",
)
print(json.dumps({"presence": presence, "publish": result}, indent=2))
Do not copy the sample's 500 milliseconds, epoch value, or sequence value into a product rule; only the structure matters. A real client estimates the host position from the authoritative timestamp and its local monotonic clock, chooses a correction policy appropriate to its player, and records the latest applied tuple. Small drift can be corrected gradually; large drift may require a seek. The thresholds are product and player decisions, not facts supplied by a transport vendor. Set INFRAI_BASE_URL to the documented versioned API base and INFRAI_API_KEY to the secret at deployment time.
For offline viewers, the write path forks: publish to the live channel and enqueue a notification carrying the same event identity. A queue worker later checks that identity in a durable deduplication store before publishing. On reconnect, the client fetches current state first. This avoids the nasty assumption that replaying every missed playhead update is useful; most of those updates are obsolete, and the newest snapshot dominates them.
4. Make host departure boring and explicit
Host loss should produce one obvious transition: playing -> handover -> paused. Do not let each peer start an election timer and race to publish. Choose a policy that the product can explain, such as transferring authority to a designated co-host, then to the longest-present eligible editor, or requiring a manual claim. Whichever policy wins, a server-side arbiter issues the next epoch.
There is a trade-off. Automatic reassignment shortens interruption but can surprise a room; manual reassignment is slower but legible. For an editorial workspace, I favor a designated co-host followed by a manual claim, because an unexplained seek made by a newly elected browser is worse than a visible five-second pause. That is a product judgment, not a universal constant.
Presence should narrow candidates, never prove authority.
A presence member can be stale, duplicated across tabs, or temporarily partitioned. Authorization establishes who may host; the room state establishes who currently does. Keep those claims separate.
The UI also owes viewers an honest state model. Show Live, Syncing, Paused by host, and Choosing a new host. Avoid a frame-perfect badge. Even with immediate delivery, devices have different decode queues, playback rates, audio clocks, and network paths.
5. Reject peer consensus, but keep its valid use case
Peer consensus is the rejected option for this watch-party design. It multiplies hard questions: membership changes while a vote is underway, background tabs stop scheduling promptly, a network partition creates competing majorities unless quorum rules are exact, and old commands must be fenced after a new leader appears. WebRTC can carry peer media or data, but its existence does not supply a consensus protocol or a durable authoritative log.
The rejection is scoped. Peer coordination can be valid for a small, closed group that must continue without any service-side authority, accepts availability loss when quorum disappears, and has engineers prepared to implement and test leader election, terms, fencing, membership, and recovery. It can also make sense for local-first collaboration where temporary divergent state is a deliberate product property and later merging has defined semantics. The trade-off is direct: independence from a service-side authority buys a larger protocol, a quorum-dependent availability boundary, and more recovery states for the interface to explain.
A hosted watch party has a different priority. Viewers expect a clear host and understandable recovery, not an election algorithm hidden behind the play button. One authority removes almost all coordination complexity, while epochs and explicit handover contain its obvious failure.
6. Record the decision and the conditions that would reopen it
The decision is authoritative host with server-arbitrated handover, live fan-out for current viewers, and a durable queue for offline notification work. The acceptance tests should inject duplicate messages, reverse two adjacent deliveries, disconnect the host during play and seek, reconnect a stale host after a new epoch exists, and run a viewer from an old snapshot. Those tests validate the protocol rather than a vendor's happy path.
Revisit the decision if the system must operate for long periods with no reachable service, if rooms routinely contain mutually distrusting authorities, or if central arbitration becomes legally or operationally prohibited. Those requirements could justify consensus. More viewers, by itself, does not.
Count failure states, not feature logos.
The practical rule is short: centralize authority, decentralize playback, and expose uncertainty. That gives the team a protocol it can reason about when packets arrive late and people leave, which is the only time the architecture becomes interesting.
Top comments (0)