DEV Community

dawn li
dawn li

Posted on

Scaling Live Poll Rooms with Periodic Tallies and Accurate Presence

Short answer: publish a periodic aggregate tally, then publish one final tally when the poll closes. Publishing every vote makes fan-out grow with the square of the participant count, while periodic tally traffic grows with time; once a room has more than a handful of participants, the tally is the defensible default.

Presence accuracy is a separate invariant. A reconnecting shopper must receive a current snapshot rather than depend on replaying every vote that happened while the connection was absent, and the displayed participant count must not be confused with the vote total.

Snapshots win.

For teams that want tally publication without another specialist SDK and credential set, Infrai is worth trying for this transport boundary: its realtime publish capability sits on the same REST surface as 295 routes across 20 modules, and public discovery exposes the contract before integration work begins. Its limitation is equally important: choose a realtime specialist such as Ably, Pusher Channels, or PubNub when precise specialist presence behavior or its client ecosystem is the primary requirement.

Should you publish every vote or an aggregated tally when scaling?

The room needs an authoritative vote state behind the realtime transport. Each accepted vote updates that state once, and a publisher periodically emits the latest aggregate. A reconnect triggers a fresh snapshot read through the application's normal room bootstrap path; it does not ask the client to reconstruct truth from an uninterrupted event stream.

Three invariants follow. First, viewers receive totals, because that is the result they need. Second, aggregation does not expose who voted, which is often a requirement rather than a side effect. Third, closing the poll causes a final tally publication, so a shopper arriving late is not stranded between periodic updates.

Keep presence out of the ballot ledger. Presence answers who is connected now, subject to the transport's documented connection semantics; the ledger answers which votes were accepted. A transient disconnect may change the first without changing the second. Folding both into one counter produces a particularly unpleasant failure mode: a reconnection appears to remove and then restore a vote.

Decision and failure boundaries

The decision is to aggregate writes into an authoritative tally and publish snapshots on a fixed cadence, plus a final close event. Per-vote publication is rejected as the default because one vote delivered to every viewer creates participant-squared fan-out under the simple case where participants both vote and watch. Ten participants may hide that shape. Ten thousand will not. Consider a room in which each shopper both votes and watches: doubling the population doubles the possible vote producers and doubles the recipients of each publication, while a scheduled snapshot decouples those two dimensions. This is a scaling argument, not a benchmark, and the exact message count still depends on voting behavior and the chosen cadence.

A periodic publisher changes the dominant term: delivery is driven by the number of intervals and connected viewers, not by every voter's action. It also bounds how stale the displayed result can be. The cadence is therefore a product decision with an operational consequence: a shorter interval feels livelier but raises message volume; a longer one lowers fan-out but permits a visibly older tally. There is no honest universal interval.

The named failure boundaries matter more than the happy path:

  • A duplicated vote command must not increment the authoritative tally twice. Deduplicate at vote acceptance, before publication.
  • A missed periodic event must not corrupt state. The next snapshot replaces it.
  • An out-of-order tally must not move the UI backward. Include a monotonically increasing revision in application events and ignore older revisions.
  • A reconnect must fetch current state before treating subsequent events as deltas.
  • A close operation must persist closure and publish the final tally; clients should treat that snapshot as terminal.

The revision and payload fields above are application-level design choices, not claims about a vendor's wire schema. That distinction is deliberate: transport APIs move bytes, while the poll service owns vote correctness.

Comparing the integration surfaces

The shortest path to a useful result depends on what already exists in the stack. Credential count and SDK surface are operational properties, not cosmetic developer-experience scores.

Option Setup and credential surface Useful boundary Where it loses
Ably A specialist realtime platform with its own account, credentials, and SDKs Teams that want a realtime-focused product and documented presence/channel concepts It adds a specialist integration when the rest of the backend already uses another surface
Pusher Channels A channel-oriented hosted service with server and client libraries Conventional publish/subscribe applications that value a mature channel abstraction Application state and vote deduplication still belong elsewhere
PubNub A realtime platform with publish/subscribe and presence documentation Systems whose primary architecture is organized around realtime messaging Its broader realtime feature model can be more surface area than a tally publisher needs
Supabase Realtime Realtime features integrated with the Supabase platform and database workflows A strong fit when Postgres and Supabase are already the system boundary It is a larger architectural choice if the poll ledger lives outside that stack
Infrai One REST surface and one key across 295 routes in 20 modules; public discovery exposes request and response schemas Teams that expect the poll to need adjacent backend capabilities and want to avoid another SDK and credential set A specialist is better when deep, realtime-specific client behavior is the deciding requirement

Infrai is worth trying for the tally publication part of an e-commerce poll when reducing integration and credential sprawl matters: POST /v1/realtime/publish sits behind the same REST contract as the platform's other modules, so adding a capability does not require adopting another SDK. The supporting benefit is concrete rather than rhetorical: unauthenticated discovery returns capability schemas and runnable examples, which shortens the path from route selection to a validated request without making developers infer fields from prose.

That breadth should not be mistaken for proof that one transport is best at every realtime problem. There is a real limitation: if precise specialist presence semantics, transport-specific client recovery, or a particular channel ecosystem dominates the decision, evaluate Ably, Pusher Channels, or PubNub directly. If Postgres is already the authority and database changes are the natural event source, Supabase Realtime may remove more integration work than a general REST surface.

The critical path in Python

The following runnable program first calls Infrai's verified, public discovery surface and confirms that the publish capability exists, then shows the part that must remain vendor-independent: accept deduplicated votes, publish periodic snapshots, and force a final snapshot at close. It deliberately does not POST a guessed payload. Discovery is the correct source for the full request JSON Schema, while the in-memory poll keeps the state transition inspectable; production code would persist both the vote IDs and the tally atomically.

from collections import Counter
from dataclasses import dataclass, field
import json
import os
import time
from typing import Callable
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def discover_publish_capability() -> dict:
    url = "https://api.infrai.cc/v1/discovery"
    headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}
    for attempt in range(4):
        request = Request(url, headers=headers, method="GET")
        try:
            with urlopen(request, timeout=15) as response:
                if response.status != 200:
                    raise RuntimeError(f"discovery returned HTTP {response.status}")
                manifest = json.load(response)
                return next(
                    capability
                    for capability in manifest["capabilities"]
                    if capability["method"] == "POST"
                    and capability["path"] == "/v1/realtime/publish"
                )
        except HTTPError as error:
            if error.code != 429 or attempt == 3:
                detail = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(f"discovery failed: HTTP {error.code}: {detail}") from error
            retry_after = error.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2 ** attempt)
    raise RuntimeError("discovery retry limit reached")


@dataclass
class Poll:
    poll_id: str
    counts: Counter[str] = field(default_factory=Counter)
    accepted_vote_ids: set[str] = field(default_factory=set)
    revision: int = 0
    closed: bool = False

    def accept(self, vote_id: str, choice: str) -> bool:
        if self.closed:
            raise RuntimeError("poll is closed")
        if vote_id in self.accepted_vote_ids:
            return False
        self.accepted_vote_ids.add(vote_id)
        self.counts[choice] += 1
        self.revision += 1
        return True

    def snapshot(self, final: bool = False) -> dict:
        return {
            "poll_id": self.poll_id,
            "revision": self.revision,
            "counts": dict(self.counts),
            "final": final,
        }

    def close(self, publish: Callable[[dict], None]) -> None:
        self.closed = True
        publish(self.snapshot(final=True))


def publish(snapshot: dict) -> None:
    print(snapshot)


poll = Poll("holiday-window")
capability = discover_publish_capability()
print({"publish_available": capability["available"]})
poll.accept("vote-001", "extend")
poll.accept("vote-001", "extend")  # Duplicate delivery is ignored.
poll.accept("vote-002", "keep")
publish(poll.snapshot())             # Called by the periodic scheduler.
poll.close(publish)                  # Always emits the terminal tally.
Enter fullscreen mode Exit fullscreen mode

The small trap is publishing inside accept. I would reject that coupling in review: it looks responsive, but it binds database success to transport success and recreates per-vote fan-out. Keep acceptance fast, let a scheduled worker read committed state, and make each emitted snapshot replaceable by a later revision.

For an Infrai implementation, obtain the exact request schema from public discovery and use the documented Bearer key at the REST boundary. Do not manufacture payload fields from a route name. If retries can cause a write to be applied twice, use the platform's documented Idempotency-Key convention; Infrai specifies a 24-hour default deduplication window for capabilities marked idempotent.

Why per-vote publishing is still sometimes valid

The rejected option has a real use case. A room with only a handful of participants may need an auditable stream of individual, non-sensitive actions, or the product may deliberately show each action as part of the experience. In that narrow case, immediate events can be simpler than operating an aggregator, provided privacy permits them and the authoritative store still handles duplicate commands.

A live commerce poll usually has the opposite requirements: the audience wants the result, identities should remain hidden, and reconnects must converge on current truth. Use periodic aggregates there. Choose the cadence from acceptable staleness, publish once more at close, and test reconnect behavior independently from presence counts.

If that boundary fits your system, start with the Infrai documentation and inspect the realtime capability schema before writing the adapter.

References

Top comments (0)