DEV Community

ApexZ69
ApexZ69

Posted on

Aggregate Live Poll Votes — Publish a Reliable Tally on an Interval

A live poll should accept each valid vote immediately, aggregate it in one authoritative place, and publish a versioned snapshot on a fixed interval. Do not broadcast after every vote. In a healthtech room, the same delivery rule also applies to ephemeral typing indicators and durable read receipts: clients may miss, repeat, or reorder fan-out messages, so every tally needs an identity and a monotonic version.

Short answer: keep the current counts and accepted vote IDs on the server, increment a version only when the tally changes, and let one timer broadcast an immutable snapshot. Clients replace their view only when the incoming version is newer. Accepted votes affect the next snapshot; reconnecting clients fetch current state instead of trusting a fragile stream history.

From event spray to snapshot publication

The tempting design is vote in, tally out. It looks fast. It also ties write traffic directly to fan-out traffic: 800 votes arriving together can create 800 nearly identical broadcasts per connected client. A slow connection then receives stale intermediate totals that nobody needs.

Use a different mental model. Before: each vote is both a state change and a publication command. After: vote handlers change authoritative state; a clock decides when that state becomes a published snapshot. The interval is a latency budget, not a durability mechanism.

Picture the path in words: browser to validation to deduplication to aggregate, then a timer takes a snapshot, increments its version, and fans it out. A reconnect travels on a separate path: browser to snapshot endpoint to the same authoritative aggregate. This split matters. WebSocket gives you a two-way connection, but it does not make application messages durable for a disconnected browser.

For a clinical education poll, avoid putting patient identifiers, free text, or other sensitive data in vote payloads. A poll-scoped pseudonymous participant ID is enough to enforce one vote per participant when that is the rule. Authentication and authorization still belong before the vote reaches the aggregate.

How can Node.js aggregate votes and publish a tally on an interval?

This in-memory example keeps the mechanics visible. It is correct for one process. The voteId makes retries idempotent within that process, and version lets clients reject an older tally that arrives after a newer one.

import express, { type Request, type Response } from "express";
import { WebSocketServer, WebSocket } from "ws";
import { createServer } from "node:http";

type OptionId = "yes" | "no" | "unsure";
type Vote = { voteId: string; optionId: OptionId };
type Tally = {
  pollId: string;
  version: number;
  publishedAt: string;
  counts: Record<OptionId, number>;
};

const app = express();
app.use(express.json({ limit: "4kb" }));
const server = createServer(app);
const sockets = new WebSocketServer({ server, path: "/live" });

const pollId = "medication-training-01";
const counts: Record<OptionId, number> = { yes: 0, no: 0, unsure: 0 };
const acceptedVoteIds = new Set<string>();
let version = 0;
let dirty = false;

function snapshot(): Tally {
  return {
    pollId,
    version,
    publishedAt: new Date().toISOString(),
    counts: { ...counts }
  };
}

function publishTally(): void {
  if (!dirty) return;
  version += 1;
  dirty = false;
  const message = JSON.stringify({ type: "tally.updated", data: snapshot() });

  for (const client of sockets.clients) {
    if (client.readyState === WebSocket.OPEN) client.send(message);
  }
}

const interval = setInterval(publishTally, 1_000);
interval.unref();

app.post("/polls/:pollId/votes", (req: Request, res: Response) => {
  if (req.params.pollId !== pollId) {
    res.status(404).json({ error: "poll_not_found" });
    return;
  }

  const vote = req.body as Partial<Vote>;
  const validOptions: OptionId[] = ["yes", "no", "unsure"];
  if (typeof vote.voteId !== "string" || !validOptions.includes(vote.optionId as OptionId)) {
    res.status(400).json({ error: "invalid_vote" });
    return;
  }

  if (acceptedVoteIds.has(vote.voteId)) {
    res.status(200).json({ accepted: true, duplicate: true });
    return;
  }

  acceptedVoteIds.add(vote.voteId);
  counts[vote.optionId as OptionId] += 1;
  dirty = true;
  res.status(202).json({ accepted: true });
});

app.get("/polls/:pollId/tally", (req: Request, res: Response) => {
  if (req.params.pollId !== pollId) {
    res.status(404).json({ error: "poll_not_found" });
    return;
  }
  res.json(snapshot());
});

sockets.on("connection", (client) => {
  client.send(JSON.stringify({ type: "tally.snapshot", data: snapshot() }));
});

server.listen(3000);
Enter fullscreen mode Exit fullscreen mode

The client rule is tiny, and it is the part many demos omit. Apply only a newer version. On reconnect, fetch the snapshot before resuming normal rendering.

type Tally = {
  pollId: string;
  version: number;
  publishedAt: string;
  counts: Record<"yes" | "no" | "unsure", number>;
};

let renderedVersion = -1;

function applyTally(next: Tally): void {
  if (next.version <= renderedVersion) return;
  renderedVersion = next.version;
  renderCounts(next.counts);
}

function renderCounts(counts: Tally["counts"]): void {
  document.querySelector("#yes")!.textContent = String(counts.yes);
  document.querySelector("#no")!.textContent = String(counts.no);
  document.querySelector("#unsure")!.textContent = String(counts.unsure);
}
Enter fullscreen mode Exit fullscreen mode

The server answers 202 after accepting the vote into its local aggregate, not after every viewer renders it. Those are different acknowledgements. A requirement that says “the voter must know everyone saw this” needs per-recipient acknowledgements and a timeout policy; this tally endpoint does not claim that guarantee.

That distinction matters.

What delivery guarantee does each signal need?

Typing indicators are presence hints. If one disappears, sending an old typing: true later is worse than dropping it. Give the client a short expiry and treat newer state as authoritative. Do not queue a historical typing stream.

Read receipts are different. They record user-visible progress and often need durable storage, idempotent writes, and a monotonic marker such as the highest message sequence read in a conversation. A repeated marker is harmless; a marker moving backward is not.

Poll tallies sit between those cases. Individual published snapshots are disposable because the next one supersedes them, but accepted votes usually are not.

Signal Write-side requirement Fan-out behavior Recovery
Typing indicator Best effort Latest state wins; expire it Wait for the next activity
Read receipt Durable, idempotent, monotonic Duplicates are acceptable Load the durable marker
Poll vote and tally Durable vote; versioned snapshot Newer version replaces older Fetch current tally

That table is the architecture. “Realtime” is not one delivery guarantee. Name the state transition first, then choose what may be lost.

What changes when one process becomes several?

The in-memory Set and counters stop being authoritative as soon as traffic reaches two application instances. Each instance sees only some votes, and each timer can publish a different total. Sticky sessions do not repair the split aggregate.

Now scale it.

Move vote identity and counting into a shared durable store. The acceptance operation must atomically establish that a voteId is new and update the selected counter. Then elect one publisher, or assign one poll to one publishing partition, so versions have a single writer. Another valid shape is an append-only vote log plus a stateful consumer that owns each poll partition. The choice depends on recovery and throughput requirements, but the invariant stays fixed: one accepted vote contributes once.

Backpressure deserves an explicit policy. The WebSocket API exposes bufferedAmount, the number of application bytes queued but not yet transmitted. If a client's buffer keeps growing, close that connection and let it recover from the tally endpoint. Keeping every stale snapshot in memory trades one slow browser for server-wide risk. For tallies, dropping superseded snapshots is correct.

The interval needs a trade-off, not folklore. One second means a vote accepted just after a tick may wait nearly one second for the next scheduled publication, plus processing and network time. A shorter interval lowers visible latency and increases serialization and fan-out work. Start from the product's freshness target, load-test the full connected audience, then choose.

How do you prove the tally is healthy?

Track the lifecycle, not merely request status. Useful metrics include accepted votes, duplicate vote IDs, rejected votes by reason, publish duration, connected clients, disconnected slow consumers, and the age of the last successful snapshot. Measure end-to-end publication lag from vote acceptance to a client acknowledgement in a controlled probe if that latency has an objective.

Logs should carry pollId, voteId, the resulting snapshot version, and a trace or request ID. Never log authentication tokens or sensitive form bodies. Keep the fields stable. An alert such as “no snapshot for two intervals” is meaningful only while the poll has pending changes; otherwise silence is normal.

Test the uncomfortable sequence. Submit the same voteId twice and assert one increment. Disconnect a client across several intervals, reconnect it, and assert that the fetched snapshot catches it up. Deliver versions 12, 14, then 13 to the renderer and verify that 13 cannot roll the UI backward. Finally, pause one consumer so its send buffer grows and confirm that the server bounds the damage. This sequence catches three separate mistakes: deduplication that happens after counting, reconnect logic that assumes missed messages will be replayed, and a renderer that trusts arrival order. The trade-off is deliberate. A snapshot client gives up every intermediate animation in exchange for bounded recovery work and a view that can be rebuilt from one response.

There is one more trap: process shutdown. Stop accepting new votes, persist or drain accepted work, publish the final dirty snapshot if the deployment contract requires it, and then close connections. The sample keeps state only in memory, so it is a teaching boundary rather than a production persistence plan.

Choose the guarantee before choosing the transport

A fixed publication interval decouples vote volume from viewer fan-out. The reliable design comes from the surrounding rules: idempotent acceptance, one authoritative aggregate, monotonic snapshots, reconnect recovery, and bounded slow-client behavior.

Keep ephemeral signals ephemeral. Persist consequential state. Make every client able to recover without replaying the entire room, and make the dashboard show where time is actually spent. Those decisions survive changes in transport and deployment topology.

Further reading

Top comments (0)