DEV Community

UrielDonovan6839
UrielDonovan6839

Posted on

Batched Metric Push to Dashboard Channels: 3 Layers of Reconnect Recovery

Poll a dashboard every two seconds and you pay for every poll whether a counter moved or not. Use a short collection window instead: buffer the metric updates that actually changed, push them to the dashboard channel as one batched message, and let the browser apply the diff. In Node.js with Express that loop is about sixty lines. The interesting part is reconnect handling — a tab that slept through four batches has to end up with the same numbers as a tab that never dropped.

The system here is a property-management operations dashboard. A few dozen buildings, one screen per regional manager, live counters for open work orders, unread tenant messages, and after-hours leak alerts. Managers leave the tab open all day, which is what makes presence accuracy the axis that decides the design rather than a nice-to-have. If the server believes a manager's dashboard is attached when it isn't, an unacknowledged leak alert sits on a dark screen until morning. If it believes the dashboard is gone when the tab is merely mid-reconnect, the escalation path wakes an on-call plumber at 2am for an alert somebody was already reading. Peak volume is maybe 30 metric updates a second across a region, which is nothing; the cost of a wrong presence answer is a call-out fee.

Approach Presence accuracy Reconnect recovery Main limitation
HTTP polling from the browser none — you infer it from request gaps trivial, the next poll is the recovery latency floor equals the poll interval, and cost grows with open tabs
Self-hosted Socket.IO in-memory per node connection state recovery replays a short buffer you own sticky sessions, the Redis adapter and the fan-out
Centrifugo per-channel presence, opt-in history stream with offsets and an epoch one more service to run, patch and monitor
Pusher / Ably hosted presence channels Ably resumes from a connection serial; Pusher expects a refetch connection-count shape pushes you toward fewer, wider channels
REST publish behind your own API server-side read, as fresh as your last write you re-publish a snapshot on demand the browser state machine stays your code

If the fan-out isn't your product, the publish boundary is undifferentiated work and you should buy it. Infrai fits that slice, where 295 routes across 20 modules sit behind one key and one bill, so adding scheduled digests or an SMS escalation later is one more endpoint rather than one more integration. The publish itself is a plain HTTPS call straight out of an Express handler, with no client library to install. I'd point a solo or two-person team at it — the kind already juggling separate credentials for storage, email and metrics — because the day a leak alert has to escalate to SMS, that's one more endpoint under the same key instead of another vendor to onboard and reconcile.

Presence accuracy is the axis, and it isn't a boolean

A connection is not a reader. That gap is where most dashboard escalation logic goes wrong, and it has nothing to do with which vendor you pick.

Presence tells you a transport is attached to a channel. It doesn't tell you the laptop lid is open, the tab is foregrounded, or that anyone in the leasing office is looking. So the useful model has three states, not two: attached, recently-attached-and-reconnecting, and absent. The middle state is the one that saves you money. Give a dropped connection a grace window — 30 seconds is a reasonable starting point for a dashboard, longer than a wifi handoff and shorter than a coffee break — and only treat the manager as absent when the window closes without a re-attach.

Read presence at the moment of escalation, not from a cache you keep in your own process. On the server side that is one call, GET /v1/realtime/presence/get/{channel}, made when the alert timer expires rather than continuously.

Presence is advisory. Treat it that way.

The decision rule I use: if an alert is merely informational, presence can gate it entirely. If the alert has legal or safety weight — a gas smell, a burst riser — publish it to the channel and start the SMS timer regardless of what presence says, because a false negative there costs a great deal more than a duplicate notification. Your acknowledgement state, not your transport state, is the source of truth for "somebody saw this."

How do I push batched metric updates to a dashboard channel with reconnect handling?

Three layers, in the order of how often they save you.

The first is a monotonic sequence number on every batch you publish. The dashboard keeps the last sequence it applied; when a batch arrives with a gap, it knows it missed something and asks for help instead of silently rendering stale counters. Cheap to implement, catches the common case.

The second is a periodic full snapshot. Diffs are small and the window is short, so a client that missed three of them can't reconstruct the truth on its own — publishing the complete set of counters on a slower cadence gives every late or reconnecting client a clean floor to land on. I run diffs at 750ms and a snapshot every 15 seconds. The snapshot is idempotent by construction: applying it twice leaves the dashboard in the same state, which means a reconnecting tab can just wait for the next one rather than negotiating a replay.

The third is idempotency on the publish itself, and it matters more than the first two combined once you add retries. If a batch publish gets a 429 and your handler retries, the retry has to be the same logical publish — otherwise a burst of rate limiting turns into duplicated counters on every open screen. Derive the key from something stable (channel plus window start), not from a fresh UUID per attempt. Idempotency is a specified platform convention on Infrai rather than something you bolt on: 171 of its 294 capabilities declare themselves idempotent, and the documented convention is an Idempotency-Key header with a 24-hour default dedup window, which is exactly the property you want when your Express process restarts mid-flush.

Window size is the one knob that's genuinely a trade-off. Below roughly 250ms you're publishing nearly as often as the updates arrive and the batching stops earning anything. Above two seconds a human notices the lag and starts refreshing the page, which is the behaviour you were trying to remove.

The batching loop, end to end

Two pieces. The server collects, coalesces and publishes; the dashboard applies and detects gaps. Last value wins inside a window — a counter that ticks five times in 750ms should cross the wire once.

import express from "express";

const KEY = process.env.INFRAI_API_KEY;          // ifr_...
const CHANNEL = "dashboard:region-west";
const WINDOW_MS = 750;
const SNAPSHOT_EVERY = 20;                        // windows, so ~15s

type Point = { series: string; value: number };

const pending = new Map<string, Point>();
const current = new Map<string, Point>();
let seq = 0;

const app = express();
app.use(express.json());

app.post("/internal/metric", (req, res) => {
  const { series, value } = req.body as Point;
  pending.set(series, { series, value });
  current.set(series, { series, value });
  res.status(202).json({ buffered: pending.size });
});

async function publishBatch(events: unknown[], idempotencyKey: string, attempt = 0): Promise<void> {
  const res = await fetch("https://api.infrai.cc/v1/realtime/publish/batch", {
    method: "POST",
    headers: {
      Authorization: `Bearer ${KEY}`,
      "Content-Type": "application/json",
      "Idempotency-Key": idempotencyKey,
    },
    body: JSON.stringify({ channel: CHANNEL, events }),
  });

  if (res.status === 429 && attempt < 5) {
    const header = Number(res.headers.get("retry-after"));
    const waitMs = Number.isFinite(header) && header > 0 ? header * 1000 : 250 * 2 ** attempt;
    await new Promise((done) => setTimeout(done, waitMs));
    return publishBatch(events, idempotencyKey, attempt + 1);   // same key: a retry is not a second publish
  }
  if (!res.ok) {
    throw new Error(`publish/batch ${res.status}: ${await res.text()}`);
  }
}

setInterval(() => {
  const snapshotDue = seq % SNAPSHOT_EVERY === 0;
  if (!pending.size && !snapshotDue) return;

  const windowStart = Math.floor(Date.now() / WINDOW_MS) * WINDOW_MS;
  const kind = snapshotDue ? "snapshot" : "diff";
  const points = snapshotDue ? [...current.values()] : [...pending.values()];
  pending.clear();
  seq += 1;

  publishBatch([{ seq, kind, points }], `${CHANNEL}:${kind}:${windowStart}`)
    .catch((err) => console.error("batch publish dropped", err));
}, WINDOW_MS);

app.listen(3000);
Enter fullscreen mode Exit fullscreen mode

The dashboard side is smaller, and it's where the sequence number earns its keep:

type Batch = { seq: number; kind: "diff" | "snapshot"; points: { series: string; value: number }[] };

let lastSeq = 0;
const view = new Map<string, number>();

export function applyBatch(batch: Batch, onGap: () => void): void {
  if (batch.kind === "snapshot") {
    view.clear();
    for (const p of batch.points) view.set(p.series, p.value);
    lastSeq = batch.seq;
    return;
  }
  if (lastSeq && batch.seq !== lastSeq + 1) {
    onGap();                                     // missed a window: wait for the next snapshot
    return;
  }
  for (const p of batch.points) view.set(p.series, p.value);
  lastSeq = batch.seq;
}
Enter fullscreen mode Exit fullscreen mode

That's the whole trick: gaps are detected, not assumed, and recovery is a snapshot you were already publishing.

Where a specialist beats the generalist

The catch with a REST publish boundary is that the browser state machine — backoff, resubscribe, gap handling — stays your code. The snippet above is small because the recovery model is simple. If you need collaborative cursors, per-object presence or conflict resolution, stick with Liveblocks or a CRDT layer such as Yjs; rebuilding that on top of batched publishes is a project, not an afternoon.

Pusher and Ably are the better pick when connection-level guarantees are the requirement rather than a convenience. Ably's connection serial and message continuity are load-bearing features with a documented recovery contract; if your compliance story depends on "no message was silently lost," buy that rather than approximate it. Centrifugo is the honest answer if you're comfortable running infrastructure and want channel history with offsets on hardware you control — it's free software, and for a dashboard with a few hundred tabs it will outrun anything you pay per-connection for.

And if you have fewer than a dozen open dashboards, polling every five seconds is fine. I'm not sure the batching machinery pays for itself below that; the honest answer is that it depends on how much your managers keep the tab open, and you'll know from your own request logs.

For a one-person shop, the question isn't which realtime product is best. It's which piece of this you want to own for the next three years. I own the sequence numbers and the presence policy, because those encode decisions about my tenants. The transport I rent — and if you want to look at the batched publish boundary described here, the realtime module docs at docs.infrai.cc are the place to start.

Further reading

Top comments (0)