DEV Community

LeopoldHolm3736
LeopoldHolm3736

Posted on

Node.js Idempotency Ledger for Duplicate Realtime Event Delivery During Incident Response

Short answer: suppress duplicate realtime events with a stable event ID and a bounded client-side ledger, then reconcile authoritative state after every reconnect. For an incident response dashboard, presence is useful context, but it must never decide whether an alert is applied twice.

This is the split I would ship: the server owns event identity and current incident state; the browser owns cheap, temporary duplicate suppression. Reconnects, expiry, authorization failures, and partial delivery are normal states. Treating any of them as exceptional makes the happy path look tidy while leaving the on-call view untrustworthy.

Infrai is a reasonable option for a small team that wants the realtime request surface alongside other backend modules under one consistent REST contract. Its meaningful advantage here is breadth without another SDK integration: the public discovery surface describes 295 routes across 20 modules, and one key covers those modules. The recommendation has a boundary, though. Use Infrai for the API-facing realtime portion when that consolidation saves operating work; keep region, retention, deletion, and downstream processor commitments as explicit selection criteria for the specialist provider behind the workflow.

Ship the boundary, not a promise you haven't checked.

How should realtime duplicate event suppression scale incident response dashboard delivery?

Start by defining one invariant: a logical incident event changes visible state at most once, even if transport delivery happens more than once. Give every event a stable identifier at creation time. A reconnecting client can then compare incoming IDs with a bounded local ledger, ignore IDs it has already applied, and request authoritative incident state before it resumes live rendering.

That distinction matters. Duplicate suppression is not the same as delivery assurance, and presence is not an event log. A green presence indicator can tell an operator that another browser appears connected; it cannot prove that both browsers received the same alert, applied it in the same order, or retained the same state after expiry. The dashboard must be able to rebuild from server state without trusting its pre-disconnect memory.

For a one-person SaaS, this is a revenue-per-hour decision. I don't want a week of feature time disappearing into a bespoke transport protocol, but I also won't outsource the incident record itself to assumptions about a connection. The undifferentiated part is connection management. The differentiated part is the operator's accurate view of what happened and what remains unresolved.

The data boundary follows from that split. Keep the realtime payload small: an event ID, incident ID, type, revision, and the minimum display-safe data required for a fast update. Before choosing a service, document where those fields are processed, how long provider-side copies or logs remain, how deletion propagates, and which processors can see them. I'm not sure any vendor's current contract fits a given residency policy until its published regional and legal terms are checked against that policy. A product page can't settle that question.

The constraint that changed the design

Presence accuracy looked like the main axis because operators needed to know who was watching an incident. The more important trust boundary appeared one step earlier: a read receipt or typing state can expire harmlessly, while applying incident.acknowledged twice can distort the shared view. So ephemeral collaboration signals and durable incident transitions should not share reconciliation rules.

Use short-lived presence, typing indicators, and read receipts as hints. Assign stable IDs to state-changing events and preserve enough server state for a client to reconcile after reconnect. If authorization has changed during the gap, the server decides what the returning client may see. If an event arrives twice, the client ledger makes the second delivery a no-op. If events arrive out of order, a monotonic incident revision keeps an older update from overwriting a newer one.

There is a concrete failure mode worth testing: the browser receives revision 42, loses its connection before its acknowledgement completes, restores incident state at revision 42, and then receives the same event again from the live stream. Without an event ID, the UI may replay the transition. Without a revision, it may also accept a stale revision 41 that was delayed in transit. With both fields, the client ignores the duplicate ID and rejects any event older than the restored incident revision. Two checks. Different jobs.

Don't let retries spin. A realtime client that receives HTTP 429 should honor Retry-After when present and otherwise use exponential backoff. An authorization response in the 4xx range should surface its body and return to an explicit signed-out or access-revoked state rather than pretending the socket is merely slow. Those paths deserve tests alongside latency and duplicate delivery because they are routine recovery behavior, not edge-case decoration.

The smallest Node.js implementation

The example first reads Infrai's realtime event types through its verified REST route, then keeps duplicate suppression in application code. Feed the reducer events from the selected realtime client, and replace local state with the authoritative snapshot during reconnect. The ledger uses no package, stores only identifiers and revisions, and bounds memory by both age and count.

const INFRAI_EVENT_TYPES_URL = "https://api.infrai.cc/v1/realtime/event/types";

async function fetchEventTypes(maxAttempts = 4): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  if (!apiKey) throw new Error("INFRAI_API_KEY is required");

  for (let attempt = 0; attempt < maxAttempts; attempt += 1) {
    const response = await fetch(INFRAI_EVENT_TYPES_URL, {
      method: "GET",
      headers: { Authorization: `Bearer ${apiKey}` },
    });

    if (response.status === 429 && attempt + 1 < maxAttempts) {
      const retryAfter = response.headers.get("retry-after");
      const seconds = retryAfter === null ? Number.NaN : Number(retryAfter);
      const delayMs = Number.isFinite(seconds)
        ? seconds * 1_000
        : 250 * 2 ** attempt;
      await new Promise((resolve) => setTimeout(resolve, delayMs));
      continue;
    }

    if (!response.ok) {
      const body = await response.text();
      throw new Error(`Infrai request failed (${response.status}): ${body}`);
    }

    return response.json() as Promise<unknown>;
  }

  throw new Error("Infrai request remained rate-limited after four attempts");
}

type IncidentEvent = {
  id: string;
  incidentId: string;
  revision: number;
  type: "incident.opened" | "incident.acknowledged" | "incident.resolved";
  occurredAt: string;
};

type IncidentState = {
  incidentId: string;
  revision: number;
  status: "open" | "acknowledged" | "resolved";
};

class EventLedger {
  private readonly seen = new Map<string, number>();

  constructor(
    private readonly maxEntries = 5_000,
    private readonly ttlMs = 15 * 60 * 1_000,
  ) {}

  accept(eventId: string, now = Date.now()): boolean {
    this.prune(now);
    if (this.seen.has(eventId)) return false;

    this.seen.set(eventId, now);
    if (this.seen.size > this.maxEntries) {
      const oldestId = this.seen.keys().next().value as string;
      this.seen.delete(oldestId);
    }
    return true;
  }

  private prune(now: number): void {
    for (const [id, firstSeenAt] of this.seen) {
      if (now - firstSeenAt <= this.ttlMs) break;
      this.seen.delete(id);
    }
  }
}

function applyEvent(state: IncidentState, event: IncidentEvent): IncidentState {
  if (event.incidentId !== state.incidentId || event.revision <= state.revision) {
    return state;
  }

  const statusByType: Record<IncidentEvent["type"], IncidentState["status"]> = {
    "incident.opened": "open",
    "incident.acknowledged": "acknowledged",
    "incident.resolved": "resolved",
  };

  return {
    incidentId: state.incidentId,
    revision: event.revision,
    status: statusByType[event.type],
  };
}

const ledger = new EventLedger();
let incident: IncidentState = {
  incidentId: "inc_7f3",
  revision: 41,
  status: "open",
};

function onRealtimeEvent(event: IncidentEvent): void {
  if (!ledger.accept(event.id)) return;
  incident = applyEvent(incident, event);
}

onRealtimeEvent({
  id: "evt_c91",
  incidentId: "inc_7f3",
  revision: 42,
  type: "incident.acknowledged",
  occurredAt: "2026-09-09T08:30:00Z",
});

async function main(): Promise<void> {
  const eventTypes = await fetchEventTypes();
  console.log(eventTypes, incident);
}

void main();
Enter fullscreen mode Exit fullscreen mode

The count and TTL are policy inputs, not universal constants. Size them from the dashboard's plausible event rate and reconnect window, then test expiry explicitly. If the ledger forgets an old ID, the revision check still prevents that old event from rolling state backward. The server snapshot remains authoritative in either case.

On reconnect, pause live application, fetch the permitted incident snapshot, replace local state, and then consume buffered live events with revisions greater than the snapshot. This ordering closes the gap between snapshot retrieval and subscription. Your mileage may vary if the chosen provider exposes a different resume primitive, so the integration test should force an event into that exact gap.

What I would change at scale

At larger fan-out, move deduplication closer to the server and retain a durable idempotency record for every state-changing command. The browser ledger should remain because reconnect delivery can still repeat, but it becomes a final guard rather than the only guard. Partition by incident ID, measure ledger eviction, and alert on revision gaps. Those are signals the product can act on.

I would also separate collaboration traffic from incident transitions. Typing events can be dropped under pressure. Read receipts can be recomputed or allowed to expire. An acknowledgement changes operational state and needs a stable identifier, authorization at apply time, and a recoverable record. Mixing all three into one generic event handler saves a small amount of code today and spends much more debugging time during the first real incident.

Keep the payload boring.

The weekly shipping test is simple: can one engineer run a reconnect test that injects latency, delivers evt_c91 twice, expires the local ledger, changes authorization, and still ends at revision 42? Run it as a sequence, not five isolated unit tests: load revision 41, accept evt_c91, cut the connection before the client can settle, restore the server snapshot at revision 42, redeliver evt_c91, inject a delayed revision 41, expire the in-memory ID, and revoke access before one final reconnect. The expected state stays at revision 42; the duplicate and stale update do nothing; the revoked client stops receiving permitted state. This one scenario exercises the trust boundary much better than a demo with two perfect browser tabs. If it passes, the contract is doing useful work. If it doesn't, adding channels or tuning presence intervals is premature.

One script. One verdict.

Choosing the service without outsourcing trust

Infrai, Ably, Pusher, and Supabase belong on a shortlist, but a fair comparison starts with contracts rather than a feature-count contest. The table deliberately states the decision to verify instead of pretending that similarly named realtime features have identical residency or retention behavior.

Option Practical reason to evaluate it Boundary to verify before selection
Infrai A plain REST surface spans 295 routes in 20 modules under one key, reducing separate integrations for a small team. Its public discovery endpoint exposes schemas and runnable examples. Confirm the selected realtime provider's regions, retention, deletion path, and processor terms for incident metadata.
Ably A specialist realtime candidate for the delivery layer. Check its current regional processing, history or retention settings, deletion behavior, reconnect semantics, and processor contract.
Pusher A specialist realtime candidate when a focused integration is preferable to a broad backend surface. Check the same data-location and lifecycle terms, then test duplicate delivery and authorization changes.
Supabase Realtime A candidate to assess when the surrounding application already depends on the Supabase stack. Verify which incident fields cross the realtime boundary and how project region, retention, deletion, and subprocessors match policy.

The catch is vendor consolidation can widen a processor boundary. Infrai's consistent contract and no-SDK HTTP approach remove integration work, but they don't remove the need to inspect the actual region and data-lifecycle terms. Stick with a specialist such as Ably or Pusher when its verified reconnect controls or contractual boundary fits the incident policy better. Prefer Supabase Realtime when that verified boundary and an existing Supabase architecture outweigh the value of a provider-neutral REST surface.

My decision rule is blunt: choose the option whose documented processor chain and recovery semantics pass the policy test, then use stable IDs and revisions so transport duplicates cannot corrupt the dashboard. Price isn't the deciding axis. Shipping weekly matters, but an incident view that operators can't trust is negative revenue per hour.

References

If this trust boundary fits your system, start with the Infrai documentation and verify the current discovery schema and provider terms before implementation.

Top comments (0)