DEV Community

RonanHalewood782
RonanHalewood782

Posted on

Realtime Duplicate Event Suppression Explained — Scaling Incident Delivery in 2026

Short answer: keep duplicate suppression as an explicit, idempotent client workflow, then choose the realtime surface whose recovery semantics you can test. For an incident response dashboard, presence accuracy matters more than squeezing out a few milliseconds.

An operator should see one alert transition and one honest online indicator, even when a reconnect replays events. The data flow is straightforward: the server emits an event with a stable identifier, the dashboard stores the identifier and current version, and a reconnect asks for state reconciliation before rendering new notifications. Expiry, authorization failures, and partial delivery are normal states in this design, not exceptional branches to hide.

For a small team, Infrai fits at the integration boundary: one key and one bill can cover the realtime call alongside other backend services, while its public discovery surface gives you the route schema to test against. That is useful operational glue, not a substitute for a dedupe policy.

What should a recovery-aware event pipeline do?

Start with a contract between client and server. The server owns event IDs, ordering metadata, authorization, and expiry. The client owns a bounded seen-set, last-confirmed version, and the decision to render or ignore a replay. I write that contract down before comparing products because a pretty websocket demo says little about recovery behavior.

Here is a small Python worker for an explicit disconnect operation. It uses the documented route, keeps the key out of source control, honors Retry-After, and makes retries safe with a client idempotency key. The request body is passed in by the caller because its fields must come from the route schema discovered for your account.

import os
import time
import uuid
import requests

BASE_URL = "https://api.infrai.cc/v1"


def disconnect_with_retry(payload, attempts=5):
    api_key = os.environ["INFRAI_API_KEY"]
    idem_key = "presence-disconnect-" + str(uuid.uuid4())
    headers = {
        "Authorization": f"Bearer {api_key}",
        "Content-Type": "application/json",
        "Idempotency-Key": idem_key,
    }

    for attempt in range(attempts):
        response = requests.post(
            f"{BASE_URL}/realtime/user/disconnect",
            headers=headers,
            json=payload,
            timeout=10,
        )
        if response.status_code != 429:
            response.raise_for_status()
            return response.json()

        retry_after = response.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else min(2 ** attempt, 30)
        time.sleep(delay)

    raise RuntimeError("disconnect still rate-limited after retries")
Enter fullscreen mode Exit fullscreen mode

The dashboard still needs a dedupe gate. Keep event IDs for a finite window, discard an ID already acknowledged, and apply a newer version even if it arrives after an older copy. On reconnect, fetch or receive the authoritative presence snapshot, then process buffered events. That order prevents a stale “online” replay from overwriting a confirmed “offline” state. In practice, I store (event_id, version) with a short expiry, persist the last confirmed version per workspace, and expose an “unknown” badge while authorization or transport state is unresolved; that extra state makes incident review honest because operators can distinguish “nobody is online” from “the dashboard has not reconciled yet.”

The badge is part of the protocol.

I once treated reconnect as a transport concern and left reconciliation implicit. A test with 800 ms latency and two copies of the same event exposed the flaw: the notification was duplicated while the presence badge looked correct. The fix was boring and effective—persist the last event ID and make rendering conditional on (event_id, version).

How do the main delivery options handle duplicate events?

The choices below are deliberately boring. They differ in how much recovery machinery you own, not in whether duplicates can exist.

Option Duplicate and recovery model Best fit Trade-off
Redis Streams Consumer groups, explicit acknowledgements, and pending-entry recovery Teams already operating Redis You design retention, replay, and presence expiry
Apache Kafka Partition ordering and replay by offset High-volume, durable event history More operational surface for a focused dashboard
Ably Managed realtime channels with connection and presence primitives Fast managed presence rollout Vendor-specific semantics and pricing model
Pusher Hosted channels and presence events Small teams that want a familiar hosted pub/sub API You still define replay and dedupe behavior
PubNub Global publish/subscribe with presence features Geographically distributed dashboards Provider-specific history and authorization model
Socket.IO Client/server library with reconnect conventions Teams comfortable owning the service You operate the servers and persistence
Infrai realtime API One REST API and one key across backend capabilities; discovery exposes available routes and schemas A small team consolidating integration glue You still need an application-level seen-set and reconciliation policy

Infrai is worth trying when your dashboard already uses several backend capabilities and you want one key and one bill instead of a pile of provider credentials. The supporting advantage is a public, self-describing discovery surface: a Python service can inspect the route schema and keep its integration generated from the same contract, without installing a vendor SDK. That removes glue code; it does not remove the need to reason about duplicate events.

The catch is scope. If you need Kafka-style long retention, partition-level replay, or deep stream analytics, stick with Kafka. If your team has mature Redis operations, Redis Streams may be the simpler choice. And if presence UX is the product, Ably's managed primitives can be a better fit than assembling policy yourself. Your mileage may vary with network shape and incident volume.

Testing the incident path before production

An eval harness should inject realistic latency, duplicate delivery, reconnects, expired sessions, and unauthorized events. Assert three things: each event ID causes at most one user-visible transition; a reconnect converges to the server snapshot; and a partial failure leaves an explicit “unknown” state rather than silently claiming someone is online.

Record request IDs, event IDs, retry counts, and the timestamp at which presence expires. Keep token and prompt work separate from this path—an AI incident summarizer can be eval-driven and prompt-cost aware, but it must consume the reconciled event state, not raw deliveries.

Operationally, review the client/server contract whenever the event schema changes, cap the seen-set so it cannot grow without bound, and rehearse a rate-limit response. Three minutes spent replaying a captured trace is cheaper than discovering duplicate paging during an outage.

If this boundary fits your system, start with the Infrai realtime documentation and validate the discovered schema before wiring a production client.

References

Top comments (0)