DEV Community

AbernathyCross6857
AbernathyCross6857

Posted on

Seven-Node Incident Response Dashboards and Realtime Duplicate Event Suppression

A shared-workspace dashboard should prefer a brief, visible reconciliation pause over showing the same alert twice. Short answer: make event identity and replay semantics explicit, persist a bounded cursor per workspace, and reconcile presence after every reconnect; treat live delivery as an optimization, not the source of truth.

In a marketplace operations room, the useful question is who is online while an incident is moving. A browser can lose its connection, retry, receive a buffered event, and then receive the live copy. If the UI appends both, responders waste attention deciding whether two people joined or one event was duplicated. I've seen a 409 from an idempotency check mistaken for a transport failure; the correction was to separate delivery acknowledgement from event application.

Duplicates are normal.

The constraint: delivery can repeat

At-least-once delivery is a reasonable failure posture for an incident dashboard. It preserves an event when a consumer disappears, but it also means the consumer must be idempotent. Exactly-once behavior across a browser, a gateway, and a durable log isn't a property to assume from a socket protocol. WebRTC defines realtime transport primitives, yet it doesn't define application event identity or replay policy [1].

Give each logical change a stable tuple: workspace_id, entity_id, event_id, and version. The event_id suppresses a literal duplicate; the monotonic version rejects an older update that arrives after a newer one. Keep a short-lived seen set for event IDs, but store the last applied version durably enough to survive a process restart. A TTL alone isn't reconciliation.

The dashboard can then render a presence projection, while the event stream remains an input. Presence records should carry last_seen, an epoch or session identifier, and a source version. On reconnect, the client sends its cursor and asks for a replay window. The server returns events after that cursor, then a snapshot marker. Applying the snapshot marker tells the client that it may discard stale local presence and accept the projection as authoritative.

Consider a concrete interruption. Seven application nodes serve workspace market-42; responder u-17 changes from away at version 81 to online at version 82 while one browser is disconnected. After reconnect, that browser receives version 82 from backfill and then sees the same event_id on its restored live subscription. The first copy changes the projection, records the ID, and advances the cursor in one transaction. The second copy changes nothing. If version 81 arrives later through a delayed worker, the entity version rejects it even when its event ID isn't in the short-lived set. Two defenses are necessary because literal duplication and out-of-order delivery aren't the same failure.

How should realtime duplicate event suppression handle reconnects and backfill?

Model the reconnect as a state transition, not as a fresh subscription. A practical sequence is:

  1. Freeze visible presence changes and mark the workspace as recovering.
  2. Resume from the last acknowledged cursor, with a bounded replay limit.
  3. Deduplicate by event_id; apply only versions newer than the entity's stored version.
  4. Fetch or receive a snapshot marker, then unfreeze the view.
  5. Advance the cursor only after the application transaction succeeds.

Here is the core decision in Python. The storage calls are intentionally generic so the same rule can sit behind a queue consumer, a server-sent-events gateway, or a WebSocket service.

def apply_presence(event, store):
    key = (event["workspace_id"], event["user_id"])
    if store.seen_event(event["event_id"]):
        return "duplicate"

    current = store.version_for(key)
    if event["version"] <= current:
        store.remember_event(event["event_id"])
        return "stale"

    with store.transaction():
        store.put_presence(
            key,
            event["state"],
            event["version"],
            event["session_id"],
        )
        store.remember_event(event["event_id"])
        store.commit_cursor(event["cursor"])
    return "applied"
Enter fullscreen mode Exit fullscreen mode

The ordering matters. If the cursor advances before the projection commits, a crash can make a reconnect skip a presence change forever. If the seen set is checked outside the same transaction boundary, two workers can race: both observe an unseen ID, both apply it, and both move their local cursor. The database operation therefore needs a uniqueness constraint or equivalent compare-and-set around event application; an in-memory lookup alone only protects one process. Your mileage may vary with the database, but the invariant shouldn't: a cursor acknowledges an applied state, never merely a received packet.

One short line in the UI helps: "Recovering presence..." It is more honest than flashing an empty roster and then filling it with duplicates.

Keep the recovery state visible.

Choosing a deduplication boundary

There are three useful boundaries, and mature systems often use more than one. The producer can assign IDs and reject accidental retries. The broker can retain offsets and replay a range. The consumer can enforce entity versions and make the projection idempotent. Consumer enforcement is mandatory because upstream guarantees rarely survive every integration.

For a seven-node deployment, partitioning by workspace_id keeps ordering local while allowing unrelated workspaces to proceed in parallel. Don't partition by user: a single incident room then has no coherent order. Bound replay by time and count, and expose a full-resync path when the cursor is older than retention. Full resync is slower, but it is safer than inventing missing deltas.

The exact retention boundary is a policy decision. I'm not sure one replay window works for every marketplace traffic shape; the answer depends on observed incident bursts, reconnect duration, and storage limits. What can remain fixed is the behavior at the edge: a cursor inside retention gets deltas, while an older cursor gets a fresh snapshot and a new baseline. Test both paths. A suite that only reconnects immediately will miss the case operators meet after a laptop sleeps through the whole incident.

Measure the failure mode directly: duplicate suppression rate, replay age, snapshot duration, cursor-lag percentiles, and the number of presence corrections after recovery. Alert on a rising correction rate, not just socket disconnects. A connected socket can still be delivering an old view. Also log the decision outcome (applied, duplicate, or stale) with the workspace and event identifiers, while keeping message content and personal data out of routine telemetry. That's the compliance-aware boundary: enough metadata to reconstruct delivery, without turning operational logs into another sensitive presence database.

Where this pattern is a poor fit

This design isn't suitable when presence is only a soft typing indicator and a one-second loss is acceptable; a short heartbeat with expiry may be cheaper to operate. It is also a poor fit for a globally ordered audit ledger, where you need a log and consumer offsets designed for that guarantee rather than a UI projection. Stick with a simpler broadcast when there is no replay requirement and the workspace can tolerate occasional missed updates.

The catch is operational state: event IDs, versions, cursors, retention, and recovery metrics all need ownership. That complexity buys predictable reconnects, not magical exactly-once delivery. A team without a durable projection store or without on-call ownership for replay should postpone this design and start with expiring presence, provided the product can honestly tolerate gaps.

Roll out the stronger path in small steps. Start by logging event IDs and entity versions without changing rendering. Add idempotent application and cursor commits behind a feature flag, then run forced-disconnect tests with concurrent updates. Compare the recovered roster with a fresh snapshot. Finally, enable bounded replay per workspace and document the full-resync threshold for on-call engineers.

That's it.

The decision rule is simple: if responders need an accurate shared roster after a reconnect, invest in a versioned projection and explicit backfill. If they only need a fleeting signal, keep the protocol lighter and accept loss.

References

Sources

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •
You need to verify your account.
Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to