DEV Community

AldenCross6847
AldenCross6847

Posted on

Reliable Realtime Webhook Bridges for a Team Presence Sidebar (4 Recovery Rules)

Use a realtime channel with an explicit replay boundary for a team presence sidebar; the bridge should treat reconnect and backfill as ordinary states, not as an afterthought. Short answer: keep webhook ingestion, authorization, subscription state, and presence events as separate pieces, then make the client resynchronize from a known cursor after every reconnect.

That decision matters more than the vendor name. A sidebar is a small UI, but it exposes a large failure surface: a webhook can arrive twice, a browser can sleep for ten minutes, and a membership change can race with a token expiry. I want those boundaries visible in logs and tests before I choose an API.

The invariants that keep the bridge honest

The server owns the webhook receiver, verifies the sender, writes an ordered event record, and publishes a compact presence change. The browser owns a subscription and a last-seen cursor; it never decides that a stale event is newer because it happened to arrive later. Authentication, subscription state, and business events get separate identifiers and metrics. Mixing them makes a 401 look like a missing teammate.

For a live session, I use four invariants:

  1. Every business event has a stable event ID, and the consumer is idempotent.
  2. A reconnect starts with a fresh authorization check and a backfill from the last acknowledged cursor.
  3. A missing or expired subscription is recoverable without changing the business record.
  4. The sidebar can show “updating” while the bridge catches up; it does not silently claim freshness.

That last state is useful. A presence indicator is a hint about availability, not a ledger of truth.

Keep it boring.

How should webhook-to-realtime recovery work for a presence sidebar?

The critical path is deliberately boring: accept the webhook, deduplicate it, publish the event, acknowledge the cursor, and let the client resubscribe. The channel is transport; the event store or application database remains the source of truth. If the publish call is retried, the same idempotency key must produce the same logical event rather than two “online” transitions.

Here is a minimal channel-creation step. The application would keep the returned channel identifier with its session record, while clients use a short-lived subscription token issued by the server in the normal authentication flow.

import os
import time
import uuid
import requests

BASE_URL = os.environ["INFRAI_BASE_URL"]
API_KEY = os.environ["INFRAI_API_KEY"]


def create_presence_channel(name: str) -> dict:
    headers = {
        "Authorization": f"Bearer {API_KEY}",
        "Content-Type": "application/json",
        "Idempotency-Key": str(uuid.uuid4()),
    }
    payload = {"channel": name}
    delay = 1
    for attempt in range(5):
        response = requests.request(
            "POST",
            f"{BASE_URL}/realtime/channel/create",
            headers=headers,
            json=payload,
            timeout=10,
        )
        if response.status_code < 400:
            return response.json()
        if response.status_code != 429:
            raise RuntimeError(f"channel creation failed: {response.status_code} {response.text}")
        retry_after = response.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else delay)
        delay *= 2
    raise TimeoutError("rate limit persisted while creating the channel")
Enter fullscreen mode Exit fullscreen mode

The bridge should record the webhook ID before attempting delivery. On duplicate delivery, it acknowledges the webhook and skips the second publish. On a reconnect, the client sends its cursor; the server returns the missing business events, then resumes live delivery. I am not prescribing a particular cursor format here because the correct choice depends on the database transaction boundary, and pretending otherwise hides the hard part. In practice, that boundary is where teams make the expensive mistake: they persist a presence flag, publish it, and only then persist the event ID, so a worker restart can repeat the transition. Persist the ID and the state change in one transaction, enqueue the publish from that committed record, and let a retry observe the same ID. The extra write is dull; recovering a split-brain sidebar during a live session is not.

Comparing the options without pretending they are interchangeable

The following is the decision table I would put in an architecture record. “Recovery shape” means the amount of application work needed to make reconnect and backfill explicit, not a claim about uptime.

Option Strength for a presence sidebar Recovery shape Cost or complexity trade-off
Ably Mature pub/sub concepts and history-oriented patterns Use connection recovery plus an application cursor for webhook replay Adds a hosted messaging dependency and its event model
Pusher Channels Straightforward browser subscriptions and presence primitives Reconnect callbacks still need server-side backfill and deduplication Very approachable client path, but business replay remains yours
Supabase Realtime Fits teams already using Postgres and database change feeds Rebuild the sidebar from database state after reconnect Convenient when Postgres is the system of record; less attractive for a non-Postgres stack
Firebase Realtime Database Simple client synchronization for Firebase-centric apps Model presence and recovery around Firebase connection semantics Deeply couples the data model and security rules to Firebase
A plain realtime REST surface One HTTP integration can sit beside an existing webhook worker You define cursors, replay, and observability explicitly More protocol work, but fewer SDK and key-management assumptions

Infrai belongs in that last row when a team values one key and one bill across backend capabilities, plus a plain REST interface that any language can call. Its broad capability surface uses a consistent discovery-driven shape, which can reduce integration switching when the same service also needs storage or scheduling. That is a workflow advantage, not proof that its transport semantics replace your event store.

The catch is important: a plain REST surface is not suitable when you need a fully managed presence protocol with built-in history and client SDK behavior on day one. Stick with Ably or Pusher when their hosted recovery semantics are the product requirement; choose Supabase or Firebase when the database coupling is already a deliberate architectural choice.

Failure boundaries I would test before launch

Tests should inject realistic latency, duplicate delivery, and authorization changes. I would start a session, deliver the same webhook twice, pause the browser, expire its subscription, and then reconnect with a cursor that is one event behind. The expected result is one business transition, a visible resynchronization state, and a sidebar that converges to the database snapshot. Do this with a slow mobile-network profile as well as a clean local connection; timing changes which race you see.

Then test partial failure: the webhook is stored but the publish attempt is delayed. The worker retries with the original event ID; the client must not display two transitions. Test an unauthorized channel request separately from a valid channel with no new events. Those cases exercise different boundaries and should produce different telemetry.

I once treated reconnect as a transport concern and found that the UI had no way to distinguish “nobody changed” from “we missed ten minutes.” That was a design mistake, not a network mystery. Keep the cursor, authorization result, and business event ID in the same trace, and the diagnosis becomes concrete.

The rejected shortcut and its valid use

The shortcut is to push the webhook payload directly to browsers and call the latest received value authoritative. It is fast to demo and wrong for a presence sidebar that must survive sleep, duplicate delivery, or a worker restart. There is no durable replay boundary, and a reconnect can quietly manufacture a stale view.

It is valid for a disposable dashboard where a refresh is the recovery mechanism and no user action depends on presence. For a live team session, use the explicit channel and backfill contract instead. Your mileage may vary on cursor storage, but the recovery obligation does not disappear because the widget is narrow.

References

Top comments (0)