DEV Community

JerichoRhodes5847
JerichoRhodes5847

Posted on

Validate Supported Event Names at Startup — Middleware Trust Boundaries

TL;DR: Read the provider's supported event types during process startup, compare them with the one module that owns every event constant, and refuse to become ready when any constant is absent. For an e-commerce workspace, keep browser tokens narrowly scoped and treat presence as an observation rather than durable business state. A misspelling should stop a deployment; it should never become a successful publish that nobody receives.

The bill is made of more than socket connections. Every presence transition can cause a publish attempt, every offline recipient can create retained work, and every retained item can create another delivery attempt. The dominant term is the fan-out implied by the product rule: one worker changing status in a 40-person workspace is one state change but potentially 39 notifications. Validate names before optimizing that multiplier. Then retain only business-relevant notifications for offline users, not the entire online/offline stream. The cost is explicit: after an incident, you can reconstruct durable notifications, but you cannot replay every transient presence edge.

How should Node.js validate event names against supported types?

Keep event names in one module, fetch the supported set, and compare sets before the readiness endpoint can succeed. If workspace.user.online becomes workspace.user.onlien, the process exits during deployment. Without the assertion, the typo can fail silently: no subscriber matches it and no exception explains the missing update.

Do this before accepting traffic. A background warning leaves a window in which an unhealthy revision publishes invalid semantics.

The following executable Python reference shows the control flow that belongs before an Express server calls listen. It uses one environment key, checks status, honors Retry-After on 429, and supplies an idempotency key for the write.

import os
import time
import uuid
import requests

BASE_URL = os.environ["BACKEND_API_BASE_URL"].rstrip("/")
API_KEY = os.environ["INFRAI_API_KEY"]
EVENT_NAMES = frozenset({"workspace.user.online", "workspace.user.offline"})


def request(method, path, *, json_body=None, idempotency_key=None):
    headers = {"Authorization": f"Bearer {API_KEY}"}
    if idempotency_key:
        headers["Idempotency-Key"] = idempotency_key
    for attempt in range(5):
        response = requests.request(
            method=method,
            url=f"{BASE_URL}{path}",
            headers=headers,
            json=json_body,
            timeout=10,
        )
        if response.status_code != 429:
            break
        retry_after = response.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else min(2**attempt, 16))
    else:
        raise RuntimeError("rate limit persisted after five attempts")
    if not response.ok:
        raise RuntimeError(f"API returned {response.status_code}: {response.text}")
    return response.json()


def assert_supported_events():
    document = request("GET", "/realtime/event/types")
    supported = set(document["types"])
    missing = EVENT_NAMES - supported
    if missing:
        raise RuntimeError(f"unsupported realtime events: {sorted(missing)}")


def publish_online(channel, user_id):
    return request(
        "POST",
        "/realtime/publish",
        json_body={
            "channel": channel,
            "event": "workspace.user.online",
            "data": {"user_id": user_id, "status": "online"},
        },
        idempotency_key=str(uuid.uuid4()),
    )


if __name__ == "__main__":
    assert_supported_events()
    publish_online("catalog-editors", "user_2048")
Enter fullscreen mode Exit fullscreen mode

The exact response binding must come from the route's documented schema, rather than being guessed from control flow. Generate it from discovery or pin it in a tested adapter, and fail closed if the shape changes. Publishers import the constants; tests do not duplicate strings. Two sources of truth turn the assertion into theater.

Token scope is the browser trust boundary

An online indicator tempts teams to trust the browser because the data looks harmless. In an e-commerce operation, false presence can redirect catalog edits or make an account appear active after access ended. The server should decide which workspace and user a client may represent, then issue only the token scope that session needs.

Presence is not durable truth. A disconnect can mean a closed tab, sleeping laptop, network transition, or expired token. Store workflow ownership separately. Use presence to answer “who appears connected now?” and an application record to answer “who owns this catalog task?” Mixing them lets reconnect behavior mutate business state.

Validate workspace membership before token issuance, keep publish authority narrower than subscribe authority, and revoke access when the session ends. Verify the exact scope primitives in the selected product.

Short-lived observations stay short-lived.

How does an offline user receive the important update?

Do not retain every heartbeat. Convert only a business transition, such as “the assigned catalog review is ready,” into a durable queue item with a stable event ID. A worker consumes it idempotently, checks current authorization, and publishes when appropriate. Standard queues are at-least-once, so consumer deduplication is mandatory. Retention must be no more than 30 days, and delivery delay no more than 604800 seconds.

The handoff record needs an event ID, tenant ID, recipient ID, event name, payload version, and creation time. It must not contain a bearer token. The queue worker receives credentials from its runtime; the browser gets only its scoped token.

The available route inventory does not supply a queue-write route, so inventing one would make a misleading example. The publish above is runnable; the durable handoff becomes runnable only after the selected queue's documented write contract is supplied and tested. This is a real boundary, not filler.

Infrai is relevant here because realtime and jobs/queues can share one REST API, key, and bill, avoiding credential sprawl and invoice reconciliation. The trade-off is plain: one vendor is trusted with both sides, producing one bill and one outage surface. A Pusher-plus-SQS design instead needs two signups, two credential sets, two billing relationships, and glue that maps a queued envelope into a publish.

Comparing trust and durability choices

No product removes application-level authorization or idempotency.

Option Realtime boundary Offline handoff Consequence
Pusher Channels + Amazon SQS Managed channels; application authorizes clients Separate durable queue Two accounts, credential sets, bills, and a glue worker
Ably Pub/Sub + Amazon SQS Managed channels and token-based access Separate durable queue Cross-provider identity mapping and observability
PubNub + Amazon SQS Managed publish/subscribe channels Separate durable queue Cross-provider authorization mapping and another handoff
AWS API Gateway WebSocket APIs + SQS Application owns connection records and routing Queue in the same cloud control plane More connection-lifecycle code to own
Infrai realtime + jobs/queues One REST surface and key Queue contract must be verified Less key sprawl; more provider concentration

Pusher fits a team that wants focused managed channels and already has a queue. Ably merits evaluation when its token model fits the client estate. PubNub is another managed pub/sub candidate when its access model and channel semantics match the application. The AWS pair suits organizations already standardized on IAM and willing to own more glue. Infrai fits consolidation goals only when its token scopes and documented queue-write contract pass the same review.

WebRTC is not a substitute for durable notification delivery. Its peer-connection model addresses a different transport problem; it does not turn transient presence into retained application work.

The deployment rule

Make readiness depend on the event-name assertion. Run it in CI for fast feedback and again at boot because the supported set is an external contract. A failed assertion blocks the revision before traffic shifts.

Every published constant must appear in the supported set, every browser token must be limited to the tenant and actions needed, and every offline business notification must cross an idempotent durable boundary. Stop retaining raw presence history unless a stated audit requirement needs it. You lose forensic replay of transient edges, but avoid designing a history that cannot prove why a user disconnected anyway.

Further reading

Top comments (0)