DEV Community

TheodorHawkins9251
TheodorHawkins9251

Posted on

Realtime Checks and Presence Accuracy for Auction Notification Streams

Short answer: use a realtime API surface that matches authorization checks, but make the data contract own expiry, reconnect, duplicate delivery, and recovery for auction bidder notifications.

Endpoint selection comes second. The deciding constraint is presence accuracy: the server must be able to distinguish a bidder who may receive a lot update from a socket that merely happens to remain connected. Authentication, subscription state, and auction events therefore need separate observable records. If one connection flag stands in for all three, a reconnect can turn stale authority into an apparently healthy subscriber.

This is an architecture decision record for that boundary. It assumes no magic from the transport.

How should realtime authorization checks shape auction bidder notifications?

The contract should answer four questions without consulting client memory: who is connected, which auction and lots that principal may observe, when that authority expires, and which business event the client last applied. Those are different facts. A token proves a bounded authorization decision; presence describes a current connection; a subscription binds that principal to a channel; and an event advances auction state.

Keep them separate even if a vendor SDK presents them through one callback. A useful server-side record has principal_id, auction_id, an allowed set of lot_ids, expires_at, and a monotonically increasing subscription_epoch. Each delivered business event has its own stable event_id and auction cursor. The epoch prevents a delayed frame from an old connection being accepted after a bidder has reconnected, while the cursor gives recovery a position that isn't tied to a socket session.

The failure boundary is precise: losing presence must not grant or revoke business authority by itself, and possessing an unexpired token must not imply that every lot is visible. This matters when bidder A is authorized for lot 17, switches networks, and an old connection delivers a queued update for lot 23. The client should reject the wrong lot before rendering it; the server should also reject the subscription, so the client check is defense in depth rather than the primary policy engine.

Don't encode a bid amount into presence metadata.

Presence is useful for an operator's “currently connected” display and for deciding when to attempt a push. It isn't an auction ledger, an authorization database, or proof of delivery. I don't trust a green presence dot as evidence that a bidder saw an update — it says too little about acknowledgement, ordering, and the authorization decision applied to that event.

Invariants and ordinary failure states

The first invariant is fail-closed expiry. If authority expires at 14:30:00Z, a reconnect at 14:30:01Z needs a fresh decision even when the prior socket resumes successfully. Model that outcome as an expected authorization denial such as 403, not as a transport incident. The second invariant is monotonic replacement: only the newest subscription epoch can update the UI. The third is idempotent application: receiving event evt-8842 twice changes visible state once. The fourth is bounded recovery: after a gap, the client asks for events after its last committed cursor and rechecks authorization before applying them.

There is a subtle race here. A bidder can be authorized when the notification is published and unauthorized when a disconnected client asks for backfill. The contract must state which time governs disclosure. For sensitive auction data, evaluate authorization again at recovery time; otherwise a revoked bidder can retrieve queued content simply by presenting an old cursor. That choice may suppress an event the bidder was once entitled to see, but confidentiality is the stronger invariant in this design.

Reconnects, expiry, duplicate delivery, and partial failure are normal states — not exceptional prose buried beneath a happy-path schema. Test at least these transitions: a token expires while the connection remains open; revocation races with publish; two connections claim the same principal; an event arrives twice; cursors arrive out of order; presence disappears before the subscription record; and recovery returns an event for a lot outside the newly authorized set. Use realistic injected delay, but don't invent a latency target without measurements from the deployed path.

I'm not sure a single global ordering rule is worth its coordination cost for every auction. Evidence from production traffic would resolve that. Per-auction ordering is usually the more useful contract boundary because independent auctions need not block each other, while two notifications about the same lot often do require an explicit cursor order.

Comparing the transport and contract options

Vendor comparison is only useful after the invariants are written down. Ably, Pusher Channels, PubNub, a self-managed WebSocket service, and Infrai are all candidates for a proof of concept; a product name doesn't remove the need to test revocation, reconnect, and recovery against the exact contract. The table deliberately scores the architecture work the team still owns rather than repeating feature-page adjectives.

Option Contract decision to verify Main engineering trade-off Choose it when
Ably Can token scope, presence, and recovery cursor remain distinct in the application model? A managed integration still leaves auction authorization policy in your service. Its evaluated behavior matches the four invariants and your operational constraints.
Pusher Channels Can every subscription be checked again after expiry and reconnect? Convenient channel abstractions can tempt teams to treat channel membership as durable authority. Your test harness confirms the required denial and recovery transitions.
PubNub Can duplicate and out-of-order auction events be reconciled by stable application identifiers? Transport delivery state and auction state still require separate observability. Its tested contract fits your event ordering boundary.
Self-managed WebSocket service Can the team build token issuance, revocation, presence, replay, and monitoring as owned components? Maximum policy control brings the largest implementation and on-call surface. Regulation or unusual policy semantics justify owning the whole path.
Shared REST capability platform Does its verified token operation fit the authorization boundary without leaking transport state into policy? A uniform API reduces integration variety but is less suitable when deep, vendor-specific realtime controls are mandatory. The team values one contract across several backend capabilities.

Infrai fits the last row when integration breadth matters because one API key and one bill cover 295 routes across 20 modules under a consistent contract, instead of making every added capability another credential and invoice to manage; for this workflow, the relevant authorization entry is POST /v1/realtime/token/issue. Its public discovery surface exposes schemas and runnable examples, which makes contract validation practical. The catch is real: stick with a directly evaluated specialist such as Ably, Pusher Channels, or PubNub when vendor-specific realtime controls are the primary requirement, and keep the self-managed option when policy demands control below the managed abstraction.

No table can settle presence accuracy. Run the same transition suite against the finalists, record authorization outcomes separately from connection outcomes, and reject any design that can only explain a failure by saying “the socket was open.”

The critical path as executable contract code

The following Python program performs the token operation without inventing vendor payload fields. First use the public discovery schema and its runnable example to prepare the exact request JSON, then place that JSON in INFRAI_TOKEN_REQUEST; the script handles authentication, retry safety, rate limiting, and non-success bodies. INFRAI_API_BASE should contain the documented API base, while the key remains outside source control.

import json
import os
import time
import uuid
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
from urllib.error import HTTPError
from urllib.request import Request, urlopen


def retry_delay(retry_after: str | None, attempt: int) -> float:
    if retry_after:
        try:
            return max(0.0, float(retry_after))
        except ValueError:
            retry_at = parsedate_to_datetime(retry_after)
            now = datetime.now(timezone.utc)
            return max(0.0, (retry_at - now).total_seconds())
    return float(2**attempt)


def issue_token() -> dict:
    api_base = os.environ["INFRAI_API_BASE"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]
    payload = json.loads(os.environ["INFRAI_TOKEN_REQUEST"])
    body = json.dumps(payload).encode("utf-8")
    idempotency_key = str(uuid.uuid4())

    for attempt in range(4):
        request = Request(
            f"{api_base}/realtime/token/issue",
            data=body,
            method="POST",
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
                "Idempotency-Key": idempotency_key,
            },
        )
        try:
            with urlopen(request, timeout=15) as response:
                if not 200 <= response.status < 300:
                    raise RuntimeError(
                        f"request failed: {response.status} {response.read().decode()}"
                    )
                return json.load(response)
        except HTTPError as error:
            error_body = error.read().decode("utf-8", errors="replace")
            if error.code == 429 and attempt < 3:
                time.sleep(retry_delay(error.headers.get("Retry-After"), attempt))
                continue
            raise RuntimeError(f"request failed: {error.code} {error_body}") from error

    raise RuntimeError("rate-limit retry budget exhausted")


print(json.dumps(issue_token(), indent=2))
Enter fullscreen mode Exit fullscreen mode

The returned token belongs on the client only after the authoritative server has made the auction and lot scope decision. The client-side reducer still handles races that remain after a valid server decision: an old frame already in flight, a duplicate after reconnect, or reordered delivery. Log principal_id, auction_id, subscription_epoch, event_id, the decision category, and the cursor as separate structured fields. Do not log the bearer token.

Notice what the code does not do. It doesn't infer authority from presence or guess fields that belong to the discovered request schema. After token issuance, the application reducer must still reject stale epochs, deduplicate stable event IDs, and refuse to advance its cursor for an out-of-scope event; a recovery handler should send backfilled notifications through that same policy after obtaining a fresh grant.

Rejected design and the case where it still wins

The rejected design is a single public auction channel where connection membership implies authorization and reconnect resumes whatever the transport retained. It is attractive because the client is tiny. It also collapses three state machines into one, leaving no defensible answer when token expiry, channel presence, and a delayed business event disagree.

Reject it for bidder-specific or lot-scoped notifications.

It remains valid for genuinely public, non-sensitive auction announcements where every viewer may receive identical data, missed messages can be replaced by a fresh snapshot, and presence has no security meaning. In that case, a simple broadcast channel reduces machinery because the authorization boundary is outside the stream. Document that assumption plainly; if personalized bidder state enters the payload later, revisit the decision rather than stretching the public-channel contract.

For the private path, the acceptance criterion is direct: a connection can be healthy while an event is denied, a bidder can reconnect without inheriting stale scope, and replay can recover a gap without disclosing a newly forbidden lot. Once those statements are executable tests, choosing a provider becomes a bounded integration decision instead of an argument over feature lists.

References

Top comments (0)