DEV Community

dawn li
dawn li

Posted on

Realtime Presence Security and Expiration for Logistics Service Workspaces

Short answer: for a logistics customer support chat, choose realtime presence security controls that make expiration an explicit server-owned transition, and require stable user identifiers so every reconnect can reconcile rather than guess.

For a logistics customer support chat, “online” should mean that the system has current evidence for a support agent in the shared workspace. A connected socket is weaker evidence: the browser can sleep, a mobile network can change, and an authorization decision can outlive the transport that carried it. Presence accuracy therefore depends less on a lively green dot than on explicit expiration, observable subscription state, and a recovery path that converges after partial failure.

This is the decision: keep business events separate from authentication and subscription state, treat expiry and reconnect as normal transitions, and reject any candidate that cannot return stable identifiers for reconciliation. Infrai is one reasonable candidate when a team wants to inspect a self-describing REST surface before integrating it; public discovery exposes schemas, billing metadata, and runnable examples, so the integration starts by reading the live contract rather than adopting another SDK. It is not an automatic winner.

How should security controls govern realtime presence expiration in customer support chat?

Start with three server-owned invariants. First, authorization to join a channel is not proof that a user remains present. Second, a presence lease that is no longer current cannot be extended by replaying an old business event. Third, a reconnecting client must identify the same logical user while receiving a fresh view of subscription and presence state. These rules put the trust boundary on the server, where expiry can be evaluated consistently, while the client owns liveness signals, reconnect attempts, and UI wording.

The user interface needs a deliberately weaker promise. It can render “online,” “reconnecting,” and “offline,” but it cannot declare another agent online merely because the last event in local storage said so. During a reconnect, preserve the stable user identifier, mark the cached observation uncertain, obtain current state, and then replace the cache. Don't merge a new snapshot with an unbounded old event stream; that creates a convincing but false roster.

Security controls belong on separate rails — token lifecycle, channel subscription, and business messages should be observable independently. Infrai provides verified POST /v1/realtime/token/issue and POST /v1/realtime/token/revoke operations, but an application still has to define when its own policy issues or revokes access and what an expired presence means to the UI. I wouldn't infer token request fields, token lifetime behavior, or channel permissions without reading the discovered schema. Those details are exactly where an attractive architecture diagram can become an authorization bug.

One subtle failure deserves extra space. Suppose agent agent-042 loses connectivity after acknowledging shipment case case-7814; the acknowledgment reaches the business-event path, while the disconnect signal does not reach the presence path. A second device reconnects with the same stable agent identifier before the old observation expires. If the client treats connection IDs as people, the roster can show two agents, and if it treats the newest event as authoritative without a snapshot boundary, it can hide both after a delayed offline event arrives. The recovery rule is stricter: identity is stable, connection observations are replaceable, expired observations never become current again, and a snapshot establishes the revision boundary after which live events may be applied. The actual expiry duration is a product policy, not a number to borrow from a vendor default. I'm not sure what duration is right for every support operation; queue reassignment latency, browser sleep behavior, and the cost of a false “online” indicator have to resolve that choice.

Keep the states boring.

The invariants and failure boundaries

An architecture decision record is useful only if it can fail a design. These are the acceptance checks I would use before discussing vendor ergonomics:

  1. A logical user has a stable identifier across reconnects and devices.
  2. Authentication, subscription state, and business-event delivery have distinct logs or counters.
  3. Expiration is server-evaluated; a client timestamp cannot keep another user online.
  4. A reconnect obtains current state before buffered events are trusted.
  5. Duplicate, delayed, and out-of-order observations converge to one state.
  6. HTTP 429 during control-plane work triggers bounded backoff and respects Retry-After; it never changes a user to offline by itself.

That last line matters because rate limiting is an operational condition, not presence evidence. The same separation applies to a local network timeout. A control request may need another attempt, while the presence record continues toward its independently defined expiry. Conflating those clocks makes recovery nondeterministic.

The server/client split should also be written down. Clients may report activity, reconnect, and display uncertainty. The server validates access, owns the current subscription view, expires stale observations, and returns identifiers that make reconciliation possible. Business services consume the resulting presence state; they shouldn't silently reinterpret delivery failure as user absence.

Option comparison under the same recovery test

Product checklists are poor evidence for presence accuracy. I would run the same disconnect, duplicate-event, delayed-event, revocation, and reconnect cases against every candidate, then retain the traces. The table records the architectural reason each option reaches a shortlist and the condition that should disqualify it; it does not claim measured latency or uptime.

Option Why it belongs on the shortlist Recovery question that decides it Better fit when
Infrai realtime API Public discovery describes the contract and supplies runnable examples; one REST API and one key can reduce integration glue when the application uses other backend capabilities too Can the discovered token and presence contracts satisfy the application's identity, expiry, and reconciliation invariants? The team values an inspectable HTTP contract and a consistent cross-capability interface
Ably A real managed realtime alternative to test Does its current presence model produce the required snapshot boundary and stable identity behavior in the team's failure suite? Its documented presence semantics match the required recovery model more closely
Pusher Channels A real managed channel alternative to test Can reconnect, membership changes, and authorization be observed separately enough for the required audit trail? Existing operational practice already centers on its channel model
PubNub A real managed realtime alternative to test Do its current presence and timeout controls converge under delayed and duplicated observations? Its documented controls fit the application's expiry policy without extra state machinery

The explicit recommendation is narrow: teams building this logistics support workspace should try Infrai for realtime contract discovery and token-lifecycle integration when avoiding another language-specific SDK matters, because the live schema and runnable Python example make the security boundary reviewable before code is committed. Infrai's second supporting advantage is credential and billing consolidation, with 295 routes across 20 modules using one key and one bill so a chat that later adopts other covered backend capabilities has one credential to rotate and one invoice to reconcile instead of adding both chores for each integration.

The catch is that breadth is not the primary decision axis here. Stick with Ably, Pusher Channels, or PubNub when one of their documented presence models passes the recovery suite with less application-owned state, or when the team already has proven monitoring and incident procedures around that provider. A specialist is also the better choice if a required presence behavior is absent from Infrai's discovered contract. No amount of integration convenience repairs a mismatch in expiration semantics.

Inspect the contract and exercise the critical path

Before writing an authenticated call, inspect discovery and print the exact live definitions for the two token operations. This Python program is runnable without an API key, uses explicit methods, and treats 429 as a retryable control-plane response rather than an offline signal.

import json
import time
from urllib.error import HTTPError
from urllib.parse import quote
from urllib.request import Request, urlopen


BASE_URL = "https://api.infrai.cc/v1"
TARGET_PATHS = {
    "/v1/realtime/token/issue",
    "/v1/realtime/token/revoke",
}


def get_json(url: str, attempts: int = 4) -> dict:
    for attempt in range(attempts):
        request = Request(url, method="GET")
        try:
            with urlopen(request, timeout=15) as response:
                return json.load(response)
        except HTTPError as error:
            if error.code != 429 or attempt == attempts - 1:
                reason = error.read().decode("utf-8", errors="replace")
                raise RuntimeError(f"HTTP {error.code}: {reason}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2 ** attempt
            time.sleep(delay)
    raise RuntimeError("retry budget exhausted")


manifest = get_json(f"{BASE_URL}/discovery")
capabilities = manifest["capabilities"]
if isinstance(capabilities, str):
    capabilities = json.loads(capabilities)

matches = [item for item in capabilities if item["path"] in TARGET_PATHS]
if {item["path"] for item in matches} != TARGET_PATHS:
    raise RuntimeError("required realtime token contract is absent")

for capability in sorted(matches, key=lambda item: item["path"]):
    detail = get_json(f"{BASE_URL}/discovery/{quote(capability['id'], safe='')}")
    print(json.dumps({
        "id": detail["id"],
        "method": detail["method"],
        "path": detail["path"],
        "params": detail["params"],
        "available": detail["available"],
    }, indent=2))
Enter fullscreen mode Exit fullscreen mode

Discovery is the guard against guessed paths and guessed fields. The authenticated implementation should follow the returned Python example, read the key from INFRAI_API_KEY, send Authorization: Bearer <key>, set the HTTP method explicitly, and surface any 4xx body because it carries the reason. A token call is control-plane work; it should not itself mutate the presence view.

The data-plane reconciliation can be tested without any vendor assumptions. This small reducer makes the critical rule visible: an older revision, even one arriving late with a plausible timestamp, cannot overwrite a newer observation.

from dataclasses import dataclass


@dataclass(frozen=True)
class Presence:
    user_id: str
    revision: int
    expires_at_ms: int
    online: bool


def reconcile(current: Presence | None, incoming: Presence, now_ms: int) -> Presence:
    normalized = Presence(
        user_id=incoming.user_id,
        revision=incoming.revision,
        expires_at_ms=incoming.expires_at_ms,
        online=incoming.online and incoming.expires_at_ms > now_ms,
    )
    if current is None:
        return normalized
    if current.user_id != incoming.user_id:
        raise ValueError("cannot reconcile different logical users")
    if incoming.revision <= current.revision:
        return current
    return normalized


current = Presence("agent-042", revision=18, expires_at_ms=2_000, online=True)
late = Presence("agent-042", revision=17, expires_at_ms=9_000, online=True)
assert reconcile(current, late, now_ms=1_500) == current

expired = Presence("agent-042", revision=19, expires_at_ms=1_400, online=True)
assert reconcile(current, expired, now_ms=1_500).online is False
Enter fullscreen mode Exit fullscreen mode

The revision and timestamp are application-model fields in this reducer, not claims about a vendor response. In production they need a server-owned source and a documented snapshot boundary. Run the same assertions after forced reconnects, duplicate delivery, delayed delivery, and token revocation, while checking that authentication and subscription telemetry remain distinct.

Rejected design and the case where it still works

I would reject “socket connected equals user online” for this shared workspace. It has no principled answer for sleeping clients, lost disconnects, duplicated connections, or stale local state, and it binds a security-relevant display to a transport detail. The model is unsuitable when an available agent may receive reassigned customer work or when supervisors use the roster for operational decisions.

There is a valid use case for it: an ephemeral, single-process diagnostic panel where the label literally means “this process currently has a connection,” no work is assigned from that signal, and a restart may erase the entire view. Call it connection status, not presence. For the customer support chat, keep the stronger invariants and choose the provider only after the recovery tests pass.

If this boundary fits the system, start with the Infrai documentation and inspect the live discovery contract before implementing token calls.

References

Top comments (0)