DEV Community

VespasianBlack3884
VespasianBlack3884

Posted on

Logistics Support Chat Leases: Python Controls for Presence Expiration and Reconnects

Use a short, server-owned lease for presence, then make every reconnect re-authenticate and backfill from an ordered event log. That decision keeps a logistics support chat from showing a driver as available after a tablet disappeared, while still preserving messages sent during a tunnel outage. Presence is a hint; authorization and message history are the controls.

Architecture Decision Record: the invariants

The system has four invariants. A presence lease expires without client cooperation. A user can read or write only rooms allowed by the current authorization check. Every accepted message gets a monotonic room sequence. Finally, a reconnect is a new session, not a continuation of a trusted socket.

Those boundaries matter more than whether the transport is WebSocket, Server-Sent Events, or a WebRTC data channel. WebRTC describes peer connection behavior, but it does not define your room membership, retention, or access policy. Keep those decisions in the application layer and validate them at every session boundary.

Here is the decision record I would put next to the design:

Option Delivery at fan-out Presence security Recovery cost Appropriate use
Broadcast directly from each socket Best effort; disconnects lose events Hard to revoke consistently High client complexity Low-risk typing indicators
Broker pub/sub without a log Fast fan-out, no replay guarantee Lease still required Medium; gaps are ambiguous Ephemeral dashboards
Ordered log plus fan-out workers Replay from a sequence, measurable lag Server validates every lease Predictable storage and worker cost Customer support rooms

The third option fits support conversations because an unanswered delivery update is an operational incident, not decoration. It also gives compliance staff an audit trail without treating presence as evidence that a person saw a message; it's a record of connectivity, and it doesn't prove attention.

How should realtime presence expiration secure a customer support chat?

Give each connection a lease such as expires_at, scoped to a user, device, and room. A heartbeat can renew it, but only after the server checks the session token and room grant. Store the lease in a monotonic-time domain, and publish an offline transition when it expires. Do not let a browser timestamp decide that transition; clock skew turns a security control into a suggestion.

Presence is disposable state.

For a logistics tenant, I use separate identities for a dispatcher and a carrier contact even when both appear in the same room. A revoked carrier grant must stop new sends immediately. Existing messages remain subject to retention and legal hold, while the presence record can disappear after a short grace period. The catch is that aggressive expiration creates false offline states on cellular networks, so the UI should label presence as “last heartbeat” and never infer read status from it.

The critical path can stay small. This Python sketch uses generic interfaces so the policy is testable without a particular broker:

from dataclasses import dataclass
from time import monotonic

@dataclass(frozen=True)
class Lease:
    user_id: str
    device_id: str
    room_id: str
    deadline: float

def renew_presence(lease: Lease, token: str, now: float, ttl: float,
                   auth, leases, events) -> Lease:
    if not auth.session_is_valid(token, lease.user_id, lease.room_id):
        raise PermissionError("session is not authorized")
    refreshed = Lease(lease.user_id, lease.device_id, lease.room_id, now + ttl)
    leases.put(refreshed)
    events.append("presence.renewed", {
        "user_id": lease.user_id, "device_id": lease.device_id,
        "room_id": lease.room_id, "expires_at": refreshed.deadline,
    })
    return refreshed

def expire_due(leases, events, now: float) -> int:
    expired = 0
    for lease in leases.due_before(now):
        if leases.remove_if_current(lease):
            events.append("presence.expired", {
                "user_id": lease.user_id, "device_id": lease.device_id,
                "room_id": lease.room_id,
            })
            expired += 1
    return expired

# A monotonic clock prevents wall-clock adjustments from extending a lease.
now = monotonic()
Enter fullscreen mode Exit fullscreen mode

The compare-and-remove operation is important. Without it, an old expiry job can delete a newer lease and flicker a user offline. I once diagnosed that as a fan-out bug because the symptom was a missing “agent available” event; the actual mistake was allowing a stale timer to win a race. The log made the sequence visible: presence.renewed, then an incorrectly applied presence.expired, then a reconnect.

Reconnects, backfill, and delivery guarantees

On connect, authenticate, authorize the room, and return the newest retained sequence before accepting live fan-out. The client sends its last applied sequence on reconnect. The server replays (last_applied, current], then attaches the socket to the live stream. If the requested sequence is outside retention, return a resync instruction and fetch a room snapshot; silently skipping the gap is not acceptable for a support transcript.

Use idempotent message identifiers at the write boundary. A retry after a mobile timeout can then produce one room sequence, not two customer-visible replies. Fan-out workers should acknowledge a sequence only after durable handoff to their queue, and metrics should distinguish publish latency, replay latency, and consumer lag. “Delivered” must have a defined recipient boundary: queued, socket-written, or client-acknowledged are different claims.

Keep presence out of the message guarantee. A contact can be offline while a message is durably stored, and a green presence dot can coexist with an unread queue. This separation prevents a privacy leak in which support staff infer a driver's location or work shift from heartbeat precision.

Failure boundaries and the rejected shortcut

The rejected shortcut is a client-only timeout: set online = false after 30 seconds and reconnect when the flag changes. It is easy to ship and unsuitable when a stolen token, revoked room grant, or shared tablet is in scope. It also fails under background-tab throttling, which can pause JavaScript timers for minutes while the radio remains connected. A second browser window can keep rendering the old green dot even after the account is disabled, and a device that wakes from a long sleep can send a burst of stale heartbeats. Use the client flag only as a presentation fallback while the authoritative lease and authorization checks run on the server; the server must be able to revoke a room without waiting for a tab to wake up. That extra check is what turns an expiration timer into a security boundary instead of a cosmetic timeout.

Testing should exercise time, not sleep. Inject a clock, advance it past the deadline, and assert one expiration event. Race a renewal against expiry and assert that the newer lease survives. Reconnect with sequence 41 after the room has reached 47 and assert replay of 42 through 47 in order. Add property tests for duplicate message IDs and authorization revocation between heartbeat and send.

Your mileage may vary on lease length. Network radio wakeups, regulatory retention, and the sensitivity of presence metadata all affect the number; measure reconnect rates and false-offline reports in your own fleet before choosing it. Stick with a simpler pub/sub design when rooms are disposable and a missed event has no business consequence.

References

Top comments (0)