Short answer: keep a short, permission-checked replay window for a delivery tracking map, and send current cursor positions through a separate ephemeral channel. The bill and the risk are both driven by retained events, so measure bytes and replay demand before choosing a longer history.
The bill starts with retained bytes
A map editor for a logistics team has two very different streams. A courier's location is useful for a while, while a collaborative cursor is useful only while someone is editing the same dispatch note. Treating both as durable history makes the storage bill grow in the least valuable place and gives a leaked client more to read.
The first estimate is mechanical: event rate multiplied by encoded event size, multiplied by retention time, plus indexes and replication. If 2,000 active editors emit one cursor event every two seconds and each event is 180 bytes after encoding, one hour is roughly 648 MB before overhead. That is a planning input, not a benchmark; compression, fan-out, and replication change the result. The important point is that retention time is a direct multiplier. I've seen teams argue about storage engines for weeks while leaving that multiplier untouched, then discover during a shift change that replay reads, not writes, were saturating the service.
I keep the cursor stream in a bounded ring or log segment. A new subscriber can replay enough state to draw the map, then receives live updates. Once a cursor has been superseded, retaining every intermediate pixel position adds forensic weight without adding product value.
Keep it short.
The deliberate loss is precise: after the window expires, an operator cannot reconstruct every cursor movement. They can still inspect the latest checkpoint and the audit records that matter for dispatch decisions. That trade is acceptable for cursors, but it would be wrong for a proof-of-delivery record or a payment event.
How should a realtime event retention policy scale delivery tracking map updates?
Start with event classes, not a single global TTL. A practical policy for this map might look like this:
| Event class | Purpose | Retention decision | Replay behavior |
|---|---|---|---|
| Cursor position | Show where an editor is working | Seconds to a few minutes | Latest position per client |
| Vehicle location | Show operational movement | Short operational window, then aggregates | Last known point plus sampled track |
| Dispatch edit | Explain who changed an assignment | Durable according to the audit requirement | Ordered replay with authorization |
| Delivery proof | Support a customer or legal dispute | Durable storage with explicit access policy | Never exposed by a cursor token |
This split protects the map from a common failure mode: a reconnecting browser asks for a huge replay, the service reads an entire day of high-frequency points, and live delivery falls behind. The fix is not merely a larger machine. Cap the replay range, return a current snapshot, and make the client request older history through a separate, slower path. The trade-off is that a reconnect may lose intermediate cursor motion; that is acceptable only because the product needs the current editing location, not a film of every mouse move.
Tokens should express that boundary. A cursor token can carry a subject, map identifier, allowed event classes, an expiry, and a maximum replay horizon. It should not be a general-purpose database credential. The server still checks the token on every subscription and filters events before fan-out; trusting a field supplied by a browser is not an authorization model.
Here is the shape I use for a small policy function. It is intentionally boring: explicit scopes make review easier.
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
@dataclass(frozen=True)
class ReplayGrant:
subject: str
map_id: str
scopes: frozenset[str]
expires_at: datetime
max_age: timedelta
def authorize_replay(grant: ReplayGrant, map_id: str, event_kind: str,
requested_since: datetime, now: datetime) -> bool:
if grant.map_id != map_id or now >= grant.expires_at:
return False
if event_kind not in grant.scopes:
return False
if requested_since < now - grant.max_age:
return False
return True
def cursor_payload(client_id: str, x: float, y: float, revision: int) -> dict:
return {"client_id": client_id, "x": x, "y": y, "revision": revision}
The coordinates are not a security boundary. The server derives the tenant and map from authenticated context, validates the revision, and limits payload size. If the client needs to share a cursor with another editor, it sends a semantic event such as “editing stop 42,” not a token that can read the whole dispatch database.
What fails when retention and trust are coupled?
Three failure modes recur in production designs.
First, a long-lived token outlives the person who received it. A stolen browser profile then becomes a replay reader. Short expiry helps, but revocation and audience checks still matter for sensitive events.
Second, deduplication is mistaken for ordering. A reconnecting client may receive an event twice or miss a transient cursor update. Give each event a monotonic revision within its map, make handlers idempotent, and let a snapshot repair a gap. Do not promise global ordering across every vehicle and editor; that promise is expensive and rarely useful.
Third, retention is measured only in storage. Egress and replay reads can dominate when a large dispatch center opens the map at shift change. Track accepted events, bytes retained, bytes replayed, reconnect rate, and time from snapshot to live delivery. Alert when replay traffic consumes the same capacity as live traffic.
WebRTC can carry peer media and data channels, but its recommendation does not remove the need for an application-level authorization and retention policy. A direct channel may reduce relay traffic in some topologies; it also makes peer membership, NAT traversal, and observability different problems. The browser still needs a trusted service to decide which map events it may receive.
A retention change needs a reversible rollout
I prefer a two-phase change. First, record what would be discarded while continuing to serve the old window. Compare the proposed cutoff with incident investigations and customer support cases. Then enforce the cutoff for one map or tenant, with a feature flag that can restore the previous window without rewriting history.
The test suite should include an expired token, a token for the wrong map, a request older than max_age, duplicate revisions, and a reconnect that starts from a snapshot. Load tests should vary fan-out and replay size independently; a system that passes a steady live stream can still fail when 500 clients reconnect together.
Keep an audit trail separate from the cursor feed. It should record authorization decisions and durable dispatch changes, not every heartbeat. Redact coordinates where a support log does not need them, and make retention deletion observable so compliance staff can verify that a policy actually ran.
The boundary I would keep
The least complex design is a short-lived, narrowly scoped cursor token, a bounded event buffer, and a snapshot endpoint for reconnects. Add longer vehicle history only when operators can name the decision it supports. Add durable dispatch history when an audit or legal requirement demands it.
This is not suitable when the map itself is the system of record, when offline clients must replay weeks of edits, or when regulators require immutable location history. In those cases, keep the realtime feed thin and choose a durable event store with an explicit retention schedule. Stick with a durable audit log when a dispatcher must prove the exact sequence of assignment changes; a bounded cursor buffer cannot provide that evidence. Your mileage may vary because the correct window depends on incident response and contractual requirements, not on a fashionable default.
The useful stopping rule is simple: stop retaining data when it no longer changes an operational decision, and stop trusting a token when its scope cannot be explained in one sentence.
References
- W3C, “WebRTC: Real-Time Communication in Browsers”: https://www.w3.org/TR/webrtc/
- RFC 9449, “OAuth 2.0 Demonstrating Proof of Possession (DPoP)”: https://www.rfc-editor.org/rfc/rfc9449
- RFC 8725, “JSON Web Token Best Current Practices”: https://www.rfc-editor.org/rfc/rfc8725
Top comments (0)