DEV Community

FitzgeraldBlake3561
FitzgeraldBlake3561

Posted on

How to Budget Realtime Score Retention Across Cellular Handoffs: API Boundary Patterns

Short answer: make the feed a replayable log with an explicit cursor, and let the client switch from live delivery to bounded backfill whenever a mobile handoff breaks the stream. The API boundary should describe event identity and retention, not promise that one connection survives every tower change.

Start with the retention bill

A sports score feed has two very different costs: delivering current events and retaining enough history to repair a missed interval. Delivery is usually the visible line item, but retention determines whether reconnects are cheap and predictable or turn into database scans and duplicate fan-out. Measure both before choosing a transport.

For each event, keep a compact envelope: game ID, sequence, event type, server timestamp, and a payload such as the score. A 90-second replay window may be enough for a casual scoreboard; a betting or officiating workflow may need several minutes and stronger audit storage. Those are product decisions, not defaults hidden inside an SDK. I write the retention number into the contract so operations can alert on it.

The arithmetic is easiest to miss when the match is busy. Suppose a game emits a burst of scoring, substitution, and clock events, while tens of thousands of phones reconnect during a train ride through a tunnel. The hot log now pays for each event once in storage and again when a gap is replayed. Keeping a compact event for 90 seconds can be cheaper than materializing a full score snapshot for every reconnect, but that only holds when the snapshot is bounded and the replay query is indexed by (game_id, sequence). If the query walks timestamps, a clock skew or duplicate timestamp turns a small recovery into a broad scan. I budget storage by events per game and replay bytes per reconnect, then load-test the busiest interval rather than averaging across an entire day. Your accounting system may classify these lines differently; the engineering decision remains the same: retention is part of the API contract, and its expiry behavior must be observable.

Keep it boring.

The catch is that a longer window consumes storage and replay bandwidth even when handoffs are rare. A shorter window lowers that bill but forces a full snapshot more often, and a snapshot can be heavier than a handful of score events. I deliberately stop keeping per-device delivery receipts in the hot log. When an incident requires forensic detail, that information lives in sampled telemetry or a separate audit store, with its own access controls and retention policy.

What should mobile network handoffs and API boundaries guarantee?

The API should guarantee ordering within a game, durable event IDs, and a clear recovery response. It should not guarantee packet delivery or an uninterrupted TCP/WebSocket session. Cellular changes can invalidate an address, pause radio traffic, or move a device between IPv4 and IPv6 paths. Treat the connection as disposable.

One workable contract uses a live stream plus a backfill endpoint. The client sends its last applied cursor; the server either returns the missing events or says the cursor is outside the replay window and includes a fresh snapshot. Keep the route names boring and stable. For example, a service can expose GET /score-feed/live for the stream and GET /score-feed/replay?game_id=...&after=... for recovery. These are illustrative interface names, not a requirement for a particular vendor.

The replay response needs an unambiguous boundary. If after=1842 returns events 1843 through 1851, the next request must use 1851, not a timestamp guessed by the device clock. A monotonic sequence also makes deduplication straightforward when the client receives 1851 twice after a reconnect.

A reconnect loop that does not poll

Here is the core of a Python client. It uses a push transport for normal operation, then performs one bounded replay request after a disconnect. There is no timer fetching the latest score every few seconds.

from dataclasses import dataclass
import asyncio


@dataclass
class Cursor:
    game_id: str
    sequence: int = 0


async def consume_game(feed, cursor: Cursor):
    while True:
        try:
            async for event in feed.live(game_id=cursor.game_id,
                                         after=cursor.sequence):
                if event.sequence <= cursor.sequence:
                    continue
                if event.sequence != cursor.sequence + 1:
                    # Ask for the gap before applying a later event.
                    missing = await feed.replay(
                        game_id=cursor.game_id, after=cursor.sequence
                    )
                    for item in missing.events:
                        apply_score(item)
                        cursor.sequence = item.sequence
                    if missing.snapshot_required:
                        snapshot = await feed.snapshot(game_id=cursor.game_id)
                        apply_snapshot(snapshot)
                        cursor.sequence = snapshot.sequence
                apply_score(event)
                cursor.sequence = event.sequence
            raise ConnectionError("live stream ended")
        except (ConnectionError, TimeoutError):
            await asyncio.sleep(1)


def apply_score(event):
    pass


def apply_snapshot(snapshot):
    pass
Enter fullscreen mode Exit fullscreen mode

The one-second delay is a backoff floor, not a polling interval: the client is waiting for a stream to become available. In production, cap exponential backoff, add jitter, and stop retrying when authentication is invalid. A 429 response should honor Retry-After; blindly reconnecting can turn a tower handoff into a rate-limit storm.

The server must make replay idempotent. A client crash after rendering sequence 1851 but before persisting it is normal. Applying 1851 again should be harmless, while applying 1853 without 1852 should be rejected or repaired. Persist the cursor with the rendered state in one local transaction when possible.

Testing the failure path before launch

Most handoff bugs hide in transitions, not in a clean Wi-Fi demo. Test radio disable/enable, airplane mode, captive portals, IPv6-only networks, and a device that sleeps for longer than the replay window. Inject a disconnect between every pair of events and assert that the final state equals a clean snapshot plus the suffix of the log.

I also test malformed and adversarial input. A sequence that jumps backward, a duplicate event, or a payload with an unexpected game ID must not corrupt another game in the cache. Authentication tokens need refresh without losing the cursor. For a logistics dispatch app that sends similar in-app alerts, the same rule prevents a driver from seeing an old route update after a newer one.

Instrument the boundary, not just the socket. Record reconnect count, replay length, snapshot fallbacks, cursor age, and the percentage of clients whose cursor falls outside retention. Keep payload sampling privacy-aware: score feeds are public, but device identifiers and notification metadata may not be.

An event log plus replay is a good fit when updates are small, ordering matters, and a client can recover from a short outage. It is not suitable when every message must be delivered exactly once to a regulated ledger, or when the payload is a large media object; use a transactional queue or object transfer with explicit acknowledgements in those cases. Stick with a simpler snapshot endpoint when the UI only needs the current score and a missed intermediate play has no user-visible effect.

WebRTC data channels can provide low-latency peer or server paths, but the W3C recommendation does not remove the need for application-level cursors and replay. Managed realtime services can reduce connection plumbing; evaluate their retention limits, authentication model, and rate semantics against this contract rather than assuming reconnect behavior is identical. Your mileage may vary because radio firmware and carrier NATs differ across regions.

The decision rule is small: define the cursor, state the replay window, specify the snapshot fallback, and publish retry semantics. Then price retention and replay traffic with real traces. A transport is an implementation detail after those boundaries are testable.

References

Top comments (0)