DEV Community

NorbertChristensen3183
NorbertChristensen3183

Posted on

Implementing Accurate Multiplayer Quiz Presence Through Client Subscription Teardown

Short answer: treat subscription cleanup as an auditable state transition, not a socket callback: stop accepting business events, record the player's intended departure, release the client subscription, and let the server confirm presence independently before the dashboard removes that player.

For a live media dashboard attached to a multiplayer quiz, the expensive part is rarely the cleanup function itself. The bill is made from whatever a provider meters while a session remains live, the events sent during reconnects, and the durable history retained for later reconciliation. Before choosing a service, express that exposure without guessing at a vendor price: live exposure = connected player-seconds + reconnect traffic, while audit exposure = departure records x retention interval. Measure those terms from your own telemetry. If stale presence stretches connected player-seconds long after viewers leave, shortening business-event retention won't cure the dominant term.

The useful change is to separate ephemeral presence from the compact audit trail needed to explain it. Keep the current roster only as long as the live game needs it; retain a departure record containing a player pseudonym, channel, client operation ID, reason, and server-observed time under a policy approved for your jurisdiction. Don't retain every heartbeat merely because storage exists. The trade-off is sharp: aggressive deletion limits data exposure and storage, but when a producer disputes who was eligible to answer, a compact transition journal can explain the decision while a vanished heartbeat stream cannot. I'm not sure what retention interval your contracts and local privacy rules permit; counsel and the incident-response requirement must settle that value, not a code sample.

How should realtime client subscription cleanup handle multiplayer quiz failures?

Give the client and server different jobs. The client owns intent: once the player taps Leave, the page hides, or a new quiz replaces the old one, it closes its local subscription exactly once and sends no more answers for the old round. The server owns authority: it authenticates the player, rejects business events outside the player's current round, observes presence separately, and writes the audit transition. A dashboard consumes the authoritative state rather than treating a browser's close callback as proof.

That distinction matters during partial failure. A client can lose its network after recording local intent but before the server observes departure; it can also reconnect while an earlier transport still appears alive. Model these as ordinary states: active, leaving, detached, and rejoining. Attach a stable operation ID to the departure intent so retries refer to one transition. This is an exactly-once mindset implemented over delivery that may duplicate: repeat observations are allowed, but only the first valid transition changes quiz eligibility.

Don't mix three ledgers. Authentication answers who may connect. Subscription state answers which channel the device follows. Business state answers whether a player may submit an answer. If a reconnect refreshes transport state, it must not silently reopen a closed quiz attempt. That invariant is more valuable than a clever disconnect handler because every recovery path can be checked against it.

One rule is enough: after leaving, no event from that subscription can alter the score.

Make the transition observable before automating deletion

Use a small transition record on the server, with a unique constraint on the client operation ID and an append-only audit entry in the same transaction as the eligibility change. The browser may request departure more than once. The handler should return the already-committed result for the same operation ID, yet reject reuse of that ID for a different player or channel. This pattern does not promise that the network delivers exactly once; it makes duplicate delivery harmless and leaves evidence for reconciliation.

Presence is a separate observation. Poll or query it from a trusted backend, compare it with the transition ledger, and expose disagreement as a metric rather than rewriting history immediately. A player marked leaving but still present may be inside an expiry or reconnect window. A player absent while still active may have lost connectivity. In either case, the dashboard can show a bounded transitional state while server authorization remains decisive.

The following program performs that trusted observation against the one verified presence route. It deliberately prints the response JSON without assuming member-field names, because a generated struct would be fiction unless it came from discovery. Every request states its method, checks status, and treats HTTP 429 as a recoverable control signal; Retry-After is honored as either seconds or an HTTP date.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func retryDelay(value string, fallback time.Duration) time.Duration {
    if seconds, err := strconv.Atoi(value); err == nil && seconds >= 0 {
        return time.Duration(seconds) * time.Second
    }
    if deadline, err := http.ParseTime(value); err == nil {
        if delay := time.Until(deadline); delay > 0 {
            return delay
        }
    }
    return fallback
}

func getPresence(ctx context.Context, baseURL, key, channel string) ([]byte, error) {
    pathTemplate := "/v1/realtime/presence/get/{channel}"
    path := strings.Replace(pathTemplate, "{channel}", url.PathEscape(channel), 1)
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, strings.TrimRight(baseURL, "/")+path, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := http.DefaultClient.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            delay := retryDelay(resp.Header.Get("Retry-After"), time.Second<<attempt)
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("presence request returned %d: %s", resp.StatusCode, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("presence request remained rate limited after 5 attempts")
}

func main() {
    baseURL := os.Getenv("REALTIME_API_BASE_URL")
    key := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || key == "" || len(os.Args) != 2 {
        fmt.Fprintln(os.Stderr, "set REALTIME_API_BASE_URL and INFRAI_API_KEY, then pass one channel")
        os.Exit(2)
    }
    body, err := getPresence(context.Background(), baseURL, key, os.Args[1])
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

Run it from a trusted environment; a browser must not receive the backend key. The raw result belongs in observation logs with the request time and channel, while any identity fields should be minimized or pseudonymized according to the retention policy. This probe does not unsubscribe a browser and should never be described as doing so. It verifies the server-side presence used during reconciliation.

Compare the operational contract, not the happy-path demo

All five options below can enter a serious evaluation, but their integration contracts differ. Ably documents presence around members entering, updating, and leaving channels. Pusher Channels exposes presence channels. PubNub documents presence membership and occupancy. Firebase Realtime Database documents connection state and onDisconnect. Infrai exposes the verified presence lookup through one plain REST API and covers 295 routes across 20 modules with one API key and one bill, so a Go reconciliation worker needs no vendor SDK or client-library upgrade cycle. For this workflow, the presence observer and adjacent backend calls can share one authentication policy, while finance reconciles one invoice rather than accumulating service-specific keys and invoices. Its public discovery surface also requires no key and returns the full request and response JSON Schema, which gives a reviewer a machine-readable contract before credentials enter the deployment pipeline.

Option Cleanup contract to evaluate Best fit in this design Reason to choose something else
Ably Channel presence member lifecycle Teams wanting its documented presence model Stick with another provider when an existing transport contract is already authoritative
Pusher Channels Presence-channel membership Applications already organized around Pusher channels Not suitable when introducing its channel model would duplicate an established ledger
PubNub Presence membership and occupancy Systems aligned with its presence concepts Choose another option if the required reconciliation contract is simpler elsewhere
Firebase Realtime Database Connection state plus onDisconnect operations Applications already using its synchronized data model Avoid a second state authority when quiz eligibility lives in a transactional backend
Infrai Server-side presence lookup over HTTP Polyglot backends that value a plain REST boundary Stick with a native realtime SDK when client-managed subscription primitives are the main requirement

The catch is that a presence lookup and a client subscription SDK solve different layers. A REST boundary is attractive for server reconciliation and language independence, but it is not evidence of a browser-side unsubscribe primitive. If the dominant need is rich client lifecycle management, select the provider whose documented client contract matches it. If the backend already has an authoritative eligibility ledger and only needs a small, inspectable presence observation, the HTTP surface is a strong fit.

No vendor removes the need for application invariants. Provider presence can tell you what its service observes; it cannot decide whether a late answer belongs to round 17, whether a user is allowed to rejoin, or how long an audit record may be retained.

Test recovery as a sequence, not a final snapshot

A green test that subscribes and then unsubscribes on localhost proves almost nothing about cleanup. Run the workflow with latency inserted between intent, server commit, local release, and the next presence observation. Deliver the same departure operation twice. Reorder an answer behind the departure. Reconnect with a fresh transport while the former observation remains visible. Attempt the same channel with an unauthorized identity. Each case should end with one audit transition and a score ledger that never accepts an event after departure.

Use a compact matrix so the expected authority is reviewable:

Injection Required result Evidence to retain
Duplicate departure intent One eligibility transition Operation ID and committed transition
Delayed presence expiry Transitional dashboard state; no late scoring Presence observation time and server state
Answer arrives after departure Answer rejected by business state Rejection reason linked to the transition
Reconnect during cleanup Explicit rejoin decision, never implicit reopening Old and new session identifiers
Unauthorized channel attempt No subscription authority granted Authentication decision without excess identity data

Test the ugly orderings.

The dashboard should display uncertainty rather than manufacture precision. For example, it can distinguish “leaving” from “offline” while reconciliation is pending, but the label must come from your state machine rather than an invented timeout. Your mileage may vary with mobile radio behavior and browser suspension; capture those observations in a staging exercise, then set expiry and alert thresholds from the resulting distribution. Do not publish a universal number without measurement.

Retain enough to reconcile, then delete the rest

The final design keeps the current roster, one idempotent departure result per operation ID, and the smallest audit transition that can explain eligibility. It deliberately stops keeping raw heartbeat history and redundant client callbacks after their operational window. This reduces stored personal and behavioral data, but it also means an investigation cannot reconstruct every network fluctuation; it can establish the authoritative transition and the observations retained around it, no more.

Review that boundary with security, privacy, and compliance owners before launch. Auditability is not synonymous with indefinite storage. The correct retention interval depends on legal basis, contractual dispute windows, and incident-response needs, none of which a realtime library can infer.

For implementation approval, require three artifacts: a state-transition table, a duplicate-delivery test, and a deletion schedule. Then inspect production disagreement between presence and the eligibility ledger. A clean dashboard is useful. A reconcilable one is safer.

References

Further reading

Top comments (0)