DEV Community

NyxenL29
NyxenL29

Posted on

Go Realtime Queues: What One Key Buys Gaming Presence During Reconnects

The page says a game room has 84 connected players while only 79 are actually active. Carrying one key across realtime ingress and queues lets on-call connect each delayed presence update to the connection lifetime that produced it; without that link, reconnects, queue lag, and state writes are visible but their causal order is not. The first action is to stop treating those events as interchangeable.

Short answer: one stable key carried through the realtime path and the queue gives every observation about a single connection lifetime the same name. For a beginner, that buys correlation, ordering within the scope represented by that key, and a practical place to reject stale work. It does not create global ordering, exactly-once delivery, or accurate presence by itself. The useful design is a room-and-session key, plus a monotonic sequence number and an explicit connection generation.

That distinction matters because presence accuracy is not a socket count. It is a claim about whether the server's current view matches the player's current connection generation within a stated time bound.

Why did the page fire after the reconnect storm?

Work backward from what the responder sees. The visible symptom is a gauge mismatch: the room snapshot includes five players whose older disconnect jobs arrived after their newer reconnect jobs, or whose stale heartbeat work extended an expired generation. A dashboard grouped only by room cannot distinguish those cases. A queue depth graph cannot either.

The earlier signal should have been the rate of rejected stale updates, partitioned by room and operation, alongside the age of the oldest unapplied presence event. Those signals expose causality damage before the user-facing count drifts far enough to page. They also separate a busy room from a wrong room; traffic is load, while an update from generation 6 attempting to overwrite generation 7 is evidence of stale work.

Use an SLO that describes the player-visible state, not the transport. For example: the proportion of sampled room views whose member set agrees with authoritative, unexpired connection generations within the chosen convergence window. The window and target are product decisions and must come from game behavior, not copied from an infrastructure template. A turn-based lobby and a short match with rapid rematches do not have the same error budget.

This is the alert-to-action trace:

  1. Page on sustained presence disagreement, not raw reconnect volume.
  2. Pivot from the affected room to the correlation key and generation.
  3. Compare accepted and rejected sequence numbers at the consumer.
  4. Inspect event age to decide whether the fault is stale ordering, delayed consumption, or expiry policy.
  5. Mitigate at the narrowest scope: one generation, one player session, or one room partition.

The key makes step two possible. Without it, responders are matching timestamps across systems and hoping clock proximity means causality. It does not.

What does one key across realtime and queues buy?

A key such as roomID:playerID:sessionID should identify the unit whose history must remain coherent. Generate the session component when a logical connection lifetime begins, retain it through transient transport recovery when the application still considers that the same lifetime, and create a new generation when the application declares a replacement. The exact lifecycle is yours to define; ambiguity is the expensive part.

Do not put mutable state in the key. A region, server address, or online flag can change and should be fields. Do not reuse a player ID alone either, because two tabs or a superseded connection can then share an ordering domain accidentally. Long keys also have a capacity cost: they appear in queue records, logs, indexes, and metric labels. Use an opaque, bounded identifier in telemetry and keep high-cardinality values out of metric dimensions.

The WebRTC specification is useful context even when the room uses another realtime transport: its connection model exposes asynchronous state changes and defines state such as RTCPeerConnection.connectionState. Application presence still has to interpret transport state. A transport transition is an input to the state machine, not proof that a player is present or absent.

Here is a deliberately small Go envelope. The same structure crosses the realtime ingress and queued worker boundary.

package presence

import "time"

type Event struct {
    CorrelationKey string    `json:"correlation_key"`
    RoomID         string    `json:"room_id"`
    PlayerID       string    `json:"player_id"`
    Generation     uint64    `json:"generation"`
    Sequence       uint64    `json:"sequence"`
    Kind           string    `json:"kind"`
    ObservedAt     time.Time `json:"observed_at"`
}
Enter fullscreen mode Exit fullscreen mode

The correlation key answers “which history?” Generation answers “which replacement?” Sequence answers “which update is newer inside that generation?” A timestamp answers none of those questions reliably enough to substitute for them; it remains useful for latency and expiry decisions.

One key is not magic. This needs saying plainly.

If producers can emit sequence numbers concurrently, define who allocates them. If consumers process the same key concurrently, serialize the compare-and-apply operation or make the state write conditional. If a queue partitions by key, that can reduce concurrency conflicts for one history, but the consumer must still defend against duplicates and late work. These are design obligations, not features implied by a field named correlation_key.

The limitation is scope: this pattern is a poor fit when the business rule needs a total order across every room, because per-session keys intentionally preserve independent concurrency and cannot answer which of two unrelated sessions happened first. A single global key would answer that narrower ordering question, but it would also create one serialization point and erase useful parallelism. For gaming presence, the practical trade-off is to order only the history that can overwrite the same player's current state, then use separate room revision rules for snapshots that combine many players.

Reject stale work where state changes

The presence store is the enforcement point because that is where an old event can become a false room member. The consumer should compare generation first, then sequence. A later generation supersedes an earlier one. Within the same generation, an event whose sequence is not greater than the stored sequence changes nothing.

package presence

type State struct {
    Generation uint64
    Sequence   uint64
    Online     bool
}

func Apply(current State, event Event) (State, bool) {
    if event.Generation < current.Generation {
        return current, false
    }
    if event.Generation == current.Generation && event.Sequence <= current.Sequence {
        return current, false
    }

    next := State{
        Generation: event.Generation,
        Sequence:   event.Sequence,
        Online:     event.Kind != "disconnected",
    }
    return next, true
}
Enter fullscreen mode Exit fullscreen mode

This example expresses the rule but omits persistence deliberately. In production, reading the current value and writing the next one must behave as a single conditional transition; otherwise two workers can both pass the comparison and commit in the wrong order. The rejected event should increment a low-cardinality counter labeled by operation and rejection reason, while its correlation key belongs in logs or traces where responders can search it.

Test the state machine with permutations, not only the happy path. Feed connected(7,1), heartbeat(7,2), and disconnected(6,9) in every relevant arrival order. Repeat events. Delay a disconnect beyond expiry. Start generation 8 before every generation 7 event has drained. Then assert the final generation, sequence, online state, and rejection reason.

Replay it twice.

If the second application changes presence or renews a lease, the idempotency boundary is incomplete. That small test catches a consequential mistake: a consumer may reject an older sequence while still extending expiry before the rejection branch, which makes the state look unchanged in ordinary assertions even though a stale delivery has prolonged the player's visible membership. Assert expiry and side effects as well as the stored generation and sequence.

Capacity planning before choosing the boundary

Presence traffic grows with connections and refresh frequency, while fan-out grows with interested recipients. Those are different terms. Estimate them separately before deciding that every heartbeat belongs on the durable queue or that every presence transition deserves room-wide broadcast.

For planning, define C as concurrent connection lifetimes, H as heartbeat events per connection per second, T as other transitions per connection per second, and B as average encoded event bytes. Ingress is approximately C * (H + T) events per second and C * (H + T) * B bytes per second before replication, indexing, and protocol overhead. This is a model, not a benchmark. Measure the omitted terms in a load test with the actual envelope and retention policy.

The key's cardinality deserves its own budget. Retaining a searchable record for every session increases storage and index work; emitting each session as a metric label can overwhelm the telemetry system's series budget. Keep aggregate SLO metrics bounded, then attach exemplars, logs, or traces for individual investigations according to the capabilities of the observability stack.

Decision Managed capability Self-operated capability SRE question
Key-based partitioning Less partition machinery to operate Direct control over hashing and movement What happens to hot rooms and during repartitioning?
Conditional state update Service-specific atomic primitive Schema and concurrency model are team-owned Can the compare-and-apply rule be tested under contention?
Replay and retention Policy exposed through a service boundary Storage, compaction, and recovery stay on-call How long must stale-generation investigations remain possible?
Correlation telemetry Integrated tracing may reduce setup Open schemas reduce coupling Can a key be followed without making it a metric label?

The buy-versus-build decision should be based on operational ownership, lock-in at the envelope and partitioning layers, and the team's willingness to rehearse recovery. Price matters, but it is downstream of the workload model and the on-call surface. A cheap queue that obscures partition behavior is expensive during a presence incident; a feature-rich service is also a poor fit if its ordering boundary cannot represent the chosen session key.

Instrument the invariant, then tune the alert

Add four observations at the consumer: accepted transitions, stale-generation rejections, stale-sequence rejections, and apply latency from observation to committed state. Add a periodic correctness sample that compares the room view with authoritative unexpired generations. The first group explains mechanism. The sample protects the outcome.

Deployment should be staged. First emit keys, generations, and sequences without changing state decisions. Verify that every path populates them and that cardinality stays within the planned telemetry budget. Next run the acceptance rule in shadow mode and record disagreements. Only then enforce rejection, starting with a small slice of rooms and preserving a rollback path that does not discard the new fields.

Alert thresholds require restraint. A single stale rejection can mean the defense worked exactly as designed, so paging on any rejection punishes correctness and trains responders to ignore the signal. Page when the user-visible disagreement consumes the presence SLO's error budget quickly enough to demand action; use stale-rejection rate and event age as diagnostic or earlier warning signals, with thresholds derived from observed normal reconnect behavior.

Too loose is bad too. If the alert waits for a large absolute mismatch, a small room can be completely wrong without crossing it. Prefer a proportion with a minimum sample rule, and evaluate large and small rooms separately. False positives spend on-call attention, erode trust in the page, and can provoke unnecessary failovers that create more reconnects. The final threshold is therefore part of the capacity and reliability design, not a dashboard cosmetic.

One shared key buys a chain of evidence and a scope for enforcing causality. Presence becomes accurate only when that key is paired with generations, sequences, conditional writes, expiry semantics, and an SLO that checks the room state players actually observe.

Further reading

Top comments (0)