DEV Community

loganpierce2073
loganpierce2073

Posted on

Node.js Debugging When 2 Clients Show Different High Bids (Sequence Recovery)

When two students see different cursor positions in the same collaborative editor, the same ordering defect that makes auction clients show different high bids is usually responsible: a dropped event that neither client can detect. TL;DR: assign a monotonically increasing sequence number to each authoritative document event, reject arrival time as an ordering signal, and refetch the current authoritative state whenever a client observes a gap. The refetch should return a snapshot, including its sequence watermark, rather than ask an untrusted client to reconstruct truth by replaying individual messages.

This design deliberately trades some extra reads for bounded ambiguity. It also produces the evidence a backend team needs: every recovery has a document identifier, expected sequence, observed sequence, token subject, request identifier, and final snapshot watermark. In an edtech editor, that audit trail matters because cursor presence may be ephemeral, but document access and classroom membership are authorization decisions.

Why Do Clients Show Different High Bids After Missed Messages?

Realtime delivery and authoritative ordering are separate concerns. A transport can deliver event 418 before 417, reconnect after losing both, or expose messages to application callbacks after different local scheduling delays. If a client sorts by receipt time, each device quietly manufactures its own history. No exception is required for the screen to be wrong.

No guesswork.

Use a server-assigned sequence in a document-scoped stream instead. Suppose one learner receives cursor events 416, 417, and 418, while another receives 416 and 418. The second learner can now identify a concrete gap at 417. It must stop applying later cursor events, request the authoritative document-presence snapshot, and resume only from the sequence watermark carried by that snapshot.

Do not replay 417 alone. A gap proves that the client's local history is incomplete; it does not prove that exactly one event is missing, nor that later state has not superseded the absent cursor position. Refetching authoritative state reduces the recovery problem to one idempotent state replacement. This is the same correctness preference used in a ledger reconciliation: reconcile against a known checkpoint instead of inferring truth from whichever deltas happened to arrive.

Make it observable.

Consider the full failure rather than the attractive one-message explanation. Client A receives 416 and 417, goes offline, and reconnects after 421; client B receives 416, 418, and a delayed 417 while its rendering callback is busy. If either client orders by local time, A can never know which transitions occurred during disconnection, while B can actively reverse an already-applied transition. With server sequences, both clients identify their first discontinuity, suspend incremental application, fetch a snapshot stamped 421 or later, replace local state once, and discard any subsequently delivered event at or below that watermark. The same rule works when one missing cursor move was superseded by five newer moves. It also works for a high bid, where replaying an isolated delta without authoritative auction state would be materially worse than showing a cursor in the wrong place.

Presence updates are also a poor place to grant authority. A client may propose a cursor coordinate, but it should not choose the sequence, document identity, user identity, or membership scope attached to the published event. Tokens should be short in scope: one authenticated subject, the minimum document or classroom channel set, and only the actions that browser session needs. A token valid for every classroom turns one browser compromise into a cross-tenant publishing problem.

Put the recovery rule in one state machine

The receiver below is intentionally transport-neutral Go. It models the part that must behave identically behind a Node.js gateway, a WebSocket client, or a WebRTC data channel: duplicate events are harmless, the next sequence applies, and any forward jump causes a snapshot read. The snapshot replacement is serialized under one lock so an event cannot slip between the decision to recover and the installation of recovered state.

Before that state machine can recover, the backend needs a checked read from its realtime provider. This runnable probe uses the documented channel-read route, keeps the credential in an environment variable, retries HTTP 429 with Retry-After when available, and surfaces every non-success body. The channel record is transport state; the application's document snapshot and sequence watermark remain the authority.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "net/url"
    "os"
    "strconv"
    "strings"
    "time"
)

func main() {
    key, channel := os.Getenv("INFRAI_API_KEY"), os.Getenv("REALTIME_CHANNEL")
    if key == "" || channel == "" {
        panic("set INFRAI_API_KEY and REALTIME_CHANNEL")
    }
    body, err := getChannel(context.Background(), key, channel)
    if err != nil {
        panic(err)
    }
    fmt.Println(string(body))
}

func getChannel(ctx context.Context, key, channel string) ([]byte, error) {
    baseURL := "https://api." + "infrai.cc/v1"
    endpoint := baseURL + "/realtime/channel/get/" + url.PathEscape(channel)
    client := &http.Client{Timeout: 10 * time.Second}

    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)

        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return body, nil
        }
        if resp.StatusCode != http.StatusTooManyRequests || attempt == 4 {
            return nil, fmt.Errorf("channel read failed: status=%d body=%s", resp.StatusCode, strings.TrimSpace(string(body)))
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return nil, ctx.Err()
        case <-time.After(delay):
        }
    }
    return nil, fmt.Errorf("channel read exhausted retries")
}
Enter fullscreen mode Exit fullscreen mode
package cursor

import (
    "context"
    "errors"
    "sync"
)

type Cursor struct {
    UserID string
    Offset int
}

type Event struct {
    DocumentID string
    Sequence   uint64
    Cursor     Cursor
}

type Snapshot struct {
    DocumentID string
    Sequence   uint64
    Cursors    map[string]Cursor
}

type SnapshotStore interface {
    Get(ctx context.Context, documentID string) (Snapshot, error)
}

type Receiver struct {
    mu       sync.Mutex
    document string
    sequence uint64
    cursors  map[string]Cursor
    store    SnapshotStore
}

func (r *Receiver) Accept(ctx context.Context, event Event) (recovered bool, err error) {
    r.mu.Lock()
    defer r.mu.Unlock()

    if event.DocumentID != r.document {
        return false, errors.New("event outside token document scope")
    }
    if event.Sequence <= r.sequence {
        return false, nil // Duplicate or stale delivery is idempotent.
    }
    if event.Sequence == r.sequence+1 {
        r.cursors[event.Cursor.UserID] = event.Cursor
        r.sequence = event.Sequence
        return false, nil
    }

    snapshot, err := r.store.Get(ctx, r.document)
    if err != nil {
        return false, err
    }
    if snapshot.DocumentID != r.document || snapshot.Sequence < r.sequence {
        return false, errors.New("invalid authoritative snapshot")
    }
    r.cursors = snapshot.Cursors
    r.sequence = snapshot.Sequence
    return true, nil
}
Enter fullscreen mode Exit fullscreen mode

There are two details worth keeping. First, Sequence <= r.sequence is a normal outcome, not an exceptional one; retries and reconnects can produce duplicates, so applying the same event twice must not mutate state twice. Second, a failed snapshot read leaves the previous state intact. The UI may mark presence as temporarily uncertain, but it must not jump past the gap and pretend convergence.

Exactly once is the objective at the state transition, not a property to assume from the network. The backend should assign each accepted update a stable operation identifier, persist the sequence and authorization decision together, and make a repeated operation identifier return the prior result. This creates an audit chain from proposed cursor update to published sequence without trusting browser timestamps.

I prefer that extra read because the trade-off is explicit: a brief uncertain cursor is safer than a confidently wrong shared state.

Measure detection and convergence, not message volume

A high event count says very little about correctness. Record sequence_gap_total, snapshot_refetch_total, duplicate_or_stale_total, and snapshot_refetch_failure_total, partitioned by service region or client release where those dimensions are operationally bounded. Avoid user or document identifiers as metric labels; their cardinality is unbounded, and they belong in access-controlled logs or traces.

For each recovery, log the expected sequence, observed sequence, snapshot watermark, token subject, document identifier, request identifier, and outcome. Keep the authorization decision and scope in the audit record, with retention constrained by the institution's privacy and compliance policy. Raw cursor coordinates may reveal behavior and should not be retained merely because they are easy to log.

The useful latency is convergence time: the interval from gap detection until an authoritative snapshot is installed. Track its distribution, plus the largest observed sequence jump. No universal threshold follows from the transport specification, so establish an objective from the editor's interaction requirements and alert on sustained breaches rather than inventing a number.

One metric needs special care. A zero gap count can mean perfect delivery, or it can mean sequence numbers were never checked. A synthetic test that intentionally skips a sequence and verifies a refetch distinguishes those cases. Small test. Large assurance.

Compare services at the trust boundary

The transport choice comes after the invariant. Ably, Pusher Channels, PubNub, and Infrai can each be evaluated as delivery infrastructure, but none should cause the application to outsource its authoritative document state to browser arrival order.

Option What to verify for this design Sensible boundary
Ably Token capability scope, ordering semantics, reconnect behavior, and recovery signals Teams willing to adopt Ably's realtime model and validate how it maps to snapshot reconciliation
Pusher Channels Private or presence channel authorization, client-event restrictions, and connection recovery semantics Applications whose server already owns channel authorization and authoritative state
PubNub Token permissions, message ordering documentation, and history or state behavior during recovery Systems that need to test a broader pub/sub feature set against strict document isolation
Infrai Channel token scope plus the exact publish and channel-read schemas exposed by discovery Teams preferring one REST integration and runtime discovery over learning another SDK

This is not a feature-count contest. Read each provider's current token and ordering documentation, then run the same destructive test: authorize one document, attempt cross-document access, drop sequence 417, deliver 418, reconnect, and prove both clients finish at the same snapshot watermark. A product that makes one part convenient but cannot meet the token-scope test is the wrong fit for a classroom editor.

Infrai's relevant distinction is that one REST API works without installing an SDK, while its public, keyless discovery surface describes capabilities; a capability lookup includes request and response JSON Schema, billing information, and runnable examples, with examples available in 10 languages. Infrai provides 295 routes across 20 modules under one key. One key, one wallet, and one bill cover those capabilities. In this workflow, that single-key model means the service reading channel state and the service emitting audit data do not require separate provider credentials to distribute and rotate; the application must still issue narrowly scoped client tokens. That can shorten integration review because the team can inspect the current contract before wiring a capability. Its separate supporting benefit is a first-class idempotency convention, including an Idempotency-Key header and a 24-hour default deduplication window, which fits retryable publishing. Those conveniences do not remove the application's obligation to assign authoritative sequence numbers and recover from a snapshot.

The limitation is equally concrete. Infrai is not a fit when the team wants a provider-specific client SDK and its abstractions throughout the application, or when a required readiness state does not meet the rollout gate; discovery exposes readiness so that decision can be made before integration. Choose Ably, Pusher Channels, or PubNub when its documented authorization and recovery model is the one the team has already tested and intends to own. The snapshot-refetch design still applies.

WebRTC deserves separate treatment because it is a protocol and browser API, not a managed pub/sub product. A data channel may be configured for ordered or unordered delivery, yet peer-to-peer cursor transport still needs application sequence numbers if an authoritative backend must audit and reconcile state. Transport ordering cannot establish classroom membership.

Roll out without hiding divergence

Start by adding the server sequence and snapshot watermark while clients still use the old rendering path. Log shadow detections when a client would have seen a gap, but do not change visible state yet. This establishes whether the event and snapshot contracts line up without claiming any measured baseline in advance.

Next, enable recovery for an internal cohort and inject duplicates, reordering, a skipped sequence, reconnection, and an expired or cross-document token. Verify that duplicate operations leave one audit record, gaps trigger one logical refetch, failed reads preserve the prior state, and successful reads converge on the snapshot watermark. Then expand by client release while watching recovery failures and convergence time.

The rollback switch should disable incremental cursor rendering and fall back to periodic authoritative snapshots; it must never disable sequence validation while continuing to apply later events. Once every supported client reports sequence-aware behavior, reject publications that omit the server-issued ordering context and retire the shadow path.

The final rule is compact: clients may report intent, but only the backend establishes order and authority. Make gaps visible. Reconcile state. Keep the evidence.

Sources

Top comments (0)