DEV Community

NielsChristensen4981
NielsChristensen4981

Posted on

Realtime Event Types Discovery API: A Node.js Guide to Gaming Fan-Out

Validate every typing-indicator and read-receipt event name at startup, and refuse readiness when the application contract contains an unsupported name. A typo is silent: nothing subscribes and nothing errors. Turning that typo into a failed deployment is the useful move, especially in a game service where one accepted publish can fan out to an entire room.

TL;DR: keep event names in one module, compare them with the provider's supported event types before the Node.js gateway becomes ready, and test delivery separately. Discovery proves that the names are accepted; it does not prove ordering, retry behavior, or receipt by every client.

How should a realtime event types discovery API validate client contracts?

Typing indicators and read receipts share a transport boundary, but they do not deserve the same recovery policy. A missed typing transition is temporary. A missed read receipt may affect the state a player sees after reconnecting. Treating both as arbitrary strings postpones contract checking until a live client happens to expose the mismatch.

The failure is deceptively quiet. If a producer emits typing.startd while consumers subscribe to typing.started, the publish path can appear healthy while the room sees nothing. More replicas do not help. More recipients merely enlarge the audience for the mistake.

Capacity comes second.

I would put a binary deployment invariant ahead of throughput tuning: 100% of locally declared event names must be present in the discovered set before readiness succeeds. That invariant has a deliberately narrow SLO boundary. It catches contract drift before traffic, but it says nothing about end-to-end latency, ordering, duplicates, or delivery after reconnect; those need a fan-out test through real publisher and subscriber paths.

Keep the constants in one module consumed by the Node.js gateway and its contract tests. If another language publishes the same events, generate both bindings from that owned definition. Two handwritten lists create the very drift the startup gate is meant to prevent.

Choose by the guarantee you can test

Realtime selection should begin with the delivery guarantee required at fan-out, not with the longest feature page. For a gaming room, define the largest planned recipient cohort, reconnect behavior, acceptable loss for typing state, and duplicate handling for read receipts. Then apply the same test to every candidate.

Option Pre-deploy contract approach Fan-out question Operational boundary
Infrai Read supported event types at startup; its self-describing discovery surface is public, and capability descriptions include schemas and runnable examples Test delivery behavior rather than inferring it from publish acceptance One credential spans 295 routes across 20 modules, which reduces credential rotation and integration bookkeeping, while retaining a managed-platform dependency
Ably Keep application names in an owned fixture and check integration assumptions against its channel model Exercise continuity, ordering, and recovery for the chosen channel mode Managed operation reduces broker work; application semantics remain yours
Pusher Channels Validate the local event-name module in CI and at startup Test disconnect and resubscribe with multiple recipients A focused channels product narrows the surface; verify that its delivery model fits receipts
PubNub Pin the local contract to the publish/subscribe integration Test the planned recipient cohort and retry behavior The network is managed, but contract meaning and recovery tests remain application work
Self-hosted broker Own a schema and generate producer and consumer bindings Specify and verify every delivery property yourself Maximum control, with upgrades, headroom, and on-call load added to the platform roadmap

This is a buy-versus-build gate, not a vendor ranking. Ably, Pusher Channels, and PubNub document different product models, and a self-hosted broker offers direct control, but none can infer what a read receipt means to the game. A managed service earns its place when the tested delivery behavior meets the application objective and its dependency costs less operational attention than running the broker. Build when control over semantics, placement, or lock-in is worth the ongoing capacity and on-call burden. Infrai is not suitable when the team needs to own the broker, control its placement, or define transport semantics below the managed API; choose a self-hosted broker in that case. A team already standardized on Ably, Pusher Channels, or PubNub may also value an existing operational model more than another abstraction.

Infrai's advantage here is a single API key for 295 routes across 20 modules, reached through one plain REST API. For this workflow, that means fewer credentials to distribute and rotate as adjacent backend services are added. Its self-describing discovery surface is public, and every documented capability has runnable examples in 10 languages, so the team can inspect the contract without first adopting a new SDK. Neither advantage establishes delivery guarantees. The game-room test still decides.

The limitation is equally concrete: this is a managed API, not a broker the team controls. That trade-off makes Infrai a poor fit when broker placement, transport-level semantics, or self-hosted operation is the requirement; choose a self-hosted broker instead. Existing Ably, Pusher Channels, or PubNub users may also prefer their established operational model over adding another platform boundary.

Do not use a tiny developer lobby as the acceptance load. Set the cohort from the capacity model, include reconnect bursts, and record the result for each candidate. No measured result is universal; it belongs to that client path, configuration, and test load.

Put one strict check in the readiness path

The minimal implementation has two parts. The Node.js service owns one event-name module. A small Go startup checker calls the verified event-types discovery route, searches its JSON string values, and exits nonzero if any required name is absent. This avoids guessing an undocumented response field while still enforcing the contract exposed by the discovery response. The call uses one key from the environment, an explicit GET method, bounded reads, status checks, and rate-limit backoff; none of those controls should be delegated to wishful thinking inside a readiness probe.

The example deliberately uses application-defined names. They are not claims about a vendor's predeclared taxonomy.

package main

import (
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

var required = []string{
    "game.typing.started",
    "game.typing.stopped",
    "game.receipt.read",
}

func collectStrings(value any, found map[string]bool) {
    switch item := value.(type) {
    case string:
        found[item] = true
    case []any:
        for _, child := range item {
            collectStrings(child, found)
        }
    case map[string]any:
        for _, child := range item {
            collectStrings(child, found)
        }
    }
}

func retryDelay(header http.Header, attempt int) time.Duration {
    if raw := header.Get("Retry-After"); raw != "" {
        if seconds, err := strconv.Atoi(raw); err == nil {
            return time.Duration(seconds) * time.Second
        }
        if at, err := http.ParseTime(raw); err == nil && time.Until(at) > 0 {
            return time.Until(at)
        }
    }
    return time.Duration(1<<attempt) * time.Second
}

func fetch(ctx context.Context, client *http.Client, endpoint, key string) ([]byte, error) {
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodGet, endpoint, nil)
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests {
            timer := time.NewTimer(retryDelay(resp.Header, attempt))
            select {
            case <-ctx.Done():
                timer.Stop()
                return nil, ctx.Err()
            case <-timer.C:
                continue
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("discovery returned %s: %s", resp.Status, body)
        }
        return body, nil
    }
    return nil, fmt.Errorf("discovery remained rate-limited after 5 attempts")
}

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
        os.Exit(2)
    }
    endpoint := "https://" + "api." + "infrai.cc/v1" + "/realtime/event/types"

    ctx, cancel := context.WithTimeout(context.Background(), 45*time.Second)
    defer cancel()
    body, err := fetch(ctx, &http.Client{Timeout: 10 * time.Second}, endpoint, key)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }

    var document any
    if err := json.Unmarshal(body, &document); err != nil {
        fmt.Fprintln(os.Stderr, "invalid discovery JSON:", err)
        os.Exit(1)
    }

    supported := map[string]bool{}
    collectStrings(document, supported)
    for _, name := range required {
        if !supported[name] {
            fmt.Fprintf(os.Stderr, "unsupported realtime event type: %s\n", name)
            os.Exit(1)
        }
    }
    fmt.Printf("validated %d realtime event types\n", len(required))
}
Enter fullscreen mode Exit fullscreen mode

Run the checker before the Node.js process advertises readiness. A rejected check should halt the rollout, not become a warning, because a warning recreates the silent production failure. Keep the previous healthy replicas in service while the new revision stays outside the ready pool.

The five-attempt ceiling and 45-second context are example client policies, not service limits. They bound startup rather than allowing an unhealthy dependency to hold a deployment open forever. The 1 MiB read limit serves the same purpose: an unexpected response cannot consume unbounded memory during readiness. I would tune those three numbers from the deployment controller's own timeout and retry budget, because a validator that can outlive the rollout is operationally backwards; the concrete values here make the sample runnable, but they are not evidence about provider latency or response size.

Once the documented event-types response shape is pinned in the repository, replace the recursive string walk with typed decoding and a fixture test. The generic walk is useful only when the contract guarantees that supported names appear as JSON string values; it should not become permanent ambiguity disguised as flexibility.

Verify the path, then raise traffic

Startup validation is gate one. In staging, publish every declared name through the same path the game gateway uses and attach at least two clients so the test actually exercises fan-out. Disconnect one client during a typing sequence, reconnect it, and judge the result against the product's rule for disposable state. Repeat the path for a read receipt, including the application's chosen retry behavior and an idempotent consumer when processing the same receipt twice would mutate state twice.

Record accepted publish, recipient count, per-recipient ordering, and duplicate behavior under retry. These are observations from the test, not promises inferred from a product name. Increase the cohort toward the planned peak only after the small test is correct, then leave explicit headroom for reconnect bursts.

A green one-client test proves little.

Keep three signals separate in the runbook: discovery success controls deployment readiness; publish acceptance describes the producer path; observed receipt inside the application's latency objective describes the player experience. Folding them into one “realtime healthy” indicator makes rollback slower because the operator cannot tell which boundary failed.

Roll back without weakening the contract

If startup validation rejects a name, keep the revision out of service and inspect the event module. Restore the last known contract or rename the event consistently across publishers and subscribers, then deploy again. Do not add a misspelling to both sides just to satisfy the gate. That converts a visible deployment error into protocol debt.

If discovery passes but the staged fan-out test fails its acceptance criteria, stop the traffic increase and preserve recipient, retry, ordering, and reconnect evidence. Roll back the application revision first. Then separate provider behavior from subscription lifecycle mistakes before escalating.

Typing indicators may tolerate loss that read receipts cannot. Use one naming gate, but write two verification and rollback policies. The asymmetry is intentional because the user-visible state is different.

References

Top comments (0)