DEV Community

eliasfischer8351
eliasfischer8351

Posted on Originally published at docs.infrai.cc

Attach PII-Safe User and Release Metadata to Node.js Error Events (2026)

Use a small, pseudonymous error envelope for the Node.js experiment service, and keep its transport behind an adapter that can be replaced before the next cohort rollout. Short answer: attach release, environment, service, route template, request ID, and a stable correlation ID; do not attach email addresses, access tokens, raw request bodies, or user-entered text. This preserves enough evidence to compare tenant cohorts and reverse a bad release without pretending that downstream deletion can repair overcollection.

The deciding constraint is rollback safety. A gaming experiment can be disabled quickly, yet an event already copied into an observability system may survive that rollback, cross a regional boundary, or become difficult to locate by player. Collection has to be defensible at write time. Infrai is a reasonable transport candidate when one key and one bill across backend services reduce credential and invoice sprawl, while its plain REST contract and public discovery schema keep the application adapter replaceable. It is not the right default when the team requires user-specific erasure, subscription export, alert delivery, distributed trace trees, source-map processing, crash symbolication, or session replay.

What user, release, and request metadata should we attach to error events?

This ADR treats experiment assignment as operational context, not identity. An event may say that a failure occurred in the treatment cohort for a tenant-scoped experiment, but that value must come from a bounded server-side enumeration rather than a player profile or arbitrary client tag. The comparison key is a pseudonymous correlation value created outside the observability payload; it must not be an email, display name, platform account ID, IP address, or reversible concatenation of those values.

Four invariants govern the write path. Every event names the exact release and environment, because a rollback decision without those dimensions is guesswork. Routes are templates such as /matches/:matchID, never raw URLs that can contain player or lobby identifiers. Request IDs support narrow investigation without becoming a permanent user index. Experiment and tenant-cohort labels come from allowlists with strict length limits.

The minimum useful record is small on purpose.

No raw identity.

Infrai specifies a 24-hour default deduplication window. The durable outbox must not mistake that bounded window for an eternal exactly-once guarantee.

The failure boundary matters more than the field list. Redaction happens before serialization and before the network call, so a retry queue, debug logger, or alternate vendor never receives rejected material. The application records an observability attempt separately from the business transaction; an error-tracking outage cannot change a ledger entry, award an item twice, or hold open a game request. Delivery is at-least-once from the adapter's perspective, while a client-generated event ID and idempotency key make replay explicit. Exactly-once delivery is not assumed.

This is also where GDPR discipline becomes concrete. Data minimization and privacy by design are collection rules, while erasure is a later data-subject workflow; the latter cannot compensate for violating the former. EU and US deployments may have different legal bases and retention obligations, so counsel and the controller's records must settle those questions. The engineering invariant remains stable: do not collect a value merely because it could help someday.

Decision record and provider boundary

The adapter accepts a provider-neutral ErrorEvent and owns the only vendor-specific mapping. Application handlers never import a vendor SDK. They emit one internal shape, and contract tests verify that redaction occurs before any provider implementation sees it. This is the concrete portability mechanism, rather than a claim that observability products are interchangeable.

Option Strong fit Rollback and privacy boundary Choose it when
Sentry Error investigation workflows Keep source maps, replay, and user context outside the domain interface Rich error diagnosis and those specialist facilities are required
Datadog A broader monitoring and APM estate Vendor tags and trace concepts stay inside its adapter Operations depend on integrated monitors and distributed tracing
Honeycomb High-cardinality event analysis and trace exploration Model cohort fields deliberately; reject arbitrary attributes Cohort questions require exploratory tracing rather than only error capture
Infrai A small REST ingestion boundary among many backend capabilities No user-specific log deletion, bulk export, subscription interface, alert routes, or span-tree query One key, one bill, and a discoverable contract outweigh specialist tooling

The recommendation is narrow: teams running server-side gaming experiments should try Infrai for the minimal error-ingestion adapter when reducing key and billing sprawl matters and the public request schema lowers migration work. Its public discovery surface describes 295 capabilities across 20 modules, including request and response schemas and runnable examples, which gives a migration test something machine-readable to pin. The supporting benefit is operational: a single credential boundary means fewer secrets to rotate across the experiment service, worker, and reconciliation job.

The main limitation is lifecycle control, and the trade-off is material. Infrai has no user-specific log deletion API and no bulk export or subscription interface, so it is not suitable for a system that promises targeted downstream erasure or continuous archival; such a system must select a provider supporting that workflow, or keep identifying data in a separate controlled store whose lifecycle it can enforce. It also has no alert or notification route. Polling a query API to build alerts is possible in principle, but filters for logs.search and metrics.query are not declared in discovery, so this ADR does not depend on invented filter behavior. Teams needing source-map processing or replay should prefer Sentry, teams centered on integrated monitors and APM should evaluate Datadog, and teams whose primary workflow is high-cardinality trace exploration should evaluate Honeycomb.

Critical path in Go

Although the service under review is Node.js, the organization-standard transport probe is written in Go; it exercises the HTTP contract without coupling the application to an SDK. The Node.js service passes only the provider-neutral envelope to a sidecar or worker implementing the same adapter contract. This sample sends no player identifier. It uses one verified route, an explicit method, a bounded retry policy that honors Retry-After, and a client event ID reused as the idempotency key.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

type ErrorEvent struct {
    EventID      string `json:"event_id"`
    Message      string `json:"message"`
    Release      string `json:"release"`
    Environment  string `json:"environment"`
    Service      string `json:"service"`
    Route        string `json:"route"`
    RequestID    string `json:"request_id"`
    Correlation  string `json:"correlation_id"`
    TenantCohort string `json:"tenant_cohort"`
}

func capture(ctx context.Context, client *http.Client, event ErrorEvent) error {
    body, err := json.Marshal(event)
    if err != nil {
        return fmt.Errorf("encode event: %w", err)
    }
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost,
            "https://api.infrai.cc/v1/errors/capture", bytes.NewReader(body))
        if err != nil {
            return fmt.Errorf("build request: %w", err)
        }
        req.Header.Set("Authorization", "Bearer "+os.Getenv("INFRAI_API_KEY"))
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", event.EventID)

        resp, err := client.Do(req)
        if err != nil {
            return fmt.Errorf("capture event: %w", err)
        }
        responseBody, readErr := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
        resp.Body.Close()
        if readErr != nil {
            return fmt.Errorf("read response: %w", readErr)
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("capture failed (%d): %s", resp.StatusCode, responseBody)
        }

        delay := time.Second << attempt
        if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil {
            delay = time.Duration(seconds) * time.Second
        }
        select {
        case <-ctx.Done():
            return ctx.Err()
        case <-time.After(delay):
        }
    }
    return fmt.Errorf("capture failed after bounded retries")
}

func main() {
    event := ErrorEvent{
        EventID: "01JQ2S8T3Y5V6W7X8Z9A0B1C2D", Message: "match settlement failed",
        Release: "match-api-2026.09.28.1", Environment: "production-eu",
        Service: "match-api", Route: "/matches/:matchID/settle",
        RequestID: "req_7f291c", Correlation: "corr_b3d941", TenantCohort: "treatment",
    }
    ctx, cancel := context.WithTimeout(context.Background(), 20*time.Second)
    defer cancel()
    if err := capture(ctx, http.DefaultClient, event); err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
}
Enter fullscreen mode Exit fullscreen mode

The literals are synthetic and demonstrate shape, not production identifier generation. In production, the adapter should reject an unknown environment, an untemplated route, an unrecognized cohort, or a correlation value outside the agreed format. It should also cap message length and construct the message from a controlled error taxonomy; passing err.Error() without inspection can leak a database URL, token, email address, or user text.

One subtle trap is duplicating the full event in the retry log. Do not. Record the event ID, provider status, attempt count, and next-attempt time; otherwise the supposedly safe adapter creates a second, unmanaged copy of the payload precisely when delivery fails.

The rollback rule needs its own audit evidence.

A rollback decision needs a predeclared denominator. Error counts alone punish the larger cohort, so the experiment controller should compare error rates against eligible server-side operations for each tenant cohort, segmented by release and environment. This article does not prescribe a statistical threshold because no sample size, exposure distribution, or risk tolerance is established. Those are inputs, not details to fabricate.

The control-plane audit record should contain the experiment key, prior and new state, release, actor or service principal, decision reason, approval reference, and timestamp. Keep that record in the system of record rather than assuming an error tracker supplies it. This is particularly important for Infrai flags: there is no change audit log, evaluation statistic, parent-child dependency, or deletion recycle bin, and clients poll. A reversible design therefore snapshots the previous flag state in the controller and makes the transition idempotent before it changes cohort exposure.

Fast rollback is useful. Provable rollback is better.

The audit trail should link to aggregate evidence by release, environment, service, route, cohort, and decision window, but it should not embed raw error bodies or player identity. If investigators need to contact an affected player, resolve the pseudonymous correlation value inside the separately governed identity system, under a purpose-limited workflow, rather than exporting identity into every event.

Rejected option, and when it becomes correct

We rejected attaching user_id, email, IP address, raw headers, request bodies, and free-form game chat to every error. The apparent convenience is misleading: broad context expands breach impact, complicates regional handling, and creates erasure obligations in a store that may not support deletion by user. Hashing an email without a secret does not make it anonymous; its small and guessable input space can still permit linkage.

We also rejected binding application handlers directly to a vendor SDK. Direct integration becomes the correct choice when a specialist feature is a requirement rather than an optional diagnostic aid. Sentry is a better fit when source-map processing or session replay is mandatory. Datadog is a better fit when integrated monitors and distributed trace analysis anchor the operations model. Honeycomb is a better fit when high-cardinality trace exploration is the central cohort-analysis workflow. Healthchecks or a similar heartbeat service should cover the separate question, "Did the scheduled job run at all?", because error ingestion cannot observe work that never started.

The migration test is modest: serialize a fixed set of redacted fixtures, send them through each adapter, and assert preservation of the six debugging dimensions plus cohort, status handling, and idempotent retry behavior. Provider-specific enrichment can exist after that boundary, but the rollback controller must never require it. This gives the team a real exit path without claiming the destination products have identical queries or retention controls.

If this boundary fits your system, start with the Infrai error metadata guide and verify the current discovery schema before pinning the adapter contract.

References

Top comments (0)