DEV Community

SterlingVance2196
SterlingVance2196

Posted on

React Error Tracking — A Go Backend Without a Browser SDK

A browser error report is untrusted, potentially duplicated input, and in a gaming client it may arrive after the player has already retried the action. That constraint changes the design. Short answer: combine a React error boundary with window.error and unhandledrejection, send a small normalized event to a Go backend, and attach a client-generated fingerprint so retries and repeated crashes can be grouped. This is a sound lightweight choice for basic JavaScript and runtime errors. It is not a substitute for source-map deobfuscation, symbolication, session replay, or polished crash analysis.

Rollback safety should drive the rollout. Put the collector behind a release switch, accept old and new payload versions concurrently, and make ingestion idempotent before enabling it for every player. An error collector must never turn one frontend exception into several ledger-like records merely because a mobile connection retried.

How should a React error boundary send frontend JavaScript errors?

React error boundaries cover failures raised while rendering descendants, but they do not replace the two global listeners. window.error catches uncaught JavaScript errors; unhandledrejection covers rejected promises that no caller handled. Together, those three sources form the minimum useful collection surface.

Normalize them into one envelope: event ID, schema version, error name, message, stack, page URL, application release, browser, timestamp, source, and a client-generated fingerprint. Include a user ID only when it is appropriate and lawful. For a game, a pseudonymous player reference can help correlate repeated failures, yet it also turns an operational event into personal data with deletion consequences.

Keep the payload small. A stack may already expose paths or query strings, so do not attach application state, chat content, access tokens, payment details, or an indiscriminate dump of browser storage. Logs are a poor primary store when GDPR deletion by user is required because this route set has no per-user deletion API.

Duplicates happen.

The fingerprint deserves special care. Derive it from stable fields such as error name, normalized message, and the first useful stack frame; do not include the event ID or timestamp, which would make every occurrence unique. The event ID answers, "Have I ingested this exact report?" The fingerprint answers, "Does this report belong to an existing group?" Those are separate correctness questions.

Make the Go boundary idempotent

The browser-facing endpoint should acknowledge quickly, enforce a strict body limit, validate the schema, and deduplicate by event ID. It can then forward the normalized report to the selected error store. Before writing that forwarding code, retrieve the live contract. The following complete program asks the public discovery surface for the verified errors.capture capability, checks the response, and prints its request schema and runnable examples; this avoids copying a payload shape into an article only to let it drift from the service:

package main

import (
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
)

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    apiKey := os.Getenv("INFRAI_API_KEY")
    if baseURL == "" || apiKey == "" {
        log.Fatal("INFRAI_BASE_URL and INFRAI_API_KEY are required")
    }

    req, err := http.NewRequest(http.MethodGet, baseURL+"/v1/discovery/errors.capture", nil)
    if err != nil {
        log.Fatal(err)
    }
    req.Header.Set("Authorization", "Bearer "+apiKey)

    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        log.Fatal(err)
    }
    defer resp.Body.Close()

    body, err := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
    if err != nil {
        log.Fatal(err)
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        log.Fatalf("discovery failed: status=%d body=%s", resp.StatusCode, body)
    }

    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

Use the returned Go example for the actual capture call, with INFRAI_API_KEY supplied through the environment and Authorization: Bearer authentication; preserve one event ID across retries and use the platform's idempotency convention. A production deployment also needs a shared deduplication store with a retention window so restarts and multiple replicas do not forget accepted IDs. That is the same exactly-once mindset used around a payment command: delivery may be at least once, while the durable effect is applied once. The transport may deliver at least once. The durable effect should happen once.

There is another failure mode. If the collector returns an error, the browser must use bounded exponential backoff, honor Retry-After for HTTP 429, retain the same event ID, and stop after a small retry budget. Never let telemetry retry forever on the gameplay path.

Grouping is useful; crash analysis is different

Release, browser, URL, and fingerprint make basic triage workable: operators can ask whether a new build introduced a cluster, or whether one browser dominates it. Minified production stacks may still be opaque. Without source-map deobfuscation, crash symbolication, Electron minidump parsing, or session replay, the system cannot reconstruct the rich debugging context supplied by a mature crash-analysis product.

This boundary is decisive. Use the lightweight path when the question is "Which runtime errors increased after release X?" Choose a specialized service when the question is "Which original source line failed, what happened immediately beforehand, and can an on-call engineer be notified now?"

Infrai uses one key and one bill across 295 routes in 20 modules, which reduces credential inventory and reconciliation work when this backend later needs adjacent capabilities. Its public discovery surface describes request and response schemas, billing, and runnable examples, so integration begins by reading a capability rather than adopting another client SDK; every documented capability has runnable examples in 10 languages. The platform's idempotency convention also keeps retry behavior consistent. The limitation is material: for this use case there is no alert or notification route, no source-map analysis, and no session replay. Query polling and a separate alerting component would be required.

Compare the operational shape, not the logo

The fair comparison is about the investigation workflow and compliance boundary, not raw feature count.

Option Best fit Constraint that changes the decision
Custom Go collector plus a plain REST destination Basic runtime collection through a self-describing capability, especially when avoiding another browser SDK matters No source-map deobfuscation, session replay, notification route, per-user log deletion API, or distributed span-tree query
Sentry Teams that require a dedicated error-tracking workflow and should evaluate source maps, grouping, and replay against their release process Adds a specialized product and data-governance surface; verify current retention and deletion behavior before sending player identifiers
Datadog Teams that want to evaluate browser errors beside a broader monitoring estate The trade-off is a wider operational platform; validate SDK-free intake, deletion, and browser analysis against the current documentation
Grafana Teams already assembling an observability stack and willing to own more of its integration Validate the chosen components for grouping, source maps, replay, alert delivery, and retention rather than assuming the dashboard supplies them
Better Stack Teams comparing an integrated operational workflow around logs and incidents Confirm current browser error analysis, data location, and user-deletion controls before adopting it for player data
Healthchecks Detecting that a scheduled task or heartbeat did not run Complements exception capture; it does not replace JavaScript error grouping

Sentry, Datadog, Grafana, and Better Stack deserve a proof of concept when readable production stacks, integrated alerting, or a broader observability estate are requirements. The custom REST design is not suitable when the team expects those workflows to arrive fully assembled. Healthchecks solves a different observability gap: silent absence. A game backend can need both, because an exception report proves that something failed, while a missed heartbeat shows that expected work never began.

Do not infer traces from fields alone. Although logs can carry trace_id and span_id, this capability set has no distributed-trace query or span tree. That distinction prevents an error store from quietly becoming an incomplete tracing system.

Roll out so rollback remains boring

Start with one release cohort and one percent of sessions. Measure collector request volume, accepted-versus-duplicate counts, payload rejection, and latency added to shutdown or navigation paths; then widen the cohort only after the fingerprints group usefully and the payload has passed privacy review. The one-percent value is a rollout policy, not a performance claim.

Keep schema version 1 readable while version 2 is introduced. A rollback may leave newer clients in the wild for days, so the backend must tolerate that overlap instead of coupling deploy order to the game client. Preserve the event ID across retries, record configuration changes in your own audit trail, and rehearse disabling collection without shipping a new client.

Finally, test the ugly paths: a thrown render error, an asynchronous rejection, a network timeout, a 429 response, two identical submissions, a malformed oversized body, and a release rollback. The decision rule is compact: use this design for small, auditable runtime-error intake; move to a dedicated crash-analysis product when readable minified stacks, replay, integrated alerting, or user-level deletion is a requirement rather than a preference.

Sources

Top comments (0)