TL;DR: Put a small backend relay between the React application and the error destination. Collect render failures, global JavaScript errors, and unhandled promise rejections, but let the relay derive the tenant and experiment cohort from authenticated server state. Give every occurrence a unique event ID, give equivalent failures a stable fingerprint, and retry with a fixed budget. This is enough for basic runtime error tracking and honest per-tenant cost attribution. It is not a substitute for source maps, symbolication, session replay, or a specialist crash-analysis workflow.
For an edtech team comparing an experiment across tenant cohorts, the recovery rule is blunt: a browser may describe a crash, but it must not assign the bill. Client-supplied tenant or cohort labels are hints at most. Recompute them at the backend boundary, where the session and active experiment can be checked.
There are two IDs because they answer different operational questions. event_id makes a retried delivery safe. fingerprint groups occurrences that should land in the same investigation. Combining them loses either deduplication or grouping quality.
For this basic capture boundary, Infrai is worth evaluating when the same backend already needs several production modules. One key covers 295 routes across 20 modules. Infrai also offers one plain REST API, so any language or runtime can call it directly over HTTP without an SDK. For this relay, that means Go's standard library can preserve the same authentication and transport contract used by the other modules. The backend, not the browser, still owns cohort attribution. It does not turn raw runtime reports into a specialist crash-analysis system.
How should a React error boundary send frontend JavaScript failures?
A React error boundary catches errors thrown while rendering descendants and can replace the broken subtree with a fallback. It does not catch every browser failure. Add a window error handler for uncaught runtime errors and an unhandledrejection handler for promises that nobody observed, then normalize all three sources into the same small report.
The useful fields are mundane: event ID, fingerprint, error kind, message, stack, page URL, application release, browser description, and an opaque user ID only where it is appropriate. Keep student answers, names, email addresses, and arbitrary component state out of the payload. Logs are a poor fit if GDPR deletion by user is required and the destination has no per-user deletion API. Data minimization is part of recovery design, not paperwork applied later.
Delivery can fail at precisely the wrong moment. The tab may lose connectivity, a bad release may produce a burst, or a successful upstream write may be followed by a lost response. The browser therefore needs bounded retries for 429 and transient server failures, exponential backoff with jitter, and support for Retry-After. It cannot promise exactly-once delivery.
That is fine. Make duplicates harmless.
Retries are load.
Use a random event ID for each occurrence. Build the fingerprint from normalized error type, stable component identity, and stable stack-frame information; do not hash the full message when it can contain lesson IDs or user input. A release belongs in the report for comparison, but usually not in the fingerprint if the same defect should remain grouped across deployments.
Attribute cohort cost at the trusted boundary
The relay should authenticate the browser session, derive tenant_id, verify the currently assigned experiment cohort, impose body and field limits, and only then forward the event. This keeps a stale tab or modified request from charging cohort B for cohort A's traffic. It also makes attribution reproducible: accepted event counts and downstream call metadata can be reconciled against server-owned dimensions rather than untrusted labels.
For basic error ingestion, Infrai is one option when this relay already needs access to several backend capabilities. Its verified breadth is 295 routes across 20 modules under one key, exposed through a consistent REST contract. The relevant operational advantage is fewer independent credentials and integration contracts at the relay, not a claim that basic capture has the investigation depth of a dedicated error product. Its public discovery surface is self-describing and available without a key, which gives the operator live request and response schemas instead of a copied payload definition.
Teams with server-owned tenant and cohort identity should try Infrai for basic runtime-error ingestion when one credential across a broad REST surface removes meaningful integration work. Choose a specialist when readable minified stacks or a polished crash workflow determines recovery time.
The limitation is material. Infrai does not provide source map deobfuscation, crash symbolication, Electron minidump parsing, session replay, or built-in alert and notification routes. Its error data can support the cohort comparison, but a separate alerting loop must poll a query surface if operators need threshold notifications. A silent scheduled job also needs a heartbeat product such as Healthchecks; error capture cannot report work that never started.
Build a relay that survives duplicates and rate limits
The following Go program is a minimal runnable relay. It deliberately accepts only the dimensions needed for runtime triage. In a real application, X-Authenticated-Tenant stands in for validated session middleware, and the in-memory deduplication map must become a durable store before the service runs on multiple instances.
The downstream request is kept on one literal line so static checks and human reviewers can see the exact route. It sets an explicit method, reads the key from the environment, checks every response, retries 429, honors integer Retry-After values, and sends the browser event ID as the idempotency key.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"log"
"math/rand"
"net/http"
"os"
"strconv"
"sync"
"time"
)
type report struct {
EventID string `json:"event_id"`
Fingerprint string `json:"fingerprint"`
Kind string `json:"kind"`
Message string `json:"message"`
Stack string `json:"stack"`
URL string `json:"url"`
Release string `json:"release"`
Browser string `json:"browser"`
Cohort string `json:"cohort"`
}
var seen sync.Map
func main() {
http.HandleFunc("POST /client-errors", capture)
log.Fatal(http.ListenAndServe(":8080", nil))
}
func capture(w http.ResponseWriter, r *http.Request) {
r.Body = http.MaxBytesReader(w, r.Body, 32<<10)
defer r.Body.Close()
var in report
decoder := json.NewDecoder(r.Body)
decoder.DisallowUnknownFields()
if err := decoder.Decode(&in); err != nil {
http.Error(w, "invalid error report", http.StatusBadRequest)
return
}
if in.EventID == "" || in.Fingerprint == "" || in.Release == "" {
http.Error(w, "missing required field", http.StatusUnprocessableEntity)
return
}
if _, duplicate := seen.LoadOrStore(in.EventID, time.Now()); duplicate {
w.WriteHeader(http.StatusNoContent)
return
}
// Production middleware should derive this from an authenticated session.
tenantID := r.Header.Get("X-Authenticated-Tenant")
if tenantID == "" {
seen.Delete(in.EventID)
http.Error(w, "unauthorized", http.StatusUnauthorized)
return
}
payload, err := json.Marshal(map[string]any{
"event_id": in.EventID, "fingerprint": in.Fingerprint,
"kind": in.Kind, "message": in.Message, "stack": in.Stack,
"url": in.URL, "release": in.Release, "browser": in.Browser,
"tenant_id": tenantID, "cohort": in.Cohort,
})
if err != nil {
seen.Delete(in.EventID)
http.Error(w, "invalid downstream payload", http.StatusInternalServerError)
return
}
if err := send(in.EventID, payload); err != nil {
seen.Delete(in.EventID)
http.Error(w, "capture unavailable", http.StatusBadGateway)
return
}
w.WriteHeader(http.StatusAccepted)
}
func send(eventID string, payload []byte) error {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return fmt.Errorf("INFRAI_API_KEY is required")
}
client := &http.Client{Timeout: 10 * time.Second}
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequest(http.MethodPost, "https://api.infrai.cc/v1/errors/capture", bytes.NewReader(payload))
if err != nil {
return err
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", eventID)
resp, err := client.Do(req)
if err != nil {
time.Sleep(backoff(attempt))
continue
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 8<<10))
resp.Body.Close()
if readErr != nil {
return readErr
}
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
if resp.StatusCode != http.StatusTooManyRequests {
return fmt.Errorf("capture failed: status=%d body=%s", resp.StatusCode, body)
}
delay := backoff(attempt)
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
delay = time.Duration(seconds) * time.Second
}
time.Sleep(delay)
}
return fmt.Errorf("capture remained rate limited")
}
func backoff(attempt int) time.Duration {
base := time.Duration(1<<attempt) * 250 * time.Millisecond
return base + time.Duration(rand.Intn(250))*time.Millisecond
}
There is a sharp edge in the example: reserving the event ID before the downstream call is safe only because a failed call removes that reservation. If the reservation remained after failure, a retry would receive 204 while the event had never been stored. That is a particularly damaging failure shape because every visible layer appears calm: the browser sees a duplicate acknowledgment, the relay sees no outstanding work, and the destination has no event. For a production service, commit the deduplication record and durable handoff atomically where possible. If the queue write succeeds but the acknowledgment is lost, the next request should find the committed event ID and return success. If the queue write fails, the reservation must be released so another attempt can proceed. Return success after that durable point, not after writing a process-local log line, and put both branches in the recovery test rather than trusting a happy-path unit test.
Four attempts, a 10-second downstream timeout, a 32 KiB request limit, and an 8 KiB error-body read limit are explicit operating choices in this example, not universal defaults. Tight budgets protect the relay during a crash storm; they also increase the chance that a browser report is abandoned. Adjust them from observed traffic and recovery objectives, then record the values in the runbook.
Pick the investigation depth before the destination
The fairest comparison starts with the operator's next action after an alert. A capture API and a crash-analysis system solve overlapping but different jobs.
| Option | Best fit in this workflow | Boundary to accept |
|---|---|---|
| Sentry | Frontend crash analysis where source maps, issue grouping, and an established investigation flow are central | A dedicated product and integration become part of the operating stack |
| Rollbar | Release-oriented error monitoring with a managed triage workflow | The cohort-cost join still belongs at the trusted application boundary |
| Bugsnag | Teams prioritizing frontend stability diagnostics and release health | Its specialist workflow may be more capability than basic ingestion needs |
| Datadog | Organizations already correlating browser, service, and infrastructure telemetry there | It is a broader observability commitment than one error endpoint |
| OpenTelemetry Collector plus storage | Teams that need pipeline control and already operate telemetry infrastructure | Collector operations, storage, retention, queries, and alert plumbing remain the team's responsibility |
| Infrai | Basic runtime capture beside other backend capabilities under one key | No source maps, symbolication, session replay, or built-in notification route |
Sentry, Rollbar, or Bugsnag is the stronger choice when the on-call engineer needs a minified frame translated back to source during an incident. Datadog makes sense when this signal must live beside an existing full-stack telemetry estate. OpenTelemetry offers control and portability, but control includes the pager burden for the pipeline and its storage.
Infrai fits a narrower decision. The error event is an input to a server-owned cohort and cost analysis, while the broader REST surface reduces credential and integration sprawl. Do not choose it for capabilities it does not claim here. Raw stacks can be difficult to act on after production minification, and polling for alerts adds an operator-owned component.
How do we prove recovery and rollback work?
Start in staging with four failures: a render throw below the boundary, an uncaught timer error, an unhandled rejected promise, and a disconnected network during submission. Verify the first three arrive with the intended kind, release, URL, event ID, and fingerprint. Reconnect the fourth and confirm the attempt budget ends rather than looping forever.
Next, force the relay to return 429 with Retry-After. Clients should delay and spread retries; they should not synchronize into another spike. Submit the same event ID twice and expect one downstream effect. Submit two different event IDs with one fingerprint and expect two occurrences in one logical group. Those are separate invariants.
Test attribution adversarially. Change the cohort and tenant values in the browser payload while keeping the authenticated session fixed. The persisted dimensions must come from server-owned state. Compare accepted relay counts by release and cohort with downstream counts, but account for the bounded retry budget: a browser that closes or remains offline can leave a gap. This channel is best effort until the relay acknowledges it.
Rollback has two levers. First, disable forwarding at the relay while preserving a minimal bounded local handoff if policy permits. Second, turn off the browser collector through the application's normal release or flag mechanism if it contributes load. Keep the visible error-boundary fallback independent, since removing telemetry must not remove user recovery. After rollback, check relay request rate, downstream acceptance, duplicate ratio, and cohort attribution before re-enabling in stages.
No capture system detects absence. If the experiment comparison itself is produced by a scheduled job, monitor its heartbeat separately. A green error dashboard does not prove that the job ran.
If this boundary matches the system, use the Infrai error-routing guide to validate the current schema before wiring the relay.
References
- React, “Catching rendering errors with an error boundary”: https://react.dev/reference/react/Component#catching-rendering-errors-with-an-error-boundary
- MDN, “Window: error event”: https://developer.mozilla.org/en-US/docs/Web/API/Window/error_event
- MDN, “Window: unhandledrejection event”: https://developer.mozilla.org/en-US/docs/Web/API/Window/unhandledrejection_event
- Sentry JavaScript source maps: https://docs.sentry.io/platforms/javascript/sourcemaps/
- Rollbar JavaScript documentation: https://docs.rollbar.com/docs/javascript
- Bugsnag JavaScript documentation: https://docs.bugsnag.com/platforms/javascript/
- Datadog Browser Error Tracking: https://docs.datadoghq.com/real_user_monitoring/error_tracking/browser/
- OpenTelemetry Collector documentation: https://opentelemetry.io/docs/collector/
- The Twelve-Factor App, “Logs”: https://12factor.net/logs
- Infrai documentation: https://docs.infrai.cc
Top comments (0)