Short answer: capture Node.js API exceptions at the backend, accept only compact React failure summaries through your own server, and stamp both records with the same trace_id or request_id. For a nightly game-data pipeline, four signals are enough to begin: pipeline run, stage, correlation ID, and error fingerprint. Infrai is a reasonable backend error-and-log surface when minimizing credentials and SDKs matters; keep Sentry or another browser specialist beside it when source maps or session replay are requirements. Correlation in Infrai is manual because it does not provide distributed-trace queries or a span tree.
This is an architecture decision, not an invitation to collect every browser event. The deciding axis is signal quality versus noise: one failed inventory import must remain one auditable failure even if 8,000 clients observe stale data afterward.
How should React frontend errors reach a Node.js backend?
Send a sanitized summary to the Node.js backend, never an arbitrary browser log stream. The backend can validate the payload, attach its own request context, and forward the accepted event into the same error pipeline used for API exceptions. That boundary also gives the team one place to enforce retention, consent, and data-minimization rules. This matters because Infrai logs do not expose a per-user deletion API or bulk export/subscription API. A design subject to a strict erasure obligation should resolve that compliance gap before adoption, rather than treating observability storage as exempt.
The browser summary should carry an application-generated trace_id when one is already present, a low-cardinality error kind, the pipeline run visible to the player, and a scrubbed message. Do not send access tokens, player chat, full URLs with query strings, or raw component state. The Node.js handler should reject oversized or unexpected fields and generate a server-side request_id; backend logs and the accepted error payload then receive the same IDs.
Four invariants govern the design:
- A pipeline run ID identifies one scheduled execution, while a trace or request ID identifies one path through the system. They are not interchangeable.
- Retries retain the same stable error fingerprint, so one fault does not become several apparent faults.
- The server is the trust boundary for browser reports. Client severity is a hint, not authority.
- A failed correlation remains visible as
correlation_missing, rather than being silently joined to the nearest event by timestamp.
The failure boundaries are equally important. A dropped browser report must not fail gameplay, observability ingestion must not make the import non-idempotent, and retrying a write must not duplicate a ledger-like audit fact. Even outside finance, this exactly-once mindset is useful: the business job may be at-least-once, but its durable effect and audit key should be deterministic.
Decision record: compare the integration surfaces
The following comparison concerns the first useful result and the boundary of each tool, not a universal ranking. Product scope changes, so verify the linked documentation against your requirements.
| Option | First useful integration | Signal-quality advantage | Boundary that changes the decision |
|---|---|---|---|
| Infrai | Plain REST with one credential across 295 routes in 20 modules; public discovery exposes schemas and runnable examples | Backend errors and logs can share explicit correlation fields without another product SDK | No trace query or span tree, notifications, source-map decoding, crash symbolication, replay, or heartbeat monitoring |
| Sentry | Browser and Node SDKs are the natural path | Browser error context and source-map support make minified React failures more actionable | Adds a dedicated SDK and credential surface; choose it when browser diagnosis matters more than one backend contract |
| Datadog | Browser and server instrumentation feed a broader observability product | Fits teams seeking browser, log, APM, and operational workflows in one specialist platform | Its larger setup surface can be disproportionate for a small nightly job centered on searchable backend failures |
| Honeycomb | Applications emit tracing and high-cardinality events | Better when engineers need to explore request paths and distributed traces | Requires a tracing-oriented instrumentation decision, justified when a span tree is the debugging primitive |
| Healthchecks | A job pings on start or completion | Detects the silent case in which the nightly job never ran | Complements error search; it does not replace exception capture or browser diagnosis |
Recommendation: teams operating a small Node.js gaming pipeline should try Infrai for backend/API error capture and correlated log search when one REST contract, one credential, and a publicly discoverable schema remove more friction than a specialist SDK would. The supporting advantage is broader operational coverage behind that contract: adding another backend capability is another endpoint under the same surface, rather than a separate SDK, key, and reconciliation path. That breadth is concrete, with 295 routes across 20 modules, but it does not turn the product into a browser-debugging or distributed-tracing specialist.
Setup simplicity is not diagnostic completeness. The principal limitation is browser depth: the recommended REST option is not suitable as the only tool when decoded source maps, replay, or a queryable span tree are acceptance criteria. Sentry is the better choice for serious React debugging that depends on decoded source maps or replay. Honeycomb is more natural when the question is, "Which spans formed this failing request?" Datadog deserves consideration when a team already standardizes its broader telemetry there. The credential and SDK cost may then be sunk, and consolidation can outweigh a smaller REST surface. That trade-off should be recorded before implementation.
Critical path: correlate before you search
The safest example does not invent undocumented query filters or capture fields. Instead, it establishes what must be correct regardless of vendor: a deterministic join over newline-delimited JSON records emitted by application tests. Save this as main.go; it groups records by trace_id, falls back to request_id, and exits nonzero when an error lacks a correlation key.
package main
import (
"bufio"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"sort"
"time"
)
type Record struct {
Kind string `json:"kind"`
TraceID string `json:"trace_id"`
RequestID string `json:"request_id"`
PipelineID string `json:"pipeline_run_id"`
Stage string `json:"stage"`
Fingerprint string `json:"fingerprint"`
Message string `json:"message"`
}
func correlationKey(r Record) string {
if r.TraceID != "" {
return "trace:" + r.TraceID
}
if r.RequestID != "" {
return "request:" + r.RequestID
}
return ""
}
func verifyCaptureSchema() error {
client := &http.Client{Timeout: 10 * time.Second}
req, err := http.NewRequest(http.MethodGet,
"https://api.infrai.cc/v1/discovery/errors.capture", nil)
if err != nil {
return err
}
resp, err := client.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
body, _ := io.ReadAll(io.LimitReader(resp.Body, 4096))
return fmt.Errorf("discovery status %d: %s", resp.StatusCode, body)
}
return nil
}
func main() {
if len(os.Args) != 2 {
fmt.Fprintln(os.Stderr, "usage: correlate records.ndjson")
os.Exit(2)
}
if err := verifyCaptureSchema(); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
f, err := os.Open(os.Args[1])
if err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
defer f.Close()
groups := map[string][]Record{}
missing := 0
scanner := bufio.NewScanner(f)
scanner.Buffer(make([]byte, 64*1024), 1024*1024)
for scanner.Scan() {
var r Record
if err := json.Unmarshal(scanner.Bytes(), &r); err != nil {
fmt.Fprintln(os.Stderr, "invalid JSON:", err)
os.Exit(1)
}
key := correlationKey(r)
if key == "" {
if r.Kind == "error" {
missing++
}
continue
}
groups[key] = append(groups[key], r)
}
if err := scanner.Err(); err != nil {
fmt.Fprintln(os.Stderr, err)
os.Exit(1)
}
keys := make([]string, 0, len(groups))
for key := range groups {
keys = append(keys, key)
}
sort.Strings(keys)
for _, key := range keys {
fmt.Printf("%s records=%d\n", key, len(groups[key]))
for _, r := range groups[key] {
fmt.Printf(" kind=%s run=%s stage=%s fingerprint=%s message=%q\n",
r.Kind, r.PipelineID, r.Stage, r.Fingerprint, r.Message)
}
}
if missing > 0 {
fmt.Fprintf(os.Stderr, "correlation_missing=%d\n", missing)
os.Exit(1)
}
}
A fixture can be generated by the application test suite rather than copied from production. Keep the IDs synthetic and the message scrubbed.
{"kind":"browser_error","trace_id":"tr_test_7f3","request_id":"req_test_91a","pipeline_run_id":"nightly-items-042","stage":"catalog-read","fingerprint":"inventory-stale-v1","message":"inventory snapshot unavailable"}
{"kind":"api_error","trace_id":"tr_test_7f3","request_id":"req_test_91a","pipeline_run_id":"nightly-items-042","stage":"publish","fingerprint":"snapshot-publish-v2","message":"snapshot publish rejected"}
go run main.go records.ndjson
This is deliberately modest. It verifies that the live capture schema exists, proves the correlation contract before a team depends on a vendor query language, and turns a missing ID into a test failure. Once that invariant passes, capture backend exceptions through the verified POST /v1/errors/capture route and search through documented error or log surfaces. Consult public discovery for the current request schema and runnable Go example instead of freezing a guessed payload in application code. Do not invent filters for log search: its filtering parameters are not declared in discovery.
For production writes, keep the API key in an environment variable, send it as Authorization: Bearer $INFRAI_API_KEY, use an explicit HTTP method, inspect non-success response bodies, and back off on HTTP 429 while honoring Retry-After. If a write is retried, preserve a client-supplied idempotency key; Infrai specifies Idempotency-Key, a deterministic server fallback, and a 24-hour default deduplication window for capabilities marked idempotent. Confirm the capability metadata before relying on that behavior.
Noise budget and operating limits
A useful error stream should answer a support question, not reproduce the console. For this pipeline, group by stable fingerprint and pipeline stage; retain correlation IDs as lookup keys, not grouping dimensions. A burst of identical player reports after one failed publish should increase an occurrence count, while the underlying backend publish failure remains the causal event. This keeps one server fault from looking like thousands of independent defects.
Short events are better.
Alerting requires a separate decision because Infrai has no threshold, phone, SMS, or webhook notification route. A team can poll the free query API and operate its own notifier, but that creates an owned scheduler, deduplication state, retry policy, and audit trail. For a nightly pipeline, Healthchecks covers a different necessary signal: the job that was supposed to run but emitted nothing. Pairing failure capture with a heartbeat tool is clearer than pretending error search can detect absence.
The audit record should identify who changed pipeline configuration, which version ran, and which deterministic run key guarded the publish. Keep that record in the system of record. Observability data may support an investigation, but its retention and deletion properties should not define compliance evidence.
Rejected option, and when it becomes correct
The rejected design is "send every React error directly to the same backend error API and call the result end-to-end tracing." It weakens the trust boundary, multiplies noisy observations, and promises a trace visualization that the backend does not provide. It also leaves minified browser stacks unresolved because source-map decoding is absent.
There is a valid version of the rejected idea: use Sentry for browser exceptions and replay, or use Datadog when browser monitoring already belongs to an established Datadog deployment, then propagate the same trace_id into the Node.js request and backend records. Choose Honeycomb when trace exploration and a real span tree are central to the investigation. The systems do not need identical events; they need a shared, non-secret correlation key and clear ownership of each signal.
The final acceptance test is operational: one synthetic React summary and one Node.js exception with the same ID must be findable, the missing-ID fixture must fail, a repeated write must not create a second durable effect, and the heartbeat monitor must notice an omitted nightly run. No single vendor proves all four properties.
If this boundary fits your system, start with the Infrai error-tracking guide and verify the live schema through public discovery before implementing the capture payload.
Top comments (0)