DEV Community

KiernanBerg3867
KiernanBerg3867

Posted on

Live-Game React Frontend Error Tracking — A Backend Collector Example

A live game creates plenty of browser noise: extensions throw, stale tabs linger, and one broken asset can generate thousands of copies of the same exception. The operational constraint changes the design. Short answer: for React frontend error tracking, send window.onerror and unhandledrejection through a small backend collector, scrub PII in the browser, attach the stack, release, and environment, then group repeats by release after each deployment. This retains enough evidence to reconstruct many customer incidents without pretending a basic feed is full client observability.

The trade-off is sharp. Keep the exception type, a bounded message, the minified stack, page origin and path, browser family, game build, deployment environment, and a random incident ID. Drop query strings, player names, email addresses, chat text, access tokens, and raw rejection objects. For an indie team, fewer fields with known meaning beat an event swamp that nobody can safely inspect.

How should a React frontend send errors to a backend collector?

Start from the reconstruction question: which build failed, in which browser, on which route, and did the same signature appear after the deployment? A raw stack helps even when minified, but it may identify only a bundle offset. Without a separate build-time source-map workflow, it will not lead back to the original TypeScript line.

That is the first boundary. A basic backend capture path can accept error events and later retrieve and group them, including repeated crashes by release. It does not perform source-map deobfuscation, crash symbolication, Electron minidump parsing, or session replay. No recording will reveal the player's final clicks.

I would also resist attaching a stable player ID "just in case." This class of error and log storage has no user-specific deletion workflow suitable for a GDPR forgotten-user request. The safer design is to avoid collecting the identifier at all. Use a random per-page incident ID when correlation is necessary, and expire it in memory when the page closes.

Put the privacy boundary before the network boundary

The browser should construct an allow-listed event rather than serialize the thrown value. The following client module handles both global error channels, limits high-volume duplicates, removes URL queries and fragments, and sends only fields the collector contract names.

type SafeCrash = {
  incidentId: string;
  kind: "error" | "rejection";
  message: string;
  stack?: string;
  page: string;
  browser: string;
  build: string;
  environment: "production" | "staging";
  occurredAt: string;
};

const incidentId = crypto.randomUUID();
const recent = new Map<string, number>();

function bounded(value: unknown, limit: number): string {
  return String(value ?? "Unknown error").slice(0, limit);
}

function pageWithoutSecrets(): string {
  const url = new URL(window.location.href);
  return `${url.origin}${url.pathname}`;
}

function submit(kind: SafeCrash["kind"], value: unknown): void {
  const error = value instanceof Error ? value : undefined;
  const message = bounded(error?.message ?? value, 500);
  const stack = error?.stack ? bounded(error.stack, 8_000) : undefined;
  const fingerprint = `${kind}:${message}:${stack?.split("\n")[1] ?? ""}`;
  const now = Date.now();

  if (now - (recent.get(fingerprint) ?? 0) < 30_000) return;
  recent.set(fingerprint, now);

  const event: SafeCrash = {
    incidentId,
    kind,
    message,
    stack,
    page: pageWithoutSecrets(),
    browser: navigator.userAgent.slice(0, 300),
    build: import.meta.env.VITE_APP_BUILD,
    environment: import.meta.env.PROD ? "production" : "staging",
    occurredAt: new Date().toISOString(),
  };

  void fetch("/browser-errors", {
    method: "POST",
    headers: { "Content-Type": "application/json" },
    body: JSON.stringify(event),
    keepalive: true,
  });
}

window.addEventListener("error", (event) => submit("error", event.error ?? event.message));
window.addEventListener("unhandledrejection", (event) => submit("rejection", event.reason));
Enter fullscreen mode Exit fullscreen mode

The server owns the credential and the vendor adapter. This minimal forwarding function targets the verified capture route, reuses the page-scoped incident ID as an idempotency key, honors Retry-After on rate limits, and surfaces the real response body on failure. INFRAI_BASE_URL should be configured server-side for the API origin; keeping the origin out of browser code also prevents accidental credential exposure.

async function forwardCrash(event: SafeCrash, attempt = 0): Promise<unknown> {
  const apiKey = process.env.INFRAI_API_KEY;
  const baseUrl = process.env.INFRAI_BASE_URL;
  if (!apiKey || !baseUrl) throw new Error("Missing server-side error API configuration");

  const response = await fetch(new URL("/v1/errors/capture", baseUrl), {
    method: "POST",
    headers: {
      Authorization: `Bearer ${apiKey}`,
      "Content-Type": "application/json",
      "Idempotency-Key": event.incidentId,
    },
    body: JSON.stringify(event),
  });

  if (response.status === 429 && attempt < 4) {
    const retryAfter = Number(response.headers.get("Retry-After"));
    const delayMs = Number.isFinite(retryAfter)
      ? retryAfter * 1_000
      : 500 * 2 ** attempt;
    await new Promise((resolve) => setTimeout(resolve, delayMs));
    return forwardCrash(event, attempt + 1);
  }

  const body = await response.text();
  if (!response.ok) {
    throw new Error(`Error capture failed (${response.status}): ${body}`);
  }
  return body ? JSON.parse(body) : null;
}
Enter fullscreen mode Exit fullscreen mode

This is deliberately boring. The allow list is the feature. It prevents an enthusiastic developer from later adding location.href, an entire game-state object, or a rejected HTTP response containing credentials.

The 30-second duplicate window is a starting policy, not a universal constant. It cuts a render-loop storm on one page while retaining recurrence across sessions. Record how many events the client suppresses; if suppression is high, inspect the top signatures before widening the window. I chose client-side suppression because it stops a render loop before it consumes network and intake capacity, but I would keep independent server limits because hostile or obsolete clients can bypass browser policy. The trade-off is deliberate: a repeat inside one page may disappear, while the same crash from another page session remains visible. Also cap request size at the backend, validate the exact schema, and return quickly so error reporting cannot become a second gameplay failure.

Noise compounds.

Group around deployments, not individual players

Grouping should answer whether a deployment changed the crash population. A practical key combines normalized exception type, normalized message, and the first useful stack frame; the build remains a dimension, not part of the key. That lets one group show a clean before-and-after split instead of fragmenting into a new group for every release.

Normalization needs restraint. Removing generated line and column numbers can join equivalent minified crashes, but stripping every number may merge distinct asset, level, or status-code failures. Sample merged events periodically. False joins are quieter than duplicate groups and more dangerous because they distort impact.

For each deployment, watch four measurements: unique incident IDs per group, total occurrences, affected browser families, and first-seen build. Those are counts, not invented severity. A crash seen 4,000 times in one looping tab deserves different treatment from the same signature across 200 independent page sessions.

Infrai fits the narrow baseline when a small team wants error capture and grouping alongside many other backend modules through one REST API and one key: its discovery surface reports 295 routes across 20 modules, and documented capabilities include runnable TypeScript examples. That breadth means another backend capability can use the same contract and credential instead of adding another SDK integration. The supporting advantage here is consistent discovery metadata, which lets an adapter validate the current request schema instead of freezing guessed fields into client code. Keep the API key in that server-side adapter, never in React.

Choosing the signal system

The honest comparison is between operating models, not feature-count theater.

Option Best fit Evidence advantage Boundary to accept
Sentry Teams that want a dedicated client-observability workflow JavaScript SDK, source-map handling, and session replay are part of the product surface Another specialized SDK and service to operate
Bugsnag Release-oriented stability work Error monitoring emphasizes releases, stability, and source maps A dedicated integration is justified only if that richer workflow will be used
Rollbar Teams centered on grouped exceptions and deploy tracking JavaScript error telemetry and source maps support diagnosis from minified builds Still a separate observability product and client integration
Datadog Teams already operating a wider monitoring platform Browser errors can sit beside real-user monitoring and backend telemetry The broader platform is heavier than a narrow crash feed
OpenTelemetry Teams standardizing portable telemetry pipelines Vendor-neutral browser telemetry building blocks Error-product workflows depend on the collector and backend assembled around it
Basic REST capture Solo teams needing reconstructable crash evidence now Small contract, direct control of minimization, and grouping by build No automatic deobfuscation or replay; alerting must be built separately

Sentry, Bugsnag, and Rollbar are the stronger choices when original source locations or richer client context are requirements. Datadog makes more sense when browser failures need to share an operating surface with real-user monitoring and backend telemetry. OpenTelemetry is attractive when portability across a broader telemetry pipeline matters more than getting a finished error-triage interface immediately. The REST baseline earns its place when the requirement is narrower: preserve privacy-conscious evidence, identify deployment regressions, and avoid adding another client SDK before the signal proves useful.

Do not stretch that baseline into jobs it does not cover. There is no notification route for thresholds, phone, SMS, or webhooks, so a team needs to poll query results and own alert delivery. There is no distributed trace query or span tree, even though log records can carry trace_id and span_id. Scheduled-job silence also belongs in a heartbeat product such as Healthchecks, not in browser exception capture.

Measure this before copying the design

Run the collector for one deployment window and measure intake volume, payload bytes, duplicate suppression, groups per build, events missing stacks, and the share of groups that an engineer can actually act on. Inspect a sample for personal data. Do it manually at first.

The decision rule is simple: keep the basic path if grouped build-level evidence reconstructs the incidents you care about and review load stays low. Move to Sentry, Bugsnag, or Rollbar when minified frames repeatedly block diagnosis or replay becomes necessary. Choose OpenTelemetry when the browser signal must join a vendor-neutral telemetry architecture and the team can own the additional pipeline work.

Signal quality wins. More events do not.

Sources

Top comments (0)