Capture handled exceptions, unhandled rejections, and uncaught exceptions in one small server-side adapter, and attach the request, release, environment, tenant, and experiment cohort needed to reproduce them. For a developer-tool SaaS comparing an experiment across tenant cohorts, that is the useful baseline. Choose a full APM product instead when the real question is how a failure moved through several services.
Short answer: centralize capture, redact request data at the boundary, preserve the original stack, and make delivery bounded. Error tracking should improve signal quality without turning every rejected promise into five noisy events. The revenue-per-hour test is blunt: if an integration cannot tell which tenants and release are affected within a few minutes, it is stealing feature time.
How should a NodeJS Express error tracking API capture failures?
The hard part is not listening for uncaughtException. Node exposes that event. The hard part is retaining enough context to distinguish a bad cohort rollout from an ordinary application bug, while avoiding secrets and duplicate reports.
Suppose an experiment has control and candidate cohorts. A stack trace alone may show the same failing function for both. Add tenantId, cohort, release, method, and route, and the errors become comparable. Use a route pattern such as /api/projects/:id, not the raw URL; identifiers in raw paths create needless cardinality and may expose customer data. Do not copy cookies, authorization headers, request bodies, or query strings into an error event by default.
There is another constraint: fatal-process reporting cannot become a shutdown strategy. An uncaughtException leaves the process in an undefined state. Capture it with a short deadline, then exit and let the process supervisor restart the service. I would also separate handled operational errors from programmer errors: a rejected request caused by known validation should normally become a structured 4xx response, not an exception report. My first instinct is to capture every thrown value because missing evidence feels worse than excess data. That instinct loses once routine 4xx cases bury the failures that can actually stop a tenant from working. The capture path is for unexpected failures, and that choice cuts noise before any vendor-side grouping algorithm sees the data.
Keep it boring.
The smallest working implementation
This TypeScript example targets Node.js 20 or later and Express 5. It calls Infrai through one verified error-capture route, uses a stable idempotency key across retries, attaches explicit request metadata, and applies bounded exponential backoff for HTTP 429. Set INFRAI_API_BASE to the API base URL and keep the key in INFRAI_API_KEY; neither belongs in source control. The two-second request deadline and three-attempt ceiling make the failure budget visible instead of leaving it to a library default.
import express, { ErrorRequestHandler, Request } from "express";
import { randomUUID } from "node:crypto";
const app = express();
app.use(express.json({ limit: "100kb" }));
type ErrorContext = {
environment: string;
release: string;
request?: {
method: string;
route: string;
};
user?: {
tenantId: string;
cohort: "control" | "candidate";
};
};
const sleep = (ms: number) =>
new Promise<void>((resolve) => setTimeout(resolve, ms));
function retryDelay(response: Response, attempt: number): number {
const retryAfter = response.headers.get("retry-after");
if (retryAfter) {
const seconds = Number(retryAfter);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const dateDelay = Date.parse(retryAfter) - Date.now();
if (Number.isFinite(dateDelay)) return Math.max(0, dateDelay);
}
return 250 * 2 ** attempt;
}
async function captureError(error: unknown, context: ErrorContext): Promise<void> {
const baseUrl = process.env.INFRAI_API_BASE;
const apiKey = process.env.INFRAI_API_KEY;
if (!baseUrl || !apiKey) throw new Error("Missing error API configuration");
const normalized = error instanceof Error ? error : new Error(String(error));
const idempotencyKey = randomUUID();
for (let attempt = 0; attempt < 3; attempt += 1) {
const response = await fetch(new URL("/v1/errors/capture", baseUrl), {
method: "POST",
headers: {
"Authorization": `Bearer ${apiKey}`,
"Content-Type": "application/json",
"Idempotency-Key": idempotencyKey,
},
body: JSON.stringify({
message: normalized.message,
stack: normalized.stack,
...context,
}),
signal: AbortSignal.timeout(2_000),
});
if (response.ok) return;
if (response.status === 429 && attempt < 2) {
await sleep(retryDelay(response, attempt));
continue;
}
const detail = await response.text();
throw new Error(`Error capture failed (${response.status}): ${detail}`);
}
}
function requestContext(req: Request): ErrorContext {
return {
environment: process.env.NODE_ENV ?? "development",
release: process.env.APP_RELEASE ?? "local",
request: {
method: req.method,
route: req.route?.path ?? "unmatched",
},
user: req.header("x-tenant-id")
? {
tenantId: req.header("x-tenant-id")!,
cohort: req.header("x-experiment-cohort") === "candidate"
? "candidate"
: "control",
}
: undefined,
};
}
app.get("/api/projects/:id", async (_req, res) => {
throw new Error("Project index unavailable");
});
const errorHandler: ErrorRequestHandler = (error, req, res, _next) => {
void captureError(error, requestContext(req)).catch((captureFailure: unknown) => {
console.error("Error reporting failed", captureFailure);
});
res.status(500).json({ error: "internal_error" });
};
app.use(errorHandler);
const processContext = (): ErrorContext => ({
environment: process.env.NODE_ENV ?? "development",
release: process.env.APP_RELEASE ?? "local",
});
process.on("unhandledRejection", (reason) => {
void captureError(reason, processContext()).catch(console.error);
});
process.on("uncaughtException", async (error) => {
try {
await captureError(error, processContext());
} catch (captureFailure) {
console.error("Fatal error reporting failed", captureFailure);
} finally {
process.exit(1);
}
});
app.listen(3000);
The request handler responds immediately rather than holding the customer request open for telemetry. This is a deliberate trade. A sudden process death can lose an in-flight handled-error report, but adding a durable local queue would make this example much larger. The fatal path waits because the process is about to exit anyway, while AbortSignal.timeout prevents an indefinite hang.
One subtle trap sits in the cohort header. The sample accepts only one explicit value and maps everything else to control. In a real service, derive experiment membership from trusted server-side state. A caller-controlled header can corrupt the comparison.
The five-minute cohort triage drill
Start with groups, then inspect individual events. Basic grouping is useful for a small admin error page, but it is not a statistical experiment system. For each release, compare the number of affected tenants in candidate with the number in control; do not treat raw event totals as equivalent to tenant impact. One retry loop can emit many failures from one account.
The event should retain the original message and stack alongside the cohort context. Release is the pivot for deciding whether a new deployment lines up with the first occurrence, while environment prevents staging noise from contaminating production triage. Then look at one representative event from the dominant group and read the stack before building another chart. Avoid claiming causality from these counts because traffic and tenant behavior may differ between cohorts. Error tracking answers, “Where should I inspect?” A proper experiment analysis answers, “Did the treatment cause the change?” Those are different jobs. For a one-person SaaS, the basic admin view can stay deliberately small: unresolved groups ordered by recency, affected-tenant count, release, and the latest event. Fancy charts can wait until they change a decision.
Ship that.
Product choice comes after the signal test
The right product depends on the debugging depth you need, not the length of its feature page.
| Product | Strong fit | Boundary that matters here |
|---|---|---|
| Sentry | Application error monitoring where source maps, releases, tracing, and replay belong in one workflow | Broader client and workflow surface than a minimal backend capture adapter |
| Rollbar | Error grouping and occurrence telemetry with established Node and Express integration | Another dedicated SDK and service to operate in the application |
| Honeybadger | Focused exception monitoring plus uptime and check-in monitoring | Less compelling if the goal is consolidating many unrelated backend capabilities behind one contract |
| Datadog | Logs, APM traces, infrastructure signals, and errors need to be queried together | More platform and instrumentation than a small service may need for basic triage |
| Infrai | Basic backend error capture is one capability among 295 routes across 20 modules under one key and a consistent REST contract | Grouping is basic; there is no source-map reverse mapping, crash symbolication, session replay, distributed trace query, or span tree |
Sentry is the clearest match when browser stack traces and session evidence dominate the investigation. Datadog makes sense when an error must lead directly into service traces and infrastructure telemetry. Rollbar and Honeybadger sit closer to focused error monitoring, with mature product-specific integration paths. The consolidated API option is attractive when outsourcing undifferentiated backend integrations matters more than deep error tooling; its public discovery surface also exposes request schemas and runnable examples rather than requiring guesswork.
That is the trade. Breadth behind a simple surface saves integration time, while specialist products provide richer debugging workflows. A weekly shipping cadence favors the simple surface until missing depth begins to delay incident resolution.
Scale changes the delivery path
First, I would put a bounded queue between request handling and network delivery. The queue needs a size limit and a drop policy; “never drop telemetry” is not credible if telemetry can exhaust application memory. I would also compute a local fingerprint from normalized error type plus stable stack frames so obvious duplicates can be dampened before transmission.
Second, I would add an explicit data policy. Tenant identifiers may be necessary for impact analysis, but they still need retention rules and access control. This matters especially because the consolidated option has no per-user log deletion interface and no bulk export or subscription interface. Confirm that limitation against your deletion obligations before sending user-linked data.
Alerting needs a separate decision. The basic API has no threshold, phone, SMS, or webhook notification route, so automated alerting requires polling the query surface and owning the notification logic. Silent scheduled-job failures need a heartbeat product such as Healthchecks.io because there is no synthetic or heartbeat monitor. Those gaps are manageable for a small service, but only if they are written into the design instead of discovered during an incident.
Finally, trace IDs and span IDs in logs provide loose correlation, not distributed tracing. They cannot reconstruct a span tree or answer a trace query. Once a request crosses enough services that stack plus request context no longer explains the failure, move to OpenTelemetry instrumentation and a tracing backend. That is the point where a compact error page stops earning its keep.
The decision rule is simple: use basic capture for server-side exceptions when release and tenant-cohort context isolate the problem. Pick Sentry for rich application debugging, Datadog for joined APM and infrastructure analysis, or another specialist when its workflow removes more investigation time than its integration consumes. Revisit the choice when your missing signal has a name.
Top comments (0)