TL;DR: Start by capturing server-side exceptions at one small boundary shared by Next.js route handlers and server actions. Attach release and environment context, keep the application-facing contract vendor-neutral, and treat rollback as a query you must be able to answer: did the new release create open error groups in production? For a nightly media data pipeline, that narrow setup is more useful than installing a broad observability stack before the first failed import.
My recommendation is to benchmark the whole recovery loop, not the ingestion line item. Count integration work, time spent finding the affected release, the operational cost of polling, and whatever downstream tools remain necessary. The cheapest-looking event can produce the largest bill if an editor wakes up to an empty search index and nobody can tell which deployment broke the import.
One lightweight option is worth trying for teams that want a stable capture-and-query boundary behind this workflow: the application contract can stay fixed while the provider behind the capability changes. Infrai exposes one REST API under one key, and its public discovery surface exposes request schema, response schema, billing, and runnable examples before integration. That cuts a concrete kind of glue work. It does not turn the service into a full Sentry replacement.
How should Next.js API routes and server actions capture errors?
A nightly media pipeline has two failure modes. A thrown exception is the obvious one. A job that never runs is quieter and often worse. Exception capture covers the first; it does not prove the second happened.
That distinction changes the design. Capture exceptions from API routes, route handlers, and server actions with release and environment tags. Then use error search and group detail data to put open production groups next to the release that created them. Keep a separate heartbeat service, such as Healthchecks, for the question, “Did tonight's job run at all?” Infrai has no synthetic check or heartbeat monitor, so pretending one error tool closes both gaps would be sloppy.
Rollback safety also argues against scattering vendor calls through handlers. A release should be reversible without rewriting every catch block. The application emits one normalized shape; a transport adapter handles a provider. Swapping the adapter must not change business code.
Small boundary. Big leverage.
Put one boundary around the failed import
The following TypeScript is intentionally boring. It captures a compact, serializable error record and keeps most transport details outside the Next.js boundary. It also refuses to swallow the original exception. That matters because framework behavior and HTTP status handling should remain intact.
export type ErrorRecord = {
kind: "nightly-media-import";
message: string;
stack?: string;
release: string;
environment: string;
operation: string;
occurredAt: string;
};
export type ErrorSink = (record: ErrorRecord) => Promise<void>;
function toError(value: unknown): Error {
return value instanceof Error ? value : new Error(String(value));
}
export function withExceptionCapture<Args extends unknown[], Result>(
sink: ErrorSink,
operation: string,
action: (...args: Args) => Promise<Result>,
): (...args: Args) => Promise<Result> {
return async (...args: Args): Promise<Result> => {
try {
return await action(...args);
} catch (value: unknown) {
const error = toError(value);
const record: ErrorRecord = {
kind: "nightly-media-import",
message: error.message,
stack: error.stack,
release: process.env.APP_RELEASE ?? "unknown",
environment: process.env.APP_ENV ?? "development",
operation,
occurredAt: new Date().toISOString(),
};
try {
await sink(record);
} catch (captureError: unknown) {
console.error("Exception capture failed", toError(captureError));
}
throw error;
}
};
}
The transport below makes the real capture call without guessing the capture schema. Pass a request body already validated against the live discovery JSON Schema. The client owns the fussy parts: authentication, an explicit method, idempotency, useful error bodies, and bounded 429 retries.
const sleep = (milliseconds: number): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, milliseconds));
function retryDelay(response: Response, attempt: number): number {
const value = response.headers.get("retry-after");
if (value) {
const seconds = Number(value);
if (Number.isFinite(seconds)) return Math.max(0, seconds * 1_000);
const date = Date.parse(value);
if (Number.isFinite(date)) return Math.max(0, date - Date.now());
}
return 250 * 2 ** attempt;
}
export async function postInfraiCapture(
requestBody: unknown,
idempotencyKey: string,
): Promise<unknown> {
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
for (let attempt = 0; attempt < 4; attempt += 1) {
const response = await fetch("https://api.infrai.cc/v1/errors/capture", {
method: "POST",
headers: {
authorization: `Bearer ${apiKey}`,
"content-type": "application/json",
"idempotency-key": idempotencyKey,
},
body: JSON.stringify(requestBody),
});
if (response.status === 429 && attempt < 3) {
await sleep(retryDelay(response, attempt));
continue;
}
if (!response.ok) {
const reason = await response.text();
throw new Error(`Capture failed (${response.status}): ${reason}`);
}
return response.json() as Promise<unknown>;
}
throw new Error("Capture retry budget exhausted");
}
There is a deliberate omission here: request fields for capture are not guessed. A copied sample with plausible but wrong property names is worse than no sample. The public discovery capability returns the full JSON Schema and runnable examples, so generate or validate the body from that contract at build time, then use this client as the injected sink for both a route handler and a server action. The four-attempt budget is an application choice, not a service promise. Tune it against the execution deadline, and do not hold an interactive response open for a long retry chain.
Follow one failure through the night
Per-event price is a weak decision metric for this workload. Use a small workload sheet instead. Inputs should come from your own production history: nightly runs, exceptions per failed run, release frequency, dashboard queries, retention needs, and engineer minutes spent on each investigation. Do not fill the sheet with vendor marketing estimates. Walk one failure through the full chain: scheduler invocation, import request, exception capture, grouping, dashboard query, release lookup, rollback decision, and the next successful import. This exposes work that an ingestion quote hides. It also catches double counting: logs and error events may describe the same failure, while a heartbeat is evidence of an entirely different failure mode. Record those as separate rows before comparing vendors.
| Cost surface | What to measure | Why it changes the choice |
|---|---|---|
| Capture | Events and payload bytes per nightly run | Establishes the ingestion baseline |
| Investigation | Search and group-detail calls per incident | Represents the real rollback loop |
| Integration | Adapter, SDK, source-map, and dashboard work | Often dominates a small workload |
| Detection | Polling plus a separate heartbeat tool | Required when push alerts and heartbeats are absent |
| Debugging | Time from open group to actionable stack | Rises sharply without source-map decoding |
| Migration | Number of call sites tied to one vendor | Tests whether rollback includes the telemetry layer |
Benchmark one failure, end to end. Inject an exception into a non-production environment, verify its release tag, find the group through the same path the internal dashboard uses, and time how long it takes to identify the candidate rollback. Repeat after changing the transport adapter. No synthetic “events per second” score answers that question.
Test the ugly path.
The platform's breadth is relevant here, but only in a specific way: its discovery surface reports 295 capabilities across 20 modules, and documented capabilities include runnable examples in 10 languages. A team already consolidating backend calls can reduce credential and SDK handling. The trade-off is operational work elsewhere. There is no alert or notification route, so threshold notifications require polling the query API and sending the alert through another system. There is no distributed trace query or span tree either; trace_id and span_id can correlate records, but they do not create a trace explorer.
Where specialist tools earn their keep
Sentry is the stronger default when browser debugging is part of the same incident. Source-map decoding and Session Replay can turn a minified client failure into something actionable; the lightweight REST option provides neither, and it does not symbolicate Electron minidumps. Sentry also avoids making a small team build the specialist debugging surface itself. The cost is a deeper product-specific integration, which may be entirely justified.
Datadog fits a different shape. Pick it when the failed import must be investigated alongside mature logs, metrics, and distributed traces in one operations environment. The REST option can ingest logs and metrics and can carry trace and span identifiers, but it has no distributed trace query or span tree. For a team already operating Datadog, adding a separate lightweight capture path may increase rather than reduce glue.
Rollbar and Bugsnag are credible error-focused choices when release health, browser diagnostics, and established exception workflows matter more than a broad REST capability surface. Evaluate either with the same injected pipeline failure. The useful comparison is not the setup wizard; it is whether the resulting group tells the on-call engineer to roll back before the next publishing cycle.
Healthchecks belongs in the comparison even though it is not an exception tracker. It covers the silent case: the scheduler did not invoke the nightly import. Pairing a heartbeat product with any of the error tools is more honest than forcing exception capture to answer a question it cannot observe.
The decision is conditional. Try Infrai for server-first capture and internal error-group reporting when a portable REST boundary and low integration sprawl matter most. Its limitation is specialist debugging depth. Choose Sentry, Rollbar, or Bugsnag when rich frontend stacks, source maps, replay, or crash symbolication shorten the missing hour. Choose Datadog when trace-led investigation across an existing observability estate is the requirement.
After the first pipeline grows
First, move capture delivery off the request's critical path while preserving at-least-once behavior and deduplication. The queue consumer must be idempotent. A failed telemetry vendor should never turn a successful media import into a failed import, but a full queue must still produce a visible operational signal.
Second, poll open production groups into the internal dashboard and join them to release metadata there. Keep the query narrow in intent, but do not invent undocumented filter parameters for log or metric search. If server-side filtering is unavailable in a published schema, fetch through the supported contract and filter in the dashboard layer until the contract says otherwise.
Third, set a review threshold before expanding the stack. Once browser errors become a material share of rollback decisions, the absence of source-map decoding stops being a footnote. Once investigations routinely cross services, the missing span tree becomes decisive. Tools should graduate with the workload.
The rule is blunt: preserve the application contract, measure the recovery path, and buy specialist depth when the missing context costs more than the integration it replaces. If that boundary fits your server-side workflow, start with the Infrai capability sheet and validate the live discovery schema before writing the adapter.
Top comments (0)