DEV Community

AbernathyCross6857
AbernathyCross6857

Posted on

Backend Error Capture API: Lightweight Next.js Routes Without Sourcemaps or Replay

Short answer: choose the smallest exception-capture path that can reconstruct one failed financial agent run across every model call, tool call, retry, and outward notification. A lightweight capture API is enough only when it preserves that causal chain, records latency and cost inputs consistently, and fails without delaying the route. If it cannot, the operationally simpler choice is the fuller tracker, even when sourcemaps and replay are deliberately out of scope.

This decision is about incident evidence, not feature count. For a Next.js backend route driving an AI agent loop, a stack trace alone cannot explain whether a timeout happened before a payment-risk tool returned, after a model retry, or while an OTP notification was queued. Those boundaries change both the customer response and the accounting.

That distinction matters.

What must survive a failed agent run?

Treat the capture contract as an architecture invariant. Every exception event needs a stable run identifier, route and operation names, the failing step, attempt number, duration, exception class, and a cost input such as token counts. It also needs a deployment identifier and an outcome. Never attach raw prompts, OTPs, account numbers, authorization headers, or unrestricted response bodies. Compliance starts at the event boundary.

Three failure boundaries matter. The business operation may fail. The capture call may fail. The route may be cancelled while either is in flight. The first must remain visible, the second must not replace the first, and the third must not produce a falsely successful event.

Keep cardinality bounded. Prometheus instrumentation guidance explicitly warns against labels with unbounded sets and recommends labels over procedural metric-name generation. The same discipline belongs in exception metadata: route=agent_run is useful; a customer ID as a metric label is not. Store high-cardinality correlation data in the event record, subject to retention and access controls, rather than in metric dimensions.

One detail is easy to miss: notification delivery is a separate outcome. An exception that prevents an OTP from being queued should carry a notification stage, but delivery status must come from the delivery system. Conflating queued with delivered creates a clean dashboard and a bad incident timeline.

Consider a run with three ordered steps. The first model call records duration and usage, a risk tool then rejects its request, and the loop retries with a narrower request before queuing an OTP. An isolated exception from the first tool attempt does not reveal whether the retry ran, whether another model call added latency and usage, or whether the notification entered its queue. A useful record links each attempt to one run while keeping the outcomes separate. It also distinguishes recorded usage from an inferred monetary total; rates and billing rules can change outside the route. During reconstruction, the engineer should be able to prove that the financial decision preceded the notification, not infer that order from neighboring timestamps. This is the concrete trade-off: a smaller capture API reduces the integration surface, but only by moving correlation, redaction, retention, and alert policy into systems the team already owns. Choose that burden only because those systems are already dependable, not because the ingestion call looks short.

Should Next.js backend routes use a lightweight error capture API?

The options differ less by ingestion syntax than by how much reconstruction work the team accepts. This table is the decision record, not a ranking.

Option Good fit Incident boundary Main trade-off
Lightweight capture API A small backend surface with an existing log, metric, and trace correlation model The team owns normalization, deduplication, retention, access control, and alert routing Low integration surface, more operational policy in application code
Full exception tracker Several services or teams need shared grouping, triage, release context, and alert workflow The tracker owns more of the investigation path More capability to configure and govern
Structured logs plus metrics Exceptions are already reconstructed reliably from a central log pipeline Correlation and grouping remain pipeline responsibilities Fewer event paths, weaker ergonomics if stack and retry grouping are ad hoc

No sourcemaps narrows the question: server-side source context must come from deployed build metadata and captured stack frames. No replay removes a browser investigation aid that is usually irrelevant to a route-only incident. Neither constraint removes the need for redaction, grouping, rate control, or durable correlation.

A useful acceptance test is concrete: given only retained telemetry, can an on-call engineer order all attempts for one run, identify the first failing dependency, calculate observed model latency from recorded durations, and reconcile the recorded usage inputs? If the answer depends on searching by customer email or guessing which retry produced the charge, the design fails.

How does the critical path stay honest?

Capture once at the route boundary, but emit step records inside the loop. Normalize errors before transport, bound every field, and make the transport deadline shorter than the route's remaining budget. The following Python sketch shows the contract; the adapter behind emit_event can target an event endpoint, a queue, or the existing telemetry pipeline.

from dataclasses import asdict, dataclass
from time import monotonic
from typing import Callable
from uuid import uuid4


@dataclass(frozen=True)
class ExceptionEvent:
    run_id: str
    operation: str
    step: str
    attempt: int
    duration_ms: int
    error_type: str
    deployment: str
    outcome: str


def run_step(
    operation: str,
    step: str,
    attempt: int,
    deployment: str,
    action: Callable[[], object],
    emit_event: Callable[[dict], None],
    run_id: str | None = None,
) -> object:
    correlation_id = run_id or str(uuid4())
    started = monotonic()
    try:
        return action()
    except Exception as error:
        event = ExceptionEvent(
            run_id=correlation_id,
            operation=operation,
            step=step,
            attempt=attempt,
            duration_ms=round((monotonic() - started) * 1000),
            error_type=type(error).__name__,
            deployment=deployment,
            outcome="failed",
        )
        try:
            emit_event(asdict(event))
        except Exception:
            pass  # Preserve the original business exception.
        raise
Enter fullscreen mode Exit fullscreen mode

The deliberately absent fields matter as much as the present ones. There is no prompt text, bearer token, phone number, free-form exception message, or tool response. Add allowlisted diagnostic fields only after deciding who may read them and how long they remain available.

The swallowed transport error is also intentional, but silence must not become invisibility. The capture adapter should maintain a bounded counter for accepted, rejected, timed-out, and dropped events. A backend route must not turn a telemetry outage into a second customer-facing failure.

Keep the original failure.

Operate the evidence, not the dashboard

Deploy the schema before switching alert paths. During the overlap, compare counts by bounded dimensions such as operation, outcome, and exception class; do not expect byte-for-byte equality because grouping and retries can differ. Sample successful steps if volume requires it, but preserve failed runs and the steps needed to interpret them under a documented policy.

Then rehearse three cases: a model timeout followed by a successful retry, a tool failure after usage has been recorded, and a notification queue rejection after the financial decision is complete. The reconstructed timeline should show which effects happened and which did not. Sharp edges appear here.

Latency and cost are measurements, not labels. Record monotonic durations around each operation and retain the usage quantities returned by the component that performed the work. Aggregate them later. Do not put run IDs, exception messages, or account identifiers into metric labels; unbounded cardinality makes the measurement system itself harder to operate.

For browser performance, Core Web Vitals uses percentile-based field measurement, including a 75th-percentile assessment. That does not define a backend-agent service objective, but it is a useful reminder that one average hides tail behavior. Choose backend latency percentiles and thresholds from the route's own user and risk requirements, then document them.

Why reject direct logging as the default?

Plain structured logging is valid when a mature pipeline already provides durable ingestion, access controls, correlation, grouping, and alerts. In that environment, another capture channel can duplicate evidence and complicate deletion policy.

It is rejected as the default here because those capabilities cannot be assumed. A log line that says agent failed may be cheap to emit, yet expensive to investigate when retries interleave and notification effects live elsewhere. The deciding test stays narrow: preserve a redacted, ordered, deployment-aware incident record without extending the customer request's failure surface. Pick whichever capture depth passes that test with the least policy your application team must invent and maintain.

References

Top comments (0)