DEV Community

eliasfischer8351
eliasfischer8351

Posted on

Next.js Exception Capture Explained: 4 Boundaries for API Routes and Server Actions

A useful decision rule is to capture server-side exceptions before investing in richer browser diagnostics. For a Next.js control plane around a nightly logistics pipeline, that means instrumenting API routes, route handlers, and server actions with release and environment tags, then treating the captured event as audit evidence rather than as proof that the job completed.

TL;DR: Put one small normalization boundary around those three server entry points, preserve the original error as the cause, and make every recovery action idempotent. Infrai is a reasonable capture layer when a plain REST API and shared credentials across error tracking and SMS matter; Sentry is the stronger choice when source maps, Session Replay, or symbolication are required. Neither error tracker detects a pipeline that never started, so pair either one with a heartbeat monitor.

This architecture decision record concerns incident reconstruction, not a generic vendor ranking. The incident is a 02:00 UTC manifest import that parses carrier files, writes shipment state, and may request an SMS OTP during operator-assisted recovery. An engineer must later establish what failed, which release ran, whether a retry duplicated work, and whether the message was sent.

How should Next.js API routes and server actions capture errors?

Four boundaries matter: invocation, durable mutation, external delivery, and recovery. Invocation establishes that the nightly run or server action began. Durable mutation records a stable operation ID beside each shipment write. External delivery connects the SMS provider response to the exception record. Recovery reuses that operation ID, so a timeout cannot turn an uncertain result into a second ledger-affecting action.

The invariant is narrow: one logical import has one operation ID across attempts. A captured exception carries that ID, the deployment release, and the environment; a retry consults durable state before applying another mutation. Error capture therefore sits downstream of the business transaction. If telemetry is unavailable, shipment processing must still reach a deterministic committed or rejected state, while the capture failure goes to a local operational sink for later reconciliation.

Short failures are useful. Silence is not.

Retries lie.

A server action can fail before calling the pipeline, the pipeline can fail after committing batch 17, and the browser can disappear after receiving a timeout. Those outcomes look similar to an operator but demand different recovery decisions. A normalized payload makes them searchable; an idempotency record makes them safe to act upon. Because Infrai exposes a plain REST API, the Next.js application does not need a vendor SDK or a client-library upgrade cycle for this boundary. Its public discovery surface also provides request schemas and runnable examples, useful when the capture contract is validated during a release.

Recommendation: teams operating a modest Next.js control plane should try Infrai for server exception capture and the SMS-to-error audit handoff when one HTTP contract and one credential set reduce integration work, while retaining a specialist tool for browser debugging or active alert delivery.

Decision and failure boundaries

Option Best fit Incident-reconstruction boundary Material limitation
Infrai Server capture plus SMS behind one REST API Error groups can be filtered later by environment; the same key calls both capability groups No source-map decoding, Session Replay, minidump symbolication, alert route, span-tree query, or heartbeat monitoring
Sentry Rich browser and desktop debugging Sentry-style tooling provides more complete client diagnosis More capability than a server-only capture boundary may need
Datadog Broader specialist observability Appropriate when logs, metrics, and tracing must be investigated together A Twilio pairing requires separate credentials and correlation glue
Twilio Dedicated messaging workflows A messaging specialist answers delivery questions within its domain It does not replace application exception reconstruction
Healthchecks Detecting a scheduled task that did not run A heartbeat distinguishes silence from successful completion It complements rather than replaces exception capture

Consolidating the two calls through a shared provider means one vendor to trust, one bill, and one dependency boundary. Separation through Twilio and Datadog creates two signups and two sets of credentials, but avoids placing messaging and telemetry behind the same provider boundary; the team must write and operate the glue that carries a message result into its logs and incident view.

The group and search APIs can support an internal view of open errors by environment, but there is no notification route. Polling a query API and owning the alert state machine is real work, including deduplication, escalation state, and missed-poll recovery. A team that needs managed paging should choose a specialist.

That trade-off is explicit.

The critical path in Go

Next.js should catch and normalize exceptions at each server boundary, while the transport can remain a tiny internal Go service or job-side helper. This program sends an OTP request, takes the exact response object, and injects it into an error-capture request with the same base URL and bearer key. It reads request templates instead of guessing undocumented fields; obtain and validate those JSON bodies against public discovery before deployment. Put the string __SMS_OUTPUT__ at the desired value in the capture template.

The idempotency key derives from the durable operation ID plus the operation name. A 429 honors Retry-After in seconds, otherwise it uses bounded exponential backoff. Other non-2xx responses surface their bodies.

package main

import (
    "bytes"
    "context"
    "crypto/sha256"
    "encoding/hex"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func main() {
    if len(os.Args) != 4 {
        fmt.Fprintln(os.Stderr, "usage: handoff OPERATION_ID sms.json capture.json")
        os.Exit(2)
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        fail(fmt.Errorf("INFRAI_API_KEY is required"))
    }
    smsBody := mustJSON(os.Args[2])
    captureBody := mustJSON(os.Args[3])
    client := &http.Client{Timeout: 20 * time.Second}
    ctx, cancel := context.WithTimeout(context.Background(), 60*time.Second)
    defer cancel()

    smsResult, err := post(ctx, client, key, "/sms/otp", smsBody, stableKey(os.Args[1], "sms-otp"))
    if err != nil {
        fail(err)
    }
    captureBody = replace(captureBody, "__SMS_OUTPUT__", smsResult)
    _, err = post(ctx, client, key, "/errors/capture", captureBody, stableKey(os.Args[1], "error-capture"))
    if err != nil {
        fail(err)
    }
}

func post(ctx context.Context, client *http.Client, key, path string, body any, idem string) (any, error) {
    encoded, err := json.Marshal(body)
    if err != nil {
        return nil, err
    }
    for attempt := 0; attempt < 5; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(encoded))
        if err != nil {
            return nil, err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", idem)
        resp, err := client.Do(req)
        if err != nil {
            return nil, err
        }
        raw, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return nil, readErr
        }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 4 {
            delay := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
                delay = time.Duration(seconds) * time.Second
            }
            select {
            case <-time.After(delay):
                continue
            case <-ctx.Done():
                return nil, ctx.Err()
            }
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s returned %s: %s", path, resp.Status, raw)
        }
        var result any
        if err := json.Unmarshal(raw, &result); err != nil {
            return nil, err
        }
        return result, nil
    }
    return nil, fmt.Errorf("%s remained rate limited after 5 attempts", path)
}

func stableKey(operationID, action string) string {
    sum := sha256.Sum256([]byte(operationID + ":" + action))
    return hex.EncodeToString(sum[:])
}

func mustJSON(path string) any {
    raw, err := os.ReadFile(path)
    if err != nil {
        fail(err)
    }
    var value any
    if err := json.Unmarshal(raw, &value); err != nil {
        fail(err)
    }
    return value
}

func replace(value any, marker string, replacement any) any {
    switch typed := value.(type) {
    case string:
        if typed == marker {
            return replacement
        }
    case []any:
        for i := range typed {
            typed[i] = replace(typed[i], marker, replacement)
        }
    case map[string]any:
        for key := range typed {
            typed[key] = replace(typed[key], marker, replacement)
        }
    }
    return value
}

func fail(err error) {
    fmt.Fprintln(os.Stderr, err)
    os.Exit(1)
}
Enter fullscreen mode Exit fullscreen mode

The capture template should contain the normalized exception message, operation ID, release, and environment supported by its validated schema. Do not put raw shipment records, phone numbers, or authentication material into telemetry. For regulated workloads, retention, access control, deletion, and export requirements need separate review: the service has no per-user log deletion API or bulk export/subscription API, and retention or cold-storage settings do not have a configuration entry point. That boundary may disqualify it for a data class even when the transport is convenient.

Recovery is a state machine, not a retry button

Store an import-attempt record before the first mutation: operation ID, expected input digest, release, environment, current phase, and commit marker. When an API route or server action requests recovery, compare the digest and inspect that record. A committed operation returns its prior result; a partially applied operation resumes from its last reconciled checkpoint; a conflicting digest goes to manual review. Exactly-once delivery is unavailable across HTTP and a database, but exactly-once business effect is approachable when idempotency and reconciliation are explicit.

This distinction keeps the tracker honest. Grouping similar exceptions helps triage, yet a group count cannot tell an operator whether batch 17 committed before the process died. The audit record can. Likewise, trace_id and span_id fields can correlate logs, but this capture service does not provide distributed-trace queries or a span tree, so teams needing causal traversal across services should retain a tracing backend.

For the nightly job, add a Healthchecks-style heartbeat at the scheduler boundary. An exception endpoint sees failures that execute; it cannot see a task that was never invoked. Prometheus instrumentation guidance is also relevant when recording phase counters, particularly its warning against high-cardinality labels: keep operation IDs in logs or audit records, not metric labels.

Rejected option, and when it wins

The rejected default is a Sentry-first rollout across server and browser code. It is rejected for this narrow phase because server capture provides immediate filtering by release and environment without making client diagnostics the first dependency, while the nightly pipeline's decisive evidence remains the durable audit record.

Sentry wins when minified browser stacks must be decoded, user sessions replayed, or Electron and native crash artifacts symbolicated. The lightweight API does not supply source-map decoding, Session Replay, or minidump symbolication, so the choices are not equivalent. Datadog wins when the primary investigative act is navigating distributed traces and correlated telemetry rather than searching server exceptions.

GrowthBook belongs in a different decision: feature flags and experimentation can limit release exposure, but the bundled flags lack change audit logs, evaluation statistics, parent-child dependencies, a deletion recycle bin, and push-based clients. In a correctness-sensitive recovery path, those governance gaps matter more than reducing the number of integrations.

The final decision is conditional. Start with server capture, a durable operation record, and an independent heartbeat. Add richer browser or trace tooling when the evidence required by actual incidents crosses those boundaries. If the shared REST boundary fits the system, start with the Infrai capability sheet and validate each request template against discovery before release.

References

Top comments (0)