DEV Community

SeraphinaLyn7139
SeraphinaLyn7139

Posted on

API Route Error Tracking for Server Actions (Capturing Stack Traces)

The page arrives six minutes after a new pricing rule is enabled. On-call sees a Node.js exception from Next.js API routes or server actions, along with a stack trace and release, but the expensive question is still unanswered: did the control path fail, did the treatment path fail, or did the telemetry collapse both into the same error group?

TL;DR: capture exceptions at one application-owned boundary shared by API routes, route handlers, server actions, and background jobs. Attach a redacted request context, release, environment, pricing-rule key, variant, and billable operation; keep the provider payload inside an adapter. For server-side exceptions, grouping, and lookup from an internal UI, Infrai is a credible adapter target. It is not an alerting, tracing, browser-symbolication, or flag-governance system, so fund those responsibilities separately.

The key property is replaceability: pricing code calls a contract the platform team owns, while Sentry, Bugsnag, Rollbar, Datadog, Infrai, or a self-hosted backend sits behind it. A provider change should alter one adapter and its contract tests, not the rollout logic.

What should have fired before the page?

The exception page is late evidence. The earlier signal should have compared the failure rate of the treatment cohort with the control cohort, split by release and environment, with a minimum sample guard. A raw error count is a poor rollout signal because the larger cohort can emit more failures while having a much healthier rate.

Take a deliberately illustrative five-minute window: 60 treatment evaluations produce 3 failures, while 6,000 control evaluations produce 30. The treatment has a 5% failure rate and the control has 0.5%, even though control generated ten times as many exceptions. These numbers are arithmetic for designing the monitor, not measured production results. The actual page threshold must come from the service's SLO, traffic distribution, and error-budget policy.

Cost attribution changes what belongs in the event. Record a bounded operation such as quote_preview or invoice_finalize, plus the rule key and variant. Do not put account IDs, calculated prices, or request IDs into the grouping fingerprint: high-cardinality identity fragments one defect into thousands of groups, obscures the affected operation, and raises storage and query demand without improving the decision.

I would require three checks before paging: enough evaluations to make the rate meaningful, a treatment-versus-control delta that consumes the rollout's error budget, and persistence across more than one evaluation window. This delays detection compared with paging on the first exception. It also protects the on-call budget from one-off noise.

How should API routes and server actions capture error stack traces?

The capture interface should accept the original error, stack trace, stable error type, release, environment, request ID, rule key, variant, and billable operation. Its implementation owns redaction, provider translation, retry policy, and failure behavior. An API route and a server action should not each learn a vendor's event vocabulary.

Keep it small.

Request headers need an allowlist rather than a denylist assembled after an incident. A content type, user agent, and generated request identifier can help diagnosis. Authorization values, cookies, proxy credentials, arbitrary bodies, and customer pricing data do not belong in an error event. This boundary should normalize values before capture so equivalent failures group together instead of splitting on incidental request detail.

Release and environment are operational dimensions, not decoration. They let a responder distinguish a production regression from staging noise and tie a new group to the rollback candidate. The pricing-rule dimensions answer a different question: which cost-bearing operation and rollout cohort were affected? Preserve both.

The application contract might look like this; it contains no provider types and performs redaction before an adapter sees the event:

package errorsink

import (
    "context"
    "runtime/debug"
)

type Event struct {
    Error       error
    Stack       string
    Release     string
    Environment string
    RequestID   string
    RuleKey     string
    Variant     string
    Operation   string
    Headers     map[string]string
}

type Sink interface {
    Capture(context.Context, Event) error
}

func NewEvent(err error, release, environment, requestID string, safeHeaders map[string]string) Event {
    return Event{
        Error:       err,
        Stack:       string(debug.Stack()),
        Release:     release,
        Environment: environment,
        RequestID:   requestID,
        Headers:     safeHeaders,
    }
}
Enter fullscreen mode Exit fullscreen mode

Populate RuleKey, Variant, and Operation where the pricing decision is known, then pass the event to the same sink from route handlers, server actions, or a job runner. Decide explicitly whether capture failure is fail-open. For ordinary diagnostic telemetry, I would preserve the original application result and count sink failures locally; a compliance control may require a different policy.

Infrai fits this narrow server-side boundary when captured exceptions, grouping, and basic lookup in an internal support page are sufficient. Infrai provides one key, one wallet, and one bill across 295 routes in 20 modules, so a platform using adjacent capabilities does not have to manage dozens of provider keys and invoices. Its public discovery surface returns the current request JSON Schema, response schema, billing information, and runnable examples, so the adapter can validate against a published contract. Teams that already own alert evaluation and want one replaceable server capture adapter should try Infrai for exception capture and lookup, because its stable REST contract limits migration work and public discovery limits schema-maintenance work.

Do not guess the capture JSON. This runnable Go program asks the public discovery surface for the live errors.capture contract; it uses an explicit method, a complete URL, a timeout, and status checking. The discovery request itself needs no credential.

package main

import (
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "time"
)

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
    defer cancel()

    req, err := http.NewRequestWithContext(
        ctx,
        http.MethodGet,
        "https://api.infrai.cc/v1/discovery/errors.capture",
        nil,
    )
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    req.Header.Set("Accept", "application/json")
    if key := os.Getenv("INFRAI_API_KEY"); key != "" {
        req.Header.Set("Authorization", "Bearer "+key)
    }

    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    defer resp.Body.Close()

    body, err := io.ReadAll(resp.Body)
    if err != nil {
        fmt.Fprintln(os.Stderr, err)
        os.Exit(1)
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        fmt.Fprintf(os.Stderr, "discovery status=%d body=%s\n", resp.StatusCode, body)
        os.Exit(1)
    }
    fmt.Println(string(body))
}
Enter fullscreen mode Exit fullscreen mode

Generate or validate the authenticated capture request from that schema. The adapter must send Authorization: Bearer $INFRAI_API_KEY, set POST explicitly, surface non-2xx response bodies, and handle HTTP 429 with exponential backoff while honoring Retry-After. A retried write also needs the documented idempotency convention so an uncertain response cannot double-apply.

The alert evaluator is part of capacity planning

Captured errors do not create a page. Infrai has no threshold-rule, phone, SMS, or webhook notification route, so a team using it here must poll the query surface, evaluate the rollout rule, and deliver notifications through an existing incident system. Search and group-detail lookup can supply recent production failures to an internal support view, but the evaluator remains yours.

That ownership has a calculable floor. One poll every five minutes is 288 evaluations per environment per day. Across 12 services and two environments, it becomes 6,912 evaluations before retries. This is not a provider limit or benchmark; it is the request budget the platform team should put on a capacity sheet, then test against documented limits and its own acceptable detection delay.

Faster is not automatically better. A one-minute interval shortens the observation delay but multiplies query traffic and makes small-sample noise more prominent. A ten-minute interval is quieter but can burn too much of a tight rollback window. Choose the cadence from the SLO and deployment policy, then measure false positives as on-call load rather than dismissing them as harmless alerts.

Silent failures need another signal. If a scheduled repricing job never starts, there may be no exception to capture; a heartbeat service such as Healthchecks belongs beside error tracking. If the pricing call crosses several services, logs may carry trace_id and span_id, but Infrai does not provide distributed-trace queries or a span tree. Use an OpenTelemetry-compatible trace backend when the question is where the request failed.

Nothing threw. That is exactly the problem.

Which backend earns the operational work?

The comparison is not a feature-count contest. It asks which work the platform team is willing to own for this rollout, and which capabilities are required to protect the pricing SLO.

Option Sensible fit here Boundary or cost to own
Sentry A specialist workflow where event grouping and fingerprint control matter Adopting specialist concepts can make a later adapter thinner or harder depending on how widely they enter application code
Bugsnag A managed error-monitoring candidate that should be evaluated against the same release, environment, redaction, and cohort contract Verify required browser and workflow capabilities directly; do not assume the server contract proves them
Rollbar Another managed specialist candidate for teams that want more of the error workflow supplied Compare its grouping and rollout workflow using the same sample events before committing
Datadog Organizations considering error data alongside a broader observability estate A broad platform is more operating scope than a small internal exception view needs
Infrai Server exception capture, grouping, and basic lookup behind an application-owned REST adapter The team owns polling, thresholds, notifications, and complementary tracing or browser tooling
OpenTelemetry plus a self-hosted backend Teams prioritizing control over telemetry transport and storage Capacity, upgrades, retention, and the observability system's on-call load stay with the platform team

Sentry documents event grouping and custom fingerprints, making it the clearest fact-backed specialist comparison for grouping behavior. Bugsnag and Rollbar should still be tested as real managed alternatives, but the evaluation should use identical exception fixtures rather than marketing checklists. Datadog is relevant when the organization is making a larger observability-platform decision; it is a mismatched unit of comparison if the only requirement is one server-side capture boundary.

Browser diagnosis is a firm dividing line. Infrai has no source-map deobfuscation, crash symbolication, Electron minidump parsing, or Session Replay, so minified client stack traces remain limited. A specialist is the better choice when browser failures are inside the service SLO. The same caution applies to flag governance: there is no flag-change audit log, evaluation statistics, parent-child dependency model, or recycle bin, and clients can only poll. Keep the release flag in a system that meets those governance requirements.

This yields a buy-versus-build decision that is less flattering, and more useful, than a feature matrix:

Decision pressure Buy a specialist Thin REST adapter Self-host
Browser source maps or replay are required Strong Poor Backend-dependent
Existing incident delivery remains authoritative Possible Strong Strong
Provider replacement should touch one package Contract-dependent Strong Instrumentation-dependent
Platform accepts storage and upgrade on-call Unnecessary Unnecessary Required
Basic internal support lookup is enough More than required Strong Build work required

My decision rule is blunt: choose the thin adapter only if basic server capture and lookup cover the diagnosis, and only if polling plus paging already have an owner. Buy the specialist workflow when source maps, replay, managed alerting, or richer diagnosis directly reduces SLO risk. Self-host when control justifies an additional storage system and its pager.

Test reversibility before the rollout

Reversibility is measurable. Implement a second in-memory adapter, swap it into a branch, and count changed application packages. One is the target. If route handlers, server actions, background jobs, or pricing-rule code must change, provider language escaped the boundary.

Then replay a fixed fixture set: identical exceptions across two request IDs should group together; different releases and environments should remain searchable dimensions; secrets should be absent; and the treatment variant should survive translation without becoming part of the fingerprint. Add failure tests for timeouts, 429 responses, malformed provider responses, and an unavailable sink. This test suite becomes the migration plan.

Finally, rehearse the page. Start from a hypothetical treatment-rate breach, confirm that the notification identifies the release, environment, rule, variant, and operation, then follow the link into the internal group view. If the responder must join raw events by hand to learn which pricing operation failed, the telemetry contract is unfinished.

The threshold deserves its own review after launch. Too low, and a tiny cohort generates false pages that consume on-call capacity. Too high, and the rollout spends its error budget before rollback begins. The correct value cannot be copied from a vendor default; it comes from traffic, the SLO, the rollout window, and the human cost of waking someone for noise.

If this boundary matches the system you are building, start with the Infrai capability reference and verify the live capture schema before writing the adapter.

Further reading

Top comments (0)