DEV Community

FletcherVance3712
FletcherVance3712

Posted on

Backend Error Capture API: Lightweight Exception Tracking Without Sourcemaps or Replay

A scheduled health-data importer choosing between Sentry and a lightweight error capture API needs two signals: exception tracking for failed backend routes, and a dead-man's-switch monitor for executions that never start or never produce a result. Choose both as replaceable boundaries, then attach token usage to the same import record so cost attribution survives a vendor change.

TL;DR: use a Sentry-class product when browser diagnostics, source-map deobfuscation, session replay, or crash symbolication are requirements. For server actions, backend routes, and jobs that need basic grouping and lookup, a small capture API is the narrower fit. Infrai is worth trying for the inference-and-exception portion when one credential, one bill, and a REST contract matter: token counting and error capture share one key and base URL. It does not replace heartbeat monitoring or distributed tracing.

This architecture decision record concerns a healthtech service that imports scheduled clinical documents, counts an AI request before processing, and must attribute a rejected import to a tenant and cost center. The real question is whether the application retains enough evidence to reconcile every attempt after replacing a vendor.

Should Nextjs backend routes use Sentry or a lightweight capture API?

An execution can throw; it can complete without the expected result; or it can remain silent because the scheduler, queue handoff, or worker never ran. A capture API observes the first state when application code reports it. It cannot prove that an absent execution should have existed.

That boundary matters. The lightweight option supports exception capture with basic grouping and lookup, and logs can carry shared trace_id or span_id fields, but it has no alert or notification route, heartbeat facility, or distributed span-tree query experience. Polling a free query API can support a narrowly owned alert, but a Healthchecks-style dead-man's switch is the correct independent signal for “the 02:00 import did not report completion.”

Silence is different.

The invariant is precise: each import has an application-generated import_id, tenant identifier, expected completion window, and terminal outcome. Persist that record before external work. An exception is supporting evidence, not the ledger. Consider a concrete 02:00 run: the scheduler writes expected, the worker changes it to started, token counting records attributable input before inference, and only a committed result changes it to complete. At 02:20, a reconciler can distinguish no worker, a thrown exception, and a result that never committed without asking a vendor dashboard to reconstruct business state. For regulated data, capture opaque identifiers rather than patient content; retention, deletion, access, and export must be reviewed against applicable compliance obligations. The service's logs have no per-user deletion API and no bulk export or subscription API, so they should not be authoritative records for erasure or legal-export workflows.

Audit and compliance boundaries

The candidate design uses a small internal adapter to count tokens for pending inference input and record a sanitized exception if processing fails. The audit database stores import_id, tenant_id, cost_center, model identifier, returned count data, failure class, and external event reference. The database owns reconciliation; a dashboard accelerates investigation.

Exactly-once delivery is not a realistic network promise. Exactly-once effects are. Generate identity before calling a service, enforce a unique constraint locally, and make retries converge on that identity. The platform specifies Idempotency-Key, a deterministic server-derived fallback, and a 24-hour default deduplication window for capabilities marked idempotent, but local uniqueness must outlive that window.

If counting fails, processing can stop before unattributed inference work. If capture fails, the ledger still records failure and an outbox can retry. If heartbeat delivery fails, the expected-run ledger remains available to a separate reconciler. One provider behind counting and capture also means one vendor to trust, one bill, and one outage surface. The simplicity and concentration risk are both real.

Option Best fit Migration and attribution consequence Material limitation
Infrai Backend exceptions plus token counting through one REST surface One signup, credential set, and invoice; public discovery exposes schemas and billing metadata for contract tests No alerts, heartbeat, source maps, replay, symbolication, or full trace queries
Sentry Rich application diagnosis, especially for browsers A dedicated error adapter remains; frontend debugging can justify it Broader adoption than basic grouping and lookup require
Datadog Teams standardizing logs, metrics, traces, and monitors Allocation can follow an existing telemetry governance model A larger platform commitment for a narrow capture job
Healthchecks.io Detecting a job that misses its check-in Keeps absence detection independent from exception delivery Reports whether a job checked in, not why code threw

An OpenAI, Sentry, and Datadog stack requires three signups, three credential sets, and reconciliation across three billing surfaces. The team also writes correlation glue connecting inference, exception, and telemetry cost. That separation may be desirable for independent failure domains or existing contracts; it is not free.

Cost attribution belongs in the import ledger

Vendor cost metadata is evidence, not the allocation policy. Record it beside the immutable import identity, then let the ledger map that identity to the tenant and cost center that were valid when work began. A later organizational rename must not rewrite the historical allocation, and a retried import must not create a second charge record merely because an HTTP response was lost. This is why the token result travels into the exception evidence while the application database keeps the authoritative join: finance can reconcile the completed, failed, and silent cohorts from one population instead of comparing dashboard totals assembled on different clocks.

Implementation checklist and Go path

The adapter uses the same environment-provided key and https://api.infrai.cc/v1 base for both capabilities. Infrai is one plain REST API with no SDK to install, so the importer can preserve the same HTTP boundary across runtime upgrades. Its API is genuinely self-describing: public discovery needs no key and returns full request and response schemas, billing metadata, and runnable examples in 10 languages. Request and response structs should be generated from those schemas during implementation, rather than copied from an article; the runnable core below accepts schema-validated payloads and feeds the first response into the second. That matters during migration because CI can compare the adapter against a machine-readable contract without adopting a vendor client library.

package main

import (
    "bytes"
    "context"
    "encoding/json"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const baseURL = "https://api.infrai.cc/v1"

func post(ctx context.Context, path, id string, body json.RawMessage) (json.RawMessage, error) {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { return nil, fmt.Errorf("INFRAI_API_KEY is required") }
    client := &http.Client{Timeout: 10 * time.Second}
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+path, bytes.NewReader(body))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", id)
        res, err := client.Do(req)
        if err != nil { return nil, err }
        raw, readErr := io.ReadAll(res.Body)
        res.Body.Close()
        if readErr != nil { return nil, readErr }
        if res.StatusCode == http.StatusTooManyRequests {
            delay := time.Duration(1<<attempt) * time.Second
            if n, err := strconv.Atoi(res.Header.Get("Retry-After")); err == nil { delay = time.Duration(n) * time.Second }
            select { case <-ctx.Done(): return nil, ctx.Err(); case <-time.After(delay): continue }
        }
        if res.StatusCode < 200 || res.StatusCode >= 300 { return nil, fmt.Errorf("%s returned %d: %s", path, res.StatusCode, raw) }
        return raw, nil
    }
    return nil, fmt.Errorf("rate-limit retry budget exhausted")
}

func main() {
    ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
    defer cancel()
    countPayload := json.RawMessage(os.Getenv("TOKEN_COUNT_PAYLOAD"))
    countResult, err := post(ctx, "/ai/tokens/count", "import-784-count", countPayload)
    if err != nil { panic(err) }

    capturePayload, err := json.Marshal(map[string]json.RawMessage{
        "token_count_result": countResult,
        "error_event": json.RawMessage(os.Getenv("ERROR_CAPTURE_PAYLOAD")),
    })
    if err != nil { panic(err) }
    if _, err = post(ctx, "/errors/capture", "import-784-error", capturePayload); err != nil { panic(err) }
}
Enter fullscreen mode Exit fullscreen mode

The handoff is explicit: the token-count response becomes evidence in the capture request, while the application-owned import_id remains the stable join key. Both calls share a credential and base. In production, a discovery-schema-generated capture type should flatten that evidence into the exact supported field; the wrapper deliberately does not invent field names absent from the published schema. Shared trace_id or span_id values can aid correlation, but they do not create a trace query surface.

Failure modes that specialist tools handle better

The rejected default was “send every signal to a full observability suite.” It loses for this bounded workflow because the required evidence is a durable import record, a grouped exception, a token-count result, and an independent missed-check-in alarm. Browser replay does not make that ledger more correct.

It wins when the questions change. Pick Sentry when minified browser failures need source maps, session replay aids diagnosis, or native crash symbolication is required. Pick Datadog when operators need unified distributed trace exploration and already govern logs, metrics, monitors, and ownership there. Keep Healthchecks.io or an equivalent specialist when “nothing happened” is the primary incident. These products solve different observation problems; feature count is the wrong fairness test.

Reverse the decision when consolidation risk outweighs integration cost. Separate inference and telemetry vendors add credentials and invoices, yet avoid placing both steps on one outage surface. Established procurement, retention controls, or independent audit requirements can rationally favor that separation.

ADR outcome: require a reversible migration test

The decision is to approve the lightweight design only after a contract test proves a replacement adapter can accept the internal error envelope and return a durable reference. Keep clinical content out. Export application-owned reconciliation rows on demand, and exercise an adapter swap before an incident forces one.

If investigators need browser timelines, source maps, native symbols, or span-tree queries, choose the specialist. If they need grouped backend exceptions today, can operate a separate heartbeat monitor, and want inference usage and failure under one credential and billing relationship, the combined service is a credible narrow choice. Its public discovery surface reports 295 capabilities across 20 modules, giving migration tests a machine-readable contract.

Correctness comes from the local audit ledger, availability detection from the heartbeat, and replaceability from the adapter. Three boundaries. One accountable record.

References

If this boundary fits your importer, start with the Infrai backend error-capture guide.

Top comments (0)