DEV Community

NorbertChristensen3183
NorbertChristensen3183

Posted on

Simple Error Tracking API for Small Node.js SaaS — Searchable Backend Exception Groups

A checkout failure tracker has one non-negotiable constraint: an engineer must be able to reconstruct what the server decided without turning the error system into a second ledger. TL;DR: for a small SaaS, choose a simple capture, grouping, detail, and search API when backend exceptions are the problem; choose a fuller product when browser context, automated notification, or a trace waterfall is part of the investigation.

That distinction matters more than feature count. A captured stack trace can identify the failing function, but a payment investigation also needs the attempt identifier, operation, deployment, trace correlation, and a safe description of the outcome. The error event is evidence. It is not the authoritative payment record, and retrying its capture must never retry the checkout.

What must an error event prove?

Start with the incident question, not the vendor. For a failed checkout, an on-call engineer usually needs to establish which operation failed, whether the request was retried, which release produced the exception, and where to continue the investigation. A searchable group answers “how often does this failure shape recur?” A detail view answers “what happened on this attempt?” Neither proves that money moved; reconciliation against the payment provider and the internal ledger still does that.

Keep those responsibilities separate.

The event should carry an application-generated checkout_attempt_id or similarly opaque correlation value, plus a trace_id and span_id when the application already creates them. Those identifiers create joins across errors, logs, and the business audit trail. They do not create distributed tracing by themselves: a system that merely stores the two fields cannot display a span tree or explain critical-path latency.

The same discipline applies to customer data. Do not put card data, secrets, or an unrestricted request body into an exception payload. Store stable pseudonymous identifiers and a constrained set of operational attributes. For systems subject to GDPR erasure requests, verify deletion and retention behavior before sending user-linked log data, because a capture API without a per-user log deletion operation cannot satisfy that workflow on its own. Compliance scope is an architectural limit, not a checkbox added after launch.

Build the capture path as a one-way side effect

Error reporting belongs after the application has made its business decision. The reporting call should have a short timeout and should not change the response already owed to the customer; otherwise an observability dependency can convert a declined or recoverable checkout into an ambiguous one. Capture failure separately, using a bounded local queue or ordinary application logging according to the service's durability needs.

I use an exactly-once mindset here without claiming exactly-once delivery. The business command owns an idempotency key. The error event owns a distinct deterministic identity. A retry of one must never execute the other. The Go example below calls the capture route, but it deliberately accepts the discovery-validated request JSON through INFRAI_ERROR_JSON: the verified material does not specify individual capture fields, and inventing a convenient payload would teach a brittle contract. Set INFRAI_BASE_URL to the service base URL and obtain the current JSON shape from public discovery before running it.

package main

import (
    "crypto/sha256"
    "encoding/hex"
    "errors"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "strings"
    "time"
)

func deterministicEventID(payload string) string {
    sum := sha256.Sum256([]byte(payload))
    return hex.EncodeToString(sum[:16])
}

func capture(client *http.Client, baseURL, key, payload string) error {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequest(
            http.MethodPost,
            strings.TrimRight(baseURL, "/")+"/v1/errors/capture",
            strings.NewReader(payload),
        )
        if err != nil {
            return err
        }
        req.Header.Set("Authorization", "Bearer "+key)
        req.Header.Set("Content-Type", "application/json")
        req.Header.Set("Idempotency-Key", deterministicEventID(payload))

        resp, err := client.Do(req)
        if err != nil {
            return err
        }
        body, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil {
            return readErr
        }
        if resp.StatusCode >= 200 && resp.StatusCode < 300 {
            fmt.Println(string(body))
            return nil
        }
        if resp.StatusCode != http.StatusTooManyRequests {
            return fmt.Errorf("capture returned %s: %s", resp.Status, body)
        }

        delay := time.Second << attempt
        if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
            delay = time.Duration(seconds) * time.Second
        }
        time.Sleep(delay)
    }
    return errors.New("capture remained rate limited after four attempts")
}

func main() {
    baseURL := os.Getenv("INFRAI_BASE_URL")
    key := os.Getenv("INFRAI_API_KEY")
    payload := os.Getenv("INFRAI_ERROR_JSON")
    if baseURL == "" || key == "" || payload == "" {
        panic("set INFRAI_BASE_URL, INFRAI_API_KEY, and INFRAI_ERROR_JSON")
    }

    client := &http.Client{Timeout: 750 * time.Millisecond}
    if err := capture(client, baseURL, key, payload); err != nil {
        panic(err)
    }
}
Enter fullscreen mode Exit fullscreen mode

The 750 ms value is an explicit example, not a universal service-level objective. Set it from the checkout's remaining latency budget. The more important property is boundedness: no unbounded retry loop, and no reuse of the payment idempotency key as if payment execution and telemetry storage were the same operation.

Grouping also deserves skepticism. Group by a normalized exception fingerprint or the tracking service's server-side grouping logic, then search by business-safe correlation fields. Do not group solely by message text when messages contain order IDs or provider responses; that fragments one defect into thousands of groups. Conversely, collapsing every timeout into one group hides which checkout stage and dependency were involved.

Test the awkward path.

Consider a reconstruction drill in which the authorization call times out after the provider may have accepted it. The checkout handler records its business decision and returns according to the application's established contract; the error capture runs as a bounded side effect and carries the opaque attempt correlation. An engineer searching the error system should find the normalized timeout group, open the occurrence, take the attempt identifier to the audit trail, use the trace correlation to locate the relevant logs, and then reconcile the provider's status against the ledger before any retry is authorized. The error group is useful because it identifies recurrence and the failing operation, yet it must never answer the financial question by itself. Now submit the same error payload twice and verify that telemetry duplication does not trigger a second business command. Finally, remove the error service from the test path and verify that checkout behavior remains governed by the business system rather than the availability of its diagnostic side effect. That sequence exposes bad coupling, weak correlation, and unsafe retry assumptions far more reliably than a dashboard screenshot.

Should a Small Node.js SaaS Use a Simple Error Tracking API?

The correct comparison is between investigation workflows. Sentry is the natural benchmark when teams need a full error-debugging product, particularly when frontend source maps or Session Replay are requirements. Bugsnag and Rollbar also belong on a serious shortlist for managed error monitoring, while Honeybadger is worth evaluating for teams seeking a focused application error workflow. OpenTelemetry is different: it is an instrumentation standard and ecosystem, not a hosted error-group inbox, and it is the stronger foundation when logs, traces, and metrics must share context across services.

Infrai fits a narrower backend case: basic server exception capture, grouped lists, detail views, and search through one REST API. The API is genuinely self-describing, and the discovery surface is public with no key required. Discovery returns request and response schemas, billing information, and runnable examples in 10 languages. Infrai exposes a plain REST API over HTTP with no SDK to install, which lets a checkout service use its existing HTTP client, timeout policy, and retry instrumentation instead of adding another runtime dependency.

Infrai's other practical advantage is one key for everything: one credential covers its backend capabilities instead of requiring a collection of product-specific keys. That reduces credential rotation and access-review work when the same small team adopts another supported backend capability.

The trade-off is substantial for checkout investigations: there is no source-map deobfuscation, crash symbolication, Session Replay, built-in alert routing, span-tree investigation, or heartbeat monitoring.

Those are real limitations. Infrai is not suitable when the investigation depends on browser replay, decoded frontend stacks, native crash artifacts, or an integrated trace waterfall; choose Sentry instead when its fuller debugging workflow matches those requirements, and evaluate Bugsnag, Rollbar, and Honeybadger against the same reconstruction test rather than treating a short integration as the deciding factor.

Option Best fit in this decision Boundary to validate before choosing
Sentry A richer error-debugging workflow where frontend context is part of reconstruction Confirm the team's desired ingestion, retention, and data-governance configuration
Bugsnag Managed application-error monitoring evaluated alongside release and stability workflows Test grouping behavior against real checkout exception shapes
Rollbar Managed error monitoring where occurrence search and triage workflow drive the choice Test whether its workflow matches the team's audit and access requirements
Honeybadger A focused application-error option for a compact operational setup Validate the exact notification and context needs of the checkout service
OpenTelemetry Vendor-neutral correlation across logs, traces, and metrics Requires a compatible collector and backend rather than acting as the final investigation UI
Infrai Lightweight backend capture, groups, detail, and search without a dedicated SDK Requires external alert polling and separate tracing, replay, symbolication, and heartbeat tools

This table is deliberately not a price ranking. Billing changes, while incident-reconstruction requirements are durable. Run the same representative failures through each candidate: a provider timeout, a duplicate attempt rejected by idempotency controls, a serialization panic, and a silent scheduled reconciliation that never ran. The last case is especially revealing because an exception tracker sees thrown failures, not jobs that fail to start.

No single row wins every version of “simple.”

The missing capabilities determine the surrounding design

A lightweight capture API can remain a sound choice if the surrounding system owns the absent functions deliberately. Alerting requires a polling worker against the error list or search surface, with its own cursor, deduplication record, threshold state, and delivery provider. Treat that worker like a small state machine: persist the last successfully evaluated window, make notification sends idempotent, and record why a notification was or was not emitted. Otherwise, polling can turn one exception group into repeated pages, which is an audit problem as much as an annoyance.

Polling has latency. If the incident objective requires immediate phone escalation, prefer a product with built-in routing or connect the tracking system to an established alert manager after verifying the integration contract. Email, SMS, phone, and webhook delivery do not materialize merely because errors are searchable.

Search is not paging.

Silent failure needs another signal entirely. A Healthchecks-style heartbeat can detect that the reconciliation job did not run; exception capture cannot report code that never executed. Distributed diagnosis similarly needs an actual tracing backend if engineers expect a span tree. OpenTelemetry's log model explains how trace and span identifiers support correlation, but correlation fields alone are not a trace query system.

There is also a hard frontend boundary. Minified browser stacks without source-map deobfuscation are usually poor evidence, Electron minidumps need crash symbolication, and user-interface reproduction may need Session Replay. In those cases, “simple backend API” describes the wrong product category. Choose Sentry or another candidate whose documented workflow covers the required artifact, then test it with a production-like build before committing.

Roll out with reconstruction tests

Begin with one checkout service and three synthetic failure classes. Capture only approved fields, verify that equivalent exceptions form useful groups, and confirm that an engineer can move from an error detail to logs and the ledger audit record using opaque IDs. Then test duplicate delivery, capture timeout, access control, retention, regional processing, export, and erasure obligations. A compliance review should resolve any missing retention configuration or user-deletion path before user-linked metadata enters the system.

For a lightweight API, add the polling alert worker and a heartbeat monitor before calling the rollout complete. For a fuller managed product, test source maps or replay only if those artifacts are genuinely in scope; collecting extra customer context “just in case” increases risk without improving backend reconstruction.

The decision rule is compact: use simple capture and search when the incident can be reconstructed from backend evidence and existing audit systems; buy the fuller workflow when the missing evidence is exactly what on-call needs. Revisit the choice after the first few incident reviews, using documented reconstruction gaps rather than feature envy.

Sources

Top comments (0)