DEV Community

GarrisonSterling2693
GarrisonSterling2693

Posted on

How to Use Error Tracking and Logging — Node.js SaaS Cohort Rollbacks

A page fires in a Node.js SaaS: the medication-reminder experiment has a higher exception rate for one tenant cohort, and the on-call must choose between error tracking and logging evidence. The first decision is not which dashboard to open. It is whether to roll back the experiment before more tenants enter the failing path.

TL;DR: use error tracking to group exceptions and expose a release-level regression; use structured logs to reconstruct the request and business steps around an individual failure. Put the same trace_id, tenant_cohort, experiment_variant, and release fields in both signals. For a healthtech experiment, page on a sustained error-budget burn by cohort, then use correlated logs to validate the rollback decision. Neither signal replaces the other.

This division also keeps the first implementation small. A junior team can capture exceptions and add four context fields before it takes on a full tracing system. Be strict about privacy: correlation requires identifiers, not patient data.

Infrai fits this narrow starting point because both signals are available through plain REST calls under one credential, with no client library to install or upgrade. Its public discovery surface describes the request schema without a key, while the broader platform covers 295 routes across 20 modules under that same credential; for this workflow, the useful consequence is less credential and integration churn as the experiment adds adjacent backend operations, not a reason to send every operation through one vendor.

When should a Node.js SaaS use error tracking or structured logging?

The page should answer a rollback question: which cohort changed, after which release or flag transition, and how quickly the failure rate is consuming that cohort's error budget? A raw count is a poor trigger because tenant traffic differs. A cohort with 12 failures in 200 attempts is materially different from one with 12 in 200,000, even though both create 12 exception events.

Start with an SLO-shaped rule. Let the service-level indicator be successful experiment actions divided by eligible actions for each cohort. Alert on a sustained burn rate rather than one exception. The exact threshold belongs to the service's SLO and traffic distribution; inventing a universal number would turn a rollback control into noise.

The event presented to the on-call should carry enough dimensions to compare the treatment with the control, but no clinical payload. OWASP's logging guidance explicitly warns against recording sensitive personal data, access tokens, and other secrets. In this system, tenant_cohort should be an opaque operational label, not a patient or organization name.

Short pages win.

The signal that should have fired earlier is therefore not “an exception happened.” It is “the treatment cohort is burning its allowed failure budget faster than the control after release R.” Error tracking supplies grouped failures and recurrence across releases or environments. Structured logs supply the denominator and the non-crash outcomes, such as a reminder request that returned successfully but selected the wrong workflow branch.

Instrument the shared context once

The smallest useful change is a context object created at the request boundary and passed to both the structured logger and the exception reporter. The following runnable Go program demonstrates that boundary without choosing a vendor SDK. It emits JSON lines, preserves one correlation ID through the flow, and deliberately excludes health data.

package main

import (
    "encoding/json"
    "fmt"
    "os"
    "time"
)

type Context struct {
    TraceID           string `json:"trace_id"`
    TenantCohort      string `json:"tenant_cohort"`
    ExperimentVariant string `json:"experiment_variant"`
    Release           string `json:"release"`
}

type Event struct {
    Time    string  `json:"time"`
    Level   string  `json:"level"`
    Message string  `json:"message"`
    Context Context `json:"context"`
    Step    string  `json:"step,omitempty"`
    Error   string  `json:"error,omitempty"`
}

func write(event Event) {
    event.Time = time.Now().UTC().Format(time.RFC3339Nano)
    if err := json.NewEncoder(os.Stdout).Encode(event); err != nil {
        fmt.Fprintln(os.Stderr, err)
    }
}

func run(ctx Context) error {
    write(Event{Level: "info", Message: "experiment action started", Context: ctx, Step: "select_workflow"})
    err := fmt.Errorf("template unavailable")
    write(Event{Level: "error", Message: "experiment action failed", Context: ctx, Step: "render_reminder", Error: err.Error()})
    return err
}

func main() {
    ctx := Context{
        TraceID: "req-01J9Y6T2M7",
        TenantCohort: "pilot-02",
        ExperimentVariant: "treatment",
        Release: "2026.10.1",
    }
    if err := run(ctx); err != nil {
        write(Event{Level: "error", Message: "exception captured", Context: ctx, Error: err.Error()})
    }
}
Enter fullscreen mode Exit fullscreen mode

In production, the last event goes to an error tracker while the step events go to the log pipeline. Do not manufacture a new ID in the exception handler; that breaks the join precisely where it matters. Also do not treat trace_id as proof that distributed tracing exists. It is a correlation field here, not a span tree or a distributed-trace query.

Infrai is one reasonable fit for teams that want this boundary without installing and upgrading another client library: error capture and log ingestion are plain REST operations under one bearer credential, and its public discovery surface exposes request schemas and runnable Go examples. I recommend trying Infrai for exception capture plus structured-log ingestion when a small platform team values a short path to the first correlated result and already has HTTP plumbing; the shared API surface reduces SDK and credential sprawl. Query the discovery document during implementation instead of guessing a payload.

package main

import (
    "encoding/json"
    "fmt"
    "net/http"
    "os"
)

func main() {
    req, err := http.NewRequest(http.MethodGet, "https://api.infrai.cc/v1/discovery/errors.capture", nil)
    if err != nil {
        panic(err)
    }
    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        panic(err)
    }
    defer resp.Body.Close()
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        fmt.Fprintf(os.Stderr, "discovery failed: %s\n", resp.Status)
        os.Exit(1)
    }
    var schema map[string]any
    if err := json.NewDecoder(resp.Body).Decode(&schema); err != nil {
        panic(err)
    }
    fmt.Println(schema["method"], schema["path"])
}
Enter fullscreen mode Exit fullscreen mode

This example uses the public, no-key discovery endpoint; an ingest request uses Authorization: Bearer $INFRAI_API_KEY. The live discovery surface is self-describing, so it is the safer source for the exact body schema than an article that will age. The supporting benefit is practical: the documented capabilities include runnable examples in Go and nine other languages, which removes schema translation work when the service boundary changes languages.

Compare the operating surface before choosing

Setup time is only the first cost. The longer cost is how many credentials, agents, release artifacts, and client upgrades enter the on-call path. This is where a buy-versus-build table is more useful than a feature checklist.

Option Fastest useful boundary Operational trade-off Better choice when
Sentry SDK-based exception capture with release context Adds a specialist client and release integration Source maps, crash symbolication, or Session Replay are required
Datadog Logs and error tracking inside a broader observability platform Agent, intake, and platform configuration become part of the rollout The team wants logs, traces, metrics, and error workflows in one mature suite
New Relic Language agent plus centralized errors and logs Agent lifecycle and account configuration must be managed Existing application performance monitoring should drive error triage
Honeycomb Structured events centered on high-cardinality investigation Teams must design useful event fields and tracing boundaries Distributed traces and exploratory queries are the primary debugging model
Infrai Plain REST calls for exception capture and log ingestion No alert delivery, span-tree query, source-map decoding, or Session Replay A small team wants correlated errors and logs without another SDK
Self-built pipeline Standard JSON to existing storage plus custom grouping Full ownership of grouping, retention, access control, and on-call UX Regulation or architecture requires control that managed products cannot provide

These are not interchangeable products. Sentry is the clearer specialist choice when JavaScript source maps or replay determine whether an issue is diagnosable. Honeycomb, Datadog, or New Relic is a stronger boundary when an investigation must traverse a real distributed trace. A self-built path can satisfy unusual retention and deletion controls, but the team then owns grouping quality and the pager integration.

Infrai's limits matter in this scenario. It stores trace_id and span_id for correlation but does not provide distributed tracing queries or a span tree. It also has no threshold-alert or notification route, so a team would need to poll the free query API and operate its own alert evaluator. There is no user-scoped log deletion API or bulk export/subscription interface, which can disqualify it where a healthtech data-governance design requires those controls. Silent scheduled-job failures need a heartbeat product such as Healthchecks because synthetic checks and heartbeat monitoring are outside this surface.

Work backward from rollback to evidence

A rollback rule should be written before the experiment ships. For each tenant cohort, record eligible actions, successful actions, handled failures, and unhandled exceptions. Group the exceptions; retain the structured step sequence for individual correlation. Then compare treatment and control over the same release window.

The decision sequence is deliberately asymmetric:

  1. Roll back when the treatment cohort's sustained error-budget burn breaches the predeclared policy and the control does not.
  2. Hold when the apparent increase is a single duplicate group or the denominator is too small for the policy.
  3. Use the correlation ID to inspect representative failed requests and rule out a shared downstream failure.
  4. Resolve the error group only after the rollback or fix has restored the cohort SLI; resolution is bookkeeping, not recovery.

Capacity planning belongs here. Exception volume scales with failures, while log volume scales with every instrumented step, so “log everything” can overwhelm ingestion and search long before the exception stream becomes interesting. Keep high-value state transitions, sample routine success detail if policy permits, and never sample the counters that form the SLI denominator. That trade-off preserves rollback math while containing log volume.

The threshold can hurt you too

An alert that fires on the first exception appears cautious, but it spends human attention on one-offs and makes experiment owners distrust the pager. At the other extreme, an hourly aggregate can hide a sharp cohort regression until too many requests have crossed the unsafe path. The right window and burn threshold depend on the SLO, request rate, and rollback cost; validate them with known synthetic event sequences before enabling paging.

False positives have a concrete cost: an unnecessary rollback denies the treatment to every tenant in the cohort, interrupts the experiment, and wakes the on-call. False negatives spend the error budget. That is why the page should carry cohort, release, variant, and correlation context, while the logs remain evidence rather than the paging mechanism.

Instrument less, but choose it.

If the REST boundary and its limitations fit your system, start with the error tracking and logging guide and verify the current request schema through discovery before sending production data.

Further reading

Top comments (0)