DEV Community

GarrisonSterling2693
GarrisonSterling2693

Posted on

Node.js SaaS Error Tracking API: Backend Heartbeats, Optional React Source Maps

For a Node.js SaaS app, use an error tracking API for diagnosis, but detect a scheduled healthtech import that produces no result with a durable completion heartbeat for every expected backend run. The deciding constraint is rarely whether the tracker can capture a stack trace. It is whether the on-call engineer can distinguish an intentionally empty import from a scheduler, queue, credential, or downstream failure, while the platform owner can assign the resulting telemetry cost to the tenant and import class that caused it.

TL;DR: keep exception events for diagnosis, but make a small, backend-owned run ledger the source of truth for silence. Record one terminal outcome per scheduled run, preserve the expected schedule separately, and evaluate lateness after a grace window. Browser source maps can remain optional because they do not establish whether a backend import completed.

Why can't a Node.js SaaS app error tracking API catch silence?

An exception tracker answers a valuable but narrower question: which captured execution failed, and where? A missing-result alert asks whether an execution that should have happened ever reached a terminal state. If the scheduler never dispatched work, the worker process died before initializing its SDK, or a queue message never arrived, there may be no exception event to capture.

Silence wins.

The opposite case matters too. A correctly completed import can produce zero records because the source had no new data. Alerting on records_written == 0 turns a valid business outcome into noise and trains responders to ignore it. The terminal state therefore needs at least succeeded, failed, and an explicit count; absence of a terminal state is inferred by the monitor, not reported by the worker.

Define the service-level indicator before selecting the collection system:

  • Eligible runs are scheduled runs not administratively disabled before their due time.
  • A good run reaches a terminal state before scheduled_at + grace_period.
  • An empty successful run is good.
  • A late success can restore current health, but it still consumes the reliability budget for the evaluation window.

This prevents a serious accounting error: measuring only jobs that started. If the denominator excludes runs lost before dispatch, the dashboard can show perfect reliability while customers receive stale data. The expected schedule must be independent of worker execution.

Model the heartbeat before choosing its destination

Use a stable run identity derived from the tenant, import definition, and scheduled time. Retries must update or append against that logical run rather than inventing another expected run, or a retry storm will inflate apparent workload and telemetry spend. Keep patient data and source payloads out of this record.

A compact event carries tenant_id, import_id, scheduled_at, finished_at, status, records_read, records_written, attempt, and telemetry_class. The tenant identifier should be opaque. The telemetry class should be a low-cardinality workload label such as scheduled_import, useful for allocation without putting unbounded source names into metrics.

This Go example sends a terminal heartbeat through a generic interface. The receiver may write a relational ledger, emit an OpenTelemetry log record, or do both, but the application contract stays small.

package heartbeat

import (
    "context"
    "errors"
    "time"
)

type Outcome string

const (
    Succeeded Outcome = "succeeded"
    Failed    Outcome = "failed"
)

type Event struct {
    RunID          string    `json:"run_id"`
    TenantID       string    `json:"tenant_id"`
    ImportID       string    `json:"import_id"`
    ScheduledAt    time.Time `json:"scheduled_at"`
    FinishedAt     time.Time `json:"finished_at"`
    Status         Outcome   `json:"status"`
    RecordsRead    int64     `json:"records_read"`
    RecordsWritten int64     `json:"records_written"`
    Attempt        int       `json:"attempt"`
    TelemetryClass string    `json:"telemetry_class"`
}

type Sink interface {
    Record(ctx context.Context, event Event) error
}

func Finish(ctx context.Context, sink Sink, event Event) error {
    if event.RunID == "" || event.TenantID == "" || event.ImportID == "" {
        return errors.New("heartbeat identity is incomplete")
    }
    if event.FinishedAt.Before(event.ScheduledAt) {
        return errors.New("finish time precedes schedule time")
    }
    if event.Status != Succeeded && event.Status != Failed {
        return errors.New("terminal status is invalid")
    }
    return sink.Record(ctx, event)
}
Enter fullscreen mode Exit fullscreen mode

The hard engineering lives around delivery semantics. A database transaction that marks the import complete and inserts an outbox row gives the publisher a durable handoff; a relay can deliver telemetry with retries. If completion state and heartbeat are sent independently over the network, either can succeed alone. That split creates false pages or invisible failures during degraded conditions.

OpenTelemetry's logs data model provides a standard representation for timestamped event records and permits trace context to correlate a log with other signals. Correlation helps diagnosis, but it does not replace the run ledger or its expectation model. A trace sampled away cannot be the only evidence that a contractual run completed.

Attribute cost without creating cardinality debt

Cost attribution needs two layers. The durable ledger retains per-tenant facts because a responder must answer which tenant is late. Aggregated metrics should usually avoid a tenant label when the tenant population is large; publish totals by bounded dimensions such as environment, import class, outcome, and region, then query the ledger for affected tenants after the monitor detects a breach.

That distinction is operational. Per-tenant metric series multiply across status, region, importer, and deployment labels. Ledger rows grow with executions instead, can follow an explicit retention policy, and support allocation reports without forcing the paging path to scan raw exception payloads.

Use a showback equation before debating storage vendors:

attributed telemetry units = terminal events + retry deliveries + diagnostic events + retained bytes

The units are provider-neutral. Measure them by tenant and workload class, then apply the internal rate card outside the collection path. Do not put a changeable currency estimate in every event. Separate mandatory reliability evidence from verbose diagnostics: sampling debug logs may be acceptable, while sampling terminal heartbeats corrupts the SLI denominator.

A capacity plan starts with scheduled executions, not monthly active users. If 8,000 imports are due every hour and each produces one terminal record, the baseline is 192,000 terminal records per day before retries. This is an example calculation, not a benchmark. Recompute it from the real schedule distribution, expected retry rate, record size, retention, index overhead, and regional replication. Inspect the peak minute too: hourly schedules aligned at :00 create a burst hidden by a daily average.

Decision Managed ingestion Self-hosted ingestion Required control
Cost ownership Usage export must preserve tenant and workload dimensions Team builds metering and allocation queries Reconcile accepted events against ledger rows
On-call load Service operations shift outward, but integration failures remain yours Storage, upgrades, scaling, and recovery remain yours Assign an owner and path SLO
Data location Verify processing and storage regions contractually Choose placement and operate replication and backups Test routing and retention by region
Exit cost Export fidelity and proprietary fields shape migration Schema and operational tooling shape migration Keep the application event contract neutral

There is no universal winner. A small platform team may rationally buy the ingestion plane because another datastore adds on-call burden. A regulated deployment with strict placement or a mature data platform may rationally operate it. The heartbeat-ledger approach also has a clear limitation: it is not suitable as the sole diagnostic system when engineers need stack traces, release correlation, or browser context, so retain a separate error tracker for those jobs. In every deployment, require an exportable event schema, a documented retention boundary, and a way to reconcile dropped or rejected events. Low ingestion cost cannot compensate for missing evidence during an incident.

Implement detection as a reconciler

The monitor compares expected runs with terminal runs after the grace period. Run it independently of the scheduler being monitored, because co-locating both behind the same failure domain makes the alarm disappear with the workload. Each evaluation should be idempotent and group pages by a useful boundary, such as importer and region, rather than opening one page per tenant during a shared outage.

A practical state progression is expected -> due -> late -> terminal. Administrative disablement is separate and auditable. When a late run becomes terminal, resolve the active symptom, preserve the late classification for SLO accounting, and attach the final attempt and counts. Do not rewrite history to green.

Error events still earn their keep. Capture backend exceptions with the logical run ID so responders can move from a late-run alert to the causal failure. For a Node.js API and React client, backend capture is required for this job; client capture and source maps improve diagnosis of browser failures but cannot observe scheduler silence. Make source-map upload a release-pipeline choice with access controls and retention review, not a prerequisite for import monitoring.

Feature toggles can reduce rollout risk, provided each toggle has an owner and retirement plan. Fowler's taxonomy is useful here: release, experiment, ops, and permissioning toggles have different lifetimes and operational consequences. A short-lived release toggle around heartbeat publication is reasonable; a permanent toggle that silently disables monitoring per tenant undermines the expectation model.

Verify and roll back without blinding on-call

Start in shadow mode. Generate expectations and heartbeats, calculate lateness, but route findings to review instead of paging. Test four cases explicitly: a successful non-empty import, a successful empty import, a worker failure after start, and a run never dispatched. Then test a delayed retry that succeeds after the grace period. The first two should remain healthy; the last three should remain visible with different diagnostic states.

Observe the observer. Track expectation generation, terminal-event acceptance, outbox age, reconciliation lag, and alert-delivery success separately. A single end-to-end status hides where evidence vanished. Synthetic scheduled runs in each operating region can verify the full route without patient data, and their fixed cadence makes missing evidence unambiguous.

Deploy by importer class or a bounded tenant cohort, with paging enabled only after shadow results match authoritative job state. Compare ledger counts with scheduler expectations and investigate every unexplained gap. Set the grace period from the actual completion-time distribution plus the response objective; an arbitrary five-minute threshold can page on normal variance or wait too long to protect freshness.

Rollback has two independent controls: disable paging first, then disable publication only if publication threatens the import path. Keep expectation generation and the ledger running during an alert-policy rollback so evidence remains available for replay and tuning. If the sink degrades, the outbox may accumulate only within a defined capacity envelope; before reaching that boundary, follow a predeclared load-shedding policy that protects the primary database.

The acceptance rule is strict: ship when missing dispatches, started failures, empty successes, and delayed retries are classified correctly; tenant attribution reconciles to the run ledger; regional routing has been verified; and the telemetry path has its own SLO and capacity envelope. An exception dashboard alone cannot meet that bar.

References

Top comments (0)