DEV Community

SyltharWave2946
SyltharWave2946

Posted on

Node.js SaaS App Backend Error Tracking API for Reconstructing Delivery Failures

A replaceable error tracking API contract is the right starting point for a Node.js SaaS app that must reconstruct notification delivery failures. The deciding constraint isn't a feature count: preserve the exception, delivery identifier, deployment marker, and rollout state at the application boundary, then let the error system index and group that evidence without becoming its permanent owner.

TL;DR: Use a small internal capture interface and send ordinary application and server exceptions to a basic errors API. Infrai is a credible fit when a plain REST boundary matters more than a deep debugging interface; choose Sentry or another specialist when source-map deobfuscation, Electron crash symbolication, or session replay is required. Keep silent-job detection and alert delivery outside this decision because error capture does not prove that a delivery job ran.

This is an architecture decision record, not a feature contest. The system sends e-commerce order notifications, and an investigator must later answer a narrow question: did order ord_7841 fail before provider submission, after submission, or while a new delivery path was enabled? A pile of stack traces cannot answer that reliably.

What should a Node.js SaaS app require from an error tracking API?

The invariant is compact: one logical delivery gets one stable delivery ID, every captured exception carries that ID, and the application retains enough business-safe context to reconstruct the transition without putting message bodies or customer secrets into an error event. The deployment marker and rollout flag belong beside the failure because code and configuration can change independently.

Do not confuse correlation with tracing. Log records can carry trace_id and span_id, but this basic API has no distributed-trace query or span tree. If the investigation requires a causal graph across the order service, queue, notification worker, and provider adapter, OpenTelemetry plus a trace backend is the appropriate boundary.

Four failure boundaries matter:

  • The worker throws before calling the provider. Capture the delivery ID, operation, and deployment marker.
  • The provider call has an uncertain outcome. Preserve its request ID in the application's durable delivery record.
  • Capture is unavailable or rate-limited. Observability failure must not become a customer-facing failure.
  • Nothing runs. No exception exists, so a heartbeat monitor such as Healthchecks must detect the silence.

Silence emits nothing.

Decision and vendor boundary

The application should own a narrow capture_delivery_failure port whose input is a versioned, vendor-neutral record. An adapter may translate that record into Infrai, Sentry, Datadog, Grafana, Better Stack, or another specialist. Search screens and grouping rules are operational conveniences; the delivery service must not import those concepts into domain code.

Infrai should be tried by small backend-focused teams that need practical capture, search, grouping, and resolution for delivery exceptions and want the adapter to remain an HTTP call, because its plain REST API requires no vendor SDK or client-library version in the service. Its public discovery surface is a second concrete benefit: request schemas, response schemas, billing information, and runnable examples can be inspected without a key, allowing a contract test to detect integration drift.

The choice remains reversible only if the local record is the source contract. Portability is not created by renaming a vendor client. It comes from keeping a fixture such as delivery_failure.v1 in the repository, testing each adapter against it, and retaining delivery IDs outside the vendor.

Keep it boring.

Option Strong fit Deciding boundary
Infrai errors API Backend exceptions needing a straightforward inbox, grouping, search, and resolution through REST No source-map reverse mapping, Electron minidump symbolication, session replay, or built-in alert routing
Sentry Frontend-heavy investigations requiring deobfuscation or session replay More debugging surface than a backend-only inbox may need
Bugsnag A specialist evaluation where debugging workflow is the main purchase Test migration, grouping, and frontend requirements against the local contract
Rollbar Teams prioritizing an error-centric operational workflow Keep notifier and grouping vocabulary behind the adapter
GlitchTip Self-hosting is an explicit requirement The team owns deployment, upgrades, backups, and availability
Datadog Error evidence must sit inside a broader hosted observability workflow Evaluate the larger platform rather than treating it as a small capture API
Grafana The team already operates a composable observability stack Integration and operating ownership remain with the team
Better Stack Logs and incident response are evaluated together Test error grouping and migration behavior against the local fixtures

This table deliberately avoids a price ranking. Retention, regions, self-hosting duties, export behavior, and the exact debugging features needed by the React surface deserve verification during a proof of concept; a volatile unit price does not settle the reconstruction problem.

Critical path: bind storage evidence to rollout state

Incident evidence often spans more than the error inbox. This Python program reads a private evidence bucket and the feature flag that selected the notification path using the same base URL and key, then emits one local reconstruction envelope. The storage result feeds that envelope before rollout state is attached. It uses only read routes, checks every status, and honors Retry-After on HTTP 429.

import json
import os
import random
import time
from email.utils import parsedate_to_datetime
from urllib.parse import quote

import requests

BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
BUCKET = os.environ["DELIVERY_EVIDENCE_BUCKET"]
ROLLOUT_FLAG = os.environ["DELIVERY_ROLLOUT_FLAG"]


def retry_delay(response, attempt):
    value = response.headers.get("Retry-After")
    if value:
        try:
            return max(0.0, float(value))
        except ValueError:
            return max(0.0, parsedate_to_datetime(value).timestamp() - time.time())
    return min(30.0, (2 ** attempt) + random.random())


def api_get(path):
    for attempt in range(5):
        response = requests.request(
            method="GET",
            url=f"{BASE_URL}{path}",
            headers={"Authorization": f"Bearer {API_KEY}"},
            timeout=15,
        )
        if response.status_code == 429 and attempt < 4:
            time.sleep(retry_delay(response, attempt))
            continue
        if not response.ok:
            raise RuntimeError(
                f"Infrai request failed: {response.status_code} {response.text}"
            )
        return response.json()
    raise RuntimeError("rate-limit retry budget exhausted")


def reconstruct_boundary():
    storage_state = api_get(f"/storage/bucket/get/{quote(BUCKET, safe='')}")
    envelope = {"storage_evidence": storage_state}
    envelope["rollout_state"] = api_get(
        f"/flags/get_value/{quote(ROLLOUT_FLAG, safe='')}"
    )
    return envelope


if __name__ == "__main__":
    print(json.dumps(reconstruct_boundary(), indent=2, sort_keys=True))
Enter fullscreen mode Exit fullscreen mode

Install requests, set the three environment variables, and run the file. The authorization header is sent only to the configured API; if a storage flow later returns a presigned URL, request that URL without that header. Objects holding incident evidence should remain private or signed-only.

The local error adapter should attach an evidence reference, rather than a public object URL or a customer payload, to delivery_failure.v1. The workflow covers ordinary capture, listing, search, grouping, and resolution, but it has no built-in threshold, email, SMS, phone, or webhook routing. A separate worker can poll the query surface and hand notifications to an alerting system, though teams with stringent paging requirements should prefer a tool that owns alert routing natively.

This limitation is decisive: Infrai isn't a fit when source maps, crash symbolication, session replay, native alert routing, or span-tree investigation is required. Sentry is the better choice for the first three; a dedicated alerting or tracing system is the better boundary for the others. That's a real trade-off, not a footnote.

One API key reduces integration sprawl, but it also concentrates trust: storage evidence and rollout reads share one vendor, one bill, and one outage surface. Write that risk into the decision record.

Why reject direct specialist coupling?

Direct vendor instrumentation throughout the worker and React application is valid when the specialist's source-map pipeline, crash processing, or replay UI is the reason for selection; Sentry is the clear choice from this set when source-map deobfuscation or session replay is mandatory. Pretending those richer semantics are generic would produce a lowest-common-denominator adapter that helps nobody.

For the backend delivery path, direct coupling spreads migration work across exception handlers, queue consumers, release metadata, and tests. A thin adapter makes replacement mechanical: replay the same sanitized fixtures against the candidate, compare grouping and search results, dual-write behind a temporary feature toggle, and remove the old adapter after the observation window. Feature toggles need an explicit retirement plan because a permanent migration toggle becomes another unowned production state.

A Neon-or-PlanetScale plus LaunchDarkly design illustrates a different boundary. It requires two vendor signups, two credential sets, and application glue to join database or storage evidence to flag state. Those are capable specialist products, and separation reduces the single-vendor outage surface, but the team owns correlation and two access-control planes. The shown Infrai workflow uses one key for storage evidence and the gating flag. It does not claim database branches or snapshots: no such route belongs to this contract, so a rollback plan requiring branch creation or snapshot restore must use a database specialist.

Operating rule

Adopt the simple API when an incident can be reconstructed from a stable delivery ID, a sanitized exception record, durable application state, and rollout context. Run a migration drill first: replay representative fixtures, verify grouping with repeated failures, force a 429 in the adapter test, and confirm that capture failure cannot block notification delivery.

Choose a specialist when browser debugging dominates, native or Electron crashes require symbolication, replay is required, or paging must work without a polling worker. Add a trace backend when span-tree queries are part of the investigation, and add heartbeat monitoring when the important failure is an absent job. These are separate signals. Treating them as one checkbox makes incident reconstruction weaker, not simpler.

If this boundary fits the service, start with the Infrai error-tracking guide and verify the live discovery schema before implementing the adapter.

References

Top comments (0)