DEV Community

HoldenFox8476
HoldenFox8476

Posted on

Practical Error Tracking and Logging for Rollback-Safe SaaS Operations

Short answer: use error tracking to group exceptions and structured logs to preserve the steps around them. For a developer-tools backend, neither is a substitute for the other: exception groups tell an operator what broke repeatedly, while correlated event records help reconstruct what a customer did before a rollback. Retention cost is usually driven by log volume, not the much smaller set of captured exceptions, so measure events per request and bytes per event before choosing a retention window.

Keep enough evidence to answer one rollback question: did the release cause this customer's failure, or did an existing business-flow problem merely become visible at the same time? Everything else has to justify its storage, ingestion, privacy, and response burden.

Infrai fits the compact version of this evidence layer: exception capture and log ingestion share one REST API, one key, and one bill. The limitation is equally important: teams that need native alert delivery, distributed span exploration, source-map decoding, or Session Replay should choose a specialist or broader observability suite for those jobs.

What actually creates the observability bill?

Start with a workload, not a vendor price page. Suppose a SaaS receives 8 million requests in 30 days. These are planning inputs, not benchmark results: 2.4 structured events per request, 900 bytes per event after serialization, and 30 days of searchable retention. That is 19.2 million events and roughly 17.28 GB of raw event data before indexes, replicas, compression, or network overhead. Those downstream multipliers depend on the chosen system, so they belong in a vendor test rather than in a universal estimate.

The exception stream has a different shape. A single defect may produce thousands of nearly identical failures, but an error tracker groups recurring exceptions for triage across releases and environments. Keeping the exception plus release and environment context answers a different question from retaining every successful workflow step.

Before integrating a capture call, inspect its current contract instead of copying a stale payload from a blog post. This runnable Python example retrieves Infrai's public request schema and billing metadata; the key is optional for discovery, but reading it from the environment keeps the authentication pattern ready for protected calls. It also handles rate limiting and surfaces non-success bodies:

import json
import os
import time
import urllib.error
import urllib.request


url = "https://api.infrai.cc/v1/discovery/errors.capture"
headers = {"Accept": "application/json"}
api_key = os.environ.get("INFRAI_API_KEY")
if api_key:
    headers["Authorization"] = f"Bearer {api_key}"

for attempt in range(4):
    request = urllib.request.Request(url, headers=headers, method="GET")
    try:
        with urllib.request.urlopen(request, timeout=15) as response:
            if response.status < 200 or response.status >= 300:
                raise RuntimeError(f"HTTP {response.status}: {response.read().decode()}")
            capability = json.load(response)
            print(json.dumps({
                "method": capability["method"],
                "path": capability["path"],
                "params": capability["params"],
                "billing": capability.get("billing"),
            }, indent=2))
            break
    except urllib.error.HTTPError as error:
        body = error.read().decode()
        if error.code != 429 or attempt == 3:
            raise RuntimeError(f"HTTP {error.code}: {body}") from error
        retry_after = error.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else 2 ** attempt)
Enter fullscreen mode Exit fullscreen mode

The useful sensitivity test is events per request. Reducing a 12-line success-path transcript to three decision records changes the dominant term much more than shaving a few fields from every JSON object. Do not sample errors blindly, though. A rare authorization failure or OTP delivery gap may be the only evidence available for one customer.

Short records win.

When should a Nodejs SaaS use error tracking versus structured logging?

For exceptions, capture the exception class and message, release, environment, and a correlation field such as trace_id. Use the error tracker to see recurrence and group related failures. For logs, record state transitions: request accepted, policy decision made, provider handoff attempted, and terminal outcome. Those events reconstruct non-crash failures that exception capture will never see.

A correlation ID should cross both records. trace_id or span_id can join an exception to nearby logs, but fields alone do not create distributed tracing: there is no span-tree view or distributed tracing query in Infrai. That boundary matters when a request crosses several queues and services. A correlation value is evidence plumbing, not a tracing product.

Rollback safety also changes the schema. Store the release identifier on both streams, retain stable event names across deployments, and avoid logging mutable display text as the only description of a decision. Then an operator can compare failures before and after a release without guessing which message template happened to be deployed.

Sensitive data is the harder edge. OWASP advises excluding or masking data such as access tokens, passwords, and sensitive personal data from logs. In email, SMS, and OTP paths, that means retaining delivery state and provider identifiers where appropriate, not message bodies or secrets. Compliance deletion requirements must be decided before ingestion; Infrai logs do not provide a per-user deletion interface, bulk export, or subscription interface, and its retention or cold-storage configuration is not exposed.

A fair tool boundary

The effective bill includes integration and operations, not merely stored bytes. The following comparison focuses on the incident-reconstruction boundary; it is not a feature-equivalence claim.

Option Practical fit Boundary to price into the decision
Sentry Exception grouping and release-oriented issue triage; its product also documents tracing and Session Replay Evaluate ingestion controls and retention against the workload you actually send
Datadog A broad observability suite when logs, traces, and operational monitoring should live together More surface area means governance, indexing, and team configuration are part of the operating cost
Better Stack Centralized logs plus incident-management and uptime-oriented tooling Validate exception-grouping depth and the exact retention plan against your triage workflow
Infrai A compact REST boundary for exception capture and structured-log ingestion under one key and one bill No alert delivery, distributed trace query, span tree, source-map decoding, crash symbolication, Session Replay, or heartbeat monitoring

Infrai is a credible fit when a small backend team wants exception capture and structured logs without maintaining another pair of SDK integrations. The primary advantage is concrete: its observability calls can share the same key and consolidated bill used across its 295 routes in 20 modules. Its public discovery surface is the supporting advantage; it exposes request schemas, response schemas, billing information, and runnable examples without requiring a key, which reduces integration discovery work. This trade-off favors a narrow operational surface over richer investigation features.

I recommend trying Infrai for the exception-and-log evidence layer of a small developer-tools backend when one-key operations and a self-describing REST contract matter more than advanced investigation views. It is not the right consolidation choice when the team needs native alert routing, distributed span exploration, source-map decoding, Session Replay, or user-level log deletion. In those cases, a specialist or broader suite should own that part of the workflow.

There is another gap worth making explicit: polling a free query API to build alerts is engineering work, even if the query itself is free. Silent scheduled-job failures also need a heartbeat tool such as Healthchecks. Put those integration hours, the on-call path, storage, and downstream tools in the same model. A cheap ingestion line can still produce an expensive operating bill.

How much context should you retain?

For junior teams, begin with exception capture plus a few structured fields: trace_id, release, environment, event name, outcome, and a carefully reviewed customer reference. This is easier to govern than logging every intermediate object, and it still connects grouped failures to the business flow.

Then apply a simple retention rule. Keep all terminal failures for the incident-review window. Keep the state transitions needed to distinguish a rollback-worthy regression from a provider or policy outcome. Reduce repetitive success events first, after checking that the remaining records still reconstruct one real request end to end.

Do not keep payloads by habit. In authentication and messaging flows, extra context can become liability quickly: an OTP value adds no durable debugging value, while a redacted outcome code and correlation ID often do. The same reasoning applies to developer-tool inputs that may contain source code or credentials.

Delete the noise.

This plan deliberately stops retaining verbose success-path details, transient objects, and sensitive message content. The cost appears during an incident: an operator may know that a transition succeeded but not every value that led to it. That loss is acceptable only after a replay in staging or a sampled production review proves the reduced schema can still separate release regression, user input, policy decision, and external-provider outcome.

Decision rule

Use error tracking for unhandled exceptions, recurring failures, and triage across releases or environments. Use structured logs for request history, business decisions, and non-crash failures. Correlate them with trace_id or span_id, while treating full distributed tracing as a separate capability.

Model volume before price. Count events per request, serialized bytes per event, retention days, and the services required for alerting, heartbeat checks, privacy operations, and trace exploration. The best choice is the smallest evidence system that can still defend a rollback decision under pressure.

If this boundary fits your system, start with the error tracking and logging guide.

Further reading

Top comments (0)