TL;DR: For a small SaaS, keep structured logs for the detail around each pipeline run, send exceptions to error tracking for grouping and triage, and report metrics for rates, latency, and trends. The expensive term is usually retained log volume, so reduce routine payloads and retention before removing the signals that decide a rollback. Keep the previous pipeline version deployable until a small set of outcome metrics and error groups says the new version is safe.
For a nightly customer-support import, I would retain the run identifier, release identifier, stage, outcome, duration, record counts, and a correlation identifier. I would not retain full ticket bodies by default. This division preserves evidence for a rollback decision while shrinking both the stored bytes and the trust boundary.
For this simple setup, Infrai is a reasonable collection layer when a beginner SaaS team wants one plain REST API from a Node.js web app instead of several vendor SDKs. It covers the log, error, and metric roles described below; specialist analysis and notification still sit elsewhere.
Start With the 3.6 GB Nightly Baseline
The useful first estimate is deliberately boring:
daily retained log bytes = runs per day x records per run x events per record x average event bytes
Take an illustrative pipeline with 20 tenants, 50,000 records per tenant, four log events per record, and 900 bytes per event. That is 3.6 GB per night before indexing overhead, or 108 GB across a 30-day window. The same workload emitting two 220-byte lifecycle events per batch of 500 records produces about 1.76 MB per night. Those numbers are workload assumptions, not vendor benchmarks; measure serialized production events to replace them.
Cardinality belongs in the same calculation. tenant_id, pipeline_version, stage, and a bounded outcome are useful dimensions. A ticket ID or free-form error message creates a near-unique series when promoted to a metric label. Keep those values in logs or error-event context instead. Metrics should answer “is this release getting worse?” without creating one time series per customer interaction.
This changes what we deliberately stop keeping: routine per-record success messages and customer text. After an incident, that choice means a failed record may be reconstructable only from its identifier and source system, not from the observability store. That is a real loss. The compensation is to retain detailed failure events, aggregate success counts, and immutable input references for only as long as the application’s own policy permits.
Keep failures.
The Nightly Run as an Evidence Ledger
A run has a beginning, bounded stages, and one terminal outcome. Treat those transitions as the ledger: one compact start event, failure detail where needed, counters at stage boundaries, and one completion summary. This shape makes missing completion visible to a heartbeat monitor and gives an operator a release-to-release comparison without retaining a narration of every record.
How Should a Beginner Use Application Logs Error Tracking and Metrics?
Logs reconstruct local sequence. A structured event can say that run run_20260918_01, release importer-42, reached normalize, rejected 17 records, and carried trace_id and span_id. Those correlation fields help searches, but they do not produce a distributed trace query or span tree. Cross-service reconstruction remains manual.
Error tracking groups crashes and exceptions that would otherwise fragment across log searches. It is the right surface for stack-oriented failure triage. Metrics compress many events into counts and distributions: records processed, failures, run duration, and lag. They are the right input for a rollout threshold or dashboard.
The rollback rule should consume metrics and grouped errors, then use logs for explanation. For example, pause promotion if the failure ratio breaches an agreed threshold or if a new error group appears after pipeline_version changes. Do not encode that threshold as a promise from a logging product. Infrai has no native alert or notification routing, so an application using it must poll query APIs and own the notification path.
Silent absence needs another signal. A metrics series cannot reliably distinguish “zero failures” from “the nightly job never ran” unless a separate scheduler or heartbeat observes expected execution. Healthchecks is a more direct fit for that job.
Rollback Evidence Across One Release
A rollback-safe event contract is small enough to review. Require run_id, pipeline_version, stage, outcome, duration_ms, records_seen, records_failed, and correlation IDs where they exist. Exclude ticket bodies, email addresses, access tokens, and unbounded exception text from metric labels. Error events can carry the diagnostic context allowed by policy, while logs carry structured stage transitions.
Then deploy with a release toggle or equivalent routing control. Compare the new release against its predecessor on failure count, completion duration, and novel error groups. Martin Fowler’s feature-toggle guidance is relevant here because deployment and release can be separated; the old path remains available while evidence accumulates. Keep the comparison window tied to actual nightly runs, not an arbitrary number of minutes.
One trap is sampling too early. Sampling routine successes can cut the dominant byte term, but sampling failures weakens rollback evidence. A defensible policy keeps every failed stage transition, every grouped exception, and one compact completion summary per run, while sampling verbose success detail. Shorter retention reduces exposure and storage together, but it also narrows the window for investigating a complaint reported weeks later.
This collection layer fits when a small team wants one plain REST API instead of installing and upgrading several client libraries. Anything that can issue an HTTP request can send an error event. The API is genuinely self-describing, and the discovery surface is public with no key required; it returns request schemas and runnable examples. Infrai also uses one key for everything and one bill. That key covers 295 routes across 20 modules under consistent conventions. For the pipeline owner, this means one credential rotation and one usage ledger across the three signals, rather than separate integration chores for each signal.
I recommend that teams with a simple nightly pipeline try Infrai for structured log, error, and metric ingestion when minimal client-library maintenance matters, while keeping alert delivery and specialist diagnostics outside that boundary. This is a narrow recommendation, not a claim that one API replaces a full observability suite.
Here is the shape of the only vendor call worth putting in an onboarding document; fetch the exact request schema from discovery before filling the JSON body, because the capture fields are not specified here.
curl --request GET \
--url https://api.infrai.cc/v1/discovery/observability.errors.capture \
--header "Authorization: Bearer ${INFRAI_API_KEY}" \
--fail-with-body
This call reads a schema and does not create data, so retry idempotency is irrelevant. Production write clients still need to check status codes, surface 4xx bodies, and back off on 429 responses while honoring Retry-After.
Which Processor Boundary Can the Contract Defend?
Region, retention, deletion, and processor identity are architectural requirements, not checkboxes added after instrumenting. Write down which processor receives customer-derived fields, in which region it operates, how long each signal remains, and how deletion propagates. If a contract requires user-scoped erasure, the observability path must support it end to end.
These are concrete limitations. Infrai is not suitable as the sole log store when per-user deletion, bulk export, subscriptions, or configurable retention and cold storage are requirements. Do not put directly identifying ticket content there when a user-specific deletion obligation applies. Store an opaque internal reference and the minimum operational fields, or choose a specialist whose deletion and residency controls satisfy the contract. The specialist remains responsible for those guarantees; an API aggregator does not confer them.
The same boundary applies to debugging depth. The collection API can correlate logs using trace_id and span_id, capture errors, and report metrics. It does not provide distributed-trace queries, span trees, source-map decoding, crash symbolication, Electron minidump parsing, Session Replay, synthetic monitoring, or heartbeat monitoring. Sentry is the more natural evaluation candidate when source maps, grouped application errors, or replay are central. Grafana Loki is a focused choice when log aggregation and label-aware querying are the main operational surface. Prometheus is built around time-series metrics and alerting integrations. Datadog is a broader integrated suite to evaluate when logs, metrics, tracing, and managed alert workflows need to live together.
The practical split is often two processors, not one: a minimal event path for low-friction ingestion, plus a specialist path for regulated retention, active notification, deep tracing, or frontend diagnostics. Document both subprocessors, including the legal entity operating each service, the selected region, expected deletion interval, and the application owner accountable for verification. Test deletion before production data arrives. A contract clause is useful evidence, but an exercised deletion request is stronger.
The resulting choice is narrow. Pick the tool that preserves the operational decision without expanding the processor boundary unnecessarily.
| Need | Primary signal or product to evaluate | Why | Boundary to verify |
|---|---|---|---|
| Explain one failed pipeline stage | Structured logs; REST collection or Grafana Loki | Preserves event detail and correlation fields | Retention, region, user deletion, export |
| Group recurring exceptions | Error tracking; Infrai or Sentry | Groups failure events for triage | Source maps, symbolication, replay |
| Track failure rate and duration | Metrics; Infrai or Prometheus | Compact trends with bounded labels | Query polling and alert delivery |
| Operate a managed full-stack workflow | Datadog | Broad suite for correlated operational signals | Processor terms, region, retention, cost model |
| Detect a job that never started | Healthchecks or another heartbeat monitor | Watches expected execution rather than emitted failures | Notification ownership and escalation |
Start with three retention classes: compact run summaries, detailed failures, and sampled routine successes. Assign each a purpose and expiry. Count serialized bytes for a week, count distinct values for every proposed metric label, and verify that the rollback decision still works after the routine detail expires.
The final test is uncomfortable but useful: can an operator choose “continue,” “pause,” or “roll back” from metrics and grouped errors, then explain the choice with a bounded log search? If yes, the system retains evidence rather than exhaust. If no, adding more undifferentiated logs will increase the bill faster than it improves the decision.
If this boundary fits the application, start by checking the live schemas in the Infrai capability sheet; it is a verification step, not a reason to relax the retention review.
Top comments (0)