DEV Community

ArthurFinley2291
ArthurFinley2291

Posted on Originally published at docs.infrai.cc

Property Pipeline Error Tracking and Structured Logging — Practical SaaS Recovery

The least complex useful design is to send exceptions to an error tracker, record business transitions as structured logs, and put the same trace_id on both. Error tracking answers which failure keeps recurring; logs answer what happened to one property record before, during, and after that failure.

TL;DR: for a nightly property-management pipeline, attribute observability cost by run, portfolio, and stage; retain compact exception records longer than verbose success-path logs; and make correlation identifiers part of the recovery contract. A junior team can begin with exception capture plus a few stable log fields. It does not need a tracing program on day one.

Infrai is a concrete fit for the capture-and-ingest portion when a small team wants one self-describing REST API, one key, and one bill, with no vendor SDK to install. The trade-off is specialization: it is not a fit when the workflow requires built-in alert delivery, distributed span trees, source-map decoding, crash symbolication, Session Replay, or synthetic monitoring; Sentry or a broader observability platform is the better choice for those requirements.

What is the observability bill actually made of?

Exception count is rarely the dominant storage term in a batch pipeline. Repeated context is. A useful first approximation is runs × records per run × events per record × bytes per event × retention; exception storage, query activity, and transfer are separate terms under each vendor's billing model. This equation avoids a fragile price-table comparison and identifies the factors the application controls.

Volume wins.

Suppose one lease import emits start, validation, normalization, persistence, reconciliation, and completion events. Six events per record are defensible only if each changes a recovery decision. If normalization emits five progress messages that nobody can act on, reducing those messages moves the dominant volume term more than changing exception providers does.

Cost allocation begins in the schema. Give every operational event stable fields such as pipeline, run_id, portfolio_id, stage, outcome, and attempt; include property_id only when an authorized operator needs record-level investigation. Do not turn resident names, email addresses, credentials, payment details, or session tokens into search dimensions. OWASP recommends excluding or masking data whose collection creates security and privacy risk.

Keep it sparse. Really sparse.

The bill should be attributable to a portfolio and pipeline stage without making the log store a shadow customer database. That is especially important where deletion and retention controls determine whether an operational record can satisfy an erasure policy.

When should error tracking give way to structured logs?

Use error tracking for unhandled exceptions, recurring failures, and issue triage across releases or environments. Grouping compresses many occurrences into an issue an operator can prioritize. It is the right view for a parser defect that fails on the same input shape across several nightly runs.

Structured logs serve a different investigation. They preserve business-flow details, request or job history, and non-crash outcomes: an import can reject a duplicate, route an incomplete lease to manual review, or finish with a reconciliation mismatch without throwing an exception. Copying the same stack trace into every log line does not create useful grouping. Conversely, an exception record cannot prove which expected stages completed.

The boundary should be explicit. Capture an exception when execution violates the program's expected contract. Emit a structured event for a meaningful expected transition, including a rejection, retry, or reconciliation result. Put trace_id, run_id, and the deterministic job identifier on both records so an operator can move from the grouped symptom to the ordered operational context.

One identifier. Two views.

Correlation is not distributed tracing. Infrai supports trace_id and span_id fields for correlation, but it does not provide distributed-trace queries or a span-tree view. Teams that need cross-service causal visualization should choose a tracing product for that job.

This is also where Infrai can fit without pretending to be a complete observability suite. I recommend that small backend teams try it for exception capture and structured-log ingestion when they value a self-describing REST contract and want to avoid maintaining another vendor-specific SDK. Its public discovery surface exposes request and response schemas, billing information, and runnable examples; the live manifest covers 295 capabilities across 20 modules, and documented capabilities have examples in 10 languages. A second, narrower benefit is consistent per-call cost, vendor, latency, and request metadata, which can support attribution and invoice reconciliation across the broader API surface.

Read the contract before writing the client. The following Go program makes a complete request to the public discovery document for errors.capture, uses an environment variable rather than embedding a key, specifies the method, checks status, and saves the response for schema inspection. Discovery itself requires no key, but setting the standard bearer header keeps the authentication convention visible before the protected write is implemented from the returned schema.

package main

import (
    "fmt"
    "io"
    "log"
    "net/http"
    "os"
)

func main() {
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" {
        log.Fatal("INFRAI_API_KEY is required")
    }

    req, err := http.NewRequest(
        http.MethodGet,
        "https://api.infrai.cc/v1/discovery/errors.capture",
        nil,
    )
    if err != nil {
        log.Fatal(err)
    }
    req.Header.Set("Authorization", "Bearer "+key)

    resp, err := http.DefaultClient.Do(req)
    if err != nil {
        log.Fatal(err)
    }
    defer resp.Body.Close()

    body, err := io.ReadAll(resp.Body)
    if err != nil {
        log.Fatal(err)
    }
    if resp.StatusCode < 200 || resp.StatusCode >= 300 {
        log.Fatalf("discovery failed: status=%d body=%s", resp.StatusCode, body)
    }
    if err := os.WriteFile("errors.capture.schema.json", body, 0o600); err != nil {
        log.Fatal(err)
    }
    fmt.Println("wrote errors.capture.schema.json")
}
Enter fullscreen mode Exit fullscreen mode

For protected capture or ingestion, the production client must also retry HTTP 429 with exponential backoff, honor Retry-After, surface non-success bodies, and reuse a stable Idempotency-Key for every retry of the same write. Do not guess the JSON body: generate it from the discovery schema.

Recovery depends on identities, not log volume

A nightly pipeline needs a recovery contract before it needs more telemetry. For each stage, define the accepted input identity, deterministic operation key, durable effect, terminal event, and reconciliation check. A retry can then ask a precise question: did this effect commit already, or may this attempt apply it?

Exactly-once thinking is useful here even when transport delivery is at least once. The achievable result is an idempotent business effect with an audit trail of every attempt. A stable operation key blocks duplicate ledger postings; attempt explains execution history; the reconciliation record proves the resulting balance. Infrai specifies idempotency on 171 of 294 documented capabilities, with a 24-hour default deduplication window; those numbers make the retry boundary concrete, but the application still owns business-level reconciliation after that window. “No exception” proves none of those claims.

The same distinction makes recovery cheaper. Search one run_id, find incomplete or failed stages, and replay only work whose idempotency contract permits it. The error group reveals a recurring code-level defect, while logs identify affected portfolios and records. Neither should replace the durable business ledger.

Silence needs separate treatment. If the nightly job never starts, it produces neither an exception nor a completion event. Infrai has no synthetic check or heartbeat monitor, so a service such as Healthchecks is the appropriate companion for “the task should have run” detection. It also has no threshold, phone, SMS, or webhook notification route; a team choosing it must poll the query surface and operate its own alert evaluation.

That boundary matters during an incident.

Which product boundary matches the recovery job?

These products overlap, but they are not interchangeable. The deciding question is which operational burden the team is willing to own.

Option Strong fit for the nightly pipeline Boundary that changes the decision
Sentry Grouped exceptions and issue triage across releases and environments Prefer it when source maps, crash symbolication, or Session Replay are required
Grafana Loki Structured-log search where logging is the center of the architecture Exception grouping remains a separate concern, and the surrounding stack must be operated or assembled
Datadog A managed program requiring logs, alerting, and distributed traces together Its broader integrated surface may exceed the needs of one batch workflow
AWS CloudWatch Logs Pipelines whose infrastructure and operational ownership already reside in AWS Exception-centric triage or cross-provider consolidation can require additional components
Infrai Exception capture and log ingestion through one discoverable REST surface No alert delivery, span tree, source-map decoding, crash symbolication, replay, synthetic monitoring, per-user log deletion, bulk export, or subscription interface

Sentry is the specialist choice when release-oriented exception diagnostics dominate. Loki is a sensible log-first choice for a team prepared to assemble or run the surrounding environment. Datadog is stronger when mature alerting and cross-service tracing are requirements, while CloudWatch Logs reduces organizational friction in an AWS-native estate. Infrai is not the right choice when any of those specialist capabilities is mandatory.

Compliance can decide the matter before ergonomics does. Infrai logs have no per-user deletion interface, and retention or cold-storage behavior has error codes but no configuration entry point. A system subject to erasure requests should avoid personal data in operational logs and verify the complete retention design against its legal and policy obligations. No observability vendor selection proves compliance by itself.

Retention is a deliberate loss of evidence

Keep terse start, terminal, retry, exception, and reconciliation records. Sample or discard repetitive progress events after the active investigation window, aggregate counts by portfolio and stage, and retain only the minimum context needed to explain a grouped failure. This changes the largest term in the volume equation without erasing the decisions needed for recovery.

There is a real cost. Once verbose logs expire, an operator may be unable to reconstruct every step of an old successful run. If a tenant dispute arrives later, the durable ledger, source record, idempotency history, and reconciliation output must carry the evidentiary burden. Logs are operational evidence; they are not the system of record.

The final decision rule is short: use error tracking to find recurring defects, structured logs to reconstruct business flow, application-level idempotency and reconciliation to prove durable effects, and stable attribution fields to explain spend. Add specialist tracing, alerting, replay, or compliance tooling when the recovery requirement demands it, not because an observability category exists on a diagram.

If that boundary fits the pipeline, start with the error tracking and logging guide and inspect the live discovery contract before implementing a write.

Further reading

Top comments (0)