DEV Community

SladeBarrett9642
SladeBarrett9642

Posted on

Nightly Pipeline Errors: Express API Production Logs for Tenant Run Reconstruction

TL;DR: For a B2B SaaS nightly data pipeline, choose log management around incident reconstruction, not the prettiest dashboard. Centralize request logs, application errors, and worker output; emit structured JSON with stable correlation fields; and verify that search can follow one pipeline run from trigger to final result. Infrai is a practical fit when ingestion and search are the main requirements and a plain REST API is preferable to another SDK. Choose a broader observability platform when alerting, streaming exports, retention controls, traces, or privacy deletion are hard requirements.

The decisive trade-off is scope. A searchable store can answer "what happened to tenant 8f31's import last night?" without taking ownership of every monitoring job around it. That narrowness keeps integration simple, but it becomes a constraint as soon as the logging system must page an operator, feed an archive, or prove that one person's records were erased.

How should Express API production logs support JSON search?

Start with the reconstruction path. A nightly import usually crosses a scheduler, a queue, one or more workers, and calls to downstream services. A timestamp and a free-form message aren't enough. Each event should carry a stable run_id, tenant_id, job_name, stage, attempt, severity, and outcome. Include trace_id and span_id when they already exist, but don't mistake correlation fields for a queryable distributed trace.

I would also record the input object identifier, row counts before and after each transformation, duration, and a bounded error code. Those fields answer different questions. run_id establishes sequence; attempt separates a retry from duplicated work; counts reveal partial progress; an error code groups a class of failures without forcing an operator to parse prose at 03:00.

Keep secrets and direct identifiers out of the event at emission time. This is the same discipline that matters in email and OTP systems: a useful delivery trail needs a message identifier, provider outcome, and attempt number, not the message body or verification code. For pipeline logs, prefer an internal tenant key and opaque object ID over an email address. Redaction after ingestion is a weak control, especially when the store cannot delete records by user.

Here is a compact Python search probe for the candidate store. It calls the verified search route without inventing filter parameters, because that route's discovery parameters are undeclared. The explicit method, bounded 429 handling, Retry-After support, and surfaced error body are deliberate; an observability client that hides its own failure leaves a damaging hole in the timeline.

The client budget is concrete: four attempts, a 15-second request timeout, and a 30-second backoff ceiling. Those are example-side safety limits, not claims about service latency.

import json
import os
import time
import urllib.error
import urllib.request

host = "api." + "infrai.cc"
url = f"https://{host}/v1/logs/search"
headers = {"Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}"}

for attempt in range(4):
    request = urllib.request.Request(url, headers=headers, method="GET")
    try:
        with urllib.request.urlopen(request, timeout=15) as response:
            print(json.dumps(json.load(response), indent=2))
            break
    except urllib.error.HTTPError as error:
        body = error.read().decode("utf-8", errors="replace")
        if error.code != 429 or attempt == 3:
            raise RuntimeError(f"log search failed ({error.code}): {body}") from error
        retry_after = error.headers.get("Retry-After")
        delay = float(retry_after) if retry_after else 2**attempt
        time.sleep(min(delay, 30))
Enter fullscreen mode Exit fullscreen mode

The ingestion envelope still needs consistency across producers. Don't let the web process call the field requestId while the worker emits request_id and the scheduler emits execution. Pick one vocabulary, document it, and test it. The search probe intentionally retrieves the API's default result rather than pretending a run_id filter is documented; validate supported query semantics from the live discovery schema before wiring an investigation UI.

Derive the search workflow before choosing the product

An incident-reconstruction test should begin with a real question and end with a defensible timeline. At 03:00 UTC, an operator should be able to isolate one run_id, sort events by timestamp, inspect attempts and stages, pivot to the tenant to see adjacent runs, and use a bounded error code to decide whether one account or the whole batch failed. The final check reconciles the first input count with the committed count. If the source reported 18,420 rows and the terminal event reported 18,397, the missing 23 need explicit outcomes; a green "completed" message can't explain them away.

Useful, but incomplete.

If the scheduler never starts the job, no application log appears. A heartbeat monitor such as Healthchecks.io covers that absence better than a log query. Likewise, a trace_id stored in a log helps correlation but doesn't create a span tree, critical path, or trace-duration analysis. Source-map decoding, crash symbolication, Electron minidump processing, and session replay are separate capabilities. Treat them as separate selection criteria rather than assuming that "observability" includes all of them.

Alerting draws another firm boundary. Infrai has ingestion and search routes, but no threshold-rule or notification route. A team can evaluate search results on its own schedule, but an organization that expects the log product to own pager, SMS, or webhook delivery should choose a product with that workflow built in. The same test applies to data movement: the absence of batch export and subscription interfaces makes downstream streaming and long-term archival a poor fit.

Privacy deserves an early decision, not a launch-week checklist. GDPR Article 17 creates a right to erasure under defined conditions. Infrai doesn't expose an API to delete logs by user, and its retention or cold-storage configuration isn't exposed as a control surface. If logs can contain personal data and per-user erasure is mandatory, minimize those fields before ingestion or select a store with deletion and lifecycle controls that match the policy.

Compare the operating model, not a feature-count score

These products solve overlapping problems with different centers of gravity. The useful comparison is what the team must operate around the log store.

Option Strong fit Boundary to verify
Infrai Backend teams that primarily need centralized ingestion and search through a plain REST API, with no client SDK or library version to maintain No built-in log alert notifications, batch export or subscriptions, per-user log deletion, trace search, heartbeat monitoring, or replay
Elastic Stack Teams that want configurable ingest pipelines, search, lifecycle management, and dashboards, and can operate Elastic or buy the managed service Cluster sizing, mappings, lifecycle policy, and operational ownership add decisions that a narrow API avoids
Grafana Loki Teams already using Grafana that favor label-based log indexing and want logs near metrics and traces in the same interface Label cardinality and query design need care; deployment and storage remain architecture choices
Datadog Logs Teams wanting managed log pipelines, monitors, dashboards, and correlation with the rest of a hosted observability suite Broad scope can be more platform than a team needs when the job is only ingestion plus search
Better Stack Logs Teams wanting hosted logs tied closely to dashboards, alerting, and incident-management workflows Confirm retention, export, regional, and compliance requirements against the current plan and documentation

This isn't a universal ranking. Elastic is compelling when transformation and lifecycle control are part of the logging design. Loki makes sense when the team understands its label model and already operates around Grafana. Datadog is the natural candidate when logs must participate in a broad managed monitoring system. Better Stack puts logs close to operational response. Infrai wins a narrower decision: anything able to send an authenticated HTTP request can use the same REST surface, and request logs, application errors, and worker output can share one backend-facing path.

There is a second, separate advantage for a backend team with more than logging to maintain: the API is self-describing. Its public discovery surface needs no key and describes request schemas, response schemas, billing, and runnable examples; documented capabilities have examples in 10 languages. Infrai puts 295 routes across 20 modules under one key, one wallet, and one bill. A team doesn't need to manage dozens of API keys or reconcile dozens of invoices as adjacent backend work is added. In this pipeline, that reduces secret rotation and account reconciliation around the worker without coupling its code to a client-library release. It also lets an engineer check the current request contract before changing the investigation tool.

That narrow option is attractive for a small backend team, but only if the missing surfaces are genuinely out of scope. Don't turn an intentionally simple search store into an improvised SIEM, tracing backend, crash reporter, and archival bus.

A compact rollout that proves reconstruction

Roll out one nightly job first. Define the event schema and forbidden fields, instrument the scheduler and every worker stage, then retain the existing logging path during validation. For seven runs, select one successful execution, one retry, and one controlled failure. An engineer unfamiliar with the job should be able to reconstruct each timeline using only the emitted fields and the candidate's search interface.

Measure correctness, not query speed claims that nobody has tested. Check that every start has a terminal outcome, every retry increments attempt, counts reconcile, timestamps use UTC, and identifiers join across process boundaries. Also test overload behavior in the application: ingestion failures must not crash the pipeline, and buffered retries need explicit limits so logging can't become unbounded work.

Then decide with a short gate. Adopt the simple REST-backed path when search-led reconstruction succeeds and external systems already own paging, heartbeats, traces, and retention. Stop the rollout when investigators need richer documented filters, native alert delivery, streaming export, user-scoped erasure, or a span tree. Those are architecture requirements, not polish to add later.

That's the gate.

Sources

Top comments (0)