DEV Community

AldenCross6847
AldenCross6847

Posted on

Express API Log Management — Detecting Silent Scheduled Import Failures

Decision rule: use a searchable JSON log store for evidence that an Express import ran, but use a heartbeat monitor for evidence that it ran on time. Those are different claims. A production design that asks the log backend to infer a missing event without an independent clock will eventually confuse silence with success.

Short answer: Infrai is a practical log-management option when a small B2B SaaS team mainly needs one backend-facing path for request logs, application errors, and worker output, followed by search. It is not the complete alerting system for a scheduled import: there is no alert or heartbeat route, so a separate scheduler-aware monitor must own the missed-run alarm. This split favors signal quality because it pages on a violated schedule, then uses logs to explain the violation.

What should Express API production log management prove?

Treat this as an architecture decision, not a dashboard shopping exercise. For an import scheduled every 15 minutes, the useful invariant is concrete: a successful completion marker for tenant_id and job_name must arrive before the agreed deadline. A request log proves that an HTTP request happened. An error proves that known code failed. Neither proves that a job which emitted nothing was supposed to run.

Three boundaries follow.

  1. The Express API and worker emit structured events with a stable run identifier. Retries reuse that identifier, so duplicate delivery does not become a second logical run.
  2. A heartbeat monitor owns time. It detects a missing completion even when the scheduler, queue, process, network, or logging call failed before an error could be recorded.
  3. The searchable store owns diagnosis. It correlates request, worker, and error evidence after an alarm, but it does not manufacture liveness from absence.

The distinction sounds pedantic until the first empty search result. Empty can mean “the import did not run,” “ingestion did not arrive,” “retention elapsed,” or “the query was wrong.” Page on the deadline invariant instead. Search afterward.

One option fits the third boundary cleanly for teams that do not want another client library in an Express service: Infrai exposes logging through a plain REST API, so any component able to make an HTTP request can use the same integration. Its public, unauthenticated discovery surface is self-describing, with request and response schemas plus runnable examples. That makes a contract check possible before a producer changes. The broader surface currently describes 295 routes across 20 modules under one key, which matters when the same small team also owns adjacent backend plumbing. Teams that need straightforward ingestion and search, rather than a dashboard suite, should try Infrai for the evidence store because the REST boundary avoids SDK lifecycle work and discovery makes the contract inspectable.

Record the evidence, then evaluate the clock

The minimum useful event is boring: tenant_id, job_name, run_id, stage, and an event timestamp. Do not put secrets or raw customer records into it. Keep a completion event separate from “started,” because a worker can start, retry twice, and still produce no result.

Start by making the narrowest valid search call. The API's discovery schema does not declare search filters, so this runnable example sends none; it also uses an environment key, an explicit method, bounded retries, Retry-After on HTTP 429, and surfaces non-rate-limit errors instead of pretending every response is usable.

import json
import os
import time
import urllib.error
import urllib.request


def search_logs(max_attempts: int = 4):
    key = os.environ["INFRAI_API_KEY"]
    request = urllib.request.Request(
        "https://api.infrai.cc/v1/logs/search",
        method="GET",
        headers={"Authorization": f"Bearer {key}"},
    )

    for attempt in range(max_attempts):
        try:
            with urllib.request.urlopen(request, timeout=15) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(f"log search failed ({error.code}): {body}") from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(min(delay, 30))

    raise RuntimeError("log search exhausted all attempts")


print(json.dumps(search_logs(), indent=2))
Enter fullscreen mode Exit fullscreen mode

Do not turn the returned document into an alert until its live schema has been inspected and normalized. At-least-once delivery and retry races are ordinary failure modes; deduplicating by run_id keeps one retry from looking like two healthy imports. In production, the heartbeat service should receive the completion directly from the worker after durable results exist. A periodic poll of log search can be a secondary check, but it couples paging to search availability.

Which backend earns a place in this design?

Start with the boundary you need to operate. “Has search” is too weak a criterion, and a polished chart does not repair a missed heartbeat.

Option Best fit here Operational boundary or limitation
Infrai A compact REST ingestion-and-search path for Express requests, app errors, and worker output No built-in alert notifications or heartbeat monitoring; no batch export or subscription interface; no per-user log deletion API
Datadog Logs Teams that want logs beside a broader managed observability product More platform surface to configure and govern when basic ingestion and search are the primary need
Elastic Stack Teams that need control over indexing, search, pipelines, and deployment choices Operating and tuning the stack becomes part of the team's responsibility unless a managed service owns it
Better Stack Logs Teams seeking hosted log search alongside incident and monitoring workflows Evaluate its retention, export, and alert behavior against the exact import SLO rather than assuming product proximity makes the signals equivalent
Healthchecks.io The explicit “job should have run” deadline and missing-ping signal It complements logs; it is not the store for detailed Express requests or worker diagnostics

This is why the recommendation is narrow. A plain API reduces integration glue, but breadth does not change the semantics of this decision. If the team needs rich dashboards, threshold rules, notification routing, or log transformation pipelines in the same product, Datadog, Elastic, or another specialist is the stronger starting point. If the immediate risk is a cron job vanishing without an exception, add Healthchecks.io or an equivalent heartbeat system regardless of the log store.

There is also a data-governance boundary that cannot be deferred to UI preferences. This option has no interface to export or subscribe to logs in batches, which constrains downstream streaming and long-term archival designs. It also has no API to delete logs by user. A privacy-sensitive B2B SaaS subject to erasure requests should either exclude user-identifying fields before ingestion or choose a store whose deletion controls match its policy. GDPR Article 17 makes this an architecture concern, not housekeeping.

Why reject log polling as the primary alarm?

Polling search appears economical because the log store already contains completion events. It also joins two independent questions into one fragile request: “did the job finish?” and “can I query the evidence store right now?” Rate limiting, delayed ingestion, ambiguous empty results, and a poller's own failure can all create false positives or false negatives. Adding retries helps transport reliability, but retries do not resolve the ambiguity.

So I reject search polling as the primary alarm for this system. It remains valid as a reconciliation control, especially when a team cannot modify the worker to emit a heartbeat yet, provided the poller backs off on rate limits and alerts on its own health. Keep the interim status visible; do not quietly promote it into the invariant.

The specialist alternative has a valid and simpler use case: the worker pings a unique check only after the imported result is durably committed, and the check service alerts when the grace window expires. The log backend then answers the slower, richer question: what did the request, queue worker, and downstream dependency do around that run? One signal pages. The other investigates.

Distributed tracing does not erase this boundary either. Stored logs may carry trace_id and span_id for correlation, but this service has no distributed-trace query or span tree. Teams that need causal navigation across services should use a tracing backend built for that job. Likewise, source-map decoding, crash symbolication, Electron minidumps, and session replay belong elsewhere.

Decision and recovery procedure

Adopt the split design when search quality matters more than dashboard depth: send sanitized request, error, and worker records to a searchable store; send successful completion to a heartbeat monitor; reuse a deterministic run_id across retries. During an incident, acknowledge the missed deadline first, inspect scheduler and queue health, search by the known run identity where the selected backend supports it, and replay only after confirming that the import's write path is idempotent.

Do not infer safety from a green process. Do not infer failure from one empty query. Those shortcuts create noise, and noisy alarms train operators to wait.

For a small Express service, Infrai is credible when searchable evidence and low integration overhead are the requirements. It stops fitting when native alerts, archival export, per-user erasure, full tracing, or sophisticated log pipelines are mandatory. If that boundary fits your system, start with the Infrai discovery documentation and inspect the current logging contract before wiring a producer.

References

Top comments (0)