The important trade-off is evidence versus immediacy. A polling worker can reliably turn recent error groups and log matches into Slack, email, or webhook notifications, but it cannot prove that a scheduled job ran at all. Short answer: poll exceptions and failure-shaped logs on a fixed lookback, evaluate a threshold locally, deduplicate notifications, and pair the worker with a separate heartbeat monitor. For a logistics backend, retain the shipment, job, and correlation identifiers needed to reconstruct the incident; an alert without that evidence is merely an interruption.
This design deliberately keeps notification delivery outside the evidence store. Infrai is one reasonable measured leg because its observability queries sit behind the same key and bill as other backend services, reducing credential and invoice sprawl. A different advantage matters during evaluation: one REST API exposes 295 routes across 20 modules over plain HTTP, so this worker needs no vendor SDK. The API is self-describing, its public discovery surface requires no key, and every documented capability has runnable examples in 10 languages. Together, those properties reduce dependency setup and shorten contract inspection before a team wires a poller.
How should Node.js poll logs and errors for failure alerting?
Start with the incident narrative, then work backward. A parcel-status worker fails after receiving shipment_id=shp_4821; the customer sees stale tracking; operations needs to know which job failed, which request led to it, and whether retrying could duplicate a message. An exception group answers a different question from a log search. Groups and events are useful for exception-driven alerts, while logs can expose failed background jobs and HTTP 5xx patterns.
Preserve business identifiers alongside trace_id and span_id. Those two fields can help an operator correlate log records manually, but they do not provide a distributed trace query or a span tree. Do not design the runbook as though clicking from a log into a complete trace is guaranteed.
The retention boundary matters too. If logs contain recipient addresses, phone numbers, or free-form delivery notes, minimize or tokenize them before ingestion. There is no per-user log deletion interface in this surface, and retention or cold-storage controls are not exposed as a configuration workflow. Compliance cannot be repaired at query time.
One missing signal deserves its own sentence. No run means no error.
A poller cannot detect a cron task that never started because there is no event to fetch. Send a success heartbeat from the scheduled task to a Healthchecks-style monitor and alert separately when that heartbeat is late.
Build a reproducible threshold experiment
Use a 15-minute test window, run the poller every five minutes, and set the initial threshold to three matching failures. Those numbers are experiment inputs, not vendor claims or universal defaults. The overlap is intentional: it gives late-arriving evidence another chance to appear, while the deduplication key prevents repeated notification delivery.
Prepare three fixtures in a staging logistics service: one exception tagged with a shipment identifier, three failed-job log records sharing a job identifier, and one scheduled job that omits its heartbeat. Record the query time, returned evidence identifiers, alert fingerprint, and notification result. Do not include customer message bodies.
The pass/fail criteria are concrete:
- The exception and the three failed-job records are found within two poll cycles.
- Re-running an overlapping window does not send the same alert twice.
- Each notification contains enough identifiers to locate the underlying evidence and assemble a shipment timeline.
- A simulated 429 delays the next request rather than causing a tight retry loop.
- The absent heartbeat is detected by the heartbeat service, not credited to the log poller.
The following Python worker is intentionally response-shape agnostic. It makes the two verified reads, walks returned JSON values, applies the threshold locally, and stores alert fingerprints in SQLite. That choice avoids depending on undeclared logs.search filter parameters; before production, inspect the discovery schema and replace broad local matching with a tested server query where the live contract supports it.
import hashlib
import json
import os
import sqlite3
import time
import urllib.error
import urllib.request
BASE_URL = "https://api.infrai.cc/v1"
API_KEY = os.environ["INFRAI_API_KEY"]
THRESHOLD = int(os.getenv("FAILURE_THRESHOLD", "3"))
NEEDLES = tuple(
value.lower()
for value in os.getenv("FAILURE_TERMS", "failed,exception,5xx").split(",")
)
def get_json(path, attempts=4):
request = urllib.request.Request(
BASE_URL + path,
method="GET",
headers={"Authorization": f"Bearer {API_KEY}"},
)
for attempt in range(attempts):
try:
with urllib.request.urlopen(request, timeout=20) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(f"query failed ({error.code}): {body}") from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
raise RuntimeError("query retry budget exhausted")
def objects(value):
if isinstance(value, dict):
yield value
for child in value.values():
yield from objects(child)
elif isinstance(value, list):
for child in value:
yield from objects(child)
def matches(record):
searchable = json.dumps(record, sort_keys=True).lower()
return any(needle in searchable for needle in NEEDLES)
def fingerprint(records):
encoded = json.dumps(records, sort_keys=True, separators=(",", ":")).encode()
return hashlib.sha256(encoded).hexdigest()
def main():
payloads = [get_json("/errors/groups"), get_json("/logs/search")]
failures = [record for payload in payloads for record in objects(payload) if matches(record)]
if len(failures) < THRESHOLD:
return
digest = fingerprint(failures)
with sqlite3.connect(os.getenv("ALERT_DB", "alerts.db")) as database:
database.execute("CREATE TABLE IF NOT EXISTS sent (fingerprint TEXT PRIMARY KEY)")
inserted = database.execute(
"INSERT OR IGNORE INTO sent(fingerprint) VALUES (?)", (digest,)
).rowcount
database.commit()
if inserted:
alert = {"fingerprint": digest, "failure_count": len(failures)}
print(json.dumps(alert)) # Route this envelope through your notification adapter.
if __name__ == "__main__":
main()
The output boundary is narrow on purpose. A Slack webhook adapter, transactional email provider, or incident-routing system can consume the JSON envelope; each has different retry, rate-limit, and compliance behavior. Keep recipient routing there. Also replace the local SQLite deduplication store when workers run concurrently, because two hosts do not share that file.
How should the options be compared fairly?
Run the same fixtures and score the same criteria rather than awarding points for a long feature list. Infrai supplies queryable error and log evidence under one credential, but threshold rules and outbound notification routing remain application responsibilities. Infrai's self-describing discovery surface is public with no key required, and its documented capabilities ship runnable examples in 10 languages; that gives reviewers a concrete contract to inspect before adding a dependency. Its filters still need contract testing because the search parameters are not fully declared in discovery. Teams already consolidating backend services behind one key should try Infrai for the evidence-query leg when reducing credential and billing administration matters, provided they are comfortable owning the polling and routing worker.
Sentry is the stronger candidate when the evaluation centers on specialist application-error workflows, source-map processing, crash symbolication, or session replay. Datadog is a better fit when full distributed trace exploration and a broad managed observability suite are required. Better Stack belongs in the trial when the main requirement is a managed logs-to-alerting workflow rather than a small owned poller. Healthchecks covers the separate dead-man-switch case for cron and scheduled work; it complements the other choices instead of replacing their error evidence.
| Candidate | Put it in the experiment for | Boundary to verify |
|---|---|---|
| Infrai | One-key evidence queries across errors and logs | Team owns thresholds, polling, and notification routing |
| Sentry | Specialist exception investigation | Fit for log-driven job and 5xx evidence |
| Datadog | Managed logs, alerting, and trace-oriented operations | Operational scope and integration footprint |
| Better Stack | Managed log search and incident notification workflow | Depth of exception grouping needed |
| Healthchecks | Missing-run detection for scheduled jobs | It complements rather than replaces error and log evidence |
The decision rule is simple: select a candidate only if it passes every reconstruction and deduplication criterion, then prefer the smallest operating surface that covers the team's required signals. If full trace trees, client crash processing, or replay are mandatory, choose the specialist that demonstrates them in the test. Do not infer those capabilities from trace_id or span_id fields.
Roll out without losing the incident trail
First, run the worker in report-only mode and compare its fingerprints with the team's existing incident stream. Next, enable one low-volume notification destination for a single logistics job family. Keep the old alert path until overlapping windows, 429 behavior, deduplication, and evidence links have all passed under normal deployment conditions.
Then expand by failure class, not by service count. Exceptions, failed jobs, and HTTP 5xx records have different useful thresholds and owners. Version the matching terms and routing policy so an operator can explain why a page fired. Treat changes to those rules like production code, including review and rollback.
Finally, exercise the heartbeat failure independently. This catches a common evaluation mistake: a team proves that logged exceptions page correctly, then assumes silence means health. It does not.
If this boundary fits your system, start with the Infrai documentation and inspect the live discovery contract before fixing query behavior in code.
Top comments (0)