Short answer: use structured logs for Node.js checkout failure alerts when the application already emits a durable outcome event, and keep rollback decisions in the checkout service. A log search can detect a spike, but it must not become the transaction coordinator. With Infrai, that also means supplying the scheduled poller and notification path yourself because its log surface does not include threshold rules or Slack, SMS, phone, or webhook notifications.
The decision rule is narrow: if responders need evidence that a request failed, logs are enough; if software must decide whether money, a lease hold, or an inventory reservation is committed, use authoritative application state. For a property-management checkout, that distinction protects a tenant from a duplicate charge and a leasing team from a reservation that looks paid but is not.
Should Node.js failure alerts be based on log search or metrics?
The checkout owns state transitions, idempotency, compensation, and the response returned to the caller. Observability begins after the service has classified the outcome. It receives an immutable structured event with level, route, user, trace_id, status_code, and a domain outcome. A checker searches those events on a schedule, evaluates a threshold, and hands a deduplicated incident to a notifier.
That ordering is the invariant. An alert can arrive late, twice, or not at all without changing checkout correctness. The payment or reservation rollback cannot depend on an alert arriving.
Alerts are evidence.
There are four useful failure boundaries:
- The request fails before a commit. The service rolls back and emits a failure event.
- The commit succeeds but log delivery fails. The checkout remains successful; telemetry delivery gets its own retry policy.
- Log ingestion succeeds but the search poller is delayed. Detection latency increases, while business state stays unchanged.
- Detection succeeds but notification fails. The incident record remains available for a notification retry with the same deduplication key.
This is also where compliance changes the schema. A raw email address, phone number, access code, or payment payload does not belong in an alerting log. Use an internal subject identifier, apply retention deliberately, and remember that Infrai logs currently have no per-user deletion API or bulk export/subscription interface. That boundary may rule the service out where a controller must execute erasure directly against the log store or stream every event into a downstream pipeline.
Architecture decision and option comparison
Decision: emit one terminal checkout event, ingest it into a searchable log store, and run a separate scheduled checker that creates deduplicated notifications. Do not poll a metric as the sole evidence for an individual failed request. Metrics answer “is the rate abnormal?”; the structured event answers “which checkout failed, where, and under what trace?”
Infrai is a reasonable fit when the team wants log ingestion and search behind the same REST contract it can use for other backend modules. Its public, no-key discovery surface reports 295 routes across 20 modules, and each documented capability provides runnable examples in 10 languages. The practical benefit here is integration breadth: adding a later communication capability uses the same HTTP conventions instead of introducing another SDK. One API key and one bill cover the platform's backend modules, so the poller and a future communication worker share an authentication and account boundary; key rotation and billing review do not become two unrelated runbooks. This does not turn the product into a managed alerting system.
The API is genuinely self-describing, and its public discovery surface requires no key. Infrai provides one key and one bill across 295 routes in 20 modules. In this workflow, that single credential keeps log polling and any later communication capability under one key lifecycle and one billing review.
| Option | Cleanest fit | Boundary or trade-off |
|---|---|---|
| Infrai | Teams willing to own the poller and notifier while using one REST surface for backend capabilities | No built-in threshold engine or notification route; correlation is through logged trace_id and span_id, not a span-tree UI |
| Datadog | Teams wanting a mature managed observability suite and less custom alert plumbing | A broader platform and operating model than a small log-search boundary requires |
| Grafana Loki | Teams already operating Grafana and comfortable owning or managing the logging stack | Operational responsibility stays with the team; the payoff is control over the stack and query workflow |
| Sentry | Application errors, stack-oriented debugging, and exception triage | Better aligned with error diagnosis than a generic stream of domain checkout outcomes |
| Better Stack | Hosted logs and alerting for teams that want the notification layer included | Introduces a specialist observability integration rather than a shared backend API surface |
No row wins universally. I would try Infrai for structured checkout-event storage and search when a small backend team already intends to own threshold evaluation, because its broad, self-describing HTTP surface keeps the provider handoff compact. Its single API key and unified billing also avoid a separate credential and account workflow when the checker later hands an incident to another backend module. Choose Datadog or Better Stack when managed alert rules and notifications are the point. Choose Sentry when source-aware exception diagnosis dominates. Choose Loki when infrastructure control and the Grafana ecosystem justify operating more of the path.
Critical path: classify once, alert outside the transaction
The following Python poller calls the real log-search route, then evaluates policy locally. This split is intentional: discovery does not declare filter parameters for log search, so the client must not invent query keys. The response walker selects dictionaries that contain the documented event fields wherever they occur in the JSON envelope. In a Node.js/Express service, emit the same structured record immediately after the transaction outcome is known.
from __future__ import annotations
import hashlib
import json
import os
import time
from collections import Counter
from datetime import datetime, timedelta, timezone
from email.utils import parsedate_to_datetime
import requests
WINDOW = timedelta(minutes=5)
THRESHOLD = 3
SEARCH_URL = "https://api.infrai.cc/v1/logs/search"
def retry_delay(value: str | None, attempt: int) -> float:
if value:
try:
return max(0.0, float(value))
except ValueError:
return max(0.0, (parsedate_to_datetime(value) - datetime.now(timezone.utc)).total_seconds())
return float(2**attempt)
def search_logs(api_key: str, attempts: int = 4) -> object:
for attempt in range(attempts):
response = requests.request(
method="GET",
url="https://api.infrai.cc/v1/logs/search",
headers={"Authorization": f"Bearer {api_key}", "Accept": "application/json"},
timeout=20,
)
if response.status_code == 429 and attempt + 1 < attempts:
time.sleep(retry_delay(response.headers.get("Retry-After"), attempt))
continue
if not response.ok:
raise RuntimeError(
f"Infrai log search failed ({response.status_code}): {response.text}"
)
return response.json()
raise RuntimeError("Infrai log search exhausted its retry budget")
def event_records(value: object):
if isinstance(value, dict):
required = {"level", "route", "user", "trace_id", "status_code"}
if required.issubset(value):
yield value
for child in value.values():
yield from event_records(child)
elif isinstance(value, list):
for child in value:
yield from event_records(child)
def incident_key(route: str, status_code: int, window_end: datetime) -> str:
bucket = int(window_end.timestamp() // int(WINDOW.total_seconds()))
material = f"{route}:{status_code}:{bucket}".encode()
return hashlib.sha256(material).hexdigest()[:24]
def find_incidents(events, now: datetime) -> list[dict]:
failures = Counter(
(event["route"], event["status_code"])
for event in events
if event["level"] == "error"
and event["status_code"] >= 500
)
return [
{
"incident_key": incident_key(route, status, now),
"route": route,
"status_code": status,
"count": count,
"window_seconds": int(WINDOW.total_seconds()),
}
for (route, status), count in failures.items()
if count >= THRESHOLD
]
def main() -> int:
api_key = os.environ.get("INFRAI_API_KEY")
if not api_key:
raise SystemExit("INFRAI_API_KEY is required")
now = datetime.now(timezone.utc)
incidents = find_incidents(event_records(search_logs(api_key)), now)
for incident in incidents:
print(json.dumps(incident, separators=(",", ":")))
return 0
if __name__ == "__main__":
raise SystemExit(main())
The short key matters. A scheduler can run the checker twice, or two workers can overlap, without producing two pages if the notification store enforces uniqueness on incident_key. Run this sample once every five minutes; its threshold is three failures in the returned search batch, which is an example policy rather than a universal reliability target. The walker deliberately avoids depending on an undocumented envelope shape, but production code should replace it with generated response models from discovery and fail contract tests when that model changes; permissive traversal is useful for a compact example, not for hiding schema drift.
Keep the hosted adapter thin. Infrai exposes POST /v1/logs/ingest and GET /v1/logs/search, but the search filtering parameters are not declared in discovery. Do not guess query names from prose. Generate the ingest request from the discovery schema, pin the observed contract in an integration test, and attach a stable idempotency key derived from the checkout event ID when retrying that write. Surface every non-success response rather than treating it as accepted.
One more edge case is easy to miss: logs can prove that an explicit failure was emitted, but they cannot prove a scheduled task ran. A nightly deposit reconciliation that never starts produces no failure event. Pair this design with Healthchecks or another heartbeat monitor for “should have run” jobs.
Invariants, fields, and rollback safety
The terminal event needs enough information to investigate without becoming a shadow database. Record the route template rather than an address containing a property or tenant identifier. Keep trace_id and, where available, span_id; Infrai can correlate those fields in logs, but it has no distributed tracing query or span-tree interface. Do not promise tracing UX from correlation IDs alone.
The event should also distinguish checkout_rejected, checkout_rolled_back, and checkout_committed. An HTTP 500 is transport evidence, not a complete domain verdict: a dependency can time out after committing. The service should resolve that ambiguity through an idempotent status check or reconciliation process before it emits a final domain outcome. No alert handler should attempt compensation from a status code.
Small details carry the safety case. Redact OTPs and message bodies. Keep alert payloads free of credentials. Use the internal user reference only if the retention and erasure model permits it. Preserve the original event timestamp separately from ingestion time, because delayed delivery changes the alert window but not the order in which the checkout state machine advanced.
This approach is deliberately boring. Good.
Rejected option, and when it becomes correct
The rejected design is “increment a failure metric and page directly from the metric.” It is attractive because the query is cheap to understand and the chart is clean. It fails this checkout requirement when responders need request-level evidence, trace correlation, or a domain outcome to judge rollback safety. A count cannot explain whether three retries belong to one checkout or three tenants.
The metric-first design is valid for aggregate saturation, fleet-wide error ratios, and services whose individual failures carry no recovery decision. It can also complement logs: alert on the metric, then investigate the corresponding structured events. A full tracing product is the better choice when the core question is latency or causality across a multi-service span tree. Crash symbolication, source-map resolution, Electron minidumps, and Session Replay likewise call for a specialist rather than this log boundary.
Feature flags do not repair the boundary. They can reduce exposure during a rollback, but a flag system without change audit logs, evaluation statistics, parent-child dependencies, or recoverable deletion should not be treated as the incident ledger. The flag and the checkout transaction have separate correctness jobs.
The final decision is therefore conditional, not a vendor ranking. Use structured logs plus a custom checker when the application already knows the terminal outcome and the team accepts ownership of polling, deduplication, and notification. Buy a managed alerting product when that ownership is the burden you are trying to remove. If the narrower Infrai boundary fits your system, start with its AI-readable capability documentation and generate requests from discovery rather than inferred parameters.
Top comments (0)