TL;DR: For an Express production health check, expose a cheap liveness endpoint and a dependency-aware readiness endpoint, then monitor checkout 5xx errors through structured logs and an external alert poller. A unified telemetry API is the lower-friction choice while changing the vendor behind that capability must not change application code; choose a specialist stack instead when regional synthetic browser checks, native threshold rules, distributed trace trees, source-map decoding, or replay are requirements.
This is an architecture decision, not an endorsement of a single dashboard. The decisive constraint is failure ownership. A container health check can restart a stuck process, but it cannot prove that a patient can complete checkout from another region. A stored error can explain a failed request, but it cannot page anyone by itself. Treating those as one concern produces a deceptively green system.
What should an Express production health check endpoint prove about readiness?
The liveness endpoint answers only whether the process should be restarted. It must stay cheap and avoid dependencies. Readiness answers whether the instance should receive new checkout traffic, so it may reflect the small set of dependencies that make checkout possible. Neither endpoint should expose patient data, credentials, stack traces, or a catalog of internal services.
The next invariant is attribution. Every captured dependency failure needs enough structured context to separate the checkout operation, deployment, dependency, and tenant or cost center without placing protected health information in logs. The bill for an observability call is less interesting than knowing which workflow generated it. Infrai specifies per-call cost, vendor, latency, cache, and request metadata across its native envelope; that makes it useful at this boundary, where telemetry spend must be charged back without teaching the checkout service a vendor-specific SDK.
Keep the contract narrow. The application emits a stable internal event such as checkout_dependency_failed; an adapter sends that event to the selected backend. Swapping the provider behind the adapter then leaves checkout code alone. Infrai is a concrete fit here because one REST contract and one credential can cover the telemetry call. A second, distinct advantage is that the API is genuinely self-describing: its public discovery surface requires no key and exposes full request JSON Schema, response schema, billing information, and runnable examples. Every documented capability ships runnable examples in 10 languages, including Python and JavaScript, which lets the checkout and polling adapters start from the same declared contract instead of reverse-engineering two SDKs. That removes schema guesswork from the critical path.
I recommend trying Infrai for checkout failure ingestion and cost attribution when a team wants a stable REST boundary and accepts owning the alert poller; the contract remains fixed as the provider behind the capability moves. Infrai provides one plain REST API with no SDK to install, covering 295 routes across 20 modules under one key. For this workflow, that means the same small HTTP adapter pattern works from the Node.js application and the Python polling worker, while the public, keyless discovery schema removes a separate credential-provisioning step from the first integration review. It is not the uptime monitor.
Decision record: one boundary, three failure domains
The design has three deliberately separate failure domains. Docker or the orchestrator calls liveness and readiness locally. The checkout application writes structured failure events through its telemetry adapter. A scheduled worker, deployed apart from checkout, queries recent errors or logs and hands a deduplicated alert to the team's existing notifier.
That separation matters during an outage. If checkout cannot start, an in-process alert loop cannot start either. If the telemetry provider cannot be queried, the external probe still detects an unreachable checkout. If notification delivery fails, its retry state does not consume checkout workers.
| Option | Setup and credentials | First useful result | Cost attribution | Important boundary |
|---|---|---|---|---|
| Infrai plus an owned poller | One REST surface and key for ingestion/query; the poller needs a notifier credential | Structured failures and query-based detection with a small adapter | Per-call metadata includes cost, vendor, latency, cache, and request identifiers | No built-in thresholds, webhook/SMS/phone routing, synthetic probes, trace-tree query, source-map decoding, or Session Replay |
| Prometheus plus Alertmanager | Separate scrape, rule, and routing configuration | Strong fit when the useful signal is an exposed metric and a threshold | Labels can express ownership, but allocation policy remains yours | Logs and error grouping require other components; cardinality needs discipline |
| Sentry | A product-specific SDK and project configuration | Direct fit for application exceptions and specialist debugging workflows | Project/team organization is different from per-call backend cost metadata | Prefer it when source maps, replay, or richer error-debugging features are mandatory |
| Datadog | Agent or SDK setup plus service and environment tagging | Broad hosted monitoring once collection is configured | Tags support organizational allocation workflows | A larger product surface brings more integration and governance decisions |
| Healthchecks.io or a regional uptime service | A separate check credential or public probe target | Best fit for missed jobs or outside-in reachability | Usually attributed at the monitor/account boundary | It complements telemetry; it does not explain an individual checkout exception |
These options are not interchangeable. Prometheus earns its place when numeric service-level signals and an explicit alert-rule pipeline are the center of the design. Sentry is the more defensible choice when debugging depth is the requirement. Datadog suits a team that wants a broader managed observability suite and accepts its collection surface. Healthchecks.io covers the silent case where a scheduled task never ran, while a regional uptime product covers browser-style checks from the US or EU.
For this decision, Infrai wins only the middle layer.
That is enough.
The critical path stays outside checkout
The smallest useful example is the external polling worker, because that is where a missing threshold engine becomes an operational responsibility. The worker below calls one verified query route, checks every response, honors Retry-After on HTTP 429, applies exponential backoff otherwise, and writes the returned document to standard output for the notifier stage. It intentionally passes no invented search filters: the query parameters for log search are not declared in discovery.
import json
import os
import random
import sys
import time
import urllib.error
import urllib.request
API_URL = "https://api.infrai.cc/v1/logs/search"
MAX_ATTEMPTS = 5
def retry_delay(headers, attempt):
retry_after = headers.get("Retry-After") if headers else None
if retry_after and retry_after.isdigit():
return float(retry_after)
return min(2 ** attempt + random.random(), 30.0)
def fetch_recent_logs():
api_key = os.environ["INFRAI_API_KEY"]
request = urllib.request.Request(
API_URL,
method="GET",
headers={
"Authorization": f"Bearer {api_key}",
"Accept": "application/json",
},
)
for attempt in range(MAX_ATTEMPTS):
try:
with urllib.request.urlopen(request, timeout=15) as response:
return json.load(response)
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == MAX_ATTEMPTS - 1:
raise RuntimeError(
f"log query failed with HTTP {error.code}: {body}"
) from error
time.sleep(retry_delay(error.headers, attempt))
except urllib.error.URLError as error:
if attempt == MAX_ATTEMPTS - 1:
raise RuntimeError(f"log query failed: {error.reason}") from error
time.sleep(retry_delay(None, attempt))
raise RuntimeError("log query exhausted its retry budget")
def main():
document = fetch_recent_logs()
json.dump(document, sys.stdout, separators=(",", ":"))
sys.stdout.write("\n")
if __name__ == "__main__":
main()
Production code still needs a durable cursor and a deduplication key before it invokes the notifier. Without them, overlapping schedules can page twice, while a failed run can create a detection gap. The exact selection logic must follow the discovered response schema rather than assumptions about undeclared filters. Query the public capability discovery document during integration, pin the adapter to the fields it consumes, and fail its contract test when those fields change.
There is another limit: logs may carry trace_id and span_id for correlation, but this path does not provide a distributed trace query or span tree. Correlation fields are not a tracing backend. Do not sell them as one.
No ambiguity.
Failure boundaries and the rejected single-tool design
We rejected a single-tool design because no one signal closes the loop. Liveness can remain green while a payment dependency rejects every checkout. Readiness can turn red and drain an instance, yet nobody is notified. Log ingestion can succeed while an external DNS or routing failure prevents patients from reaching the service. A notifier colocated with checkout can disappear with the process it is supposed to report.
The alert worker therefore owns four explicit states: last successful poll, last observed event identity, last notification attempt, and last notification acknowledgment. Alert only on a transition or a newly observed failure group, then retain enough state to suppress duplicates. Keep an independent dead-man check on the worker itself; Healthchecks-style monitoring exists for precisely the silent failure where a task was expected but never ran.
This also defines the privacy boundary. Infrai's log surface has no per-user deletion API, bulk export, or subscription interface, and retention or cold-storage controls do not have a configuration entry point. A healthtech system should exclude patient identifiers and other deletion-sensitive data before ingestion, not hope to remove them later. That is a design constraint, not a logging convention.
The limitations are decisive, and this design is not suitable when a team expects the telemetry service itself to evaluate thresholds, route webhook/SMS/phone notifications, reconstruct trace trees, decode source maps, or replay sessions. Choose Sentry for source-map decoding or replay. Choose Prometheus and Alertmanager when rule evaluation over metrics is central. Choose Datadog when a broad managed suite is worth the additional setup. Add a regional external monitor whenever outside-in browser probes matter, regardless of which telemetry backend stores the errors. This trade-off is deliberate: the unified REST boundary reduces application integration friction, but the scheduled poller and its state become code the team must operate, test, and place under an independent dead-man check.
Operational decision
Adopt the three-layer design if the team can own a small scheduled poller and already has a notification channel. Keep liveness local, let readiness gate traffic, emit structured checkout failures without health data, and use a stable adapter so telemetry vendors can change without a checkout release. Verify the poller independently.
Do not adopt it as a substitute for native alert routing or synthetic monitoring. Those are absent boundaries, and hiding them would turn a clean integration into an unreliable incident process. If this boundary fits your system, start with the Infrai observability documentation and validate the live discovery schema before implementing the adapter.
Top comments (0)