An Express production health check endpoint is useful only if it helps an operator make a safe decision during an incident. For a customer-support service running on Node.js, the real constraint is evidence retention: after a bad deployment, can the team distinguish liveness from readiness, find the relevant 5xx errors and logs, and explain which customer requests failed before and after rollback?
TL;DR: expose separate liveness and readiness endpoints in Express, keep both handlers cheap, and emit structured lifecycle and dependency-failure logs. Let a regional uptime service probe readiness from outside the cluster. Separately, run a scheduled poller that checks recent logs or error groups and routes notifications. A probe answers "can traffic reach this release now?"; retained evidence answers "what happened to customer requests?" You need both.
This separation matters for rollback safety. Liveness should say that the Node.js process can still serve its event loop. Readiness should say that this instance may receive customer traffic. A deployment controller can remove an unready instance without turning a transient dependency problem into a restart loop, while an operator can compare evidence from the new and previous releases before rolling back.
How should an Express production health check separate readiness and liveness?
A single green response collapses several different states. The process may be running while its database connection is unavailable. Conversely, a downstream SMS provider can be degraded while the support application remains able to accept a ticket, persist it, and retry delivery. Those states deserve different operational actions.
Use two small Express routes with deliberately narrow contracts. Liveness returns success when the process is responsive. Readiness returns success only when the dependencies required to accept new support work pass bounded checks. Include a release identifier and a timestamp in structured logs, not a dump of secrets or customer message bodies. The endpoints themselves should reveal as little internal topology as possible.
Keep readiness strict about required dependencies and explicit about optional ones. If Postgres is required to retain a ticket, a failed bounded check makes the instance unready. If outbound SMS can be queued durably for later delivery, an SMS-provider failure should not necessarily remove every instance from service. This is a compliance boundary as much as an availability choice: losing a notification is bad, but accepting a customer request without retaining its evidence can be worse.
Timeouts must be shorter than the probe interval. Cache a dependency result briefly if every probe would otherwise create a thundering herd. Do not perform migrations, repairs, or writes from a health handler. Boring is good here.
The minimum event vocabulary is also small: service_starting, service_ready, dependency_failed, service_draining, and service_stopped. Add release_id, service, environment, timestamp, and a correlation identifier where one already exists. A trace_id or span_id can correlate records, but fields alone do not provide a distributed trace query or span tree.
Design the evidence before the alert
An alert is a lossy summary. The retained record should be rich enough to reconstruct the customer incident after the page has stopped ringing.
For each support request, preserve the transition that matters: accepted, persisted, queued for delivery, delivered, or failed. Keep the customer identifier pseudonymous and avoid recording message content unless policy requires it. A useful error record identifies the operation, dependency, release, correlation ID, and failure class. It does not need an access token, OTP, phone number, or email address.
That distinction catches a common trap. A monitor that sees repeated HTTP 500 responses can page someone, yet it cannot establish whether a specific ticket was committed before the response failed. The application event can. During rollback, compare error groups and lifecycle events on both sides of the release boundary, then verify that the old release becomes ready before restoring full traffic.
Infrai fits the evidence-collection side when a team wants plain HTTP instead of another language-specific SDK. Its public discovery surface is self-describing: a capability response includes the request schema, response schema, billing information, and runnable examples, so integration starts by reading one endpoint. Every documented capability has runnable examples in 10 languages, and the same conventions cover 295 routes across 20 modules. In this workflow, that breadth means the poller and a later queue or notification integration do not each introduce a new SDK and credential model.
There is a firm limitation. For this use case, Infrai is an evidence store rather than a complete alerting system: there is no built-in threshold engine, webhook/SMS/phone notification routing, synthetic probing, or heartbeat monitor. Recent logs or error groups must be polled by a scheduled process. It is not suitable as the only monitoring product when the team needs managed escalation and regional probes; Better Stack or Datadog is the better fit for that job.
There are retention boundaries too. The logging surface does not expose per-user deletion, bulk export, or subscription APIs, and retention or cold-storage configuration is not exposed. That can be disqualifying when a support organization must automate erasure requests or archive evidence under its own retention policy. Source-map decoding, crash symbolication, Electron minidump processing, and session replay are outside this design as well.
Poll without creating a second incident
The poller below calls Infrai's verified log-search route with Bearer authentication, checks status codes, honors Retry-After on HTTP 429, and emits a JSON notification record when the returned collection is nonempty. INFRAI_BASE_URL must be the API origin supplied with the account; keeping it in configuration avoids placing a service URL or credential in source. The discovery parameters do not currently declare filters for log search, so the example does not invent time-window or severity query fields.
import json
import os
import random
import sys
import time
from datetime import datetime, timezone
from urllib import error, request
def retry_delay(response, attempt):
value = response.headers.get("Retry-After")
if value and value.isdigit():
return float(value)
return min(60.0, (2 ** attempt) + random.random())
def fetch_events(url, token, attempts=5):
for attempt in range(attempts):
req = request.Request(
url,
method="GET",
headers={
"Authorization": f"Bearer {token}",
"Accept": "application/json",
},
)
try:
with request.urlopen(req, timeout=10) as response:
return json.load(response)
except error.HTTPError as exc:
body = exc.read().decode("utf-8", errors="replace")
if exc.code == 429 and attempt + 1 < attempts:
time.sleep(retry_delay(exc, attempt))
continue
raise RuntimeError(f"observability query failed: {exc.code} {body}") from exc
raise RuntimeError("observability query exhausted its retry budget")
def main():
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
url = f"{base_url}/v1/logs/search"
token = os.environ["INFRAI_API_KEY"]
result = fetch_events(url, token)
events = result if isinstance(result, list) else result.get("data", [])
if events:
notification = {
"type": "support_service_evidence_detected",
"observed_at": datetime.now(timezone.utc).isoformat(),
"count": len(events),
}
print(json.dumps(notification))
return 2
return 0
if __name__ == "__main__":
sys.exit(main())
Run this from a scheduler at an interval that matches the service objective, and connect its nonzero result or JSON output to the organization's existing notifier. The scheduler must prevent overlapping runs. Store a cursor or stable event identifier after a successful notification so the next execution does not page on the same evidence again. Because this example performs no create or publish operation, it needs no idempotency key; a notifier added downstream does need an idempotent delivery key.
Polling has an unavoidable blind spot between runs.
It can also fail silently, which is why the poller's own expected schedule needs a dead-man's-switch service such as Healthchecks. External synthetic monitoring is a separate check again: use it to call readiness from the US or EU if those are customer regions. An in-cluster probe cannot establish that public DNS, TLS, routing, and the edge path work for a customer. This trade-off is easy to miss because all three mechanisms can produce a red status, but they observe different failure planes: the orchestrator sees a process, the external monitor sees the public route, and the evidence poller sees recorded application outcomes. Treating those signals as interchangeable makes rollback riskier, since an operator may restore traffic based on a healthy process while the dependency path remains unable to retain a support ticket.
Where do the real products differ?
Choose around the missing failure mode, not the longest feature list.
| Option | Best fit in this design | Boundary to account for |
|---|---|---|
| Better Stack Uptime | External HTTP checks and an operational alert path | Application events still need deliberate structured evidence and privacy controls |
| Datadog Synthetic Monitoring | Browser and API checks when a team already operates the Datadog stack | A broader platform brings more setup and governance surface than two basic probes |
| Sentry | Grouping application exceptions and investigating error events | It does not replace readiness semantics or an external availability probe |
| Healthchecks | Detecting that a scheduled poller or maintenance job did not run | It is a dead-man's switch, not a general log archive or browser monitor |
| Infrai | Self-described REST ingestion and searchable logs/error groups under one key | Polling and an external notifier are required; synthetic checks and heartbeat monitoring are absent |
These products overlap, but they are not interchangeable.
Better Stack or Datadog can observe the service from a customer-like network position. Sentry centers the exception investigation workflow. Healthchecks covers the quiet failure in which the poller never executes. Infrai provides a compact evidence API, with discovery and runnable examples reducing integration guesswork, but it should be paired with the missing probe and notification layers.
For a small team, that composition may be easier to reason about than adopting a full suite. For a regulated support operation needing automatic per-user deletion, bulk export, long-term archive controls, or end-to-end trace exploration, the boundaries point toward a platform with those controls. No amount of polling repairs a governance mismatch.
Roll out with a reversible decision rule
Start with liveness and readiness in shadow mode: record their decisions without letting readiness remove instances from traffic. Compare the decision stream with dependency-failure logs through at least one normal deployment and one controlled dependency interruption. Then enable readiness for a small slice of instances while keeping liveness limited to genuine process health.
Next, add the external regional check and the scheduled evidence poller. Page only on conditions that have an owner and a documented response. Three consecutive readiness failures may justify removing an instance; a single structured dependency_failed event may instead open an investigation. The exact threshold belongs to the service objective and traffic pattern, so copying an arbitrary number would create false confidence.
Finally, rehearse rollback. Mark the release boundary, verify that the previous version becomes ready, drain the new version, and confirm that accepted support work still has a terminal evidence event. Rollback is complete when traffic is healthy and the customer record is explainable. A green process alone is not enough.
Top comments (0)