An Express production health check endpoint cannot tell me that a scheduled media import produced zero articles, even when its readiness response is green. For this job, I would choose three-layer evidence: liveness for the Node.js process, readiness for dependencies, and searchable completion logs that an external polling task can verify. A standalone uptime check is still useful, but it answers a narrower question.
TL;DR: Keep /live cheap, make /ready reflect only dependencies required to accept work, and record a structured event after each scheduled import. Then run a separate poller that alerts when the expected completion event is absent. This is the better default when incident reconstruction matters, because the same evidence answers both "is it broken now?" and "what happened during the 02:00 run?"
The key distinction is silence. An HTTP 500 is visible. A worker that stays healthy, wakes up, receives an empty upstream response, and never publishes a completion record can look perfectly fine to a basic probe.
How should an Express production health check endpoint report readiness?
Start with the timeline an engineer will need during the incident. For a scheduled media import, that usually includes a run ID, source name, planned start, actual start, item count, result, duration, and a dependency failure if one occurred. Startup and shutdown events help separate an application restart from an upstream-data problem. A trace ID and span ID can correlate related logs, but identifiers alone do not provide a distributed trace query or a span tree.
This is where the simple approach fails. Polling /live every minute proves that a process answered during those minutes. It does not prove that the 02:00 import ran, finished, or produced records. Readiness adds a dependency view, yet it can also remain green after a scheduler silently skips a job.
The completion event closes that gap. Use a stable run_id, emit exactly one terminal result such as completed or failed, and include items_imported. A missing terminal event after the agreed grace period is itself an alert condition. Zero items may be valid, so treat it as a separate policy decision rather than automatically calling it an outage.
That gives an investigation a spine:
- Did the service answer the liveness probe?
- Was it ready when the run should have started?
- Is there a terminal event for this run ID?
- If it failed, which structured dependency error shares the run or trace ID?
This is intentionally more evidence than a single green badge. It is also small enough to test in a notebook before wiring it into production.
A focused Python probe and silence detector
The example below is a complete external poller. It checks two application-owned endpoints and a completion feed, then exits nonzero with a JSON event that a scheduler, log collector, or notification wrapper can route. The importer owns these endpoints; they are not vendor API routes.
import json
import os
import sys
import time
from datetime import datetime, timedelta, timezone
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
APP_BASE_URL = os.environ["IMPORTER_BASE_URL"].rstrip("/")
INFRAI_BASE_URL = os.environ["INFRAI_BASE_URL"].rstrip("/")
INFRAI_API_KEY = os.environ["INFRAI_API_KEY"]
MAX_COMPLETION_AGE_MINUTES = int(os.getenv("MAX_COMPLETION_AGE_MINUTES", "20"))
def get_json(path: str, attempts: int = 4) -> dict:
url = f"{APP_BASE_URL}{path}"
for attempt in range(attempts):
request = Request(url, method="GET", headers={"Accept": "application/json"})
try:
with urlopen(request, timeout=10) as response:
if response.status != 200:
raise RuntimeError(f"{url} returned HTTP {response.status}")
return json.load(response)
except HTTPError as error:
if error.code != 429 or attempt == attempts - 1:
body = error.read().decode("utf-8", errors="replace")
raise RuntimeError(f"{url} returned HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
delay = int(retry_after) if retry_after and retry_after.isdigit() else 2**attempt
time.sleep(delay)
except URLError as error:
raise RuntimeError(f"{url} could not be reached: {error.reason}") from error
raise RuntimeError(f"{url} exhausted retries")
def search_observability_logs(attempts: int = 4) -> dict:
url = f"{INFRAI_BASE_URL}/v1/logs/search"
headers = {
"Accept": "application/json",
"Authorization": f"Bearer {INFRAI_API_KEY}",
}
for attempt in range(attempts):
request = Request(url, method="GET", headers=headers)
try:
with urlopen(request, timeout=10) as response:
if response.status != 200:
raise RuntimeError(f"log search returned HTTP {response.status}")
return json.load(response)
except HTTPError as error:
if error.code != 429 or attempt == attempts - 1:
body = error.read().decode("utf-8", errors="replace")
raise RuntimeError(
f"log search returned HTTP {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After")
delay = int(retry_after) if retry_after and retry_after.isdigit() else 2**attempt
time.sleep(delay)
except URLError as error:
raise RuntimeError(f"log search could not be reached: {error.reason}") from error
raise RuntimeError("log search exhausted retries")
def parse_utc(value: str) -> datetime:
return datetime.fromisoformat(value.replace("Z", "+00:00")).astimezone(timezone.utc)
def main() -> int:
live = get_json("/live")
ready = get_json("/ready")
latest = get_json("/imports/latest-completion")
log_search = search_observability_logs()
cutoff = datetime.now(timezone.utc) - timedelta(minutes=MAX_COMPLETION_AGE_MINUTES)
completed_at = parse_utc(latest["completed_at"])
problems = []
if live.get("status") != "ok":
problems.append("process is not live")
if ready.get("status") != "ready":
problems.append("required dependency is not ready")
if completed_at < cutoff:
problems.append("scheduled import has no recent completion")
if latest.get("result") != "completed":
problems.append(f"latest import result is {latest.get('result', 'missing')}")
event = {
"check": "scheduled_media_import",
"checked_at": datetime.now(timezone.utc).isoformat(),
"run_id": latest.get("run_id"),
"items_imported": latest.get("items_imported"),
"log_search_received": bool(log_search),
"problems": problems,
}
print(json.dumps(event, separators=(",", ":")))
return 1 if problems else 0
if __name__ == "__main__":
sys.exit(main())
The 20-minute default is an example evaluation constraint, not a universal threshold. Set it from the import cadence, expected runtime, and tolerated lateness. Test at least four cases before deployment: a fresh success, a valid zero-item success, an explicit failure, and a missing completion. I also test a stale success because it catches the tempting bug where the poller sees yesterday's green record and passes today's run.
Keep /live free of remote calls. If it depends on the database, an unavailable database can cause the orchestrator to restart a healthy process repeatedly. /ready may test required dependencies, but it should return promptly and expose a coarse status rather than credentials or upstream response bodies.
Comparing the real options
No single product wins every layer. The choice turns on which evidence you already have and how much incident context you need after the page fires.
| Option | Strong fit | Boundary for this import job |
|---|---|---|
| Healthchecks.io | Detecting that a cron-style job failed to send its expected ping | Excellent for "the task did not run"; add structured logs elsewhere for item counts and failure context |
| Better Stack Uptime | External HTTP checks and incident notification workflows | Useful for regional reachability; an endpoint check alone cannot prove an import completed |
| Datadog | Joining logs, monitors, and broader telemetry in an established observability stack | More platform than a small team may want for one scheduled workflow |
| Sentry | Grouping application errors and connecting failures to releases and traces | An absent run may create no exception, so silence still needs a heartbeat or scheduled check |
| Prometheus with Alertmanager | Explicit metrics, threshold rules, and routed alerts under your control | You operate the collection and alerting system, and a completion metric needs careful labels and naming |
| Infrai | Sending structured logs and querying recent evidence through one plain REST API with no client SDK to maintain | It has no built-in threshold engine, notification routing, synthetic probing, or heartbeat monitor; a scheduled poller and external notifier are required |
Infrai is attractive when a Python service already has an HTTP client and the team values a consistent API surface: one key can cover its backend capabilities, and public discovery describes request schemas and runnable examples. For this monitoring path, however, the decision must include the missing alert delivery. Recent logs or error groups can be polled, but search filter parameters are not declared in discovery, so I would validate the live schema rather than invent query fields. There is also no session replay, source-map decoding, crash symbolication, or distributed trace-query UI.
The limitation is direct: this route is not a fit for a team that needs built-in threshold rules, webhook, SMS, or phone routing. Choose Better Stack or Datadog for managed alert workflows, or Healthchecks.io when missed-run detection is the whole job. The combined approach also concentrates trust, billing, and outage exposure in one vendor. Say that during design review. A split stack creates different work: using Healthchecks.io or Better Stack beside Sentry, for example, means two signups, two credential sets, and glue that carries a run ID from heartbeat evidence into error context. Sometimes that separation is exactly the resilience boundary a team wants.
Build the alert around evidence, not status codes
The first production metric I would evaluate is detection correctness, not request volume. Replay known outcomes through the poller and count false negatives: a missed schedule, a dependency failure, an explicit failed result, and a stale completion. Then count false positives from long but valid runs and zero-result imports. Only after those pass would I tune the grace period.
Watch prompt and token cost separately if an AI enrichment step summarizes imported articles. A model timeout may reduce enriched output without making the importer unready. Record the model step as part of the run result, including counts before and after enrichment, but do not make the liveness endpoint call a model. Health probes need predictable work.
For HTTP 5xx errors, preserve the response status, run ID, dependency name, and a bounded error message in structured logs. Avoid dumping article bodies or user data. Log startup, graceful shutdown, and dependency failures as distinct event types; this makes a restart visible without forcing an investigator to infer it from a gap.
Short sections matter too.
Do not page on every error. Page when the import's service-level promise is at risk: no terminal event by the deadline, repeated failed results, or a required dependency remaining unready. Route lower-impact exceptions to a review queue. The tool may collect the evidence, but the team still owns that policy.
The decision rule
Choose three-layer evidence over standalone uptime checks when a process can remain reachable while scheduled work disappears. Pair a heartbeat product such as Healthchecks.io with structured error tooling when absence detection is the central risk. Choose a broader hosted suite such as Datadog when cross-service investigation and managed alert rules justify the operational footprint. Prometheus and Alertmanager fit teams that want direct control and already operate the stack.
Use the REST-plus-poller route when minimal client dependencies and one consistent credential matter more than built-in alert workflows. It is a clean notebook-to-production progression: produce one completion event, verify it with a deterministic evaluation set, then schedule the same small poller outside the application. Add an external uptime monitor if US/EU browser-style probes are required, because log collection cannot manufacture network vantage points.
Measure before copying the choice: missed-run detection time, false-positive rate, reconstruction time, poller availability, and the percentage of incidents where the run ID connects the alert to a terminal event. Those numbers reveal whether the design works. A green dashboard does not.
Further reading
- Healthchecks.io documentation: https://healthchecks.io/docs/
- Better Stack uptime monitoring documentation: https://betterstack.com/docs/uptime/
- Datadog log monitoring documentation: https://docs.datadoghq.com/monitors/types/log/
- Sentry cron monitoring documentation: https://docs.sentry.io/product/crons/
- Prometheus metric naming guidance: https://prometheus.io/docs/practices/naming/
- Alertmanager documentation: https://prometheus.io/docs/alerting/latest/alertmanager/
Top comments (0)