Short answer: for a Node.js app, keep a cheap health endpoint for process liveness, then monitor every cron job with a durable heartbeat check. Put the run ID, completion time, item count, and compute units in that endpoint's evidence. Page only when a cohort misses its deadline; a green HTTP response cannot prove that yesterday's student export completed.
At 03:12, the page should name the cohort, experiment, last run, and missing evidence. Dashboards are context, not proof. In an edtech experiment, one noisy tenant can keep an aggregate check green while a smaller treatment cohort stops producing data.
How should a Node.js health endpoint and cron job uptime check work?
Work backward from the alert. A useful page says cohort=west-17 experiment=search-ranking run=2026-09-16T02:00Z, includes the last successful timestamp, and links to the heartbeat record. /healthz should be bounded and cheap: no database fan-out or third-party calls. The worker heartbeat is asynchronous and durable, emitted after the result and cost dimensions commit.
I once started with one global heartbeat. It hid a silent cohort. The fix was tenant and experiment identifiers in the event, with an allow-list so arbitrary requests cannot create unbounded metric labels.
How should the heartbeat carry cost evidence?
Use a stable event shape. Prices change, so record compute units, input rows, and duration; apply a rate card later. Make run_id idempotent so retries do not double-count.
package main
import "time"
type Heartbeat struct {
Tenant string `json:"tenant"`
Experiment string `json:"experiment"`
RunID string `json:"run_id"`
FinishedAt time.Time `json:"finished_at"`
Items int `json:"items"`
ComputeUnits float64 `json:"compute_units"`
}
Store full tenant IDs in logs or an event store, not metric labels when cardinality would explode. Keep student identifiers out of labels. Correlate the record with a trace or request ID.
Which query catches a silent cohort?
Alert on absence within the schedule's service objective, not one transient scrape. If a run should heartbeat every 30 minutes, allow one missed interval and page on the second. Page too early and queue delay becomes noise; page too late and the experiment collects biased data. Record the interval, grace period, and backfill policy in run metadata so the postmortem has an explicit decision.
Why can a status page stay green?
Aggregate availability can hide one school's stale recommendations. Synthetic polling should test a representative read path, while internal heartbeats cover scheduled work and cost attribution. Core Web Vitals (LCP, CLS, and INP) are user-facing signals with p75 thresholds; they belong beside, not inside, a cron heartbeat.
Prometheus is suited to time-series alerting, OpenTelemetry to trace context, and Healthchecks-style services to deadline checks. Those boundaries differ, and none defines your event contract. During an incident ask: did the process answer, did the run commit, and did cost evidence reconcile? Each answer needs one observable record.
| Approach | Integration | Best fit | Main limitation |
|---|---|---|---|
| Prometheus | Metrics scrape | Uptime and query alerts | High-cardinality run details do not belong in labels |
| OpenTelemetry | SDK or collector | Trace correlation | Requires a backend for retention and alerting |
| Healthchecks-style monitor | Heartbeat API | Cron deadline checks | Limited context unless you attach run evidence |
No single alternative covers every app boundary. Choose the simplest monitor your team can operate; a hosted status page is a poor substitute for tenant-level accounting, while a full telemetry stack is excessive for one isolated job.
That is the smallest architecture I trust at 3am: bounded liveness, durable cohort heartbeats, queryable counters, and a page that names missing proof. The false-positive cost of a threshold is operational attention; the false-negative cost is invalid experiment data.
Top comments (0)