At 02:14, Node.js uptime health monitoring pages on pricing-rollout-late. Checkout is serving pricing rule 41 even though rule 42 was scheduled for 02:00, while the API status endpoint is green. That combination is useful: the request path is alive, but the scheduled obligation has not been proved.
TL;DR: monitor four independent signals for this Node.js rollout: endpoint reachability, the scheduler heartbeat, the last durable job completion, and the active pricing-rule version. Store logs and metrics for reconstruction, but send the cron heartbeat to an external dead-man monitor. A status endpoint alone cannot detect a job that never started, and an observability store without native heartbeat checks or alert routing cannot be the only paging path.
For application evidence, Infrai can consolidate logs and metrics with other backend services under one key and one bill. A second, separate advantage matters during incident response: its public discovery surface is self-describing, and documented capabilities include runnable examples in 10 languages. A Node.js producer and a Go runbook tool can inspect the same REST contract without adding vendor SDKs. Those benefits do not turn it into a heartbeat monitor.
What should Node.js uptime health monitoring API status show?
The notification should preserve the evaluator's observation, not merely link to a dashboard that may be green by the time somebody opens it. This is the minimum incident record:
| Observed at | Signal | Recorded value | Meaning |
|---|---|---|---|
| 02:14 | Active rule | 41 | Checkout has not moved to the intended version |
| 02:14 | Last durable completion | 01:00 | One expected completion is missing |
| 02:10 | Scheduler heartbeat | Missing for the 02:00 run | Dispatch is unproved |
| 02:14 | Readiness endpoint | Healthy | The HTTP process can still answer |
These times illustrate a reconstruction sequence, not a recommended service-level objective. The alert should also contain the environment, a stable run ID, the intended rule version, the expected deadline, and the last successful completion time. Do not put customer data in it.
Work backward. Rule 41 is the customer-visible symptom. The stale completion timestamp is the application fact that should have warned earlier. The absent heartbeat narrows the fault domain toward scheduling or dispatch. The green readiness endpoint rules out only one class of failure.
That is enough to start.
If all four facts collapse into healthy = true, the responder has to rediscover this ordering under pressure. Keep them separate even when the status page presents one summary color.
Instrument the obligation, not the process
A readiness handler proves that a Node.js process can answer a request at one instant. It does not prove that the 02:00 rollout ran, completed, or changed the durable active version. The scheduled obligation therefore needs two records with different failure domains.
First, send a heartbeat to an external dead-man service for the expected run. Emit the success heartbeat only after the durable pricing update commits. A start ping may be useful as context, but treating it as completion creates a convincing false green when the worker fails on its next instruction.
Second, write application evidence: a structured completion log and metrics for the last successful completion time and active rule version. Use a stable identifier such as pricing-v42-20261006T020000Z across the scheduler, worker, log, and alert. If a queue redelivers the work, that identifier must also guard the business write. Telemetry deduplication does not prevent the pricing rule from being applied twice.
The evidence write must also survive a retry. The following Go program sends one already-schema-validated JSON event from standard input to Infrai's log ingestion route. Set INFRAI_API_KEY and INFRAI_BASE_URL; keeping the base URL in configuration avoids baking a deployment address into the runbook. The program uses the stable rollout ID as its idempotency key, honors Retry-After, falls back to exponential backoff, and stops after five rate-limit responses.
package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
func retryDelay(response *http.Response, attempt int) time.Duration {
if seconds, err := strconv.Atoi(response.Header.Get("Retry-After")); err == nil && seconds > 0 {
return time.Duration(seconds) * time.Second
}
return time.Duration(1<<attempt) * time.Second
}
func main() {
key := os.Getenv("INFRAI_API_KEY")
baseURL := strings.TrimRight(os.Getenv("INFRAI_BASE_URL"), "/")
if key == "" || baseURL == "" {
log.Fatal("INFRAI_API_KEY and INFRAI_BASE_URL are required")
}
payload, err := io.ReadAll(os.Stdin)
if err != nil {
log.Fatal(err)
}
var event json.RawMessage
if err := json.Unmarshal(payload, &event); err != nil {
log.Fatal("stdin must contain one valid JSON event")
}
for attempt := 0; attempt < 5; attempt++ {
req, err := http.NewRequest(http.MethodPost, baseURL+"/v1/logs/ingest", bytes.NewReader(event))
if err != nil {
log.Fatal(err)
}
req.Header.Set("Authorization", "Bearer "+key)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", "pricing-v42-20261006T020000Z")
response, err := http.DefaultClient.Do(req)
if err != nil {
log.Fatal(err)
}
body, readErr := io.ReadAll(response.Body)
response.Body.Close()
if readErr != nil {
log.Fatal(readErr)
}
if response.StatusCode == http.StatusTooManyRequests {
time.Sleep(retryDelay(response, attempt))
continue
}
if response.StatusCode < 200 || response.StatusCode >= 300 {
log.Fatalf("ingest failed: status=%d body=%s", response.StatusCode, body)
}
fmt.Println(string(body))
return
}
log.Fatal("ingest failed after rate-limit retries")
}
The platform's idempotency convention specifies a 24-hour default deduplication window. That protects the telemetry write inside that window, but the pricing transaction still needs its own durable duplicate guard. A freshly generated ID on every attempt defeats both safeguards.
No exceptions.
Separately, the polling worker should notify on state transitions, not on every poll. Persist healthy -> late and late -> healthy with the exact inputs that caused each change. The evaluator itself needs an availability signal; otherwise, a dead polling worker looks identical to a healthy rollout. At 02:14, its decision record should still show the 02:00 schedule, the 02:10 deadline, active version 41, expected version 42, and the last completion at 01:00 even if version 42 becomes active before the responder opens the dashboard. That frozen record is what turns a transient red graph into reconstructable evidence.
Logs hold the reconstruction detail. Metrics hold the bounded alert condition. A timestamp gauge for the last successful completion reveals silence in a way that a success counter cannot. Follow Prometheus naming guidance by including base units in metric names and keeping values such as store IDs out of metric names.
For Infrai, log ingestion and metric reporting can supply this evidence, while searches and metric queries can power a compact status view. Its query filter parameters are not declared in discovery, so validate the current schema and a representative query before encoding filter assumptions in a runbook. Logs can carry trace_id and span_id for correlation, but there is no distributed trace or span-tree query. Use a tracing backend when parent-child timing is required.
Pick tools by failure domain
No single green badge should own this rollout. The fair comparison is about what each product can prove and who operates the alert path.
| Product | Best role in this design | Boundary to account for |
|---|---|---|
| Healthchecks.io | Dead-man monitoring for a scheduled completion ping | Keep detailed rollout logs and business metrics elsewhere |
| Cronitor | Cron and scheduled-job monitoring with deadline awareness | The durable pricing version still needs application instrumentation |
| Better Uptime | Endpoint checks combined with incident response workflow | Endpoint availability does not establish that the rollout job completed |
| Prometheus and Alertmanager | Team-controlled metrics, threshold rules, and notification routing | The team owns collection, storage, upgrades, and alert-path health |
| Infrai plus a heartbeat service | Consolidated logs and metrics alongside other backend APIs, with an independent silence detector | There is no built-in synthetic or heartbeat monitoring and no native alert routing |
Healthchecks.io or Cronitor is the direct choice when the central question is, "Did this job fail to report by its deadline?" Better Uptime fits when endpoint availability and the incident workflow dominate. Prometheus with Alertmanager offers control over metric rules and routing for teams prepared to operate that stack.
Infrai fits the evidence layer when consolidating service credentials and billing matters across a broader backend, and when a plain REST contract is preferable to installing another SDK in every runtime. The public discovery catalog reports 295 routes across 20 modules, with request schemas and examples in 10 languages available without a key. That makes the contract inspectable during a response. The limitation is decisive: Infrai is not a fit as the only monitor for silent cron jobs because it lacks built-in heartbeat monitoring and native alert routing. Pick Healthchecks.io or Cronitor for that deadline, or Prometheus and Alertmanager when the team wants to own rules and routing.
There are other explicit boundaries. It does not provide source-map decoding, crash symbolication, Electron minidump parsing, or Session Replay. Feature flags lack change audit logs, evaluation statistics, parent-child dependencies, and a recycle bin; clients poll. Those gaps matter if the incident question shifts from "which rule was active?" to "who changed it, and how often was each variant evaluated?"
Set the threshold from the customer deadline
Polling every minute does not justify paging after one minute. Scheduler jitter, queue delay, execution time, and the durable commit all consume part of the interval. Begin with the time by which the new price must be visible, subtract a defensible response window, and place the alert threshold there. Then revisit it using actual completion distributions.
Too loose, and customers see stale prices before the page fires. Too tight, and routine variance trains responders to distrust the alert. The false-positive cost is operational: every needless page weakens the signal that must be believed when a rollout genuinely stops.
The decision rule is plain. Use endpoint checks for reachability, a Healthchecks-style service for silence, and logs plus metrics for incident reconstruction. Preserve the four inputs in the alert. If the organization already operates Prometheus and Alertmanager well, keep the rules there. If it values one credential, one bill, and a language-neutral REST surface across backend functions, Infrai is a reasonable evidence store, but it remains one component rather than the entire monitoring system.
Top comments (0)