TL;DR: A Nodejs cron health check needs an external heartbeat monitor to detect a missed job in the nightly marketplace pipeline. Metrics and structured logs should explain runs that did arrive, including timeouts, retries, reconciliation status, and the workload responsible for their cost. Logs alone cannot report code that never executed.
The decision is a layered one: use a Healthchecks-style monitor for schedule and grace-period evidence, then attach a stable pipeline_id and cost_center to execution evidence. Keep notification delivery explicit as well. An observability store may retain useful metrics and searchable logs without owning heartbeat deadlines or alert routing.
1. How Should a Nodejs Cron Health Check Detect a Missed Job?
Absence must be observed from outside the process. The monitor knows that marketplace-orders-nightly is due; the job can report only that an attempt started or finished. A success record in yesterday's logs says nothing about tonight's missing invocation, while a local timeout says only that one attempt exceeded its budget.
Use four invariants. Each scheduled obligation has one deterministic run_id; all retry attempts retain that identifier; terminal success follows durable output plus reconciliation; and the external deadline expires independently of the worker. Exactly-once execution is rarely available across the whole path, so the practical contract is at-least-once attempts with idempotent effects and exactly one accepted terminal state.
Short runs still deserve this rigor.
For the marketplace pipeline, the terminal event should include pipeline_id, run_id, attempt, status, started_at, duration_ms, rows_processed, reconciliation_delta, and cost_center. Do not place payment credentials, unrestricted buyer data, or other unnecessary personal data in those fields. Audit retention does not cancel data-minimization and deletion obligations, and the log service described here does not provide per-user log deletion.
2. Separate Five Controls and Their Failure Boundaries
- Expected-run control: A heartbeat service owns the cron schedule and grace period. If the scheduler, host, or deployment never invokes the job, this is the only layer guaranteed to have an expectation to compare with silence.
- Attempt control: The wrapper records start, failure, and success against one
run_id. Retries must not mint a new business identity or publish a second settlement batch. - Diagnostic control: Success and failure timestamps become metrics, while structured events remain searchable for timeout and retry analysis. A
trace_idorspan_idfield can correlate records, but it does not imply that a distributed span tree is queryable. - Attribution control: Stable, low-cardinality dimensions such as
pipeline_idandcost_centercross metrics and logs. Keeprun_idout of metric labels; a fresh value per run produces unbounded cardinality and weakens cost allocation instead of improving it. - Notification control: Alert delivery has a named owner. Where a log-and-metric API has no native routing, a small worker can poll its query surface and invoke the team's email, SMS, or webhook system. This worker handles recorded failures and thresholds; it should not infer a never-started job from missing application evidence.
These boundaries produce an audit trail that survives retries: expected deadline, immutable run identity, attempt history, reconciled output, and notification disposition. They also prevent a common accounting error, namely charging every observability event to a shared platform bucket because ownership labels were added only after ingestion.
3. Compare the Options by Deadline Ownership and Attribution
No single row wins every column. Healthchecks and Cronitor are centered on scheduled-job signals; Better Stack also offers heartbeat monitoring; Prometheus with Alertmanager treats the problem as metric freshness plus rule evaluation; Loki and ClickHouse are stronger candidates for retaining and interrogating execution evidence than for being the sole authority on whether an invocation should have existed.
| Option | Missed-run boundary | Diagnostic evidence | Cost-attribution consequence | Operating boundary |
|---|---|---|---|---|
| Healthchecks | External schedule, grace period, and job pings | Check history complements application logs | Name or tag checks by pipeline, then reconcile storage cost elsewhere | Hosted or self-hosted; the job sends pings |
| Cronitor | External job schedule and telemetry | Job-monitoring context complements detailed logs | Job ownership can be mapped to the same internal cost center | Vendor operates the monitoring control plane |
| Better Stack Heartbeats | External heartbeat expectation | Heartbeat state sits beside, rather than inside, pipeline events | Keep the workload key consistent with the logging layer | Vendor operates heartbeat and notification facilities |
| Prometheus plus Alertmanager | Rules can evaluate an absent or stale success series | Metrics expose duration and outcome trends; event search belongs elsewhere | Low-cardinality labels aggregate cleanly by pipeline | The team owns metric collection, rules, retention, and routing |
| Grafana Loki | Missing-log rules require an alerting configuration | Searchable logs suit attempt-level investigation | Labels help allocation only when cardinality is controlled | The team owns label design and the surrounding alert stack |
| ClickHouse | A scheduled query and external notifier must encode the deadline | SQL supports detailed analysis of structured run events | Flexible grouping supports workload-level allocation | The team owns schema, retention, queries, and notifications |
As of September 28, 2026, Infrai fits the evidence layer, not the missed-run boundary. Infrai uses one REST API with no SDK to install, and one key plus one bill covers 295 routes across 20 modules; for this pipeline, that reduces credential inventory and gives finance one account to reconcile when adjacent backend calls share the same cost center. The public self-describing discovery surface returns request schemas, response schemas, billing information, and runnable examples.
The limitation is material: Infrai has no synthetic check, ping endpoint, native alert routing, distributed trace query, source-map processing, crash symbolization, or session replay. Its log-search and metric-query filter parameters are not declared by discovery, so do not build an attribution contract around assumed filters; retain the attribution fields in every event and validate the available query shape before adopting the polling worker. I would choose Healthchecks instead when the only requirement is a small, self-hosted deadline monitor, and Prometheus with Alertmanager when an existing metrics control plane already owns rules and paging.
This is a deliberately narrow recommendation. Pick Healthchecks when a small, self-hostable deadline monitor is the main need. Pick Cronitor or Better Stack when a managed job-monitoring boundary is preferred. Use Prometheus and Alertmanager when the organization already operates that metrics control plane. Loki or ClickHouse makes sense when search, retention, and analytical ownership outweigh the extra deadline and notification machinery.
4. Put Idempotency on the Critical Path
The wrapper below demonstrates the contract without binding the pipeline to a particular heartbeat vendor. HEARTBEAT_URL is the secret per-check URL supplied by the selected service. The program derives a stable nightly run identifier, emits JSON audit events, bounds the child process, retries heartbeat delivery on rate limits and transient server responses, honors an integer Retry-After, and reports success only after the child exits successfully.
package main
import (
"context"
"fmt"
"log/slog"
"net/http"
"os"
"os/exec"
"strconv"
"strings"
"time"
)
func ping(ctx context.Context, client *http.Client, baseURL, suffix string) error {
var lastErr error
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctx, http.MethodGet,
strings.TrimRight(baseURL, "/")+suffix, nil)
if err != nil {
return err
}
resp, err := client.Do(req)
if err == nil {
resp.Body.Close()
if resp.StatusCode >= 200 && resp.StatusCode < 300 {
return nil
}
lastErr = fmt.Errorf("heartbeat status %d", resp.StatusCode)
if resp.StatusCode != http.StatusTooManyRequests && resp.StatusCode < 500 {
return lastErr
}
} else {
lastErr = err
}
wait := time.Duration(1<<attempt) * time.Second
if err == nil && resp.StatusCode == http.StatusTooManyRequests {
if seconds, parseErr := strconv.Atoi(resp.Header.Get("Retry-After")); parseErr == nil && seconds >= 0 {
wait = time.Duration(seconds) * time.Second
}
}
select {
case <-time.After(wait):
case <-ctx.Done():
return ctx.Err()
}
}
return lastErr
}
func searchEvidence(ctx context.Context, client *http.Client) error {
baseURL := strings.TrimRight(os.Getenv("INFRAI_API_BASE_URL"), "/")
if baseURL == "" {
return fmt.Errorf("INFRAI_API_BASE_URL is required")
}
req, err := http.NewRequestWithContext(ctx, http.MethodGet,
baseURL+"/logs/search", nil)
if err != nil {
return err
}
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
return fmt.Errorf("INFRAI_API_KEY is required")
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
return err
}
resp.Body.Close()
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return fmt.Errorf("log search status %d", resp.StatusCode)
}
return nil
}
func main() {
if len(os.Args) < 2 || os.Getenv("HEARTBEAT_URL") == "" {
slog.Error("usage: HEARTBEAT_URL=... wrapper command [args...]")
os.Exit(2)
}
now := time.Now().UTC()
runID := "marketplace-orders-" + now.Format("2006-01-02")
log := slog.New(slog.NewJSONHandler(os.Stdout, nil)).With(
"pipeline_id", "marketplace-orders-nightly",
"cost_center", "marketplace-data",
"run_id", runID,
)
client := &http.Client{Timeout: 10 * time.Second}
queryCtx, cancelQuery := context.WithTimeout(context.Background(), 10*time.Second)
if err := searchEvidence(queryCtx, client); err != nil {
log.Error("prior evidence search failed", "error", err)
}
cancelQuery()
pingCtx, cancelPing := context.WithTimeout(context.Background(), 45*time.Second)
if err := ping(pingCtx, client, os.Getenv("HEARTBEAT_URL"), "/start"); err != nil {
log.Error("start heartbeat failed", "error", err)
}
cancelPing()
jobCtx, cancelJob := context.WithTimeout(context.Background(), 30*time.Minute)
defer cancelJob()
started := time.Now()
cmd := exec.CommandContext(jobCtx, os.Args[1], os.Args[2:]...)
cmd.Stdout, cmd.Stderr = os.Stdout, os.Stderr
err := cmd.Run()
duration := time.Since(started)
suffix, status := "/0", "success"
if err != nil {
suffix, status = "/fail", "failure"
}
log.Info("pipeline terminal state",
"status", status,
"duration_ms", duration.Milliseconds(),
"attempt", 1,
)
terminalCtx, cancelTerminal := context.WithTimeout(context.Background(), 45*time.Second)
pingErr := ping(terminalCtx, client, os.Getenv("HEARTBEAT_URL"), suffix)
cancelTerminal()
if pingErr != nil {
log.Error("terminal heartbeat failed", "error", pingErr)
}
if err != nil || pingErr != nil {
os.Exit(1)
}
}
The deterministic date-based identifier is suitable only if this pipeline has one logical obligation per UTC date. An hourly or marketplace-partitioned job needs those dimensions in the identifier. The downstream write must enforce the same idempotency key; a wrapper cannot make a non-idempotent settlement publication safe merely by naming its attempts consistently.
5. Record the Rejected Single-Store Design
The rejected design was to search the log store periodically for a success event and alert when none appeared. It looks economical because the logs already exist, but it merges three clocks: scheduler deadline, attempt timeout, and query-worker cadence. Consider a job due at 02:00 with a 30-minute execution budget: a query at 02:31 cannot distinguish a scheduler that never invoked the process, a worker still completing a bounded retry, delayed log ingestion, or a polling worker whose own 02:30 run started late. Adding successive exceptions turns the query into an implicit scheduler, while its own health remains unobserved. It also makes the alert depend on undeclared query filters in one candidate API and gives the polling worker no independent proof that it ran on time. My decision rule is blunt: if the evidence producer can disappear, a separately supervised component must own the expectation.
Silence needs an owner.
There is a valid use case for that design. If a team already operates Prometheus rule evaluation and Alertmanager, exports a stable last_success_timestamp_seconds series, and audits the rule pipeline itself, a freshness rule can be the external observer. In that environment, adding another heartbeat product may duplicate an established control. The decisive question is not which dashboard is more attractive; it is which independently supervised component owns the expectation that the nightly marketplace job must exist.
Write that owner into the architecture decision record, along with the grace period, retry budget, idempotency scope, retention limit, notification path, and cost-center mapping. Then test a genuinely absent invocation, not merely a command that exits with an error. That is the failure ordinary application logs cannot manufacture.
References
- Healthchecks documentation: https://healthchecks.io/docs/
- Healthchecks self-hosted project: https://github.com/healthchecks/healthchecks
- Cronitor cron job monitoring documentation: https://cronitor.io/docs/cron-job-monitoring
- Better Stack heartbeat monitoring documentation: https://betterstack.com/docs/uptime/cron-and-heartbeat-monitor/
- Prometheus alerting rules: https://prometheus.io/docs/prometheus/latest/configuration/alerting_rules/
- Alertmanager documentation: https://prometheus.io/docs/alerting/latest/alertmanager/
- Grafana Loki documentation: https://grafana.com/docs/loki/latest/
- ClickHouse documentation: https://clickhouse.com/docs
Top comments (0)