TL;DR: Treat a nightly job's completion heartbeat and its diagnostic logs as two different records. The heartbeat answers whether the job finished before its deadline; structured logs explain explicit failures and establish whether a rollback is safe. Logs or metrics alone cannot detect a task that never started. For a customer-support pipeline, retain a compact, privacy-scrubbed run receipt longer than verbose processing events, and send the heartbeat only after the output commit succeeds.
This split also defines the tooling boundary. Healthchecks.io, Cronitor, Better Stack, Datadog, and Grafana Cloud offer different ways to watch a deadline; a log API remains useful for investigation. Infrai is one possible log and metrics layer when a backend team values one key and one bill across services, but it has no heartbeat monitor, threshold alert, or notification route. It therefore needs both an external watchdog and a polling alert worker.
Price the evidence before choosing the monitor
The storage bill is driven by event volume times retention. One completion receipt per run is negligible beside a success event for every transformed ticket, so the useful change is to stop retaining per-record success logs after the rollback window. Trimming a few fields from the daily receipt misses the dominant term.
Suppose a pipeline processes 250,000 records each night, emits one 700-byte event per record, and retains those events for 30 days. That is an illustrative capacity model, not a measured vendor benchmark. The same 700-byte payload, written once per night for 365 days, is tiny by comparison:
def retained_bytes(events_per_run: int, bytes_per_event: int, days: int) -> int:
return events_per_run * bytes_per_event * days
verbose = retained_bytes(250_000, 700, 30)
receipt = retained_bytes(1, 700, 365)
print(f"verbose_gib={verbose / 1024**3:.2f}")
print(f"receipt_kib={receipt / 1024:.2f}")
The outputs are about 4.89 GiB and 249.51 KiB. Keep the ratio in mind, not the sample workload. Measure actual serialized event size and daily count before setting retention.
For this customer-support job, the compact receipt should identify the scheduled input window, stable run ID, deployment ID, start and completion times, record counts, and final status. It should not contain ticket bodies, email addresses, phone numbers, or message text. That boundary matters because a log store without per-user deletion is a poor destination for data that may later be subject to an erasure request.
What do we deliberately give up? Once verbose events expire, an operator may prove that deployment support-etl-2026-09-28.3 committed a particular input window without being able to replay every transformation decision from logs. If policy requires full reconstruction, preserve the source snapshot and transformation version in a governed data store. An observability index is not the archive.
How should Node.js alert when a scheduled job misses its heartbeat?
Because absent code emits nothing.
A thrown exception can create an error event. A successful job can write a metric, a structured log, and a completion receipt. A disabled cron entry, stopped scheduler, or deployment that never launches the worker creates no event to query. Zero errors can mean either healthy or missing.
The watchdog contributes an independent clock. It knows when the nightly run is due and alerts after a grace period when no completion ping arrives. Put that clock outside the scheduler's failure domain; co-locating the checker and job in the same Node.js deployment makes them disappear together during a rollback.
Silence wins.
Use completion, not start, as the decisive heartbeat. A start ping proves invocation only. Completion proves that work reached the commit point. Services that accept start and explicit failure signals can improve the incident timeline, but neither replaces the completion deadline for detecting silence.
The grace period should exceed the job's measured high-end duration plus normal scheduler jitter. Avoid copying a universal timeout from an example. A tight deadline pages on ordinary variance; an excessively loose one delays support-search recovery.
Make rollback identity survive a lost ping
The awkward edge case occurs after commit: the pipeline writes its output, then network delivery of the heartbeat fails. Blindly rerunning the whole input window can duplicate downstream updates. A stable run ID gives the processing layer a way to recognize an already committed window while an operator compares durable receipt evidence with the missing ping.
Order the state changes carefully: derive the run ID, process idempotently, validate the output, commit, write the receipt, and finally ping success. On an exception, send a failure signal when the watchdog supports one and preserve the original error. Do not mark success before commit.
The following Python wrapper launches an existing Node.js pipeline. It uses a scheduled UTC date as the stable identity, retries the watchdog on HTTP 429, honors Retry-After when it is an integer number of seconds, checks every response, and never embeds a secret URL in source control. The three-hour process limit is only a sample operating bound; replace it with a value derived from the pipeline's runtime distribution.
import datetime as dt
import os
import json
import subprocess
import time
import urllib.error
import urllib.request
def ping(url: str, attempts: int = 4) -> None:
for attempt in range(attempts):
request = urllib.request.Request(url, data=b"", method="POST")
try:
with urllib.request.urlopen(request, timeout=10) as response:
if 200 <= response.status < 300:
return
raise RuntimeError(f"watchdog returned HTTP {response.status}")
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(
f"watchdog returned HTTP {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After", "")
delay = int(retry_after) if retry_after.isdigit() else min(2**attempt, 30)
time.sleep(delay)
except urllib.error.URLError:
if attempt == attempts - 1:
raise
time.sleep(min(2**attempt, 30))
def read_logs(api_key: str, attempts: int = 4) -> object:
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
url = f"{base_url}/v1/logs/search"
for attempt in range(attempts):
request = urllib.request.Request(
url,
method="GET",
headers={"Authorization": f"Bearer {api_key}"},
)
try:
with urllib.request.urlopen(request, timeout=10) as response:
if 200 <= response.status < 300:
return json.load(response)
raise RuntimeError(f"log search returned HTTP {response.status}")
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == attempts - 1:
raise RuntimeError(
f"log search returned HTTP {error.code}: {body}"
) from error
retry_after = error.headers.get("Retry-After", "")
delay = int(retry_after) if retry_after.isdigit() else min(2**attempt, 30)
time.sleep(delay)
raise RuntimeError("log search retry budget exhausted")
scheduled_for = dt.datetime.now(dt.timezone.utc).date().isoformat()
run_id = f"support-search-nightly:{scheduled_for}"
environment = os.environ | {"PIPELINE_RUN_ID": run_id}
watchdog_url = os.environ["HEARTBEAT_URL"].rstrip("/")
try:
subprocess.run(
["node", "dist/nightly-support-pipeline.js"],
env=environment,
check=True,
timeout=3 * 60 * 60,
)
except Exception:
ping(f"{watchdog_url}/fail")
print(json.dumps(read_logs(os.environ["INFRAI_API_KEY"])))
raise
else:
ping(watchdog_url)
This wrapper cannot make a non-idempotent pipeline safe. The Node.js processing layer must reject or reconcile a duplicate PIPELINE_RUN_ID at its commit boundary. Also decide what should happen when the failure ping itself cannot be delivered: the local exception must still reach the process supervisor, and the missing completion ping must still expire at the external watchdog.
No automatic rerun yet.
Compare operators, not feature counts
Rollback safety depends on who owns the clock, the diagnostic trail, and notification delivery. A longer feature list does not settle that ownership question.
| Product | Sensible fit | Boundary to verify |
|---|---|---|
| Healthchecks.io | A focused dead-man's-switch model for cron and background jobs | Diagnostics remain in the existing log system |
| Cronitor | Teams that want cron monitoring and execution telemetry together | Check how its integration model fits the current scheduler |
| Better Stack | Teams that want heartbeats near logs and incident-response workflows | A heartbeat-only need may pull in a broader operating surface |
| Datadog | Organizations already using Datadog monitors and logs | Adding it solely for one nightly deadline may increase operational sprawl |
| Grafana Cloud | Teams already operating from Grafana and evaluating synthetic monitoring | Confirm that the configured check represents job completion, not endpoint uptime |
| Infrai | Backends consolidating structured logs and metrics behind one REST API, key, and bill | It supplies no heartbeat, alert-rule, or notification delivery, and queries require polling |
Healthchecks.io is the narrowest conceptual match for an external completion deadline. Cronitor is worth evaluating when execution context should accompany the deadline. Better Stack can reduce context switching when its wider incident workflow is already wanted. Datadog and Grafana Cloud are strongest candidates when either is already the team's operational home.
Infrai solves a different consolidation problem: one key and one bill can reduce credential and invoice sprawl across backend services, while its self-describing discovery surface exposes request schemas and runnable examples. For this design, it remains the diagnostic side of the split. Its log and metric query parameters are undeclared, so do not invent filters in production code; validate the live contract and build the polling worker around supported behavior. This is a real limitation, not a footnote: Infrai is a poor fit when the team wants one vendor to own the heartbeat deadline, threshold evaluation, and notification delivery. Choose Healthchecks.io for a narrow external watchdog, or evaluate Better Stack, Datadog, or Grafana Cloud when an existing incident workspace should own those responsibilities. The trade-off is more credentials and billing surfaces in exchange for a complete alert path.
Keep the watchdog independent even when a broader suite can hold every signal. During a provider or deployment failure, independence is useful evidence, not needless duplication.
Set the retention and rollback contract
Before shipping, write down four decisions: the completion deadline, the stable run identity, the verbose-log expiration point, and the authority allowed to rerun or roll back. Each has a different failure mode. The deadline catches silence. Identity prevents duplicate application. Retention bounds both privacy exposure and diagnostic reach. Human authority prevents a technically successful retry from worsening customer-support data.
Then test the states separately: explicit process failure, scheduler never firing, commit followed by a lost heartbeat, duplicate invocation with the same run ID, and investigation after verbose logs have expired. These are not interchangeable tests. A green log ingestion check says nothing about the external clock.
The practical decision is simple: use structured logs for explanation, a compact receipt for rollback evidence, and an external heartbeat for absence. Keep high-volume events only as long as they can change an operational decision. The cost of that boundary is reduced forensic detail after expiration; the benefit is a smaller privacy and storage footprint without sacrificing proof that the nightly support-search build completed.
Further reading
- Healthchecks.io documentation: https://healthchecks.io/docs/
- Cronitor cron monitoring documentation: https://cronitor.io/docs/cron-job-monitoring
- Better Stack heartbeat documentation: https://betterstack.com/docs/uptime/cron-and-heartbeat-monitor/
- Datadog cron job monitoring guide: https://docs.datadoghq.com/monitors/guide/monitoring-cron-jobs/
- Grafana Cloud synthetic monitoring documentation: https://grafana.com/docs/grafana-cloud/testing/synthetic-monitoring/
- RFC 5424, The Syslog Protocol: https://datatracker.ietf.org/doc/html/rfc5424
Top comments (0)