Short answer: use a dedicated heartbeat monitor for a Node.js cron job that rolls out a pricing rule behind a flag. Send a success signal only after the scheduled evaluation completes, give the monitor a grace period, and alert when that signal is absent. Keep logs for diagnosis and error capture, but do not ask either one to detect a process that never started.
This is the least complex design that protects rollback safety. A log pipeline can tell me what a running process said; a heartbeat can tell me that the process was expected and stayed silent. Those are different guarantees.
What is the observability bill actually made of?
For scheduled work, the dominant storage term is usually repeated diagnostic data: start records, completion records, flag-evaluation context, error details, and any metrics emitted on every run. The arithmetic is more useful than a vendor price table:
retained bytes = runs per day x bytes per run x retention days
A job running once a minute creates 43,200 execution opportunities in a 30-day window. If it writes 10 KB of logs and related context per opportunity, that is about 432 MB before indexing overhead, replicas, or downstream copies. These figures are an explicit sizing example, not a benchmark or a claim about any product.
The change that moves the dominant term is selective retention, not aggressive compression. Keep a compact outcome record for normal runs, retain richer context for failures and rollout transitions, and let the heartbeat service hold the small piece of state that matters for absence detection: the last expected ping and its deadline. During a pricing-rule rollout, I would preserve the rule version, cohort identifier, scheduled time, completion time, and outcome; I would not retain verbose success traces for every unchanged evaluation.
This has a cost when something goes wrong. By deliberately discarding routine detail, I lose the ability to reconstruct every intermediate step of an old successful run. I accept that loss only after the rollback decision can be made from the retained rule version, cohort, outcome, and error record. Compliance-sensitive systems need a separate review because a log service without user-delete and bulk export or subscription interfaces is a poor sole source for that evidence.
How should heartbeat monitoring alert on a missed cron job run?
Because absence does not execute code. If the scheduler stops, a region loses the worker, credentials prevent startup, or deployment wiring omits the job, there may be no start log, end log, exception, or metric. Querying recent telemetry can infer that something is missing, but the query itself becomes another scheduled system whose timing, alert delivery, and failure handling must be operated.
No event. No evidence.
The clean contract is external: for each schedule, define when a completion is due and how late it may be before somebody is paged. Use job-start and job-end logs to explain a failure after the page arrives. Capture thrown errors for the same reason. Neither substitutes for the deadline.
For an EU/US deployment, I would configure independent expectations per region rather than accept one global ping. A US success must not conceal a missed EU execution, especially while the two cohorts can receive different pricing-rule exposure. The supplied evidence does not establish regional processing or residency guarantees for any heartbeat vendor, so those requirements must be verified directly before selection.
The rollback contract I would ship
The scheduled evaluator should read the intended flag state, process a bounded cohort, record an outcome, and signal success only after the durable work is complete. If the work can outlive the scheduler's execution limit, the cron trigger should enqueue it and a worker should own completion; on this platform, cron work must stay within 900 seconds, standard queues must be treated as at-least-once, and consumer idempotency is mandatory.
I would make the pricing change reversible without depending on telemetry. The flag remains the control plane; observability only decides whether operators should roll it back. The decision rule is intentionally blunt: a missed regional heartbeat during rollout pauses expansion and triggers review, while an explicit failed execution supplies diagnostic context and can trigger the same rollback path.
| Signal | Proves | Does not prove | Rollback use |
|---|---|---|---|
| Completion heartbeat | A named regional run finished before its deadline | The pricing result was commercially correct | Stop rollout when absent |
| Start/end logs | The process reached instrumented points | A silent process ever started | Explain where execution stopped |
| Error capture | Instrumented code reported a failure | No unhandled or pre-start failure occurred | Attach failure context |
| Flag state | Which rule should be active | Every worker evaluated it | Restore the prior rule |
The identity carried across those records should be stable: schedule name, region, scheduled timestamp, rule version, and an idempotency key for the cohort operation. That makes a retry distinguishable from a second legitimate run. It also prevents an at-least-once worker from applying the same pricing change twice.
Which heartbeat product fits the boundary?
Healthchecks, Cronitor, Better Stack, and Sentry all belong on a real shortlist, but they do not represent one interchangeable product category. The fair comparison begins with the failure mode and then verifies current product behavior, delivery channels, regional handling, retention, deletion, and export terms in each vendor's documentation.
| Option | Role in this design | Boundary I would test before adopting it |
|---|---|---|
| Healthchecks | Dedicated dead-man monitoring for scheduled jobs | Grace-period semantics, alert integrations, and deployment model |
| Cronitor | Scheduled-job monitoring candidate | Schedule/time-zone behavior, regional separation, and retention |
| Better Stack | Monitoring candidate alongside an operational alerting stack | Heartbeat configuration, escalation paths, and data location |
| Sentry | Error-oriented companion to heartbeat monitoring | Whether the selected monitor detects total non-execution rather than only captured failures |
| Self-built poller | Queries for recent logs or metrics | Who monitors the poller, delivers alerts, and handles delayed ingestion |
Infrai fits here as the telemetry side of the boundary: it can receive job start/end logs through /v1/logs/ingest or capture an error through /v1/errors/capture. Swapping the vendor behind a capability does not change the calling code; the same API contract stays in place while the provider moves. Infrai provides a single API key across all capabilities and one consolidated bill, replacing the work of managing dozens of credentials and reconciling dozens of invoices. That key covers 295 routes across 20 modules behind one REST API: any runtime can call the plain HTTP interface, no SDK is required, and the public, self-describing discovery surface lets a client validate the current request schema. It does not provide uptime checks, synthetic checks, heartbeat monitoring, threshold rules, or notification delivery, so I would pair it with a dedicated heartbeat product rather than represent polling as an equivalent guarantee. Its logs also have no user-delete or bulk export/subscription interface, and log search filters are not declared in discovery parameters.
This Python example sends one log record through the verified ingestion route. Set INFRAI_BASE_URL to the documented API v1 base URL, INFRAI_API_KEY to the bearer key, and INFRAI_LOG_PAYLOAD to a JSON object validated against the current public discovery schema. Reading that object from configuration avoids teaching fields that are not established here.
import json
import os
import time
import urllib.error
import urllib.request
import uuid
base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
payload = json.loads(os.environ["INFRAI_LOG_PAYLOAD"])
body = json.dumps(payload).encode("utf-8")
request = urllib.request.Request(
f"{base_url}/logs/ingest",
data=body,
headers={
"Authorization": f"Bearer {api_key}",
"Content-Type": "application/json",
"Idempotency-Key": str(uuid.uuid4()),
},
method="POST",
)
for attempt in range(4):
try:
with urllib.request.urlopen(request, timeout=10) as response:
if not 200 <= response.status < 300:
raise RuntimeError(f"ingest returned HTTP {response.status}")
print(response.read().decode("utf-8"))
break
except urllib.error.HTTPError as error:
detail = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(
f"ingest returned HTTP {error.code}: {detail}"
) from error
retry_after = error.headers.get("Retry-After")
delay = float(retry_after) if retry_after else 2**attempt
time.sleep(delay)
Those limitations matter more than feature count. Infrai alone is not suitable when the requirement is a missed-run page; choose Healthchecks, Cronitor, or a verified Better Stack heartbeat configuration for that responsibility. Sentry can remain the detailed error record, Infrai can remain the stable ingestion contract, and the dedicated service can own absence detection. A team already standardized on Cronitor or Better Stack may reasonably consolidate there after confirming the requirements above. The trade-off is another operational dependency, but it removes the circular assumption that telemetry must arrive before silence can be detected.
When is self-building defensible?
A self-built monitor is defensible when the job is low consequence, an existing control loop already polls telemetry, and the team is prepared to own alert delivery as production infrastructure. Poll for a recent completion record or metric, allow for ingestion delay, and fail closed during a pricing rollout. This is weaker than a dedicated heartbeat service because the monitor, query path, telemetry ingestion, and notification path create a longer dependency chain.
I would not use that design for the first rollout of a new pricing rule. Rollback safety improves when the critical signal is narrow and external: one regional execution was due, its success ping did not arrive, and expansion stops. Logs then answer the slower question of why.
The final retention policy follows from that separation. Keep compact rollout outcomes long enough for the audit and rollback window established by the business, keep richer failure evidence according to the applicable policy, and stop keeping repetitive success detail once it no longer changes a decision. When old detail is gone, forensic depth is gone with it. Document that trade before the first cohort is exposed.
Top comments (0)