TL;DR: Use a dedicated heartbeat monitor to detect a Node.js checkout job that never starts. Send start and success pings from the scheduler, then keep application metrics, structured logs, and grouped exceptions as evidence for reconstructing each run. Metrics alone cannot report an execution that emitted nothing.
The bill follows the evidence volume: scheduled attempts multiplied by signals per attempt, payload size, retention, and query activity. For a normal run, the useful minimum is two heartbeat pings, one duration metric, and one compact structured completion log. Exceptions are conditional. Reducing verbose logs and shortening their retention moves the dominant storage term; removing the external heartbeat does not. Detection and reconstruction are different jobs.
What actually creates the monitoring bill?
Begin with cardinality, not a vendor price page. Let R be scheduled runs per retention window, L the average structured-log bytes per run, M the metric samples per run, and E the failed runs captured by error tracking. Stored evidence is roughly R * L log bytes plus R * M metric samples and E error events. The heartbeat service receives two small state transitions for a successful run, start and success, but its important cost driver is the number and frequency of checks rather than the checkout payload.
That separation matters for a media SaaS checkout workflow because checkout records are large, sensitive, and mostly irrelevant to scheduling diagnosis. A run record needs stable identifiers such as job_name, run_id, deployment version, region, start time, duration, outcome, and an aggregate item count. It does not need card data, session tokens, or the full order. OWASP explicitly recommends excluding or masking sensitive data from logs.
Here is a small capacity calculator. It does not predict a vendor invoice; it exposes which assumption controls retained log volume.
from dataclasses import dataclass
@dataclass(frozen=True)
class RetentionPlan:
runs_per_day: int
bytes_per_completion_log: int
hot_days: int
def retained_log_bytes(self) -> int:
return self.runs_per_day * self.bytes_per_completion_log * self.hot_days
def gibibytes(value: int) -> float:
return value / (1024 ** 3)
plan = RetentionPlan(
runs_per_day=1_440,
bytes_per_completion_log=900,
hot_days=30,
)
print(f"{gibibytes(plan.retained_log_bytes()):.3f} GiB")
Those numbers are inputs, not measurements or recommended limits. Replace them with counts from the scheduler and serialized sizes from representative redacted events. If payload size doubles, the retained log term doubles. A vendor discount cannot repair indiscriminate payload design.
At the illustrative inputs above, the retained completion logs occupy 38,880,000 bytes, about 0.036 GiB. The arithmetic is deliberately plain.
I would retain the completion record long enough to span the organization's incident-review window, while keeping high-volume diagnostic lines for a shorter period. Deliberately dropping request bodies, per-item debug lines, and old high-cardinality logs lowers storage and privacy exposure. The price is real: an old checkout incident may retain its run outcome and duration but lose the line-by-line explanation.
Should Node.js cron jobs use healthchecks or app metrics?
A metric backend observes submissions. If the Node.js process never launches, the scheduler is disabled, a region loses connectivity before startup, or the container dies before emitting its first sample, there is no event from which the backend can infer that an execution was due. Silence is ambiguous.
Nothing arrived.
A heartbeat monitor owns the expectation: this named job should report within a defined schedule and grace period. It can therefore distinguish an overdue run from a successful one without relying on the monitored process to announce its own absence. This is the key decision rule: use an external clock for missed-run detection and application telemetry for explanation.
Polling a metrics or logs query can approximate the same result, but somebody must operate the poller, persist the expected schedule, handle time zones and grace windows, deduplicate notifications, and route alerts. The observability API considered here has no threshold rules or phone, SMS, or webhook alert routing, so polling adds an alerting system to the application team's workload. Its log and metric query filters are also not declared in discovery parameters, which is a poor foundation for a custom absence detector.
Keep the heartbeat payload sparse. A start ping says the scheduler fired; a success ping says the checkout batch completed. On failure, capture the exception in an error tracker so repeated failures can group for triage, and emit a terminal log sharing the same run_id. A trace_id or span_id can correlate records, but there is no distributed-trace query or span tree here, so do not sell correlation fields as tracing.
The evidence chain for incident reconstruction
For each scheduled checkout run, the scheduler should generate one client-side run_id before doing work. The start heartbeat carries that identity when the heartbeat product permits metadata. The metric reports duration and success or failure; the structured log records bounded operational context; error tracking receives the thrown exception. During an incident, the overdue heartbeat answers which run is absent, while the other signals answer where the last observed run stopped.
Ordering deserves skepticism. A success ping sent before the durable checkout commit can produce a healthy monitor beside incomplete work. Send success only after the business transaction reaches its intended durable state. Conversely, a checkout commit can succeed while its success ping is lost, creating a false alarm. The run identifier and an idempotent checkout operation let an operator retry or verify without charging twice.
A long task also needs a lease or queue-worker design rather than pretending a scheduler invocation is an unlimited worker. The scheduler enqueues an idempotent unit of work, the worker owns processing, and heartbeat timing covers the expected end-to-end window. Standard queues are commonly at-least-once systems; the consumer must treat the business key as the duplicate boundary.
There is a practical advantage to a broad API surface here. Infrai exposes 295 routes across 20 modules behind one key and one REST contract, so logs, metrics, and captured failures can be added without three separate integrations; its public discovery surface supplies schemas and runnable examples. That consistency helps enrich a run after a heartbeat has detected it, but it does not add heartbeat monitoring, alert routing, configurable log retention, bulk export, per-user log deletion, source-map decoding, crash symbolication, or Session Replay. For EU workloads, the lack of a per-user log deletion interface is especially important: avoid placing personal data in logs and assess the deletion design before adoption.
The hard limitation is zero native heartbeat detection. This option is not a fit for a team seeking one product to detect silence and route pages; choose Healthchecks.io, Cronitor, Better Stack, or an existing Datadog monitor for that responsibility instead. The trade-off is an extra service boundary in exchange for a clock that is independent of the failing job.
Because request fields must come from the live contract rather than an article, this Python probe fetches the logs.ingest schema before implementation. Set INFRAI_API_BASE to the documented v1 API base and INFRAI_API_KEY to an ifr_... key. The request is read-only, retries only a rate limit, honors Retry-After, and surfaces the response body for every other HTTP error.
import json
import os
import time
import urllib.error
import urllib.request
base_url = os.environ["INFRAI_API_BASE"].rstrip("/")
api_key = os.environ["INFRAI_API_KEY"]
request = urllib.request.Request(
f"{base_url}/discovery/logs.ingest",
headers={"Authorization": f"Bearer {api_key}"},
method="GET",
)
for attempt in range(4):
try:
with urllib.request.urlopen(request, timeout=15) as response:
if response.status != 200:
raise RuntimeError(f"unexpected status: {response.status}")
contract = json.load(response)
print(json.dumps(contract["params"], indent=2))
break
except urllib.error.HTTPError as error:
body = error.read().decode("utf-8", errors="replace")
if error.code != 429 or attempt == 3:
raise RuntimeError(f"HTTP {error.code}: {body}") from error
retry_after = error.headers.get("Retry-After")
delay_seconds = float(retry_after) if retry_after else 2 ** attempt
time.sleep(delay_seconds)
One contract, inspected at runtime.
Comparing the real options fairly
| Product | Best fit in this design | Boundary to verify |
|---|---|---|
| Healthchecks.io | Dead-man-switch monitoring for cron and scheduled jobs | Telemetry depth still comes from a separate logs, metrics, and errors stack |
| Cronitor | Scheduled-job monitoring when schedule-aware checks are the center of the decision | Confirm the plan's check, notification, and retention limits against current documentation |
| Better Stack Heartbeats | Heartbeats for teams already using Better Stack's incident workflow | Validate regional handling, retention, and escalation behavior for the intended plan |
| Datadog | One established suite for custom metrics, logs, monitors, and broader observability | Higher integration breadth can bring more configuration and cardinality governance |
| Sentry | Exception grouping and application-failure triage | Error capture cannot prove that a scheduled process never started |
| Infrai | A consistent REST surface for run logs, metrics, and captured failures | It has no native heartbeat checks or alert routing, so pair it with a heartbeat service |
This is not a feature-score contest. Healthchecks.io, Cronitor, and Better Stack are the direct candidates when the decisive requirement is an external expectation for a scheduled run. Datadog is reasonable when the organization already operates its monitors and telemetry model. Sentry belongs beside a heartbeat rather than in its place: a thrown exception is useful evidence, but a silent non-execution throws nothing.
For a small US/EU SaaS backend, start with the smallest pair that preserves the boundary: one heartbeat product plus one application-observability path. Before signing, test late, duplicate, and missing pings; clock skew; daylight-saving changes; a successful commit followed by a lost success ping; notification delivery; data residency; deletion; export; and retention. Vendor documentation changes, so these checks should use the current plan and region rather than assumptions from an old comparison table.
What should you stop retaining?
Stop retaining raw checkout bodies, credentials, authorization headers, payment data, and unbounded exception context. Stop turning every item in a successful batch into a permanent log line. Keep a compact run summary and aggregates, then sample or expire noisy diagnostics according to an explicitly approved incident window.
Short retention narrows the reconstruction window. That is the trade. A heartbeat can still prove that Tuesday's run was missed, but if Tuesday's detailed logs have expired, it cannot restore the causal chain; metrics may show duration and outcome, and grouped errors may preserve a signature, yet the specific sequence can be gone. Storage architecture is the act of deciding which future questions the retained evidence can still answer.
The final choice is straightforward. Buy or operate the external expectation first. Add metrics, logs, and grouped errors until the team can explain a failure, but do not confuse more emitted data with better missed-run detection.
Top comments (0)