TL;DR: A scheduled Node.js job needs two signals. Emit a structured completion or error event so operators can search what happened, and send a heartbeat to an independent monitor so it can notice when nothing happened. Logs alone can explain a crash, but they cannot prove that a scheduler ever launched the job.
For a nightly media pipeline, I would keep those signals behind a tiny application-owned interface and deploy the monitor change before the job change. That makes rollback boring: the old and new job versions can emit the same contract while the log or heartbeat vendor changes behind it. The extra boundary also keeps a solo team from wiring business code directly to five alerting SDKs.
How should a scheduled job combine failed alerts and a heartbeat?
A log search begins with an event. If the scheduler is disabled, a container never starts, or a deployment removes the schedule, there is no event to find. A metric reported at successful completion has the same blind spot. Querying for errors covers explicit failures; querying for the absence of a success event can help, but then the query runner and notification path become another scheduled system that must stay alive.
The useful split is simple. The job writes structured state for diagnosis, while a heartbeat service owns the deadline. A successful run pings the heartbeat only after the media records are committed. An exception writes an error event and leaves the heartbeat unsatisfied. If the process never launches, the missed deadline still fires.
Silence wins.
This placement matters. Pinging at startup proves only that Node.js began executing; it says nothing about whether the nightly import finished. Pinging in a finally block is worse because failed work can look healthy. Completion means completion.
For logs, keep fields stable enough to search across releases: a job name, run identifier, outcome, duration, and processed-item count. trace_id and span_id may correlate records where supported, but they do not create a distributed trace or a span tree by themselves.
A focused TypeScript implementation
The example below treats stdout as the structured-log transport and a secret environment URL as the heartbeat transport. A platform log collector can ingest the JSON without coupling the job to its search vendor. The heartbeat URL must come from the service that owns the deadline; do not put it in source control.
import { randomUUID } from "node:crypto";
type RunResult = {
processed: number;
};
function writeEvent(fields: Record<string, unknown>): void {
process.stdout.write(`${JSON.stringify({
timestamp: new Date().toISOString(),
service: "nightly-media-index",
...fields,
})}\n`);
}
async function rebuildMediaIndex(): Promise<RunResult> {
// Replace with the pipeline call; resolve only after its durable commit.
return { processed: 0 };
}
async function sendCompletionHeartbeat(url: string): Promise<void> {
const response = await fetch(url, { method: "GET" });
if (!response.ok) {
throw new Error(`Heartbeat returned HTTP ${response.status}`);
}
}
async function verifyLogSearch(baseUrl: string, apiKey: string): Promise<void> {
for (let attempt = 0; attempt < 3; attempt += 1) {
const response = await fetch(new URL("/v1/logs/search", baseUrl), {
method: "GET",
headers: { Authorization: `Bearer ${apiKey}` },
});
if (response.ok) {
await response.json();
return;
}
if (response.status !== 429 || attempt === 2) {
const detail = await response.text();
throw new Error(`Log search failed (${response.status}): ${detail}`);
}
const retryAfter = Number(response.headers.get("retry-after"));
const delayMs = Number.isFinite(retryAfter)
? retryAfter * 1_000
: 500 * 2 ** attempt;
await new Promise((resolve) => setTimeout(resolve, delayMs));
}
}
async function main(): Promise<void> {
const heartbeatUrl = process.env.HEARTBEAT_URL;
if (!heartbeatUrl) throw new Error("HEARTBEAT_URL is required");
const observabilityBaseUrl = process.env.OBSERVABILITY_BASE_URL;
if (!observabilityBaseUrl) {
throw new Error("OBSERVABILITY_BASE_URL is required");
}
const apiKey = process.env.INFRAI_API_KEY;
if (!apiKey) throw new Error("INFRAI_API_KEY is required");
const runId = randomUUID();
const startedAt = Date.now();
try {
await verifyLogSearch(observabilityBaseUrl, apiKey);
const result = await rebuildMediaIndex();
await sendCompletionHeartbeat(heartbeatUrl);
writeEvent({
event: "scheduled_job_completed",
run_id: runId,
outcome: "success",
duration_ms: Date.now() - startedAt,
processed_items: result.processed,
});
} catch (error) {
writeEvent({
event: "scheduled_job_failed",
run_id: runId,
outcome: "error",
duration_ms: Date.now() - startedAt,
error: error instanceof Error ? error.message : String(error),
});
process.exitCode = 1;
}
}
void main();
There is a deliberate trade-off here: a completed pipeline whose heartbeat request fails will exit nonzero even though its data commit succeeded. That is preferable to silently losing the liveness signal, but it can trigger a scheduler retry. The pipeline itself therefore needs idempotent writes keyed by its logical period, such as the publication date, rather than by the random diagnostic run_id. A retry may repeat the heartbeat; it must not duplicate the media import. The log-search request checks the configured backend without inventing filters that its discovery contract does not declare. In production I would run that readiness check during deployment rather than put an external read on every nightly execution. My first instinct would be to validate everything inside the job because it feels safer. On inspection, that makes provider availability part of the pipeline's critical path, so deployment-time validation is the cleaner trade-off.
Set the heartbeat grace period beyond the real completion envelope, not exactly at the cron time. A job scheduled at 02:00 that normally takes time should be judged on its completion deadline. The correct allowance comes from observed runtime distribution and scheduler delay; inventing a universal number would produce noisy alerts.
Choosing the monitor without trapping the job
The products overlap, but they do not make the same architectural choice. This is the shortlist I would use for a small service, based on each product's documented surface rather than a price table that will age quickly.
| Product | Relevant fit | Boundary to account for |
|---|---|---|
| Healthchecks.io | Purpose-built cron and heartbeat monitoring through ping URLs | Keep searchable job details in the logging system; a ping is not a replacement for diagnostic events |
| Cronitor | Cron monitoring with job telemetry and alerting | It adds a dedicated monitoring workflow, so preserve an application-owned adapter if replacement matters |
| Better Stack Heartbeats | Heartbeat checks alongside its incident-management workflow | Decide whether coupling heartbeat and incident response is useful for the team's operating model |
| Datadog | Logs and monitors in a broader observability suite | A broad suite can be more machinery than a small nightly pipeline needs |
Infrai is another possible log and metric backend when one key across capabilities is valuable and one plain REST API can be called over HTTP from any language or runtime without installing an SDK. Swapping the vendor behind a capability does not change application code because the contract stays put; for this pipeline, that keeps a provider change out of the job's rollback path.
Its API is also genuinely self-describing. The public discovery surface requires no key and exposes full request and response schemas, billing information, and runnable examples. The live catalog contains 295 routes across 20 modules, with examples in 10 languages for every documented capability, so a deployment check can inspect the contract before the nightly worker depends on it.
Those strengths do not turn it into a heartbeat product. Its observability surface can store explicit success and error signals, but scheduled-job alerts require polling queries and supplying the notification path, while missed-run detection still belongs in a Healthchecks-style service. That division fits when an inspectable, vendor-swappable API matters more than having one observability suite own every alert.
No choice removes the core rule: the heartbeat deadline must live outside the scheduled process. If the same cron invocation both decides whether it is late and reports the answer, a missed invocation performs neither task.
Roll out the contract before depending on it
Rollback safety comes from sequencing, not a brand name. First create the heartbeat check without paging anyone. Then deploy a job version that emits the stable JSON fields and pings only after commit. Observe successful runs. Enable missed-run notification last.
Keep the previous job version compatible with the same environment variable during the rollout window. If the new release is rolled back, monitoring remains valid. If the heartbeat provider changes later, update the adapter or injected URL; the pipeline's completion semantics stay put.
One contract. Two backends.
Avoid treating the log search as a hidden second heartbeat. Polling recent logs is still useful for explicit task errors when a backend does not provide threshold rules or notification routes. Run that poller independently and make its alert delivery observable. For the absence case, the dedicated deadline monitor is easier to reason about because silence is its input, not an awkward query result.
There are also data-governance boundaries to check before centralizing the events. Do not place personal media metadata in diagnostic fields unless retention and deletion behavior meet the product's requirements. A run ID and aggregate count are usually enough for this alert path.
What to measure before copying this design
Measure the gap between scheduled time, actual start, durable completion, and heartbeat receipt. Those four timestamps distinguish scheduler delay from slow processing and network delay. Track explicit failure count separately from missed deadlines; merging them into one number hides which control failed.
Also test three cases on purpose: an exception after partial work, a process that never starts, and a successful commit followed by a failed heartbeat request. Confirm that the first produces searchable error context, the second produces a missed-run notification, and the third retries without duplicating media records. Then exercise a rollback while the check remains armed.
The decision rule is compact: use searchable logs or metrics to answer what failed, and an external heartbeat to answer did it run. Pick Healthchecks.io or Cronitor for a focused cron monitor, Better Stack when its incident workflow is part of the decision, or Datadog when the team already wants a broad observability suite. Keep the job-facing contract yours. That is the part that makes tomorrow's vendor swap and tonight's rollback manageable.
Further reading
- Healthchecks.io documentation: https://healthchecks.io/docs/
- Cronitor cron job monitoring documentation: https://cronitor.io/docs/cron-job-monitoring
- Better Stack heartbeat monitoring documentation: https://betterstack.com/docs/uptime/cron-and-heartbeat-monitoring/
- Datadog log monitor documentation: https://docs.datadoghq.com/monitors/types/log/
- RFC 5424, The Syslog Protocol: https://datatracker.ietf.org/doc/html/rfc5424
Top comments (0)