DEV Community

RemingtonCross5246
RemingtonCross5246

Posted on

Node.js Pricing Cron 2026: Heartbeat Alerts When Scheduled Jobs Fail

Decision rule: for a failed Node.js scheduled job, use logs or metrics to explain the failure and an external heartbeat alert to prove the cron task ran at all. For a weekly pricing-rule rollout, attach the rollout ID and tenant cohort to the completion event so compute cost can be attributed to the change.

TL;DR: A log-only alert misses the dangerous case: the scheduler never starts the process, so there is no error to record. Pair a completion signal with a Healthchecks-style dead-man switch. Keep the signals separate because they answer different questions.

Choice Explicit crash Job never starts Cost attribution Burden
Logs or metrics alone Yes, if emitted No Strong with rollout fields Medium
Heartbeat alone Only as missed completion Yes Weak Low
Both Yes Yes Strong Medium

The paired design is the default I would ship. It covers both failure classes without turning a pricing rollout into an observability migration. A solo SaaS should spend limited engineering hours on the rule and its safeguards, then outsource the generic deadline check.

How should a Node.js heartbeat alert cover a failed scheduled job?

That question has two incompatible answers. The process might have started and thrown while evaluating a new pricing rule. An error event or log can then carry the rollout ID, cohort, stage, and failure reason. A metric can record completion or failure too. Those dimensions let me separate the cost of a canary cohort from the broad rollout instead of staring at one blended worker total.

Or the process might never have started. No execution means no emitter.

That is the trap.

A heartbeat monitor reverses the responsibility. It expects a ping by a deadline and alerts when the ping is absent. Logs explain what happened after startup; the heartbeat catches a missing start or completion.

Define success narrowly: every intended tenant in the cohort was evaluated and the durable result was committed. Ping only then. A startup ping creates a comforting green check while the pricing update can still fail halfway through.

The two criteria that matter

The first is failure coverage. Explicit exceptions belong in logs or error events. Silent misses belong in a deadline monitor. Infrai can receive logs and metrics through a plain REST API without a client SDK, expose a genuinely self-describing public discovery surface, and cover 295 routes across 20 modules with one key and one bill. A TypeScript worker can inspect runnable examples, which exist in 10 languages, while the single credential reduces secret rotation and invoice reconciliation. For this job, though, the limitations are decisive: it has no heartbeat monitor or notification route. It is not suitable as the only missed-run monitor. You must poll query APIs to build alerts, and a separate Healthchecks-style service is still required to detect a run that emitted nothing. Query filters are undeclared, so I would not build against guessed parameters.

The second is attribution. A generic pricing_job_failed counter says little. Useful dimensions are stable identifiers the application owns: rollout_id, cohort, rule_version, and a bounded outcome. They connect resource use and failures to the change being shipped. Keep customer emails and unbounded exception messages out of metric labels; logs are the better place for high-cardinality detail.

This is a revenue-per-hour decision. The watchdog is undifferentiated infrastructure, while correct cohort selection and idempotent price changes are product work. Ship the small integration this week. Revisit it when alert routing, compliance, or on-call ownership becomes the real constraint.

A minimal TypeScript completion heartbeat

This runnable wrapper treats the heartbeat URL as an opaque secret. It retries HTTP 429 responses, honors Retry-After, uses an explicit method, and surfaces other response bodies. Replace the sample pricing operation with an idempotent application operation before production. A retry must never apply a price change twice.

const heartbeatUrl = process.env.HEARTBEAT_URL;
const infraiBaseUrl = process.env.INFRAI_BASE_URL;
const infraiApiKey = process.env.INFRAI_API_KEY;
if (!heartbeatUrl || !infraiBaseUrl || !infraiApiKey) {
  throw new Error("HEARTBEAT_URL, INFRAI_BASE_URL, and INFRAI_API_KEY are required");
}

const sleep = (ms: number) =>
  new Promise<void>((resolve) => setTimeout(resolve, ms));

async function pingCompletion(url: string): Promise<void> {
  for (let attempt = 0; attempt < 4; attempt += 1) {
    const response = await fetch(url, { method: "POST" });
    if (response.ok) return;

    const body = await response.text();
    if (response.status !== 429 || attempt === 3) {
      throw new Error(`Heartbeat failed (${response.status}): ${body}`);
    }

    const header = response.headers.get("retry-after");
    const seconds = header ? Number(header) : Number.NaN;
    await sleep(Number.isFinite(seconds) ? seconds * 1_000 : 500 * 2 ** attempt);
  }
}

async function applyPricingRule() {
  return {
    rolloutId: "pricing-2026-09",
    cohort: "canary",
    tenantsEvaluated: 0,
  };
}

async function verifyMetricsContract(): Promise<void> {
  const response = await fetch(`${infraiBaseUrl}/discovery/metrics.report`, {
    method: "GET",
    headers: { Authorization: `Bearer ${infraiApiKey}` },
  });
  if (!response.ok) {
    throw new Error(`Metrics discovery failed (${response.status}): ${await response.text()}`);
  }
  await response.json();
}

async function main(): Promise<void> {
  await verifyMetricsContract();
  const result = await applyPricingRule();
  console.log(JSON.stringify({ event: "pricing_rollout_completed", ...result }));
  await pingCompletion(heartbeatUrl);
}

main().catch((error: unknown) => {
  const message = error instanceof Error ? error.message : String(error);
  console.error(JSON.stringify({ event: "pricing_rollout_failed", message }));
  process.exitCode = 1;
});
Enter fullscreen mode Exit fullscreen mode

The order matters. Commit the idempotent business operation, write the completion event, then ping. If the ping fails after the commit, the monitor may alert even though the rollout completed. That false positive is safer than pinging early and hiding a partial update. Put the rollout ID in the runbook so an operator can verify durable state before retrying.

A long task should enqueue bounded, idempotent work per cohort and heartbeat only after the coordinator verifies completion. Small batches sharpen attribution: an unexpectedly costly cohort is visible before the next weekly release.

Choosing among the real options

Healthchecks.io is a direct fit for simple ping-based cron monitoring. Cronitor and Better Stack also provide heartbeat monitoring; both merit consideration when the surrounding monitoring workflow matters more than the smallest integration. Sentry Crons fits when scheduled-job failures already belong beside application exceptions in Sentry. These products own the missing-run deadline in a way a passive log sink cannot.

Pick Healthchecks.io when the dead-man switch is the whole job. Pick Cronitor when cron-focused monitoring is central. Pick Better Stack when heartbeat checks need to sit with broader uptime and incident tooling. Pick Sentry Crons when application error context and cron status should share one operational surface. Verify notification channels, grace-period semantics, retention, and plan limits in each vendor's current documentation; those details change.

The split stack has a cost: two systems can alert on one bad run. Name the alerts differently. The heartbeat alert should say “completion missing,” while the log alert should say “execution failed” and carry rollout dimensions. Route both to one runbook and deduplicate by rollout ID in the incident process.

There are clear cases where the runner-up wins. If useful debugging data already lives in Sentry, Sentry Crons reduces context switching. If a team operates Better Stack or Cronitor, its established escalation path may outweigh a narrow heartbeat tool. If the only requirement is proof of execution and cost attribution is irrelevant, heartbeat-only is enough.

For a new one-person developer-tools SaaS, I would begin with completion heartbeats plus structured logs. It is small enough to ship weekly and precise enough to answer two questions: did the pricing rollout finish, and what did that cohort cost to process?

References

Top comments (0)