TL;DR: Use a heartbeat when the only question is whether a scheduled job arrived on time. Use a custom metrics API when a healthtech experiment must be compared across EU and US tenant cohorts and rolled back without guessing. For that second case, keep the heartbeat anyway: it catches silence, while outcome counters explain whether the run was safe. Three gates make the decision defensible: freshness, completion, and cohort outcome.
The tempting design is one success ping at the end of a Node.js cron handler. It is wonderfully small. It also compresses every tenant, region, and experiment arm into one bit. A successful ping can coexist with skipped tenants or an unhealthy treatment cohort, while a missing ping cannot explain whether the scheduler failed, the worker stalled, or the reporting call was lost.
For a solo builder, the useful question is not which monitoring category wins. It is how little telemetry can preserve a safe rollback decision. The answer is two signal paths with different jobs, plus a rule that refuses to treat missing data as success.
Should Node.js SaaS cron monitoring use healthchecks or a custom metrics API?
A heartbeat is a dead-man signal. The monitor expects a check-in inside a known window and alerts when no check-in arrives. That is a clean detector for a process that never started or never finished. It needs no knowledge of model responses, tenant identifiers, or experiment assignment.
But this job processes tenants in two residency cohorts, eu and us, while an AI-runtime change is exposed to control and treatment arms. Rollback safety depends on what happened inside those partitions. One terminal ping cannot distinguish completed tasks from attempted tasks with partial failures. Nor can it show that one cohort produced enough valid outcomes while the other did not.
Silence matters. Shape matters too.
A custom metrics endpoint can accept counters keyed by a deliberately small set of dimensions: run, region, arm, and outcome. That buys decision detail, but creates responsibilities a heartbeat avoids. The emitter needs bounded labels, authenticated transport, retry behavior, and a rule for late events. The receiver must not count a retry twice. Tenant IDs and patient-level data do not belong in metric labels.
The simple approach is not wrong; it is incomplete for this decision. Keep its narrow strength instead of stretching it into an experiment ledger.
The limitation cuts both ways.
Heartbeat-only monitoring is unsuitable when operators must compare outcomes by cohort, because arrival cannot prove completeness or safety. A custom metrics API is a poor trade-off for a housekeeping cron whose only contract is timely completion: the extra schema, receiver, retries, and cardinality controls create work without changing the response. Even the combined design has a drawback. If metrics publication sits on the critical path, a healthy workload can appear unfinished during a telemetry outage. The conservative alert is intentional, but the runbook must distinguish missing evidence from failed health work before anyone retries side effects.
Define the rollback contract before the cron runs
Start with the operator's question: can the treatment continue for both residency cohorts? Then define three gates.
| Gate | Signal | Conservative failure meaning |
|---|---|---|
| Freshness | Start and finish heartbeat timestamps | The run missed its allowed window |
| Completion | Attempted and completed counters per cohort and arm | The comparison population is incomplete |
| Outcome | Valid and failed counters per cohort and arm | Treatment is unsafe or cannot be evaluated |
The exact alert window and outcome threshold belong to the job's service objective and risk policy. A daily job and a five-minute job have different lateness budgets; a patient-facing action deserves a different rollback threshold from internal summarization. Configure those values, review them, and version them with the experiment.
There is one firm rule: unknown is not healthy. If the heartbeat is late, either arm lacks completion data, or a cohort lacks valid outcomes under the declared policy, stop exposure or hold the previous configuration. Telemetry should support that rule without becoming the source of treatment assignment.
Keep assignment and observation separate. An experimentation system can resolve the arm; the worker records that arm with coarse cohort data and evaluates aggregates. GrowthBook is one public example combining feature flags with experimentation, but the boundary matters more than the implementation: a monitoring failure must not silently change assignment.
A focused TypeScript emitter
The smallest useful implementation emits a start heartbeat, accumulates bounded counters, publishes them with a stable idempotency key, and sends a finish heartbeat only after publication succeeds.
import { createHash, randomUUID } from "node:crypto";
type Cohort = "eu" | "us";
type Arm = "control" | "treatment";
type Outcome = "attempted" | "completed" | "valid" | "failed";
type Count = { cohort: Cohort; arm: Arm; outcome: Outcome; value: number };
interface SignalSink {
heartbeat(runId: string, phase: "start" | "finish"): Promise<void>;
counts(runId: string, key: string, values: Count[]): Promise<void>;
}
const stableKey = (runId: string, values: Count[]) =>
createHash("sha256").update(JSON.stringify({ runId, values })).digest("hex");
async function runExperiment(sink: SignalSink, values: Count[]): Promise<void> {
const runId = randomUUID();
await sink.heartbeat(runId, "start");
await sink.counts(runId, stableKey(runId, values), values);
await sink.heartbeat(runId, "finish");
}
A run identifier joins both paths without exposing a tenant. The content-derived key lets a receiver recognize a repeated batch, though the receiver must still enforce idempotency.
Do not catch a publication error and send the finish heartbeat anyway. That creates a green freshness signal while rollback evidence is missing. Leave the run unfinished, surface the error through the job's normal error path, and retry the same batch under the same key.
Error grouping is separate. Sentry documents how stack traces, exception details, messages, and explicit fingerprints influence grouping. Use a stable fingerprint based on job and failure class rather than tenant identity. Counters answer how much; grouped events retain actionable context.
Make the regional comparison honest
Region is useful only if it describes the cohort rule consistently. Do not infer it from whichever worker executed the task. Resolve residency through the same application rule used by the experiment, then emit only eu or us. This prevents deployment topology from masquerading as experiment data.
Late results need an explicit cutoff. Before it, update the run with idempotent batches. After it, retain late counts for audit or later analysis, but do not quietly rewrite a rollback decision already acted upon.
Test awkward paths: exit before the start signal, timeout during metrics publication, duplicate batch delivery, a cohort with no assignments, and a delayed finish. A synthetic run using non-production fixtures can exercise the route. Verify duplicates do not inflate counts and absent cohort data blocks a healthy verdict.
This is where a custom API costs more than its HTTP request. You own schema compatibility. Additive outcomes are manageable; renaming a cohort or changing a counter's meaning mid-experiment is not. Version the envelope when semantics change.
What should you measure before copying this choice?
Measure the scheduled interval, worst acceptable lateness, and normal completion delay. Inspect how many bounded series the dimensions produce, how often batches retry, and how long evidence arrives after the heartbeat. Those observations determine the alert window and whether the custom path is reliable enough for rollback.
Also track decision coverage: could each run produce a freshness, completion, and outcome verdict for every active cohort and arm? If a metric cannot answer the rollback question, it is storage, not evidence.
Choose heartbeat-only monitoring when safe response depends solely on timely completion. Use both signals when response depends on partitioned outcomes. Extra metrics earn their keep only when they change a rollback decision.
Top comments (0)