Do not treat a successful retry as proof that a Node.js cron job is healthy. For a gaming experiment, record one durable completion heartbeat per tenant cohort and schedule window, then let an independent watchdog compare those records with explicit deadlines; rollback only the cohorts whose evidence is late or incomplete. Rollback safety is the deciding constraint, because one global green signal can hide a partial rollout that updated the control cohort but skipped a treatment cohort.
TL;DR: Give every expected run a stable key such as experiment_id/cohort/window, write the heartbeat only after the cohort's durable work commits, and keep retry attempts separate from completion state. A five-minute schedule might use a seven-minute completion deadline, but that number is a capacity decision, not folklore: derive it from the observed upper tail of queue delay plus execution time, add bounded clock-skew allowance, and revisit it when cohort size changes.
How should a Node.js cron health check detect a missed job?
Process uptime answers the wrong question. A scheduler can be alive while an invocation never starts; a worker can return success after processing only part of its input; a retry can overlap the next window; and a heartbeat emitted at job start can remain green after the actual write fails. The operational signal must therefore represent the business unit that rollback acts on. Here, that unit is a tenant cohort inside one experiment window.
This distinction gets sharp during a staged gaming rollout. Suppose an experiment has control, treatment-a, and treatment-b cohorts on a five-minute cadence. Three expected completion records are required for each window. Two records are not 67% healthy. They mean the window is incomplete, and the safest automated response is to stop advancement for the absent cohort while preserving evidence for the others.
No guessing.
Retries need their own identity. Keep attempt for diagnosis, but make completion idempotent on the stable run key. Otherwise, attempt two can create a second heartbeat, make counts look complete, or overwrite timing evidence from attempt one. A unique constraint on the run key turns duplicate completion into a harmless no-op; it does not make the underlying cohort mutation idempotent, so that mutation still needs its own transaction or deduplication boundary.
Retries lie.
Build the completion contract before the alert
The contract needs four times: the scheduled window, the worker's start, its durable completion, and the watchdog's observation. Only the first and third decide lateness. Start time is diagnostic, while observation time prevents the checker from pretending it knew something earlier than it did. Store cohort and experiment identifiers as dimensions, not inside a free-form message.
A practical state model is small: expected, started, completed, late, and failed. Do not let started satisfy the health check. Likewise, a timeout is not automatically a failure of the mutation; it is an unknown outcome until the durable store is read. That distinction prevents a blind retry from applying the experiment twice.
The watchdog below is intentionally written in Go and kept outside the Node.js worker. It accepts records through a generic store interface, computes a deadline from the scheduled time, and produces cohort-level decisions. The production adapter should use a transactional datastore with a uniqueness rule over experiment, cohort, and window.
package watchdog
import (
"context"
"fmt"
"time"
)
type RunKey struct {
Experiment string
Cohort string
Window time.Time
}
type Completion struct {
Key RunKey
CompletedAt time.Time
Attempt int
}
type Store interface {
Completion(ctx context.Context, key RunKey) (Completion, bool, error)
}
type Decision struct {
Key RunKey
Rollback bool
Reason string
}
func Evaluate(
ctx context.Context,
store Store,
now time.Time,
grace time.Duration,
expected []RunKey,
) ([]Decision, error) {
decisions := make([]Decision, 0, len(expected))
for _, key := range expected {
completion, found, err := store.Completion(ctx, key)
if err != nil {
return nil, fmt.Errorf("read completion for %s/%s: %w", key.Experiment, key.Cohort, err)
}
deadline := key.Window.Add(grace)
if !found && now.After(deadline) {
decisions = append(decisions, Decision{key, true, "completion deadline exceeded"})
continue
}
if found && completion.CompletedAt.After(deadline) {
decisions = append(decisions, Decision{key, true, "completed after deadline"})
continue
}
decisions = append(decisions, Decision{key, false, "within completion contract"})
}
return decisions, nil
}
There is a deliberate limitation here: a missing record before its deadline is not healthy or unhealthy yet. It is pending. Alerting early trains responders to ignore noise, while rolling back early can interrupt valid work. On the other side, a very generous grace period protects completion rate by spending rollback time; capacity planning must expose that trade rather than bury it in a timeout constant.
The trade is explicit: a seven-minute deadline on a five-minute cadence permits two minutes for queueing, execution variance, and bounded clock skew, but it also means rollback detection cannot be faster than that deadline. If the observed tail no longer fits, shortening the timeout only converts predictable capacity pressure into noisy failures. The honest choices are to add capacity, reduce each window's cohort work, lengthen the cadence, or accept a slower rollback objective. Record that choice beside the SLO so an operator does not "fix" an alert by widening the grace period during an incident.
The deadline wins.
The SLO should describe completed cohort windows, not watchdog availability alone. One useful formulation is the proportion of expected cohort windows completed before their deadlines, with missing and late records consuming the error budget. Keep the rollback controller conservative when the evidence store itself is unavailable: freeze rollout progression, surface an observability-degraded state, and avoid claiming either success or failure.
Choose the ownership boundary deliberately
The heartbeat mechanism is small; operating it is not. Retention, uniqueness, clock discipline, alert delivery, restore testing, and on-call ownership determine whether the signal remains credible during a release. A buy-versus-build review should score those duties against the platform team's actual staffing and lock-in tolerance.
| Boundary | Managed service | Self-hosted service | Application-owned table |
|---|---|---|---|
| Deadline evaluation | Provider-operated | Platform-operated | Team code and scheduler |
| Cohort-specific schema | May require mapping to provider fields | Fully controllable | Fully controllable |
| On-call load | Lower infrastructure burden, external dependency remains | Storage, upgrades, delivery, and recovery stay internal | Shares failure domain with the workload unless separated |
| Portability | Export and API semantics need review | Data and deployment are under team control | High at the schema level |
| Rollback evidence | Verify retention and deduplication behavior | Define and test both | Define and test both |
No row chooses for you. For a small platform team, transferring routine operation may be rational; for strict cohort semantics or isolation requirements, a self-hosted checker may earn its cost. The deciding test is whether responders can reconstruct one expected window and explain why a rollback did or did not happen without consulting ephemeral logs.
Analytical storage can help with long-range cohort comparisons, but it should not become the sole control-plane dependency by accident. ClickHouse is documented as analytical storage, so treat it as a possible analysis sink and evaluate a separate transactional path for unique completion claims. User-experience metrics such as LCP, CLS, and INP answer a different question from schedule completion. They can validate experiment impact after the pipeline is proven; they cannot prove that every cohort job ran.
Verify failure paths, then rehearse rollback
Verification starts with synthetic windows in a non-production environment, followed by controlled fault injection. Skip one cohort entirely. Delay another past the deadline. Start two attempts with the same run key. Make the completion store unreadable. In each case, assert both the decision and the evidence retained for an operator.
The minimum acceptance set is concrete:
- One on-time completion leaves its cohort eligible to advance.
- One absent completion crosses the deadline and blocks only that cohort.
- A late completion remains visible and does not erase the prior rollback decision.
- Duplicate attempts produce one completion claim for the window.
- Loss of the evidence store freezes progression instead of reporting green.
Test the schedule boundary with fixed UTC timestamps rather than waiting for wall-clock minutes. Also test overlapping windows: if a five-minute job can take eight minutes under planned capacity, either prevent overlap, partition work so windows remain independent, or admit that the selected cadence has no defensible completion SLO. More retries do not create capacity.
Rollback itself must be idempotent and scoped by the same cohort key. Persist the decision before dispatching the action, attach the experiment version being reversed, and reject a rollback aimed at a newer version. Then rehearse a restore of the heartbeat store. An alerting design that loses its decision history during recovery is a dashboard, not an operational control.
Operational decision rule
Advance an experiment cohort only when its current window has one durable, on-time completion and the evidence path is available. Freeze on uncertainty. Roll back that cohort when its deadline is exceeded or a completed record proves it finished late, while leaving other cohorts untouched unless their own contracts fail.
This rule is intentionally stricter than process health and narrower than a global experiment alarm. It gives the on-call engineer a bounded claim: which cohort was expected, which window applied, what evidence arrived, and which version may be reversed. That is enough to automate safely and modest enough to audit.
Top comments (0)