A cohort experiment cannot be rolled back safely if its scheduled evaluator can disappear without evidence. App logs alone are insufficient because a process that never starts emits no log. The practical choice is to pair structured run logs with an externally observed heartbeat and an outcome record tied to the experiment cohort. Three signals separate absence, execution failure, and bad business output without retaining every debug byte.
TL;DR: Use a deadline-based heartbeat to prove the scheduler reached the job, structured logs to explain what happened inside the run, and a durable outcome record to show which tenant cohort changed. Alert first on a missed deadline. Then use logs for diagnosis and the outcome record for rollback scope. Treat Europe and US schedules as separate monitored instances, even when they execute identical Node.js code.
How should app logs monitor cron job silent failure?
Logging observes code that executed. Silence is ambiguous: the scheduler may not have invoked the process, the host may have failed before logger initialization, credentials may have blocked startup, or there may simply have been nothing worth logging. A search returning zero errors cannot distinguish those states.
Logs cannot prove absence.
Suppose an e-commerce experiment evaluates its Europe cohort at 01:00 UTC and its US cohort at 06:00 UTC. Those are examples, not universal recommendations. If Europe never runs, a shared dashboard may still look active because US logs arrive later. Aggregate activity answers the wrong question. Rollback safety requires evidence for each expected run, region, cohort definition, and deployed revision.
Start with an identity such as cohort-evaluator:europe:2026-09-18T01:00Z. Keep the dimensions bounded. A region label with two values is useful; a tenant_id label with one value per customer creates cardinality proportional to the customer count. Put high-cardinality identifiers in searchable fields or the durable outcome record instead of metric labels.
No event is an event.
Derive three signals from the failure boundaries
At the scheduler boundary, an external monitor expects one heartbeat per schedule instance within a stated grace period. The job sends it after acquiring the intended run identity. If none arrives by the deadline, the alert means "expected execution was not observed," a narrower claim than "no logs found."
At the application boundary, emit structured start, completion, and failure records with stable fields: run_id, region, experiment_id, revision, status, duration_ms, and aggregate counts. Do not attach raw cart contents, customer data, or an unbounded error string as labels. A completion record should agree with the heartbeat identity, letting an operator move from a deadline alert to the relevant log slice without guessing timestamps.
At the business boundary, preserve a durable outcome record containing the cohort definition version, number evaluated, number changed, and revision that produced the decision. This is rollback evidence, not a verbose trace. If a new evaluator changes 412 tenants in Europe and 0 in the US, that asymmetry deserves review even when both processes exit successfully. These numbers illustrate the evidence shape; thresholds must come from the experiment's expected population and risk tolerance.
A generic heartbeat request can remain independent of logging and storage backends:
run_id='cohort-evaluator:europe:2026-09-18T01:00Z'
curl --fail-with-body --silent --show-error \
--request POST 'https://heartbeat.example.net' \
--header 'Content-Type: application/json' \
--data "{\"run_id\":\"${run_id}\",\"region\":\"europe\",\"status\":\"started\"}"
The endpoint is illustrative. In production, authenticate the request, use TLS, bound connection and request timeouts, and make monitoring failure visible without allowing it to mutate cohort decisions. A heartbeat transport outage and a job failure are different conditions; the run record distinguishes them.
This design has a real limitation: an external heartbeat confirms that a request crossed an observation boundary, not that every cohort update was correct. A start-only heartbeat can fire before a crash, while a completion-only heartbeat can leave startup failures ambiguous. Sending both states improves classification but adds events and another network dependency. The outcome record closes part of that gap, yet it cannot decide whether a surprising count is a valid market change or a software defect. That decision needs experiment bounds owned by the application team. For a low-impact housekeeping task with no rollback consequence, the three-signal contract may be excessive; a deadline check plus a compact completion record may be enough. For a cohort evaluator that changes tenant behavior, accepting the extra signal is a deliberate trade-off because the operator needs to identify both the missing execution and the affected population.
That trade-off is deliberate.
Logs, heartbeats, and outcomes answer different questions
| Evidence | Question answered | Silent-failure coverage | Retention posture | Main misuse |
|---|---|---|---|---|
| Deadline heartbeat | Did this expected run appear on time? | Detects absence from the observer's point of view | Keep compact history through deployment and rollback windows | Sending it only at completion |
| Structured app logs | What did the running code do? | Cannot prove invocation | Retain summaries longer than verbose diagnostics | Treating zero errors as success |
| Outcome record | Which cohort changed under which revision? | Reveals missing or implausible business output | Align with audit and rollback needs | Losing region or cohort-version scope |
These signals join on run_id, but should not have identical retention. Let R be runs per day, E average events per run, B average stored bytes per event after encoding, and D retained days. Approximate log storage is R x E x B x D, before replicas and indexes. Heartbeat history replaces E with a small fixed event count. Outcome records are similarly compact. This arithmetic makes sampling a reviewable decision rather than an arbitrary percentage.
Retaining every item-level debug event multiplies storage with orders processed, while one completion summary grows with scheduled runs. Sample verbose success traces when diagnostic value no longer justifies their byte volume, but keep failure records and run summaries under a policy derived from rollback needs. Sampling must happen after preserving completion evidence; otherwise a successful run can look absent by design.
Cardinality has a parallel cost. With 2 regions, 3 revisions, 4 statuses, and 6 job names, the full combination permits 144 metric series before infrastructure dimensions. Adding 10,000 tenant identifiers can expand that upper bound to 1,440,000. Not every combination will exist, but the multiplication explains why tenant IDs belong outside metric labels.
Analytical stores can support investigation over structured event data, but a storage engine does not create a missing-run signal. ClickHouse documents itself as a column-oriented SQL database for online analytical processing. That makes it relevant to retention and query design, not a substitute for an independent deadline.
Set alerts around rollback decisions
Page when a required cohort evaluator misses its heartbeat deadline or reports a failure before applying changes. Use a lower-urgency alert when it completes but produces an outcome outside approved experiment bounds. Keep log-volume and ingestion-delay alerts separate, because telemetry pipeline trouble should not masquerade as a business rollback decision.
The grace period should exceed normal schedule jitter and expected startup delay while remaining shorter than the time available for safe rollback. There is no universal value. Measure observed start delay, document the rollback window, then select and test a threshold between them. A five-minute grace period is defensible only if those system-specific constraints support it.
Europe and US execution also creates a time-boundary trap. Store timestamps in UTC and retain the configured schedule identity; display local time only for operators. Daylight-saving transitions can make local wall-clock schedules ambiguous or nonexistent. The POSIX crontab specification notes that jobs do not run during nonexistent times caused by clock changes and can run twice when a time occurs twice.
Test the negative path deliberately. Disable one schedule in staging and confirm that the deadline alert fires without an app error. Force a failure after the start record and verify that the run differs from a skipped invocation. Then feed an out-of-bounds cohort result and confirm that it blocks promotion or triggers the documented rollback decision without depending on a human reading raw logs.
Roll out the contract without multiplying telemetry
Add the run identity and one completion summary first. Next, introduce per-region heartbeat deadlines in observe-only mode and compare expected schedules with received heartbeats. After accounting for planned pauses and deployment windows, enable paging. Finally, connect outcome bounds to the experiment rollback procedure and rehearse all three failure classes.
Keep the migration compact: one bounded set of heartbeat dimensions, one structured summary per run, and one durable outcome record per cohort decision. Verbose logs remain diagnostic data, sampled and retained according to demonstrated need. Rollback safety comes from proving expected execution and scoping its effects, not from maximizing log volume.
Top comments (1)
tr.ee/dev-to