A nightly clinical-data pipeline has an awkward constraint: a Node.js app can show perfect uptime through its health monitoring API while its cron job never starts. An application endpoint can answer probes while a scheduler misses the import, and a heartbeat can continue after useful progress has stopped. TL;DR: use separate availability, readiness, and job-progress signals; make the progress signal identify one scheduled run; and alert on a missed deadline rather than treating every log line as health evidence.
This is primarily a signal-quality problem. Searching structured logs is useful for diagnosis, but a search that asks whether a run completed must be built from a small, stable event vocabulary, not inferred from arbitrary messages. The design below keeps protected clinical content out of the monitoring record, preserves an auditable run history, and makes retries converge on one logical outcome.
How should a Node.js app combine uptime health monitoring and cron evidence?
An application health endpoint answers a narrow question about the application instance at the instant it receives a request. A scheduler heartbeat answers a different question: did the expected unit of work reach a meaningful checkpoint before its deadline? Combining them into one green or red status destroys information.
Consider a pipeline expected to import one nightly batch. Its API process may be alive, its configuration may be loaded, and its database connection may accept a trivial check. None of those observations proves that the scheduler fired, that the correct batch was claimed, or that the final reconciliation completed. Conversely, a temporary API restart need not mean the independently running batch failed.
Keep the contracts separate:
| Signal | Question answered | Useful fields | Do not include |
|---|---|---|---|
| Liveness | Can this process continue running? | service, instance, observed time | dependency inventory |
| Readiness | Should this instance receive work now? | service, state, reason code | raw exception text |
| Run progress | Did this logical batch advance? | run ID, stage, attempt, counts, event time | patient data or payload fragments |
| Completion | Did reconciliation close the run? | run ID, terminal state, accepted/rejected totals | free-form records |
The distinction also sets the right failure policy. A liveness check should stay cheap and local; coupling it to every downstream dependency can turn a dependency interruption into a restart loop. Readiness can be stricter because removing an instance from work distribution is its purpose. Progress has a clock: it becomes actionable when a promised checkpoint is late.
One endpoint cannot express all three semantics without forcing callers to guess. Don't make them guess.
Silence is data.
Derive the event contract from the deadline
Start with the operational promise, not the log backend. For example, suppose the nightly batch has a scheduled time, an allowed start delay, a maximum quiet interval while processing, and a completion deadline. Those durations are configuration chosen by the pipeline owner; they are not universal constants. The monitor evaluates the latest durable event against those explicit limits.
A compact event needs enough structure to distinguish a retry from a new scheduled run. A useful identity is a stable run_id derived from the schedule occurrence or assigned before dispatch, plus an attempt number for execution retries. The pair prevents the classic ambiguity in which two workers both emit completed and a search reports two successful batches. Completion must be idempotent for the logical run.
Use a deliberately small state vocabulary such as scheduled, started, progress, completed, and failed. State names are an audit contract, so changing them deserves the same care as changing a database enum. Free-form detail can accompany an event, but alerts and reconciliation should depend on typed fields.
The high-signal queries then become precise:
- no
startedevent exists for the expectedrun_idafter its start deadline; - the latest event is
startedorprogress, but its event time is older than the allowed quiet interval; - a terminal event exists, but accepted plus rejected records does not equal the input count;
- more than one terminal outcome exists for the same logical run.
That last condition matters even when the job uses an exactly-once mindset. Distributed execution commonly delivers retries, so correctness comes from stable identity, idempotent writes, and reconciliation evidence rather than from assuming a process will execute once. The event store should reject an exact duplicate event key or make its insertion harmless. The business transaction should do the same for imported records.
A minimal Go heartbeat with audit-friendly semantics
The smallest useful implementation is not a timer that emits "still alive." It is a typed event emitted only after a meaningful checkpoint commits. In the example below, EventID is stable for the same run, attempt, stage, and checkpoint; an EventSink implementation can use it as an idempotency key. The code carries counts, not clinical payloads.
package heartbeat
import (
"context"
"crypto/sha256"
"encoding/hex"
"fmt"
"time"
)
type Event struct {
EventID string `json:"event_id"`
RunID string `json:"run_id"`
Attempt int `json:"attempt"`
Stage string `json:"stage"`
State string `json:"state"`
Accepted int64 `json:"accepted"`
Rejected int64 `json:"rejected"`
Occurred time.Time `json:"occurred_at"`
}
type EventSink interface {
PutIfAbsent(context.Context, Event) error
}
func RecordCheckpoint(
ctx context.Context,
sink EventSink,
runID string,
attempt int,
stage string,
checkpoint int64,
accepted int64,
rejected int64,
now time.Time,
) error {
if runID == "" || attempt < 1 || checkpoint < 0 {
return fmt.Errorf("invalid checkpoint identity")
}
key := fmt.Sprintf("%s:%d:%s:%d", runID, attempt, stage, checkpoint)
sum := sha256.Sum256([]byte(key))
event := Event{
EventID: hex.EncodeToString(sum[:]),
RunID: runID,
Attempt: attempt,
Stage: stage,
State: "progress",
Accepted: accepted,
Rejected: rejected,
Occurred: now.UTC(),
}
return sink.PutIfAbsent(ctx, event)
}
PutIfAbsent is intentionally an interface rather than an HTTP call to a named service. Its contract is the important part: replaying the same checkpoint must not create new evidence. A production sink also needs bounded request time, authentication, transport security, and a durable retry path. If emitting the event fails after the business checkpoint commits, the retry must reuse the same identity.
There is a subtle ordering choice. Emitting progress before the underlying checkpoint commits can produce false success. Emitting it after commit can temporarily omit true progress if the process exits between those actions. An outbox written in the same transaction as the checkpoint resolves that gap: a separate publisher retries the event, while the deterministic ID suppresses duplicates. This adds storage and a publisher, but for regulated data flows the audit trail is usually worth that machinery.
Do not put names, identifiers, diagnoses, raw rows, or exception dumps into these events. Monitoring needs control-plane facts. Access to detailed processing errors can remain in a separately governed system with retention and authorization appropriate to the data.
How should the monitor separate signal from noise?
Page on violated promises, not on activity. A single transient sink failure is retry material; a missed start deadline is an operational failure. A rejected row may be an expected data-quality outcome; a mismatch between the input count and the reconciled total is a correctness alarm.
This distinction suggests two notification levels. Deadline and reconciliation violations require prompt human attention because waiting does not make the evidence complete. Trends such as a rising rejected-record ratio belong in review or a lower-urgency channel until a documented threshold is crossed. The exact thresholds must come from clinical operations, compliance obligations, and downstream delivery commitments rather than from a generic monitoring default.
Structured-log search still plays a central role. Index only fields used for filtering and grouping, keep high-cardinality identifiers deliberate, and search by run_id when investigating one batch. A dashboard can show the count of expected, active, late, completed, and failed runs without indexing the source data itself. Retention should follow the organization's audit and privacy policies; claiming one universal retention period would be misleading because the applicable rule depends on the record and jurisdiction. A hosted heartbeat checker such as Healthchecks can evaluate missed pings, while an internal event sink can preserve richer reconciliation fields; these are alternatives with different operating burdens, and neither changes the need for stable run identity.
This design has limitations. It is not suitable as a replacement for request tracing, host metrics, or detailed application logs, and an internal event sink adds a datastore, access controls, retention work, and an on-call surface. A simple external heartbeat receiver is the better trade-off for a low-risk cron task that only needs missed-deadline detection; the richer event ledger earns its cost when operators must distinguish attempts, prove completion, and reconcile counts without exposing source records.
Test the negative paths. Disable dispatch for one scheduled occurrence and verify that the missed-start rule fires. Pause a worker after a committed checkpoint and verify that quiet-time detection identifies the correct run. Replay the same checkpoint and confirm that only one event remains. Force two attempts to race toward completion, then confirm that business writes, completion evidence, and reconciliation remain consistent.
These tests are more valuable than generating a large volume of synthetic log messages. They prove the monitor recognizes silence, duplication, and contradiction: the three conditions most likely to be hidden by a superficially active system.
Roll out the contract without losing evidence
Introduce the event schema beside the existing logs, validate it in shadow mode, and compare its run ledger with the scheduler's expected occurrences. A feature toggle can control emission or evaluation while the team observes disagreement, but the toggle must have an owner and a removal condition; long-lived toggles create additional operating states that require testing.
Next, enable durable event writes and idempotency enforcement before turning on notifications. Backfill only when a trustworthy source can reconstruct identity and terminal status; invented historical precision weakens the audit trail. Finally, route deadline and reconciliation failures through the normal incident process, record the chosen thresholds, and review them after schedule or workload changes.
The resulting design is modest: three distinct health concepts, one stable run identity, a short event vocabulary, and alerts tied to explicit deadlines. Its value comes from restraint. A quiet pipeline becomes observable when silence has a defined deadline and every claimed success can be reconciled.
Top comments (0)