Assign one immutable run identity to the scheduled obligation, then charge every attempt, log, heartbeat, and committed output back to that identity. Short answer: monitor a nightly fintech pipeline as an accountable unit of work, not as a long-lived process. Keep the application health endpoint narrow, set a deadline for the scheduled obligation, and reconcile monitoring events with a durable completion record. A green process cannot prove that a batch ran, while a success log cannot explain which retry incurred the reads, writes, and compute.
This architecture decision record uses cost attribution as the deciding axis. The invariants are blunt: one scheduled obligation has one run_id; retries increment attempt; a terminal success follows the durable output commit; controlled ownership fields appear on failures as well as successes; and the monitoring receiver sits outside the worker's failure domain. The boundary matters because scheduler delay, worker failure, monitoring loss, and log-index delay are four different conditions. Treating all four as “down” produces bad pages and worse accounting.
1. Define the billable unit before choosing a signal
The useful unit is the obligation created by the schedule: for example, the settlement export due for a particular UTC date and input partition. Generate its run_id before dispatch and preserve that ID when the worker retries. An attempt is execution history, not a new business obligation.
That distinction prevents a common attribution error. If attempt 1 reads an object partition and fails, then attempt 2 succeeds, both attempts consumed resources; charging only the successful event hides failure cost, while minting a fresh run identity makes one scheduled obligation look like two independent jobs. Store run_id, attempt, pipeline, environment, cost_center, and a bounded partition identifier on every lifecycle event. Do not use account numbers, customer identifiers, or free-form exception messages as dimensions: they are poor aggregation keys and may contain sensitive data.
Retries count.
Four records are enough to establish the control plane:
- The schedule record says what was due and when.
- Attempt events say which executions started, progressed, failed, or succeeded.
- The completion record says which output was durably committed.
- Structured logs explain the work inside an attempt.
This is deliberately smaller than a generic telemetry schema. More labels are not automatically more evidence.
2. What Should a Node.js Cron Health Monitoring API Prove?
No single signal can. An application endpoint answers whether the checked instance can serve the probe at that moment; process uptime says only that a process has remained alive; a heartbeat says that the monitor received an event; and a durable completion record says that the business output reached its commit point. The obligation closes only when the evidence required by the operating policy agrees.
| Evidence | Question it answers | Failure boundary | Attribution value |
|---|---|---|---|
| Process uptime | Has this process remained alive? | Cannot detect a disabled scheduler or stalled work | Low: no scheduled obligation |
| App health response | Can this instance answer its bounded probe now? | Cannot prove last night's batch completed | Low: service-level evidence |
| Deadline heartbeat | Did a named attempt report by the expected time? | Delivery can fail after output commit | High: timely run and attempt identity |
| Durable completion record | Was named output committed? | Does not itself page promptly | High: reconciliation authority |
| Structured logs | What happened within the attempt? | Ingestion and parsing are separate failure domains | Medium: diagnostic detail |
Under HTTP semantics, a successful response means the request succeeded; it does not certify an unrelated scheduled workflow. Keep the health handler bounded and read-only. A probe that scans object storage or counts an entire ledger couples liveness to workload size and can add load during an incident.
RFC 9110 defines the semantics of HTTP success, but no HTTP status can prove an external batch obligation that the request did not execute.
Use an absolute UTC deadline for the batch heartbeat, derived from the schedule and the accepted completion window. Record intended schedule time separately from actual worker start time. That preserves the difference between queue delay and execution time, and it avoids pretending that a periodic “still alive” pulse is proof of completion. Silence needs classification.
3. Make the critical path auditable
The write order is the hard part. Commit the business output, persist a completion record tied to the same logical run, and only then emit success. Never keep the output transaction open while waiting for a monitoring receiver. If the receiver is unavailable after commit, preserve the event for later delivery and let reconciliation distinguish “output complete, notification missing” from “output unknown.”
The following Python models the validation at the generic monitoring boundary. A Node.js producer can emit the same JSON contract; Python is used here to make the invariant checks easy to inspect. The 3 second timeout is an example, not a universal target, and must fit inside the worker's shutdown budget.
from dataclasses import asdict, dataclass
import hashlib
import json
import urllib.request
TERMINAL_STATES = {"succeeded", "failed"}
@dataclass(frozen=True)
class RunEvent:
run_id: str
attempt: int
pipeline: str
environment: str
cost_center: str
state: str
occurred_at: str
def validate(self) -> None:
if not self.run_id or self.attempt < 1:
raise ValueError("run_id is required and attempt must be positive")
if self.state not in {"started", "progress", *TERMINAL_STATES}:
raise ValueError(f"unsupported state: {self.state}")
@property
def event_key(self) -> str:
raw = f"{self.run_id}:{self.attempt}:{self.state}"
return hashlib.sha256(raw.encode("utf-8")).hexdigest()
def publish(receiver_url: str, event: RunEvent) -> None:
event.validate()
body = asdict(event) | {"event_key": event.event_key}
request = urllib.request.Request(
receiver_url,
data=json.dumps(body).encode("utf-8"),
headers={"Content-Type": "application/json"},
method="POST",
)
with urllib.request.urlopen(request, timeout=3) as response:
if response.status // 100 != 2:
raise RuntimeError(f"heartbeat rejected: {response.status}")
The deterministic event key gives the receiver a deduplication key for repeated delivery. It does not make the business write idempotent; that is a separate contract, usually enforced by the output key, transaction, or conditional write. Conflating those two kinds of idempotency is dangerous because a monitor may accept a duplicate while downstream data has already been duplicated.
Test transitions, not happy-path strings. Disable dispatch while leaving the app endpoint healthy. Terminate the worker after output commit but before heartbeat delivery. Replay the same event key. Delay log ingestion beyond the alert window. Reject the receiver request, then verify that retrying the notification keeps the original run_id and attempt. These tests produce evidence about the boundaries without requiring a fabricated production incident.
The trade-off is real.
4. Allocate failure cost without allocating blame
Cost-center labels are useful only if their meaning is stable. Own the mapping in configuration, review it like code, and version changes so a renamed team does not split a single reporting period into unexplained categories. Feature toggles can stage a new emitter or schema, but a toggle that can suppress monitoring is operational inventory: it needs an owner and a removal condition. Fowler's feature-toggle guidance is relevant because every toggle adds carrying cost and additional behavior that must be tested.
Alerts should follow missing state transitions. Page when no start appears after the scheduling grace period, when a started run has no terminal state at its deadline, or when retry policy is exhausted. Send attribution reports on a different path. Paging a team because yesterday's cost label was absent mixes urgent service restoration with governance cleanup, and neither workflow benefits.
There is another trap: failure cost is not necessarily waste. A retry can be the intended response to a transient dependency error. The ledger should expose retry consumption by pipeline, partition, and owning cost center; the decision about acceptable retries belongs in a reviewed policy, informed by workload evidence. Measurement supplies the allocation; it does not supply the judgment.
This design has limits. It is not a fit for a best-effort task with no deadline, durable output, audit requirement, or need to allocate retry cost; a structured log and a simple failure counter are easier to operate there. It is also the wrong shape for a continuously running stream processor, where lag and checkpoint age describe progress better than one nightly terminal event. Even in the scheduled pipeline, the extra event receiver, reconciliation job, schema governance, and retained completion records create operational work. Accept that work only when the team can name who owns each component and when the accounting or audit benefit exceeds the maintenance burden.
Keep cardinality bounded. Run IDs belong in logs, event records, or trace-correlated storage where exact lookup is expected, rather than in every aggregated metric label. OpenTelemetry's attribute guidance distinguishes attributes from the resources and signals that carry them, but each backend imposes practical storage and query limits; validate those limits with representative data before deployment instead of assuming that a standard semantic shape guarantees economical indexing.
5. Reject log-only proof, but keep logs for investigation
The rejected design searches structured logs for a success message and treats its presence as completion. It looks attractive for this pipeline because operators already search the nightly logs, yet the design makes the logging path an accidental transaction coordinator. Parser changes, retention boundaries, delayed indexing, dropped delivery, and duplicate retries can all change the apparent result without changing the committed business output. A text match also struggles to represent the uncomfortable middle state where output is committed but notification delivery is uncertain.
Log-only monitoring still has a valid use case: diagnosing a run after another control identifies it as late or failed. Keep step names, bounded error categories, input and output counts, timing, run_id, and attempt in structured records. Then an operator can move from the missed deadline to the exact run without asking a log query to prove durability.
The accepted design therefore has four separate responsibilities: a narrow endpoint for current application reachability, a deadline heartbeat for prompt detection, a durable completion record for reconciliation, and structured logs for investigation. This costs more design work than searching for “success,” but the result survives retries and assigns consumption to the scheduled obligation that caused it. Choose this separation when auditability and cost ownership matter; choose log-only diagnosis when no completion claim depends on the result.
Top comments (0)