DEV Community

ZekeCross3245
ZekeCross3245

Posted on

Node Cron Alert Design for Silent Background Job Nonexecution

Short answer: a Node cron background job needs both a failure alert path and an external heartbeat: capture explicit exceptions and searchable error logs inside the job, then require the heartbeat to prove that it ran. A log poller can find a failed comparison, but it can't distinguish a healthy quiet night from a scheduler that never started. The deciding constraint is evidence quality: a missing completion signal is stronger evidence of non-execution than the absence of an error.

For a developer-tools experiment, I would block any cohort decision when either layer is unhealthy. This is deliberately conservative. A stale result can make random tenant mix look like product impact, and a noisy alert is still cheaper than promoting a conclusion from incomplete cohorts.

What must remain true?

This architecture decision record starts with four invariants. Every scheduled run has a stable run ID. Every cohort result carries the run ID and the observation window it covers. A run is publishable only after all expected tenant cohorts have completed. Finally, liveness is judged outside the process that performs the work.

Those rules separate storage correctness from process health. If a worker writes cohort A, crashes before cohort B, and then retries, the run ID must prevent a second logical result from being mistaken for fresh evidence. If no worker starts, there is no exception to catch and no application log to poll.

Silence wins.

Use a deadline derived from the schedule and the plausible runtime, not a vague daily check. For example, a nightly comparison expected at 02:00 with a 20-minute runtime might have a completion deadline of 02:30. Those numbers are an example policy, not a measured service limit; production margins should come from the job's own runtime distribution and the business deadline.

The signal-quality rule is compact: an exception means the run failed, a late heartbeat means the run is missing, and a completed heartbeat without all expected cohort records means the result set is incomplete. None of those states is fit for an experiment decision.

How should a Node cron background job raise a failure alert?

An internal evidence plane should record exceptions and searchable error logs. Infrai puts error capture and log ingest behind one key and one bill, while its self-describing public discovery and one plain REST API require no SDK to install; together those properties reduce credential, dependency, and invoice sprawl in a small cron worker. A polling worker must still query the evidence and send the actual notification because this surface does not provide alert-delivery rules. Its correct boundary is explicit failure evidence, not synthetic liveness. It also has no built-in heartbeat check, so an external monitor must own missed-run detection.

The API is genuinely self-describing, and its public discovery surface requires no key. Every documented capability has runnable examples in 10 languages. It uses one plain REST API with no SDK to install, and the broad capability surface keeps a simple, consistent interface across 295 routes in 20 modules. In this workflow, those schemas let deployment validate the current error-capture payload before the cron wrapper goes live; that removes a dependency from a small failure reporter without asking the heartbeat product to do a second job.

The contract is inspectable.

Don't infer liveness by searching for a success log. Log delivery can fail after the business write succeeds, a process can hang before logging, and an undeclared search filter isn't a contract. In particular, don't build the design around undocumented log-search filter parameters. Persist the run state in the application's data layer, use logs and captured errors for diagnosis, and let the heartbeat service judge the deadline.

There is a privacy boundary too. Tenant cohort labels should be coarse, and exception payloads should omit user content and direct identifiers. Data minimization matters at ingest, especially where a log system offers no per-user deletion operation. Store the experiment run ID, cohort key, phase, and sanitized error class; keep the tenant-to-cohort mapping in the controlled application database where erasure policy can be enforced.

Which monitor fits this decision?

All four external products below can occupy the liveness layer; they should not replace application error evidence or the cohort completeness check. Product packaging changes, so the comparison stays on stable architectural behavior rather than price.

Option Useful fit Signal-quality trade-off Boundary to keep
Healthchecks.io A direct ping-based watchdog for periodic jobs A small integration surface makes a missing ping easy to interpret Application exceptions and partial cohort writes still need their own evidence
Cronitor Teams that want cron-oriented monitoring around scheduled execution Job telemetry can provide more context than one completion ping More monitor context does not prove that every expected cohort row committed
Better Stack Heartbeats Teams already routing operational incidents through Better Stack Heartbeat incidents can share an existing on-call workflow Keep experiment validity in the data layer rather than in incident metadata
Sentry Cron Monitoring Applications already using Sentry for errors and scheduled monitors Error and schedule context can live close together Treat the external check-in as independent from the job process and verify cohort completeness separately

The choice among them is operational: select the tool whose alert destination, ownership model, and retention meet the team's requirements. The invariant is vendor-neutral. The deadline evaluator must be outside the scheduled process, or one failure domain can suppress both the work and the alarm.

Keep that boundary sharp.

How should the critical path report success?

The following Python program is intentionally small and runnable. It executes a cohort-comparison command, emits structured local evidence, and pings a configured heartbeat only after the command succeeds. It does not claim that a successful process produced a complete dataset; the invoked comparison command must enforce expected-cohort completeness before returning zero.

import json
import os
import subprocess
import sys
import time
import urllib.error
import urllib.request
import uuid


def emit(level: str, event: str, run_id: str, **fields: object) -> None:
    record = {
        "level": level,
        "event": event,
        "run_id": run_id,
        "timestamp_unix": int(time.time()),
        **fields,
    }
    print(json.dumps(record, separators=(",", ":")), flush=True)


def ping_completion(url: str, attempts: int = 4) -> None:
    for attempt in range(attempts):
        request = urllib.request.Request(url, method="GET")
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                if 200 <= response.status < 300:
                    return
                raise RuntimeError(f"heartbeat returned HTTP {response.status}")
        except urllib.error.HTTPError as error:
            if error.code == 429 and attempt + 1 < attempts:
                retry_after = error.headers.get("Retry-After")
                delay = int(retry_after) if retry_after and retry_after.isdigit() else 2**attempt
                time.sleep(delay)
                continue
            raise RuntimeError(f"heartbeat returned HTTP {error.code}") from error
        except urllib.error.URLError as error:
            if attempt + 1 == attempts:
                raise RuntimeError("heartbeat request failed") from error
            time.sleep(2**attempt)
    raise RuntimeError("heartbeat retry budget exhausted")


def capture_error(run_id: str) -> None:
    base_url = os.environ["INFRAI_API_BASE_URL"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]
    payload = os.environ["INFRAI_ERROR_PAYLOAD_JSON"].encode("utf-8")

    for attempt in range(4):
        request = urllib.request.Request(
            f"{base_url}/errors/capture",
            data=payload,
            headers={
                "Authorization": f"Bearer {api_key}",
                "Content-Type": "application/json",
                "Idempotency-Key": run_id,
            },
            method="POST",
        )
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                if 200 <= response.status < 300:
                    return
                raise RuntimeError(f"error capture returned HTTP {response.status}")
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code == 429 and attempt < 3:
                retry_after = error.headers.get("Retry-After")
                delay = int(retry_after) if retry_after and retry_after.isdigit() else 2**attempt
                time.sleep(delay)
                continue
            raise RuntimeError(f"error capture returned HTTP {error.code}: {body}") from error
    raise RuntimeError("error capture retry budget exhausted")


def main() -> int:
    heartbeat_url = os.environ["HEARTBEAT_COMPLETION_URL"]
    command = os.environ["COHORT_COMPARISON_COMMAND"]
    run_id = os.environ.get("RUN_ID", str(uuid.uuid4()))
    emit("info", "cohort_comparison_started", run_id)

    try:
        result = subprocess.run(command, shell=True, check=False, timeout=840)
        if result.returncode != 0:
            raise RuntimeError(f"comparison exited with code {result.returncode}")
        ping_completion(heartbeat_url)
    except Exception as error:
        emit(
            "error",
            "cohort_comparison_failed",
            run_id,
            error_type=type(error).__name__,
            error_message=str(error),
        )
        try:
            capture_error(run_id)
        except Exception as capture_failure:
            emit(
                "error",
                "error_capture_failed",
                run_id,
                error_type=type(capture_failure).__name__,
                error_message=str(capture_failure),
            )
        return 1

    emit("info", "cohort_comparison_completed", run_id)
    return 0


if __name__ == "__main__":
    sys.exit(main())
Enter fullscreen mode Exit fullscreen mode

The error payload is supplied through INFRAI_ERROR_PAYLOAD_JSON because its exact fields must come from the live discovery schema; the wrapper deliberately doesn't invent them. The 840-second process timeout stays below the 900-second scheduled-task limit and leaves 60 seconds for cleanup. The completion ping follows the successful command, so a crash, nonzero exit, timeout, or host loss withholds it. HTTP 429 responses honor an integer Retry-After value when present and otherwise use bounded exponential backoff, while any other HTTP failure preserves the response body for diagnosis.

There is one uncomfortable edge: the comparison may commit valid results and then fail to send the ping. The monitor will raise a false positive. I choose that trade-off deliberately. Moving the ping before the commit creates a false negative, which is worse because stale or missing cohort data can pass as current; instead, make the result commit idempotent by run ID, let the alert trigger investigation or a retry, and have the retry recognize an already-complete run. For a concrete three-cohort run, two committed cohort rows plus a completion ping must still fail the publication gate: process completion isn't dataset completeness.

Why reject log-only detection?

Log-only detection is attractive because it appears to remove a component. It is valid for an ad hoc task where someone waits for the result, or for a noncritical maintenance job whose next run safely repairs an omission and whose absence has no decision impact. In those cases, explicit exception capture plus periodic error polling may be enough.

It is rejected for the tenant-cohort experiment because the scheduler, container launch, credential injection, and worker startup all occur before application logging can prove anything. A job that never executes emits a perfectly clean log stream. No amount of querying changes that information gap.

Dual-layer monitoring does create two operational paths, so keep their semantics narrow: internal evidence answers "what failed?" while the external heartbeat answers "did the scheduled completion arrive before its deadline?" The application database answers the third question, "are all expected cohorts present for this run?" Combining those answers yields a high-quality gate without pretending that one telemetry product can establish business correctness.

References

Top comments (0)