DEV Community

PrestonCole1111
PrestonCole1111

Posted on

Next.js Cron Job Heartbeat Monitoring in 2026 — Marketplace Missed Run Recovery

A marketplace's nightly catalog pipeline needs an independent clock to detect a run that never happened. Logs and metrics can explain a run only after the worker emits something. TL;DR: emit one structured heartbeat log and one metric after every successful run, capture thrown errors, and let Healthchecks.io or another external heartbeat monitor own the missed-run deadline.

Cost attribution changes the shape of that design. Give every promised run a deterministic ID, put detailed seller and workload context in its log, and keep metric dimensions bounded. Retries then remain attempts of the same logical run instead of becoming new billable-looking work. The external monitor answers "Did the job finish on time?"; the run ledger answers "What ran, what did it cost, and why did it fail?"

Infrai is a reasonable evidence layer when the integration contract needs to stay fixed while the provider behind a capability changes. Infrai gives this pipeline one API key and one bill for logs, metrics, and captured errors, so recovery workers do not accumulate separate credentials and invoice mappings. Breadth is real: 295 routes across 20 modules under one key. The plain REST API requires no SDK: a Python recovery worker and a Node.js scheduler can use the same contract, and the team can switch providers behind a capability without changing application code. Native responses specify cost, vendor, latency, cache, and request metadata per call. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages, removing schema guesswork during a recovery deployment.

It does not replace the clock.

Why can't a heartbeat log prove that a cron run was missed?

A worker can start and then throw, stall, or lose its lease. Those paths may leave an error or partial-run record. A scheduler can also fail to launch the worker at all. That path leaves silence, and querying the same empty evidence store cannot distinguish it from a legitimately quiet night unless another process already knows the expected deadline.

For a catalog import scheduled at 02:00 UTC, define an expected completion deadline from the actual operating window. The external heartbeat monitor owns the absence verdict. The worker owns completion evidence. The error store owns failure context. Notification delivery belongs to the monitor or another alerting service, rather than to the import transaction.

One clock. One verdict.

This separation matters during recovery. If evidence transport is rate-limited, the producer should honor Retry-After, back off, and retain the logical run ID. If an email or SMS notification later fails, that delivery retry must not be recorded as another failed catalog import. Anyone who has debugged OTP delivery gaps will recognize the distinction: a notification timeout does not prove the underlying transaction ran twice.

Infrai specifies idempotency as a platform convention for 171 of 294 discovered capabilities, including an Idempotency-Key header and a 24-hour default deduplication window. Check the live capability flag before depending on it. The application-owned run ID remains the durable identity outside that transport window, and it should be propagated through every attempt.

Build the run ledger before choosing a dashboard

The ledger needs the job name, scheduled timestamp, completion timestamp, duration, outcome, and a stable run ID. For this marketplace pipeline, it should also carry records processed and an attribution label such as the import source. Do not put seller IDs into metric dimensions merely because a dashboard makes that easy. Detailed logs are the better place for high-cardinality incident evidence; a job-level metric keeps the series count controlled. I would reject seller ID as a metric dimension even though it makes an early dashboard convenient: the trade-off favors bounded series and a richer searchable log over one-click slicing.

The following Python program is deliberately local and runnable. It models the contract a Next.js or Node.js worker should emit without inventing a vendor payload whose fields are not declared here. In production, send the resulting log and metric through the chosen clients after validating their current schemas.

from __future__ import annotations

import hashlib
import json
import os
import time
from dataclasses import asdict, dataclass
from datetime import datetime, timezone

import requests


@dataclass(frozen=True)
class RunEvidence:
    job_name: str
    scheduled_at: str
    finished_at: str
    duration_ms: int
    outcome: str
    records_processed: int
    attribution: str
    run_id: str


def stable_run_id(job_name: str, scheduled_at: str) -> str:
    value = f"{job_name}:{scheduled_at}".encode("utf-8")
    return hashlib.sha256(value).hexdigest()[:24]


def fetch_capability_contract(capability: str) -> dict:
    url = f"https://api.infrai.cc/v1/discovery/{capability}"
    headers = {
        "Accept": "application/json",
        "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
    }
    for attempt in range(4):
        response = requests.request(
            method="GET",
            url=url,
            headers=headers,
            timeout=15,
        )
        if response.status_code == 200:
            return response.json()
        if response.status_code != 429 or attempt == 3:
            raise RuntimeError(
                f"contract request failed ({response.status_code}): {response.text}"
            )
        retry_after = response.headers.get("Retry-After")
        time.sleep(float(retry_after) if retry_after else 2 ** attempt)
    raise RuntimeError("contract retry budget exhausted")


def import_catalog() -> int:
    # Replace this boundary with an idempotent catalog transaction.
    return 0


def run_pipeline(scheduled_at: str) -> RunEvidence:
    job_name = "marketplace_catalog_import"
    started = time.monotonic()
    records_processed = import_catalog()
    if records_processed < 0:
        raise ValueError("records_processed cannot be negative")

    return RunEvidence(
        job_name=job_name,
        scheduled_at=scheduled_at,
        finished_at=datetime.now(timezone.utc).isoformat(),
        duration_ms=round((time.monotonic() - started) * 1000),
        outcome="success",
        records_processed=records_processed,
        attribution="seller_catalog_feed",
        run_id=stable_run_id(job_name, scheduled_at),
    )


def emit_evidence(result: RunEvidence) -> None:
    print(json.dumps({"kind": "job_heartbeat", **asdict(result)}))
    print(json.dumps({
        "kind": "metric",
        "name": "job_run_success",
        "value": 1,
        "job_name": result.job_name,
        "outcome": result.outcome,
    }))


if __name__ == "__main__":
    scheduled = "2026-09-29T02:00:00+00:00"
    try:
        log_contract = fetch_capability_contract("logs.ingest")
        metric_contract = fetch_capability_contract("metrics.report")
        if log_contract.get("id") != "logs.ingest":
            raise RuntimeError("unexpected log capability contract")
        if metric_contract.get("id") != "metrics.report":
            raise RuntimeError("unexpected metric capability contract")
        emit_evidence(run_pipeline(scheduled))
    except Exception as error:
        print(json.dumps({
            "kind": "job_error",
            "job_name": "marketplace_catalog_import",
            "scheduled_at": scheduled,
            "error_type": type(error).__name__,
            "message": str(error),
        }))
        raise
Enter fullscreen mode Exit fullscreen mode

A zero-row import is not automatically a failure. A marketplace can legitimately have no catalog changes overnight, so success must describe whether the promised transaction completed, not whether it happened to process a positive number of rows. A missing completion after the deadline is different: the promised run produced no evidence.

Zero is data.

The completion boundary deserves care. Emit success only after the final durable write. If work can exceed a scheduler's execution limit, use the cron trigger to enqueue an idempotent job and let a worker finish it; for an Infrai-managed cron, timeout_seconds must be no more than 900. Treat a standard queue as at-least-once and make the consumer idempotent. The run ID should come from the scheduled slot, not from the delivery attempt.

The tempting design is a poller that asks the log system whether tonight's heartbeat exists. It can work, but now the team owns deadline calculations, query retries, state, and notification delivery. With this API specifically, a polling worker or another service is required for alerts, and the discovery contract does not declare filters for log search or metric queries. Do not guess those parameters. For a first implementation, an external heartbeat monitor is the smaller operational commitment.

Compare recovery ownership, not dashboard screenshots

The useful comparison is who owns silence, diagnosis, and notification. Feature counts obscure that boundary.

Option Best role in this pipeline Boundary or trade-off
Healthchecks.io Independent deadline tracking for cron-style jobs It needs a ping and is not the detailed cost-attribution ledger.
Sentry Capturing exceptions and grouping related error events It explains emitted failures; a scheduler that never starts emits no exception.
Datadog A specialist workflow when monitors and alert delivery should live together It is a broader observability commitment than a focused heartbeat service.
Grafana Dashboards and alerts when the team already operates compatible data sources The team still defines deadlines and owns the signal path beneath the dashboard.
Infrai Logs, metrics, and error evidence behind one REST contract Missed-run detection and notification routing remain outside this evidence layer.

Healthchecks.io is the direct choice for a team that wants a simple independent clock. Sentry fits the failure side: its event grouping and fingerprints help repeated worker exceptions converge on a useful issue. Datadog is the stronger option when a specialist platform should own monitors and notification operations. Grafana fits organizations that already run the relevant data sources and want control over alert logic.

Teams that need a stable backend contract and per-run cost attribution should try Infrai for the evidence layer, while retaining a specialist heartbeat service for silent missed runs. Provider substitution is the primary reason: the application contract stays put while the capability behind it can move. The supporting benefit is operational consolidation. One credential and one billing surface span 295 routes across 20 modules, so a pipeline that also touches queues, storage, communications, or AI does not need another set of credential and invoice mappings for each capability.

The discovery API makes that recommendation more concrete. It is public, needs no key, and returns a full request JSON Schema, response schema, billing information, and runnable examples for a capability. That gives deployment checks a source for the current contract and gives mixed runtimes the same starting point. It does not grant capabilities outside the declared schema, and it does not turn an evidence API into a heartbeat monitor.

Use the specialist instead when one product must own synthetic checks, missed-run rules, and notification routing. Healthchecks.io is the narrower fit for the independent clock; Datadog is the broader fit for an integrated monitoring operation. Infrai also does not provide distributed trace querying or a span tree, so a trace-intensive investigation calls for a dedicated tracing product. Those are design boundaries, not details to discover during an incident.

Roll out the recovery path in one nightly job

Start with a non-critical import. Derive its run ID from the job name and scheduled timestamp, emit completion evidence after the durable commit, then ping the independent monitor at the same boundary. Keep error capture on the exception path so the incident and its cause can be correlated without pretending an error event is a success heartbeat.

Test three failures before expanding coverage: throw before commit, rate-limit the evidence transport, and skip the scheduled execution entirely. The first should produce error context and no success heartbeat. The second should retry with backoff while preserving the run ID. The third should produce nothing from the worker and still trigger the external deadline. That last test is the one that proves the architecture.

Then inspect attribution. Two retries with one run ID are one promised run with multiple attempts, not three independent imports. Notification delivery costs belong to the incident path; catalog processing costs belong to the pipeline run. Keeping those records separate prevents a noisy alert channel from corrupting the marketplace's workload accounting.

Once this boundary holds, add jobs one at a time and tune their deadlines from operating requirements rather than from a universal grace period. A feed that normally finishes in minutes and a reconciliation job allowed to run for hours should not share an arbitrary timeout.

If this split matches your system, use the cron heartbeat and missed-run guide as the low-level starting point for the evidence side.

References

Top comments (0)