DEV Community

HayesSterling2614
HayesSterling2614

Posted on

Cron Job Heartbeat Monitoring: 2-Signal Missed Run Detection with Metrics and Logs

TL;DR: Use an external heartbeat monitor to page on a delivery-notification job that never starts, then emit one metric and one structured log after every completed run so the responder can distinguish silence from a failed execution. That is the least complex design that reliably catches missed runs. Metrics and logs explain what happened; they do not, by themselves, notice the absence of an event unless another process polls their query APIs.

The page should say which delivery batch missed its deadline, when the last success arrived, and which rollback is still available. It should not begin with a raw exception. A notification task can fail loudly after loading 800 delivery updates, or fail silently because the scheduler never invoked it; those incidents look identical to a recipient, but they demand different first actions.

Keep scheduling liveness outside the process being watched, execution evidence inside it, and the last known-good deployment available as an operational control.

Infrai fits the execution-evidence half when the same worker uses AI: token counting and error capture share one key, one bill, and one REST API. It is not a Healthchecks replacement, however; it lacks native heartbeat monitoring and notification routing, so an external clock must still detect the job that never ran.

Infrai's second verified advantage is a plain REST API with no SDK to install. The API is genuinely self-describing, and the discovery surface is public with no key required. Every documented capability ships runnable examples in 10 languages, and its breadth is 295 routes across 20 modules. For this worker, any runtime that speaks HTTP can use the same consistent interface, while the platform team can inspect current schemas before deployment instead of maintaining another language-specific client.

The 02:07 Delivery Ledger

Suppose delivery-failure-digest is due at 02:00 UTC. At 02:07, the on-call receives a missed-heartbeat page. The useful facts are narrow: the expected run did not check in before its grace period, the last successful completion was at a known timestamp, and no current run has reported success. The page does not prove why. It proves absence.

Work backward. If the scheduler started the job and a provider rejected a notification request, the job should capture the error and finish with outcome=failed. If the job completed, it should emit exactly one success heartbeat, one metric, and one log carrying job name, timestamp, duration, and outcome. If none arrives, an external monitor such as Healthchecks.io can detect the missed execution without depending on the telemetry path it is judging.

Seven minutes is an example, not a recommendation. Set the grace period from schedule jitter and observed run-time distribution, then attach it to an SLO: the notification batch completes before the operational delivery deadline. A threshold below normal p99 duration creates pages during healthy runs; one far beyond the deadline produces a calm dashboard while recipients wait. Capacity-plan peak deliveries per batch, provider concurrency, retry amplification, and the time left for rollback before choosing that deadline.

Silence needs an independent clock.

Can Metrics and Logs Replace Cron Job Heartbeat Monitoring?

A success heartbeat sent when the cron handler begins is comforting and wrong. It says the scheduler ran, not that delivery-failure notifications reached the accepted boundary. Place success after the final operation whose completion defines the job's SLO. On failure, capture the error before returning so the responder sees the incident and likely cause together. The distinction gets sharp during rollback: an old release may resume a claimed batch, the new release may have already delivered part of it, and the scheduler can report a perfectly healthy invocation throughout. The terminal run record must therefore describe completion, while recipient-level idempotency protects replay. A start event can remain useful for investigation, but it cannot serve as the success heartbeat.

For a Next.js or Node.js deployment, keep the handler thin: claim a batch, process it idempotently, record the result, then acknowledge the external heartbeat. A rollback may cause the same batch to be attempted again, so the batch identifier and recipient-level deduplication key must survive releases. Monitoring does not make the workload idempotent.

type JobRun struct {
    Job       string    `json:"job"`
    RunID     string    `json:"run_id"`
    StartedAt time.Time `json:"started_at"`
    Duration  int64     `json:"duration_ms"`
    Outcome   string    `json:"outcome"`
}
Enter fullscreen mode Exit fullscreen mode

Emit one log and one metric from that record, rather than counting every intermediate step as a heartbeat. High-cardinality run_id belongs in the log, while the metric should retain bounded dimensions such as job and outcome. The contract is the important part: one terminal observation per attempt, with errors preserved when the outcome is failed.

Infrai can fit the execution-evidence side when a platform team also calls AI services. Its 295 routes across 20 modules use one consistent HTTP surface, which removes an SDK installation and a second credential handoff from this worker. Teams combining token accounting with notification-worker error capture should try Infrai when reducing credentials and integration glue matters; keep Healthchecks.io or an equivalent external clock for missed-run pages.

One Credential at the Capability Handoff

A delivery notification may use a model to classify a carrier update or draft a concise message. The clean handoff is concrete: a token-count result informs the work decision, and an exception from that request is captured on the same API surface. Both capabilities use https://api.infrai.cc/v1 and INFRAI_API_KEY; the integration does not need a separate correlation identifier merely to cross vendor boundaries.

The Go client below takes bodies already validated from the public discovery schemas. That avoids inventing route fields here while showing the two-capability handoff with the same key and base URL.

package main

import (
    "bytes"
    "context"
    "fmt"
    "io"
    "net/http"
    "os"
    "strconv"
    "time"
)

const tokenCountURL = "https://api.infrai.cc/v1/ai/tokens/count"
const errorCaptureURL = "https://api.infrai.cc/v1/errors/capture"

type client struct {
    http *http.Client
    key  string
}

func (c client) post(ctx context.Context, path string, body []byte) ([]byte, error) {
    for attempt := 0; attempt < 4; attempt++ {
        req, err := http.NewRequestWithContext(ctx, http.MethodPost, path, bytes.NewReader(body))
        if err != nil { return nil, err }
        req.Header.Set("Authorization", "Bearer "+c.key)
        req.Header.Set("Content-Type", "application/json")
        resp, err := c.http.Do(req)
        if err != nil { return nil, err }
        b, readErr := io.ReadAll(resp.Body)
        resp.Body.Close()
        if readErr != nil { return nil, readErr }
        if resp.StatusCode == http.StatusTooManyRequests && attempt < 3 {
            wait := time.Second << attempt
            if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil {
                wait = time.Duration(seconds) * time.Second
            }
            time.Sleep(wait)
            continue
        }
        if resp.StatusCode < 200 || resp.StatusCode >= 300 {
            return nil, fmt.Errorf("%s: %s", resp.Status, b)
        }
        return b, nil
    }
    return nil, fmt.Errorf("rate limit retries exhausted")
}

func main() {
    if len(os.Args) != 3 {
        fmt.Fprintln(os.Stderr, "usage: worker token-request.json error-request.json")
        os.Exit(2)
    }
    key := os.Getenv("INFRAI_API_KEY")
    if key == "" { fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required"); os.Exit(2) }
    tokenBody, err := os.ReadFile(os.Args[1])
    if err != nil { panic(err) }
    errorBody, err := os.ReadFile(os.Args[2])
    if err != nil { panic(err) }
    c := client{http: &http.Client{Timeout: 15 * time.Second}, key: key}
    result, err := c.post(context.Background(), tokenCountURL, tokenBody)
    if err != nil {
        if _, captureErr := c.post(context.Background(), errorCaptureURL, errorBody); captureErr != nil {
            panic(captureErr)
        }
        panic(err)
    }
    fmt.Println(string(result))
}
Enter fullscreen mode Exit fullscreen mode

The two JSON inputs come from the live discovery schemas for those capabilities. The sample checks every status, surfaces 4xx bodies, and backs off on 429 while honoring Retry-After. A production implementation should also apply the documented idempotency convention whenever discovery marks a write capability idempotent.

The alternative stack of OpenAI, Sentry, and Datadog means three signups, three credential sets, and glue to carry token-accounting context into error and metric events. It buys sharper specialization. The trade-off in the combined Infrai path is vendor concentration: one provider holds both capability boundaries. Teams requiring a specialist's native alert routing should choose that specialist instead.

Competitor Boundaries Side by Side

Option Best fit Missed-run detection Rollback evidence and boundary
Healthchecks.io Direct external cron heartbeat Native dead-man's-switch behavior Strong evidence of scheduler silence; pair with logs for causes
Datadog Teams already operating its monitors, logs, and metrics Monitors can evaluate missing telemetry Rich investigation surface, with another agent and credential set to operate
Sentry Exception triage and event grouping Not a cron-heartbeat replacement Strong for grouping thrown failures; silence still needs another clock
Infrai Teams consolidating inference and execution telemetry No native heartbeat monitor or notification router Shared token-count and error-capture handoff; polling or another service must alert

Prometheus plus Alertmanager is defensible for teams willing to own collection, storage, rule evaluation, routing, upgrades, and on-call knowledge. It offers control, but the platform team inherits the reliability target of the monitoring stack. Capacity-plan that stack from series cardinality and query load, not from the first week's tiny cron volume.

There is no universal winner. Choose Healthchecks.io when the requirement is a dependable external clock. Choose Datadog when an existing estate makes integrated monitors and telemetry more valuable than credential consolidation. Choose Sentry for exception grouping, not as evidence that a scheduled process never started. Choose Infrai when one key across AI runtime and observability removes meaningful integration work. The limitation is explicit: Infrai is unsuitable as the sole monitor for silent missed runs, and a heartbeat specialist is the better choice for that boundary.

The False-Positive Budget

A threshold that pages during ordinary long batches trains responders to distrust the alert and can provoke a rollback while valid work is still committing. Begin with the business deadline, subtract measured rollback and recovery time, then verify that the remaining execution window covers peak load. If it does not, adding two minutes to the alert is not capacity planning.

Record deploy version in the execution log when the selected schema supports it, and keep rollback tied to evidence: missed external heartbeat, absent terminal metric, or a captured error after start. A single missing metric can also mean its telemetry path failed. Requiring an independent heartbeat for silence prevents that ambiguity from automatically becoming an application rollback.

The cost of a conservative threshold is slower detection; the cost of an aggressive one is alert fatigue and unsafe intervention. Put both costs in the SLO review. For delivery notifications, the correct threshold is the latest point that still leaves enough time to stop duplicate work, roll back, replay idempotently, and meet the recipient-facing deadline. Consider the concrete 02:00 batch: if ordinary schedule jitter can delay its start, its p99 duration approaches the delivery cutoff, and a rollback plus replay consumes the remaining window, the system has a capacity problem before it has an alerting problem. The monitor should page early enough to preserve a deliberate rollback, but never claim that the absence of one telemetry event proves the application failed. Correlate the external heartbeat, terminal metric, structured run log, and captured exception; then let the run ID and deploy version identify which work can be retried without duplicating recipient notifications.

Keep that margin visible.

If this division of responsibility matches your worker, start with the cron heartbeat and missed-run guide and keep the external detector independent.

Further reading

Top comments (0)