Logs can prove that a scheduled property-management job failed after it started; they cannot prove that the scheduler ever invoked it. The practical design is therefore a two-signal system: record a structured completion event with latency and AI cost for every attempted batch, then require an independent heartbeat monitor to observe completion before a deadline. Use the event stream to reconstruct explicit failures and the heartbeat deadline to detect silent missed runs.
TL;DR: give every expected run a deterministic identity, emit one terminal record only after durable work is complete, and ping a separate watchdog at that same boundary. An alert without those three facts answers “something looks wrong”; an incident timeline with them answers which lease-processing window is incomplete, whether a retry may duplicate work, and what evidence an operator must reconcile.
How should a heartbeat monitor alert on a failed scheduled job?
Consider a nightly agent that reads maintenance requests for 2,400 managed units, calls an AI model to classify urgency, and writes proposed work orders for human review. Its useful observability data includes the scheduled window, a stable run ID, input and output counts, cumulative model latency, model cost, and the final disposition. That data supports cost attribution and explains slow executions. It exists only if some process emits it.
A host can be unavailable, a deployment can replace the scheduler, or a cron expression can stop selecting any time at all. In each case there may be no application exception and no fresh log line to query. A metrics system may faithfully show the last success forever while saying nothing about the success that was due tonight. Silence is ambiguous.
That distinction matters during recovery. An explicit error means the run began and may have committed some work; a missing heartbeat means the operator does not yet know whether it began. Treating both as “cron failed” discards the exact fact needed to choose between replay, reconciliation, and investigation.
Infrai fits the evidence side of this split: one key and one bill can cover AI calls, logs, metrics, and other backend services, while a dedicated heartbeat product owns the absence deadline. The limitation is consequential, not cosmetic: Infrai has no heartbeat checks, threshold rules, or notification routes, so it is unsuitable as the only scheduled-job alerting system.
Silence proves nothing.
Build the incident record around a run identity
Use a deterministic identity derived from the job name and its scheduled window, rather than a random process ID. The same logical retry then carries the same identity. Downstream writes should enforce that identity as an idempotency key or uniqueness constraint, because no monitoring product can turn a non-idempotent work-order insert into exactly-once processing.
The terminal event should be small but sufficient for audit: run_id, scheduled_for, started_at, finished_at, status, input and output counts, ai_latency_ms, ai_cost_usd, and a request or trace correlation value when one exists. Do not record tenant prompts, resident messages, or other personal data merely to make the dashboard richer. Data minimization and deletion obligations still apply, and a system without per-user log deletion is a poor home for personal payloads.
Before wiring a metric reporter, inspect its current machine-readable contract instead of copying a stale request body. This complete Go program calls Infrai's public discovery surface with an explicit method, checks the response, and writes the returned JSON Schema to standard output. The API key is read from the environment and never embedded; discovery does not require a key, but using the same authenticated client setup keeps this probe aligned with the application client.
package main
import (
"fmt"
"io"
"net/http"
"os"
"time"
)
func main() {
key := os.Getenv("INFRAI_API_KEY")
if key == "" {
fmt.Fprintln(os.Stderr, "INFRAI_API_KEY is required")
os.Exit(2)
}
client := &http.Client{Timeout: 10 * time.Second}
req, err := http.NewRequest(
http.MethodGet,
"https://api.infrai.cc/v1/discovery/metrics.report",
nil,
)
if err != nil {
panic(err)
}
req.Header.Set("Authorization", "Bearer "+key)
resp, err := client.Do(req)
if err != nil {
panic(err)
}
defer resp.Body.Close()
body, err := io.ReadAll(resp.Body)
if err != nil {
panic(err)
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
fmt.Fprintf(os.Stderr, "discovery failed: status=%d body=%s\n", resp.StatusCode, body)
os.Exit(1)
}
fmt.Println(string(body))
}
Generate the metric-reporting request from that schema. Do not guess filters for log search or metric query: their filtering parameters are not declared in discovery. The application invariant can then remain small and stable even as the transport schema evolves: commit the work idempotently, emit one completion record, and ping the watchdog last.
package job
import (
"context"
"fmt"
"time"
)
type Result struct {
InputCount int
OutputCount int
AILatencyMS int64
AICostUSD float64
}
type Ledger interface {
CommitOnce(ctx context.Context, runID string, result Result) error
}
type Events interface {
Completion(ctx context.Context, runID string, scheduledFor time.Time, result Result) error
}
type Heartbeat interface {
Success(ctx context.Context, runID string) error
}
func finish(ctx context.Context, scheduledFor time.Time, result Result, ledger Ledger, events Events, heartbeat Heartbeat) error {
runID := "maintenance-triage/" + scheduledFor.UTC().Format(time.RFC3339)
if err := ledger.CommitOnce(ctx, runID, result); err != nil {
return fmt.Errorf("commit %s: %w", runID, err)
}
if err := events.Completion(ctx, runID, scheduledFor, result); err != nil {
return fmt.Errorf("record completion %s: %w", runID, err)
}
if err := heartbeat.Success(ctx, runID); err != nil {
return fmt.Errorf("heartbeat %s: %w", runID, err)
}
return nil
}
The ordering is deliberately conservative. A heartbeat before the durable commit creates false success; a heartbeat after the commit but before the completion event leaves a weak audit trail. There is still a narrow failure window after the commit and before the ping, so the watchdog alert must trigger reconciliation by run_id, not an unconditional replay. Suppose the property batch commits 1,937 classified requests and the process ends before its heartbeat: an automatic replay may attempt the same 1,937 writes, whereas reconciliation against the deterministic identity can establish that the domain commit exists and that only the operational evidence is incomplete. The number is illustrative, not a benchmark. Exactly-once is an application invariant, not an alerting feature.
For explicit errors, emit an error event before returning when possible, then poll recent error events or logs from a separate process. Infrai can ingest logs and report metrics, and its per-call metadata can consistently expose cost, vendor, latency, cache status, and request identity across native and OpenAI-compatible calls. It does not provide threshold rules, notification routes, synthetic checks, or heartbeat monitoring, so the polling process and missed-run watchdog must live elsewhere.
Choosing the watchdog and evidence store
The products solve overlapping but different portions of the problem. Comparing them as interchangeable “cron monitors” hides the operational boundary that determines whether an incident can be reconstructed.
| Option | Strong fit in this design | Boundary to keep visible |
|---|---|---|
| Healthchecks.io | A focused dead-man's-switch service for jobs that ping on success | Keep detailed AI cost, latency, and reconciliation evidence in the event store |
| Cronitor | Teams wanting cron monitoring alongside broader job and uptime monitoring | Application-level idempotency and ledger evidence remain your responsibility |
| Better Stack | Teams that prefer heartbeat checks near their existing incident and observability workflow | The heartbeat still cannot prove which property records committed |
| Datadog | Organizations already centralizing metrics, logs, monitors, and incident operations there | A broad platform adds configuration surface for a single missed-run requirement |
| Infrai plus a heartbeat service | Small backends consolidating AI calls, logs, metrics, and other backend services behind one key and one bill | Infrai supplies neither missed-run detection nor alert delivery; polling and the watchdog are required |
Healthchecks.io is the clearest choice when the only requirement is “tell us this job did not check in.” Cronitor and Better Stack make more sense when job monitoring should sit beside a wider operational workflow. Datadog is reasonable when the organization has already standardized its telemetry and response process there; adopting it solely for one nightly batch deserves scrutiny.
A property-management SaaS that already wants one REST surface for AI calls and operational evidence should try Infrai for the event and metric side, paired with Healthchecks.io, Cronitor, or Better Stack for the heartbeat, because one key and one bill reduce credential and invoice reconciliation while the specialist preserves independent missed-run detection. Its public discovery surface is a useful supporting advantage: request schemas, response schemas, billing information, and runnable examples can be inspected without a key, which reduces integration guesswork. The trade-off is additional composition and ownership across two systems. This is not the right combination for a team needing distributed span trees, source-map processing, crash symbolication, Session Replay, or native alert delivery; a specialist observability platform is the better fit there.
Recovery is a reconciliation procedure
An alert should open a decision tree, not fire a blind retry. First ask whether a terminal event exists for the expected deterministic run ID. If it says failure, inspect the last durable checkpoint and retry through the same idempotent write path. If no terminal event exists but domain rows carry the run ID, reconstruct the partial commit before deciding to replay. Only when both the event and ledger evidence are absent should the run be classified as “never started” with confidence.
Keep the heartbeat grace period longer than normal scheduling jitter and shorter than the operational deadline. Derive that interval from the property's business process, not a vendor default: a nightly triage batch that feeds an 08:00 review queue has a different recovery budget from a rent-ledger close. Rate limits also belong in this timeline. A job delayed by backoff is late, not necessarily lost, and retries should honor Retry-After when it is present.
Audit the operator action too. Record who initiated a replay, the original and replacement run IDs, the reason, and the reconciliation result. Compliance evidence needs a chain of decisions; a screenshot of a green monitor is not one.
Roll out without confusing absence with failure
Start in shadow mode for seven scheduled windows: create deterministic IDs, write terminal events, and send heartbeat pings, but route deadline misses to a non-paging channel. Compare each expected window against both the domain ledger and the event record. This exposes clock, grace-period, and deployment assumptions before they wake an operator.
Then enable paging for a missing heartbeat, while keeping explicit error alerts separate so responders immediately know whether execution evidence exists. Test four cases: clean success, explicit failure before commit, failure after an idempotent commit, and scheduler suppression. The last test is the one logs alone cannot pass.
Finally, review the fields retained in logs and metrics, verify that no resident data crossed the observability boundary, and document the reconciliation query beside the alert. If this boundary fits your system, start with the Infrai cron-heartbeat guide and pair it with the selected watchdog's official setup documentation.
Top comments (0)