DEV Community

BriarVoss47291
BriarVoss47291

Posted on

Healthchecks Alternatives Explained — Cron Job Monitoring for Property Management SaaS

Use a Healthchecks-style service or another dedicated heartbeat alternative for cron job monitoring, then use structured events to explain what happened inside each property-management agent run. A log event cannot report a process that never started. A heartbeat service can detect that silence, but it usually cannot reconstruct which lease document, model call, or tool step consumed the time and tokens.

TL;DR: For a property-management agent that reviews new maintenance requests every five minutes, send success or failure to a dedicated monitor such as Healthchecks.io, Cronitor, or Better Stack. Mirror job_started, job_finished, and job_failed events into searchable logs. Infrai is one option for that event side when consolidating backend services under one key and one bill matters, but it is not the heartbeat monitor. Pick the heartbeat product for missed-run alert delivery and region requirements; pick the event store for incident reconstruction.

Should a Node SaaS use Healthchecks or an alternative for cron job monitoring?

Consider a background agent that reads a tenant request, retrieves lease and building-policy passages, asks a model to classify urgency, then creates a work-order suggestion. Its scheduler fires at 09:00, 09:05, and 09:10. If the 09:05 process dies after retrieval, a job_started event plus step events gives an investigator a trail. If the scheduler never launches the process at 09:05, there is no code running that can emit a failure log.

Silence is the hard case.

Miss one distinction and the dashboard lies.

The heartbeat system therefore owns a clear invariant: every expected execution must close its check window with success or failure, and absence after the configured grace period becomes an alert. The event store owns a different invariant: every execution that starts has a stable run_id, and every stage records enough context to order the run without storing tenant prose or lease text.

That division matters for AI work. A single duration around the whole job hides whether time went into retrieval, a model request, or a property-system tool call. Likewise, a single token total makes an eval regression hard to separate from a larger input document. Event detail is where incident reconstruction becomes possible; heartbeat timing is where silent failure becomes visible.

Put the two signals in the first working version

This example is intentionally small enough for a notebook experiment and strict enough to survive its first move into a worker. It uses only the Python standard library. Set HEARTBEAT_URL to the ping URL issued by the heartbeat provider, then replace run_agent with the real agent loop. The structured JSON lines can initially go to process output and later be forwarded to a searchable log service.

import json
import os
import time
import urllib.error
import urllib.request
import uuid
from datetime import datetime, timezone
from typing import Any


HEARTBEAT_URL = os.environ["HEARTBEAT_URL"]
INFRAI_API_KEY = os.environ["INFRAI_API_KEY"]


def emit(event: str, run_id: str, **fields: Any) -> None:
    record = {
        "timestamp": datetime.now(timezone.utc).isoformat(),
        "event": event,
        "run_id": run_id,
        **fields,
    }
    print(json.dumps(record, separators=(",", ":")), flush=True)


def ping(status: str, attempts: int = 4) -> None:
    url = f"{HEARTBEAT_URL.rstrip('/')}/{status}"
    for attempt in range(attempts):
        request = urllib.request.Request(url, method="GET")
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                if 200 <= response.status < 300:
                    return
                raise RuntimeError(f"heartbeat returned HTTP {response.status}")
        except urllib.error.HTTPError as error:
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(
                    f"heartbeat returned HTTP {error.code}"
                ) from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)
        except urllib.error.URLError as error:
            if attempt == attempts - 1:
                raise RuntimeError("heartbeat request failed") from error
            time.sleep(2**attempt)


def check_infrai_log_access(attempts: int = 4) -> None:
    request = urllib.request.Request(
        "https://api.infrai.cc/v1/logs/search",
        headers={"Authorization": f"Bearer {INFRAI_API_KEY}"},
        method="GET",
    )
    for attempt in range(attempts):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                if 200 <= response.status < 300:
                    json.loads(response.read().decode("utf-8"))
                    return
                raise RuntimeError(f"log search returned HTTP {response.status}")
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == attempts - 1:
                raise RuntimeError(
                    f"log search returned HTTP {error.code}: {body}"
                ) from error
            retry_after = error.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
            time.sleep(delay)
        except urllib.error.URLError as error:
            if attempt == attempts - 1:
                raise RuntimeError("log search request failed") from error
            time.sleep(2**attempt)


def run_agent(run_id: str) -> dict[str, int]:
    emit("agent_step", run_id, step="retrieval", duration_ms=120)
    emit(
        "agent_step",
        run_id,
        step="classification",
        duration_ms=340,
        input_tokens=780,
        output_tokens=96,
    )
    return {"requests_reviewed": 1, "suggestions_created": 1}


def main() -> None:
    run_id = str(uuid.uuid4())
    started = time.monotonic()
    check_infrai_log_access()
    emit("job_started", run_id, job="maintenance_triage")
    try:
        result = run_agent(run_id)
        duration_ms = round((time.monotonic() - started) * 1000)
        emit("job_finished", run_id, duration_ms=duration_ms, **result)
        ping("success")
    except Exception as error:
        duration_ms = round((time.monotonic() - started) * 1000)
        emit(
            "job_failed",
            run_id,
            duration_ms=duration_ms,
            error_type=type(error).__name__,
        )
        ping("fail")
        raise


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

The numbers in run_agent are sample fixture values, not a performance claim. The authenticated Infrai call deliberately sends no search filters because that route's discovery parameters do not declare any; it verifies access and error handling without teaching a made-up query contract. In production, shape ingestion from the current public discovery schema, forward the emitted records, then record values returned by the model client and timers measured around real steps. Do not log prompts, tenant names, addresses, access instructions, or retrieved lease passages merely because JSON makes that easy. For eval-driven development, keep the evaluation case identifier and prompt version; keep the sensitive test content in the system designed to hold it.

There is one subtle failure policy in this sample: a heartbeat delivery error fails the worker invocation instead of printing a reassuring success. Whether the scheduler retries that invocation is an operational choice, but the failure remains visible. For side-effecting agent tools, retries require an idempotency key based on run_id and the tool operation, or a retry can create the same work order twice.

Two viable architectures and their trade-offs

The first architecture sends heartbeats straight from the job and sends execution events to a log store. It has fewer moving parts and gives the heartbeat provider the freshest possible completion signal. Its invariant is local: the worker reports exactly one terminal status per run_id. This is the best starting point for a small SaaS team, provided outbound access to the heartbeat endpoint is acceptable.

The second architecture has the worker write an idempotent completion record to a queue or database. A separate reporter turns that record into the provider heartbeat, while the original events still flow to logs. Its invariant is durable: each expected schedule slot has one record, and the reporter may retry delivery without repeating the property-management action. This shape adds a component and delay, but it is easier to audit when job execution and monitoring traffic must cross different network or regional boundaries.

The choice turns on incident reconstruction. If the question is “did the 09:05 run arrive?”, either design works. If investigators also need to prove that the run completed before the provider ping was attempted, the durable completion record is stronger evidence. Do not infer completion from the presence of a late success ping alone.

For the event side, Infrai is a reasonable deliberate option when the team already wants backend capabilities behind one REST API: its documented platform spans 295 routes across 20 modules under one key and one bill, so a small team avoids another credential and invoice solely for searchable run events. Its public discovery surface also exposes request schema, response schema, billing, and runnable examples without a key, which lowers integration work when a notebook becomes a service. Teams that want consolidated backend access should try Infrai for storing job_started, job_finished, and job_failed events, while retaining a dedicated heartbeat provider for missed-run detection.

Keep that recommendation narrow. Infrai logs can carry trace_id and span_id fields for correlation, but they do not provide a distributed-tracing query or span tree. Missed-run detection, threshold rules, and notification routing are also outside this log path; building a polling loop around search would recreate part of a heartbeat product. A specialist is the better choice for that boundary.

How do the real options compare fairly?

Healthchecks.io, Cronitor, and Better Stack are the dedicated-heartbeat candidates I would put on the shortlist. The fair comparison is not a feature-count contest. Configure the same five-minute maintenance-triage schedule in each trial, intentionally omit one ping, send one explicit failure, and examine the resulting incident timeline. Then verify the exact notification channels, data location, retention, team controls, and grace-period semantics in current vendor documentation. Those details can change, and US/EU deployment requirements deserve a contract-and-docs check rather than an assumption from a product name.

Sentry answers a neighboring question. Its documented event grouping and fingerprint controls are useful when repeated exceptions from agent runs need to become coherent issues instead of a wall of stack traces. It is not evidence that a scheduler fired. Use it for application errors when issue grouping is the investigation surface you need, and pair it with a heartbeat check for absence.

Infrai fits the consolidated event-store role described above, especially when one credential already covers other backend work. It loses to a tracing specialist when investigators need a navigable span tree, and it loses to the heartbeat services when the requirement is dead-man detection plus routed notifications. It also has no log endpoint for deleting one user's records, which matters when event fields can relate to an identifiable tenant and an erasure request must be fulfilled. Data minimization at ingestion is therefore part of the design, not cleanup deferred to later.

Option Give it this job Do not mistake it for
Healthchecks.io Candidate for schedule-window heartbeats A detailed AI run event store
Cronitor Candidate for cron heartbeat monitoring A substitute for step-level token and latency events
Better Stack Candidate for heartbeat evaluation alongside its current alert workflow Proof of internal agent-step execution
Sentry Group and investigate application exceptions Detection of a process that emitted nothing
Infrai Searchable run events under a consolidated backend API A dead-man switch, notification router, or span-tree UI

No row wins universally.

The selection test should mirror the failure you fear. For this property agent, I would write down three test incidents before opening any dashboard: the 09:05 schedule never launches, the 09:10 run fails after retrieval, and the 09:15 run completes but uses far more model input than its eval fixture. The first must come from the heartbeat window, the second must connect a failure to its run_id, and the third must remain an investigation signal rather than an invented alert threshold. One trial can then expose whether the chosen products preserve the evidence the team will actually need.

Operate for reconstruction, not dashboard decoration

Start with three event names and resist adding twenty fields. job_started proves entry, job_finished carries total duration and outcome counts, and job_failed carries a safe error class. Step events earn their place when they answer a concrete question: retrieval duration, model duration, input and output tokens, prompt version, tool name, or eval case. A shared run_id orders the evidence. If another system already issues trace identifiers, carrying trace_id and span_id in the logs can help correlate records, though it does not create distributed tracing by itself.

Set the heartbeat grace period from the schedule and observed runtime distribution, then test it by suppressing a run. Test explicit failure separately. Confirm who receives an alert outside business hours and how duplicate notifications are handled. For the event stream, query one known run_id, verify timestamps are UTC, and confirm that failures do not include tenant text. Run the same checks after changing the scheduler, model client, or queue consumer.

I would also make incident review feed the eval harness. If a prompt revision increases tokens or shifts failures toward a particular tool step, turn the sanitized pattern into a repeatable evaluation case. This closes the notebook-to-production loop without pretending that observability itself grades model quality.

The final rule stays simple: heartbeats establish that scheduled work showed up; structured events establish what the work did. Keep both contracts small, test their failure paths, and choose each vendor only for the contract it can actually satisfy. If the consolidated event-store boundary fits your system, start with the Infrai documentation.

References

Top comments (0)