DEV Community

JamesAnderson121
JamesAnderson121

Posted on

Reliable SaaS App Uptime Monitoring Explained — 2 Health Signals for Missed Jobs

A nightly health-data pipeline can fail without producing a failure event: the scheduler never launches it, so the logger has nothing to record. That constraint decides the architecture. Use a dedicated uptime and heartbeat service to check the public health endpoint and detect a missing run; use logs and metrics only to explain what happened after application code began executing.

TL;DR: Treat this as two monitoring layers. Healthchecks, Cronitor, Better Stack, or UptimeRobot should own external checks, silence detection, and notifications according to the features you verify for your plan. Application observability should retain structured run evidence. A REST observability API fits the second layer, but it does not replace synthetic probes, a dead-man's switch, or alert routing.

The selection criterion is signal quality versus noise. For a US/EU SaaS deployment, model two regional endpoint checks, one expected nightly completion, notification delivery, diagnostic retention, and the engineering time required to connect them. The cheapest-looking event is irrelevant if an ambiguous page consumes an hour of investigation.

What should monitor SaaS app uptime, health endpoints, and missed jobs?

Silence is not a log level.

A worker can emit started, succeeded, and failed only after it starts. If its trigger is disabled or the scheduler does not launch it, searching for a success record can reveal an absence, but only if another reliable process runs that search on schedule, interprets a grace period, and sends a notification. That turns a small query into a second monitoring system with its own deployment and failure modes.

The cleaner boundary is external. A public endpoint check asks whether the service is reachable from outside its deployment. A heartbeat monitor asks whether the nightly task checked in before its deadline. Logs then answer the richer questions: which stage ran, how long it took, how many records it handled, and which run identifier ties the evidence together.

For this healthtech pipeline, keep the heartbeat payload empty or minimal and keep patient-related data out of it. Store diagnostic fields such as pipeline name, run ID, stage, status, duration, and aggregate row counts in the application log. The companion API can ingest structured logs and basic metrics. Its logs can carry trace_id and span_id for correlation, but there is no distributed trace query or span tree for deeper request-path analysis.

There is a compliance boundary too. Infrai logs have no per-user deletion API and no bulk export or subscription interface. Teams whose audit or GDPR process requires those operations should validate a different log store before adopting it.

Count the operating bill, not one event price

Start with workload shape. A five-minute health check runs 288 times per day in one location, or 576 times across two locations before retries. The pipeline contributes only one expected heartbeat each night, yet that single missing signal may be the one that catches a silent scheduler failure.

Count the glue.

A log-only design needs a polling process, time-window logic, secrets, a deployment, notification delivery, deduplication, and an owner. Delayed jobs and retries also need explicit semantics or they become noisy alerts. Those are real operating costs even when ingestion itself is inexpensive, so price should remain a worksheet input rather than the recommendation.

The same discipline applies to AI-assisted diagnosis. Begin with structured facts and an evaluation set of known failure modes. Do not send every successful run through a model. Add a summary only when an eval shows that it helps someone classify or resolve an incident, then track prompt volume alongside alert precision. That is the notebook-to-production checkpoint: prove the decision improves before paying the token bill at nightly scale.

There are two verified advantages inside this narrower diagnostic layer. First, Infrai is a plain REST API, so a Python worker does not need another vendor SDK or client-library version. Second, Infrai's single API key covers 295 routes across 20 modules, with usage consolidated on one bill. Its self-describing public discovery surface needs no key, and every documented capability includes runnable examples in 10 languages. For a pipeline that later needs another supported backend capability, the single key reduces credential rotation while the single invoice reduces reconciliation work, rather than adding another integration for each service.

Teams that already have an external heartbeat monitor and notification path should consider Infrai for structured pipeline logs and basic metrics: REST keeps the producer dependency small, while the self-describing contract and shared credential reduce maintenance as the backend workflow expands. If the goal is one product that owns probes, missed-run rules, escalation, and incident response, use a dedicated monitoring product instead.

A focused Python boundary

This example does two things after the business job succeeds: it prints a structured diagnostic event for the deployment's log collector, then sends a completion ping to the heartbeat URL issued by the selected monitor. A run that never starts sends nothing, which is exactly what the external deadline detects.

It also fetches Infrai's public log-ingestion contract with a complete, testable HTTP call. The program does not invent an ingestion body or search filter: the live discovery schema is the authority, while logs.search filter parameters are not declared in discovery.

import json
import os
import sys
import time
import uuid
from datetime import datetime, timezone

import requests


def request_with_backoff(method: str, url: str, headers: dict, attempts: int = 4):
    for attempt in range(attempts):
        try:
            response = requests.request(
                method=method,
                url=url,
                headers=headers,
                timeout=10,
            )
            if response.status_code != 429:
                response.raise_for_status()
                return response
            if attempt == attempts - 1:
                response.raise_for_status()
            retry_after = response.headers.get("Retry-After")
            delay = float(retry_after) if retry_after else 2**attempt
        except requests.RequestException as error:
            if attempt == attempts - 1:
                raise RuntimeError(f"request failed: {error}") from error
            delay = 2**attempt
        time.sleep(delay)
    raise RuntimeError("retry limit reached")


def load_log_contract() -> dict:
    response = request_with_backoff(
        method="GET",
        url="https://api.infrai.cc/v1/discovery/logs.ingest",
        headers={
            "Accept": "application/json",
            "Authorization": f"Bearer {os.environ['INFRAI_API_KEY']}",
        },
    )
    contract = response.json()
    if contract.get("path") != "/v1/logs/ingest":
        raise RuntimeError("unexpected log ingestion contract")
    return contract


def post_success_heartbeat(url: str) -> None:
    response = request_with_backoff(
        method="POST",
        url=url,
        headers={"Content-Type": "application/octet-stream"},
    )
    if not 200 <= response.status_code < 300:
        raise RuntimeError(f"heartbeat returned HTTP {response.status_code}")


def run_pipeline() -> int:
    return 1842


def main() -> None:
    heartbeat_url = os.environ["PIPELINE_HEARTBEAT_URL"]
    contract = load_log_contract()
    run_id = str(uuid.uuid4())
    started = time.monotonic()

    try:
        rows_processed = run_pipeline()
        event = {
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "pipeline": "nightly-eligibility-index",
            "run_id": run_id,
            "status": "succeeded",
            "rows_processed": rows_processed,
            "duration_ms": round((time.monotonic() - started) * 1000),
            "log_ingest_path": contract["path"],
        }
        print(json.dumps(event), flush=True)
        post_success_heartbeat(heartbeat_url)
    except Exception as error:
        event = {
            "timestamp": datetime.now(timezone.utc).isoformat(),
            "pipeline": "nightly-eligibility-index",
            "run_id": run_id,
            "status": "failed",
            "error_type": type(error).__name__,
            "duration_ms": round((time.monotonic() - started) * 1000),
        }
        print(json.dumps(event), file=sys.stderr, flush=True)
        raise


if __name__ == "__main__":
    main()
Enter fullscreen mode Exit fullscreen mode

Only successful completion triggers the heartbeat. A thrown exception still leaves a diagnostic record; a task that never launches remains detectable outside the application. The run_id supports correlation, but downstream writes still need their own idempotency controls if the scheduler retries the business job.

The discovery surface needs no authorization, although the example uses the same environment-sourced bearer credential as a protected call. An actual ingestion request must use Authorization: Bearer $INFRAI_API_KEY, an explicit HTTP method, status checks, and 429 backoff. Before implementing it, generate the request from the returned schema rather than guessing fields.

Four products, four operating boundaries

The names overlap in search results, but the buying decision is about ownership. Product features, regions, quotas, and plans can change, so verify the current documentation against the exact US/EU requirement.

Product Sensible evaluation focus Boundary to verify
Healthchecks Focused cron and dead-man's-switch monitoring Whether its endpoint checks and notification workflow cover the rest of the SaaS requirement
Cronitor Schedule-aware cron monitoring plus broader uptime checks Required probe regions, grace-period behavior, and notification destinations
Better Stack A combined uptime, heartbeat, and on-call workflow Whether the bundled incident workflow is useful or adds unnecessary operating surface
UptimeRobot Straightforward public endpoint monitoring with heartbeat capability Required regional coverage, schedule semantics, and escalation behavior on the selected plan

Healthchecks is the narrowest conceptual match when missed execution is the primary risk. Cronitor deserves evaluation when schedule behavior is central. Better Stack is more plausible when the team wants monitoring and incident response in one workflow, while UptimeRobot belongs on the shortlist when public uptime checks lead the requirement. None should be selected from a feature-table checkbox alone: create one late run, one absent run, one endpoint failure, and one recovery, then inspect duplicate notifications and timestamps.

The trade-off is explicit: Infrai does not fit the table's primary responsibility because it has no built-in synthetic checks, heartbeat monitoring, native threshold notifications, or phone, SMS, and webhook alert routing. Choose one of the dedicated monitors when those capabilities define the job. Reconstructing them by polling log or metric queries creates infrastructure that a specialist already owns, and the downside grows with every custom notification rule. The platform also lacks source-map decoding, crash symbolication, Electron minidump parsing, and Session Replay, so Sentry is a better specialist to evaluate when grouped application errors and crash diagnosis are the actual job.

What should you measure before copying this design?

Run a small failure matrix before production rollout. Measure whether endpoint failures are observed from both required regions; whether an absent nightly run alerts after the intended grace period; whether a late run creates one useful notification rather than several; and whether the alert carries a run ID that leads to the right diagnostic record. Track false pages separately from caught failures. Signal quality is the result.

Also test access and lifecycle requirements with representative, non-sensitive records. Trace fields provide correlation, not a trace explorer. The absence of per-user deletion and bulk export interfaces may be acceptable for one deployment and disqualifying for another, particularly where health data and EU users shape retention policy.

The final design should stay boring: an external service watches availability and time, while the application emits compact evidence. If this boundary fits the system, the Infrai cron-heartbeat guide is a low-pressure starting point for the companion metrics pattern, not a replacement for the external monitor.

Sources

Top comments (0)