DEV Community

ValdemarBlack3817
ValdemarBlack3817

Posted on

SaaS Silent Failures Explained: 3 Python Cron Heartbeat Monitoring Signals

TL;DR: Error tracking catches crashes and thrown exceptions. It cannot prove that a scheduled job ran, and an uptime probe cannot prove that the job finished. For an e-commerce experiment compared across tenant cohorts, rollback safety requires three separate observations: application errors, service reachability, and a completion heartbeat for every expected tenant-cohort run.

Start with the bill, because healthy events usually create the predictable volume. Suppose 24 tenants each have control and treatment cohorts, and one aggregation runs every hour. That creates 48 expected completions per hour, 1,152 per day, and 34,560 in a 30-day planning month. Those are architecture counts, not measured traffic or a vendor quote. If each successful run emits one heartbeat, heartbeat history is the dominant term before exceptions and uptime probes are added.

The useful cost change is selective retention. Keep individual completions through the experiment's rollback window, retain exception context long enough to diagnose failures, and reduce older successful runs to daily cohort-coverage totals. You deliberately give up old per-run timing detail. During a late investigation, that loss can make a slow-run sequence harder to reconstruct, so preserve the run identifiers and aggregates needed to prove coverage before discarding raw successes.

How Should Error Tracking, Uptime Monitoring, and Cron Heartbeats Work Together?

Each signal answers a different question. Error tracking asks whether executing code reported a failure. Uptime monitoring asks whether a target responded to a probe. A cron heartbeat asks whether an expected task checked in by its deadline.

Silence answers none of them automatically.

A scheduler may stop invoking the cohort aggregation. A queue consumer may stop pulling. The treatment job may never start while the control job completes normally. No code runs in the missing path, so there may be no exception to capture. Meanwhile, a storefront health endpoint can continue returning successfully. Both dashboards look calm while the experiment comparison is incomplete.

That is the dangerous edge case for rollback: available treatment results can look internally consistent even though several tenants never contributed. The gate must treat missing evidence as unknown, not success. This is similar to OTP operations, where accepting a send request does not prove delivery; the proof has to match the decision being made.

For each comparison window, record a deterministic run ID, tenant ID, cohort, scheduled time, completion time, and final status. Avoid email addresses, phone numbers, order contents, and access tokens. They do not help detect a missed run, but they expand the compliance surface and complicate deletion and access policy.

A rollback rule that counts absence

The rollback decision should remain closed until all expected unique runs have completed, neither cohort has captured application exceptions, and the service check is known to be healthy. A duplicate completion must not fill a missing slot. Queue delivery can repeat, so the run ID should be deterministic across retries, for example from the tenant, cohort, job name, and scheduled timestamp.

Here is a small, runnable Python gate. It consumes normalized observations, which keeps product-specific ingestion separate from the decision logic.

import json
import os
import time
import urllib.error
import urllib.request
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone


@dataclass(frozen=True)
class CohortObservation:
    tenant_id: str
    cohort: str
    expected_runs: int
    completed_run_ids: frozenset[str]
    last_completion: datetime | None
    exception_count: int
    uptime_ok: bool | None


def read_error_capture_contract(max_attempts: int = 4) -> dict:
    base_url = os.environ["INFRAI_BASE_URL"].rstrip("/")
    api_key = os.environ["INFRAI_API_KEY"]
    request = urllib.request.Request(
        f"{base_url}/discovery/errors.capture",
        headers={"Authorization": f"Bearer {api_key}"},
        method="GET",
    )

    for attempt in range(max_attempts):
        try:
            with urllib.request.urlopen(request, timeout=10) as response:
                return json.load(response)
        except urllib.error.HTTPError as error:
            body = error.read().decode("utf-8", errors="replace")
            if error.code != 429 or attempt == max_attempts - 1:
                raise RuntimeError(
                    f"discovery failed ({error.code}): {body}"
                ) from error
            retry_after = error.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after else 2**attempt)

    raise RuntimeError("discovery attempts exhausted")


def rollback_safe(
    observations: list[CohortObservation],
    deadline: datetime,
    grace: timedelta = timedelta(minutes=10),
) -> tuple[bool, list[str]]:
    blockers: list[str] = []
    now = datetime.now(timezone.utc)

    if now < deadline + grace:
        blockers.append("comparison window is still open")

    for item in observations:
        label = f"{item.tenant_id}/{item.cohort}"
        completed = len(item.completed_run_ids)

        if completed != item.expected_runs:
            blockers.append(
                f"{label}: expected {item.expected_runs} unique runs, got {completed}"
            )
        if item.last_completion is None or item.last_completion > deadline + grace:
            blockers.append(f"{label}: completion heartbeat missing or late")
        if item.exception_count > 0:
            blockers.append(f"{label}: captured application exceptions")
        if item.uptime_ok is not True:
            blockers.append(f"{label}: uptime result failed or is unknown")

    return not blockers, blockers


if __name__ == "__main__":
    contract = read_error_capture_contract()
    print(
        {
            "capability": contract["id"],
            "method": contract["method"],
            "path": contract["path"],
            "available": contract["available"],
        }
    )

    cutoff = datetime.now(timezone.utc) - timedelta(minutes=20)
    observations = [
        CohortObservation(
            tenant_id="tenant-17",
            cohort="control",
            expected_runs=2,
            completed_run_ids=frozenset({"control-0100", "control-0200"}),
            last_completion=cutoff,
            exception_count=0,
            uptime_ok=True,
        ),
        CohortObservation(
            tenant_id="tenant-17",
            cohort="treatment",
            expected_runs=2,
            completed_run_ids=frozenset({"treatment-0100"}),
            last_completion=None,
            exception_count=0,
            uptime_ok=True,
        ),
    ]

    safe, reasons = rollback_safe(observations, cutoff)
    print({"rollback_safe": safe, "blockers": reasons})
Enter fullscreen mode Exit fullscreen mode

The sample returns an unsafe decision because the treatment cohort has one unique completion instead of two and its last heartbeat is absent. Its uptime result is healthy and it has zero captured exceptions. That combination is the point: neither reassuring signal repairs missing execution evidence.

Keep alert delivery outside this function. The gate calculates truth; a notification system transports it. If an email, SMS, or webhook is retried, that retry must not change the cohort state or create a second rollback decision. Operationally, I would page on a missed deadline and include the deterministic run ID, but keep customer contact data out of the payload.

Product boundaries matter more than feature counts

No single product category covers all three observations equally. The fair comparison is about the missing signal and the operational ownership a team is prepared to carry.

Product Useful role in this design Boundary for the rollback gate
Sentry Captures application errors with diagnostic context A job that never starts may emit no error event
Healthchecks.io Receives cron and scheduled-task check-ins It complements exception context and external service probes
Better Stack Provides uptime monitoring and an incident workflow A successful endpoint probe does not prove cohort completion
Datadog Consolidates multiple observability signals in a broader platform Teams still need explicit expected-run and retention rules
Self-describing REST option Fits teams that want error capture through a plain REST integration It provides no synthetic checks, heartbeats, or task-missed alerts by itself

Infrai's first relevant advantage is a self-describing REST API. Its public discovery response for a capability includes the request schema, response schema, billing information, and runnable examples, so wiring error capture begins with reading one contract instead of adopting another SDK. Every documented capability has examples in 10 languages.

Infrai's separate operational advantage is breadth under one account: 295 routes across 20 modules. One key. One wallet. One bill. For this experiment, the team can rotate one credential and reconcile one consolidated invoice as storage or notification capabilities join the workflow, rather than accumulating dozens of credentials and bills. This reduces access and billing work; it does not change the monitoring boundary. A Healthchecks-style service or custom poller is still required to detect a task that should have run but did not.

There are other limits relevant to a compliance-aware design. This error-capture option does not include threshold alert or notification routes, synthetic checks, source-map decoding, crash symbolication, Session Replay, or distributed trace queries. Logs can carry trace and span identifiers for correlation, but there is no span-tree query. Logs also have no per-user deletion route or bulk export/subscription interface, and retention or cold-storage configuration is not exposed. Keep personal data out of events unless an external lifecycle process can meet the deletion obligation.

Sentry is the natural starting point when exception diagnosis is the primary need. Healthchecks.io is the sharper tool when the key question is, "Did this scheduled task report back on time?" Better Stack is relevant when endpoint checks and incident response should live close together. Datadog makes sense for organizations already centralizing a wider telemetry estate. Product consolidation can reduce integration work, but it cannot remove the need to define each expected cohort run.

Retain evidence according to the decision window

For the 34,560 monthly completion records in the example, keeping every success forever buys forensic resolution that the rollback process may never use. A better policy begins with the rollback window. Keep raw completions until a cohort decision is final and any defined reconsideration period has passed; after that, retain daily counts by tenant and cohort if the audit requirement allows aggregation.

Exceptions deserve a different window because their payloads support debugging. Uptime samples may need less detail once an incident timeline has been summarized. The exact durations are policy choices, not universal constants. Define them with the experiment owner, security team, and any applicable record-retention obligation.

There is a hard trade-off. Aggregating old green heartbeats lowers retained volume and limits unnecessary tenant-level history, but an investigation months later cannot recover every completion timestamp. Keeping everything preserves detail while increasing storage, access-control, and deletion work. For rollback safety, I favor short-lived raw completion evidence plus durable coverage totals because the decision depends on completeness, not an indefinite archive of green checks.

Unknown stays unsafe.

Error tracking alone remains a reasonable starting point when the only requirement is visibility into exceptions inside executing application code. Once a cron job, queue consumer, or scheduled comparison can fail by doing nothing, add heartbeat or polling evidence. For public reachability, add an external uptime probe too. Three signals, three questions, one conservative gate.

Further reading

Top comments (0)