DEV Community

KasimirBerg5341
KasimirBerg5341

Posted on

Rollback Safe Critical Error Tracking with Cron API Poll Alerts Explained

A detector for stopped property imports should alert on missing successful outcomes, not merely on new exceptions, and its state changes must be reversible. Short answer: poll the error source and the import-result ledger on a fixed schedule, persist a high-water mark only after notification delivery is accepted, and attach an idempotency key to every Slack, email, or generic webhook dispatch. That ordering makes a failed deployment or notification retry recoverable without silently advancing past an incident.

An error tracker answers one half of the question: did a new critical error appear? Property operations supply the harder half: did the scheduled rent-roll, unit, or resident import produce anything at all? A crashed worker may create an exception, but an expired upstream credential, an empty file, a scheduler failure, or a stuck queue can leave no new error event. Silence is data.

How should cron poll an API for new critical error alerts?

Define the service-level event before writing the poller. For each property and import type, record the scheduled window, the last accepted result, the expected result count policy, and the latest critical error cursor. An import is late only after its declared completion deadline plus a bounded grace period. An import that completes with zero rows is different: it may be valid for one feed and critical for another, so the rule belongs in per-feed configuration rather than a global count == 0 test.

This distinction prevents a common category error. Error counts describe failures that were observed; freshness describes expected work that did not become visible. The Google SRE monitoring model separates errors from latency and traffic for the same reason: one signal cannot stand in for the others. For an import pipeline, the useful signals are last successful completion time, rows accepted, run duration, and critical errors since the last committed cursor.

Use stable identifiers. A property display name can change, while a property ID, import kind, scheduled window, and detector rule version can form a durable incident key. That key should survive retries and process restarts.

Names drift.

Deriving the state machine from rollback safety

The tempting implementation is fetch, save cursor, send message. It is unsafe. If message delivery fails after the cursor advances, the next run treats the critical error as old and nobody is notified. Reverse those last two operations: evaluate a snapshot, dispatch the notification with an idempotency key, then commit the cursor and incident state after the receiver accepts the request.

No cursor commit yet.

A small state machine is enough: healthy, pending, alerting, and recovered. pending absorbs the grace period. alerting remains active across repeated polls but does not create a fresh incident each time. recovered is emitted only after a later successful result passes the same acceptance rule that declared the feed healthy in the first place.

Rollback complicates rule changes. Suppose version 8 lengthens a grace period and is then rolled back to version 7. If both versions overwrite one shared state record, the rollback can reinterpret an old deadline and reopen or suppress an incident. Store rule_version beside the incident key, and deploy a new rule in shadow mode before it is allowed to notify. Keep the previous evaluator readable until the migration window ends. This costs a little state; it buys deterministic reversal.

Design choice Failure mode Rollback-safe decision
Save cursor before dispatch Delivery failure permanently skips an event Commit after accepted delivery
Alert on exceptions alone Scheduler or empty-output failures stay invisible Join errors with result freshness
Use one mutable rule state Rollback reinterprets prior decisions Version rule and incident state
Retry with a new message ID Receivers create duplicate pages Reuse a deterministic idempotency key
Treat every poll as a new incident Alert storms hide the original fault Maintain an open incident until recovery

The limit should be explicit: an HTTP acceptance response proves that the receiver accepted the request, not that a person read a Slack message or email. If human acknowledgement is required, model acknowledgement as another state transition rather than claiming delivery means resolution.

Consider a property feed scheduled for 02:00 UTC with a 20-minute grace period and a five-minute polling interval. At 02:19, no result is late, even if the ledger is empty. At 02:21, one detector may open the incident; a second overlapping invocation must lose the conditional write. If notification delivery is accepted at 02:21 but the process exits before committing, the 02:26 run reuses the same key. That duplicate attempt is deliberate. Moving the cursor earlier would produce a quieter system, but the quiet would hide an unreported incident, which is the wrong trade-off for a rent-roll import whose downstream work starts on the assumption that the ledger is current.

A minimal poller with a durable commit boundary

The example deliberately leaves the transport and persistence adapters generic. ErrorSource returns critical events strictly after a cursor; ResultLedger returns the latest accepted import result; Notifier accepts a deterministic key; and StateStore.compare_and_set prevents overlapping cron invocations from committing over each other. Those contracts need integration tests against the chosen systems.

from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
from hashlib import sha256
from typing import Protocol, Sequence


@dataclass(frozen=True)
class CriticalEvent:
    cursor: str
    summary: str


@dataclass(frozen=True)
class ImportResult:
    completed_at: datetime
    accepted_rows: int


@dataclass(frozen=True)
class Checkpoint:
    revision: int
    error_cursor: str | None
    incident_open: bool


class ErrorSource(Protocol):
    def critical_after(self, cursor: str | None) -> Sequence[CriticalEvent]: ...


class ResultLedger(Protocol):
    def latest(self, property_id: str, import_kind: str) -> ImportResult | None: ...


class Notifier(Protocol):
    def send(self, *, key: str, subject: str, body: str) -> None: ...


class StateStore(Protocol):
    def load(self, key: str) -> Checkpoint: ...

    def compare_and_set(
        self, key: str, expected_revision: int, value: Checkpoint
    ) -> bool: ...


def stable_key(*parts: str) -> str:
    return sha256("|".join(parts).encode("utf-8")).hexdigest()


def check_import(
    *, property_id: str, import_kind: str, expected_by: datetime,
    grace: timedelta, rule_version: str, errors: ErrorSource,
    results: ResultLedger, notifier: Notifier, states: StateStore,
    now: datetime | None = None,
) -> bool:
    observed_at = now or datetime.now(timezone.utc)
    state_key = f"{property_id}:{import_kind}:{rule_version}"
    checkpoint = states.load(state_key)
    events = errors.critical_after(checkpoint.error_cursor)
    result = results.latest(property_id, import_kind)
    late = observed_at > expected_by + grace
    fresh = result is not None and result.completed_at >= expected_by
    should_alert = bool(events) or (late and not fresh)

    next_cursor = events[-1].cursor if events else checkpoint.error_cursor
    transition = "open" if should_alert else "recovered"
    if should_alert or (checkpoint.incident_open and fresh):
        window = expected_by.astimezone(timezone.utc).isoformat()
        notification_key = stable_key(
            property_id, import_kind, window, rule_version, transition
        )
        reasons = [event.summary for event in events]
        if late and not fresh:
            reasons.append("no accepted result by the freshness deadline")
        notifier.send(
            key=notification_key,
            subject=f"Import {transition}: {property_id} {import_kind}",
            body="; ".join(reasons) or "a fresh accepted result is visible",
        )

    updated = Checkpoint(
        revision=checkpoint.revision + 1,
        error_cursor=next_cursor,
        incident_open=should_alert,
    )
    return states.compare_and_set(state_key, checkpoint.revision, updated)
Enter fullscreen mode Exit fullscreen mode

A cron scheduler can call this function once per evaluation interval, but the interval is an operational choice, not a magic constant. It must be shorter than the tolerated detection delay, while the grace period must cover normal completion jitter. The state store needs conditional writes; without them, two overlapping invocations can both notify and then overwrite one another.

The notifier adapter may fan out to Slack, email, and a generic webhook. Fan-out needs its own delivery ledger per destination, because a Slack acceptance and an email acceptance are independent outcomes. Never mark the whole notification complete after the first destination succeeds. Cap retry duration, retain the original idempotency key, and route exhausted deliveries to an operator-visible queue.

Failure modes worth testing before cron runs

Start with time. Test a result one microsecond before the deadline, exactly at the deadline, and immediately after the grace period; use timezone-aware UTC values throughout the evaluator. Then run two detector processes against the same checkpoint and verify that only one conditional commit wins. The losing process should reload state and reevaluate, not force a write.

Next, inject failures after notification acceptance but before checkpoint commit. The following poll will send the same idempotency key, so the receiver or delivery ledger must collapse the duplicate. Inject the opposite boundary too: fail before acceptance and prove the cursor remains unchanged. These two tests are the heart of the design.

Other cases are mundane and expensive when missed: pagination that returns several critical events, a deleted or malformed cursor, a result ledger that is temporarily unavailable, clock skew, an import that reports completion before its transaction is visible, and a recovery followed immediately by another failure. Do not convert an unavailable dependency into a claim that the import is healthy. Alerting on detector health should be separate from alerting on import health, otherwise the monitor can manufacture false property incidents during its own outage.

Three metrics reveal most operational trouble: evaluation lag, oldest uncommitted notification age, and consecutive detector failures. Retain structured records containing the incident key, rule version, evaluated deadline, observed result timestamp, error cursor, delivery outcome, and commit revision. Avoid copying resident payloads or credentials into alert text; identifiers and reason codes are usually enough for triage.

Comparison after the constraints are known

Polling is appropriate when the source exposes ordered reads and the tolerated delay exceeds the polling interval. A push webhook reduces routine reads and can lower detection latency, but it still needs signature verification, replay protection, durable receipt, and a reconciliation poll because delivery can fail. Queue consumption gives stronger coordination when the producer and consumer share a durable broker contract, although retention and redelivery semantics then become part of the incident design.

This design has limits. Scheduled API polling is a poor fit when the required detection delay is shorter than a safe request interval, when the source cannot provide stable ordering or pagination, or when API quotas cannot absorb every tenant check. In those cases, use a durable event stream or signed push delivery, then retain a slower reconciliation read. The trade-off moves: push reduces routine polling delay, but replay protection and durable receipt become mandatory parts of the failure surface.

There is no free transport.

Mechanism Useful boundary Primary risk Rollback implication
Scheduled polling Existing read API and minute-scale detection Cursor gaps, pagination, overlapping runs Preserve old cursor reader during schema changes
Push webhook Source can deliver events promptly Missed delivery or replay Keep reconciliation active during rollback
Durable queue Producer can publish to a shared log Retention and poison messages Version consumers and message schemas

Cost belongs in the constraint set, but it is not the architecture. Higher poll frequency increases API calls; retaining every raw event increases storage; indexing all payload fields can add a separate ingestion and indexing charge in commercial log systems. Estimate calls per day, retained state, notification volume, and operator load with the actual tenancy count before choosing an interval. No universal interval follows from a pricing page.

Roll out without betting the alert channel

Run the evaluator in shadow mode for at least one complete import schedule: write decisions and metrics, send nothing. Compare those decisions with the result ledger, then enable one internal destination for a small property cohort. Expand by cohort only after late, empty, critical-error, duplicate-run, and recovery cases behave as specified.

Keep the prior rule version and checkpoint reader deployable while the new cohort is active. A rollback should stop new evaluations under the new version without deleting its state; deleting state destroys the evidence needed to explain duplicate or missing notifications. Once the observation window closes and no rollback is plausible, archive the old rule state according to the system's retention policy.

The deciding principle is compact: a monitor that cannot replay safely cannot be trusted to detect silence. Make freshness a first-class signal, commit progress after accepted delivery, and version the state that gives each alert its meaning.

Sources

Top comments (0)