DEV Community

TheodorHawkins9251
TheodorHawkins9251

Posted on

Postgres Uptime Monitoring — Better API Healthcheck and Cron Signals for Small SaaS

A logistics notification service should choose monitoring by the evidence required for a safe rollback, not by the length of a feature checklist. TL;DR: use an external request probe for reachability, a cron heartbeat for scheduled delivery work, and an application-level check backed by durable attempt records; then require the old and new release to emit compatible signals during the rollback window. Pingdom, UptimeRobot, and Healthchecks can sit at different points in that design, but none of those names answers the important question: can an operator distinguish a dead endpoint, a stalled worker, and a delivery provider rejection before deciding to reverse a release?

The distinction matters in logistics. A green /health response says very little if a shipment-delay message remains queued, a retry is duplicated, or yesterday's cleanup job never ran. Rollback safety makes the constraint sharper: monitoring introduced with release N must still explain the system after traffic returns to N-1. If the dashboard depends on a field or status that only N understands, the supposed safety net disappears during the operation that needs it most.

What makes uptime monitoring better for an API and cron jobs?

Start with the failure decision, then work backward to signals. For a delivery notification service, an operator usually needs to know whether the public API accepts work, whether workers make progress, whether scheduled reconciliation runs, and whether each notification attempt reaches a terminal state. These are separate claims with separate clocks. Combining them into one Boolean hides causality.

The durable record is the anchor. Store a notification identifier, an attempt identifier, the release that created the attempt, a coarse outcome, and timestamps for creation and completion. Keep labels bounded: carrier, region, customer, shipment, and raw error text belong in logs or queryable storage, not metric label sets. Prometheus explicitly warns that every unique label combination creates a new time series and recommends avoiding high-cardinality labels such as user IDs; shipment IDs have the same shape.

A rollback-compatible health contract should be additive. Release N may add a field, but it should not reinterpret ok, remove an outcome the previous release writes, or require a migration that prevents N-1 from updating existing rows. Schema expansion comes first, application changes second, and destructive cleanup waits until rollback is no longer allowed. This is slower than editing a response in place. It is also inspectable.

Use at least three signal classes:

  • An outside-in probe confirms DNS, TLS, routing, and the response contract from a real network path.
  • A worker-progress signal measures accepted work reaching a durable terminal state within a stated time budget.
  • A scheduled-job heartbeat proves that reconciliation or retry recovery completed, not merely that a scheduler attempted to start it.

Three clocks. Three failure domains.

Model evidence before selecting a service

The health endpoint should report bounded facts that both release versions can compute. It must not scan an unbounded attempts table on every probe, and it must not declare success merely because the process has a listening socket. A small aggregate table, refreshed by the worker transaction or a separate bounded query, keeps probe cost predictable.

The following Python sketch uses a generic database interface and returns only stable, low-cardinality fields. The threshold is an explicit policy input rather than a magic property of a monitoring product.

from dataclasses import asdict, dataclass
from datetime import datetime, timedelta, timezone


@dataclass(frozen=True)
class HealthResult:
    ok: bool
    component: str
    reason: str


def delivery_health(db, max_stall: timedelta) -> dict:
    row = db.fetch_one(
        """
        SELECT oldest_pending_at, last_terminal_at
        FROM notification_progress
        WHERE shard_name = %s
        """,
        ("default",),
    )
    now = datetime.now(timezone.utc)

    if row is None:
        return asdict(HealthResult(False, "delivery_worker", "no_progress_record"))

    oldest = row["oldest_pending_at"]
    if oldest is not None and now - oldest > max_stall:
        return asdict(HealthResult(False, "delivery_worker", "pending_work_stalled"))

    return asdict(HealthResult(True, "delivery_worker", "progress_within_budget"))
Enter fullscreen mode Exit fullscreen mode

This code deliberately omits shipment identifiers and provider response text. Those details are valuable during diagnosis, but exposing them as probe fields expands the contract and may leak operational data. The endpoint's job is to make a bounded assertion; the durable attempt table and structured logs supply the trail behind it.

Cron monitoring requires another contract. Send the heartbeat after reconciliation commits its result, not at process start. A start-only ping can turn a hung task green. If the task can legitimately run longer than its interval, record start and completion separately in storage and alert on overdue completion; otherwise overlapping executions can look like healthy frequency while both compete for the same work.

The comparison is about boundaries, not a winner

The names in the original selection question suggest three tools, yet a defensible evaluation cannot infer current regional coverage, retention, alert routing, or account limits from a brand name. Those details change and must be verified against current documentation and a trial from the actual US and EU deployment paths. The useful comparison is the monitoring role each candidate is being asked to fill. The main limitation of an HTTP-only design is blunt: it can prove that a response arrived, but it cannot prove that queued delivery work advanced or that a cron run committed its result.

Evidence needed Evaluation exercise Failure mode it must expose Rollback concern
Public API reachability Probe the same stable endpoint from required US and EU paths DNS, TLS, routing, timeout, or invalid response Both releases must honor the response contract
Scheduled reconciliation completion Withhold one completion heartbeat in staging Missed, hung, or late cron execution Heartbeat identity must survive a version reversal
Notification progress Queue a synthetic record and observe its durable terminal state Worker stall or a growing pending age Old code must read and update the expanded schema
Incident evidence Trigger one controlled failure and inspect timestamps and severity Alert without enough context to decide Evidence must remain queryable after deploy reversal

Evaluate Pingdom, UptimeRobot, and Healthchecks against that same worksheet. Record observed behavior, documentation links, and the date of the test. Do not award points for an integration the team will not operate, and do not treat a successful HTTP request as proof of background progress. A tool may cover more than one row; the architecture should not assume that it does until the exercise demonstrates it.

For a small SaaS, operational load belongs in the decision as well. Count the number of checks, expected alert events, credential rotations, ownership handoffs, and independent failure paths the team must maintain. Price can be recorded after the design fits, but it is a weak primary axis because an inexpensive checker that cannot distinguish backlog from endpoint failure makes rollback decisions riskier.

There is another limit: geography is evidence, not a checkbox. A US and EU probe can establish that those particular paths completed at those times. It cannot prove every customer path was healthy, and it cannot establish notification delivery by itself. Document precisely what each green state means.

Make alerts support the rollback decision

An alert should state the violated invariant and carry enough timing information to choose an action. pending_work_stalled is more useful than service unhealthy, but it still needs the first observed time, deployment version, and a link or query key for the underlying records. Avoid putting unique notification IDs into metric labels. Use them as trace or log correlation fields.

Severity also needs discipline. RFC 5424 defines standardized severity levels from Emergency through Debug, but assigning a level remains an operational policy decision. A failed synthetic probe need not be an emergency; repeated external failure combined with stalled notification progress may justify a higher response. Map conditions to paging policy explicitly, and test the mapping.

False agreement is dangerous. Two HTTP probes that traverse the same DNS and edge path are not independent confirmation, while an external probe and a transactionally updated progress record observe different parts of the system. Conversely, conflicting signals are useful: an available API plus increasing pending age points toward worker progress, whereas external failure plus fresh terminal records points toward ingress or probe-path trouble. Design for disagreement, because disagreement narrows the rollback question.

Before production, run four controlled cases: stop a worker, suppress a cron completion, return an invalid health payload, and deploy an additive schema version followed by the previous application version. The test passes only when each condition produces the expected signal and the rollback leaves monitoring interpretable. Do not fabricate a universal timeout. Derive it from the notification service's declared delivery objective, retry schedule, and longest valid job duration, then keep that value in versioned configuration.

Roll out the contract in reversible steps

First, add the durable progress fields and keep old readers working. Next, deploy producers that write both the established fields and any new additive fields. Then expose the bounded application check and run it without paging while comparing it with stored attempts. Add external probing and completion heartbeats after their failure simulations behave as expected.

Only then attach paging rules.

During the rollback window, retain the previous alert interpretation and dashboard queries alongside the new ones. Remove old fields, heartbeat identities, or queries only after the release policy says reversal is no longer permitted and stored evidence confirms that no active records depend on them. The final selection should therefore be the candidate set that passes the team's failure exercises, preserves evidence across versions, and creates an operational burden the team can sustain. Product labels do not alter that rule.

Sources

Top comments (0)