DEV Community

LunarBreeze4173085
LunarBreeze4173085

Posted on

How to Build Startup Uptime Monitoring: 4 Postgres Healthcheck API Signals

Build the startup uptime monitoring API around four separate signals: an external probe, an internal healthcheck endpoint, a cron deadline, and a durable per-run cost-and-latency record for the e-commerce AI agent loop. The deciding constraint is rollback safety, because a green storefront endpoint says very little about a recommendation agent that is slow, expensive, or quietly missing scheduled work.

TL;DR: keep monitoring outside the deployment being judged, expose a shallow health response, store bounded run records in Postgres, and make status publication a derived view rather than part of the request path. Deploy the schema and writer before any alert depends on them; during rollback, old and new application versions must both be able to emit valid records.

This is an architecture decision record, not a vendor ranking. A startup can compare hosted services later, but the useful comparison begins with failure boundaries and data ownership, not a headline monthly price.

Rollbacks are the test.

What must remain true during a rollback?

The first invariant is independence: the component deciding that checkout or the agent is unavailable cannot share the same process, release, or failure domain as that workload. A probe running inside the API container may faithfully report that its own event loop is alive while DNS, routing, TLS, or the public load balancer is broken.

The second invariant is compatibility. An application rollback must not require a database rollback, and a monitor configuration rollback must not erase observations already accepted. Use additive schema changes, tolerate unknown event attributes, and retain a stable minimum record containing a run identifier, timestamps, outcome, latency, and cost expressed in an application-defined minor unit. Do not call a dependency from /health merely to make the endpoint look comprehensive; a dependency flap can then remove healthy instances from service and amplify the original fault.

The third invariant is boundedness. Prometheus instrumentation guidance warns that every unique label set creates another time series and specifically advises against high-cardinality labels such as user IDs. Order IDs, session IDs, prompt hashes, and agent run IDs belong in logs or durable records, not metric labels. Aggregate metrics can carry route, outcome class, and deployment revision only when each has a controlled value set.

Four signals cover different failure boundaries:

Signal Proves Does not prove Rollback rule
External probe Public DNS, TLS, routing, and shallow response work from another failure domain The agent completes useful work Keep probe configuration independent of the app release
Readiness check This instance can accept its intended traffic A background schedule is firing Preserve the endpoint contract across adjacent releases
Cron heartbeat A named job reached a deadline checkpoint Every item in the job was correct Version schedules separately and allow a grace window
Run record One agent loop's latency, outcome, and attributed cost were recorded The public endpoint is reachable Add columns before readers require them

No single green light is enough.

Record the critical path without coupling it to alerts

The critical path should write one terminal record for each agent run, with monotonic time used for elapsed duration because wall clocks can move. The example below uses only Python's standard library for the measurement and a generic database connection interface for persistence. cost_minor_units is deliberately supplied by the caller: currency, token accounting, and allocation policy are business definitions and should not be guessed by monitoring code.

from __future__ import annotations

import time
import uuid
from collections.abc import Callable
from typing import Any


def run_agent(
    connection: Any,
    operation: Callable[[], Any],
    deployment_revision: str,
    cost_minor_units: int,
) -> Any:
    run_id = str(uuid.uuid4())
    started_at_ns = time.time_ns()
    timer_start_ns = time.monotonic_ns()
    outcome = "succeeded"

    try:
        return operation()
    except Exception:
        outcome = "failed"
        raise
    finally:
        duration_ms = (time.monotonic_ns() - timer_start_ns) // 1_000_000
        with connection.cursor() as cursor:
            cursor.execute(
                """
                INSERT INTO agent_run_observations (
                    run_id,
                    started_at_ns,
                    duration_ms,
                    outcome,
                    deployment_revision,
                    cost_minor_units
                ) VALUES (%s, %s, %s, %s, %s, %s)
                ON CONFLICT (run_id) DO NOTHING
                """,
                (
                    run_id,
                    started_at_ns,
                    duration_ms,
                    outcome,
                    deployment_revision,
                    cost_minor_units,
                ),
            )
        connection.commit()
Enter fullscreen mode Exit fullscreen mode

There is an uncomfortable failure mode here: if persistence fails in finally, it can obscure the operation's original exception. Production code should capture the application exception and telemetry exception separately, increment a bounded failure counter, and apply a declared policy. For checkout-adjacent recommendations, the defensible default is usually to preserve the customer response and report telemetry loss out of band; for regulated accounting, failure to record may instead be a hard failure. That choice belongs in the service contract.

Idempotency matters too. The primary key on run_id makes a retry harmless, but it does not make two different generated IDs represent the same logical job. Scheduled work should derive or persist its execution identity before retrying. Without that boundary, a timeout followed by a retry can create two plausible records and double the apparent cost.

Deploy in four reversible steps

Start by adding the observation table without changing the current application. The old release continues to work, which is the first rollback checkpoint. A minimal migration can look like this:

MIGRATION_SQL = """
CREATE TABLE IF NOT EXISTS agent_run_observations (
    run_id uuid PRIMARY KEY,
    started_at_ns bigint NOT NULL,
    duration_ms bigint NOT NULL CHECK (duration_ms >= 0),
    outcome text NOT NULL CHECK (outcome IN ('succeeded', 'failed')),
    deployment_revision text NOT NULL,
    cost_minor_units bigint NOT NULL CHECK (cost_minor_units >= 0)
);

CREATE INDEX IF NOT EXISTS agent_run_observations_started_at_idx
    ON agent_run_observations (started_at_ns);
"""
Enter fullscreen mode Exit fullscreen mode

Second, deploy the writer behind a release control and verify that write failures are visible through a low-cardinality counter. Keep the previous code path available. Third, deploy readers and alert evaluation only after records from the new writer have arrived and passed basic validity checks. Fourth, enable status-page projection from evaluated service state; status publishing must consume observations asynchronously so its outage cannot slow checkout or the agent.

For every stage, write down the rollback action before rollout. The database table remains during application rollback. Readers must tolerate absent new fields. Alert rules should use a revision dimension only if its value count is bounded, and they should avoid firing on a single missing sample when ingestion delay is plausible.

Test the transitions, not merely the steady state: new writer with old reader, old writer with new reader, duplicate delivery, a database timeout after successful agent completion, a cron run that starts but never reaches its terminal heartbeat, and a status publisher that is unavailable for longer than its retry window. These tests expose ownership mistakes that endpoint polling cannot.

How should a startup evaluate an uptime monitoring API and healthcheck endpoint?

Use a short evidence-gathering trial with synthetic traffic, then compare contracts rather than feature counts. The service must support the required probe regions, data-processing terms, retention controls, export path, incident history, and heartbeat semantics. For an EU GDPR assessment, document the categories of personal data sent, processing locations, subprocessors, deletion behavior, access controls, and the lawful basis selected by the organization; a vendor's location or a generic compliance badge does not complete that assessment.

A concrete trial configuration might probe every 60 seconds, allow a 5-minute heartbeat grace period, and retain 30 days of run detail. Those values are examples, not universal recommendations: a five-minute grace period is unsuitable for a job whose missed deadline immediately blocks order fulfillment, while a 60-second public probe may be wasteful for an internal batch service with a daily objective. Record why each threshold follows from a customer or operational deadline, then inject a missed heartbeat and a failed deployment to verify the behavior. This explicit trade-off matters more than nominal check volume because an alert that arrives inside its configured window can still arrive too late for the business.

Cost belongs in the decision, but it is an input rather than the architecture. Model the startup's actual number of checks, check frequency, retained events, team seats, notification volume, status-page needs, and expected growth. Then ask what happens at each limit: sampling, rejection, overage, delayed evaluation, or silent truncation. Prices and packaging change, so record the date and source of every quote in the decision log instead of embedding a fragile number in application code.

The rollback test is sharper: can checks keep running while the monitored release is reverted; can observations be exported before termination; can alert configuration be versioned; and can the team restore the previous evaluation rules without rewriting history? Reject any option whose answer depends on an undocumented behavior.

This design has limitations. Postgres run records add write load and retention work, shallow probes cannot validate the full agent result, and asynchronous status publication accepts a period in which the public view trails internal evidence. Choose a deeper synthetic transaction when correctness across the whole purchase path is the requirement, but keep it separate from instance health so its failure does not trigger a deployment cascade.

Rejected option: synchronous all-in-one health evaluation

We rejected a deep /health handler that calls Postgres, the model dependency, the heartbeat store, and the status publisher on every probe. It combines unrelated failure modes, increases probe latency, and lets one dependency remove an otherwise useful instance from rotation. It also turns an external observation into application work precisely when the system is under stress.

That design still has a valid use case: a protected diagnostic endpoint invoked on demand during investigation, with strict timeouts and results that identify each dependency separately. It should not be the public liveness contract and should not publish status synchronously.

The durable decision is therefore modest: separate signals by what they can prove, keep identifiers out of metric labels, persist per-run detail behind an additive schema, and make every rollout stage compatible with the release immediately before it. This does not promise perfect detection. It gives operators evidence that survives the rollback they are trying to judge.

References

Top comments (0)