DEV Community

YukiKobayashi880
YukiKobayashi880

Posted on

Rollback-Safe Health Check Endpoint Readiness and Liveness (Notification Delivery)

Use liveness to answer whether the process can still execute, readiness to answer whether this instance should receive new notification work, and external uptime monitoring to answer whether a caller can traverse the deployed path. For rollback safety, keep those contracts stable across releases and make dependency policy explicit: a failed PostgreSQL check should usually stop new work that requires durable state, while a failed Redis check should do so only when the delivery path cannot degrade correctly without Redis.

Short answer: expose separate /live and /ready endpoints, keep liveness free of network dependencies, give readiness checks strict time budgets, return a small versioned response, and monitor the public path from outside the service. A single all-dependencies endpoint couples process recovery, traffic routing, and incident detection; during a bad release, that coupling can turn a limited dependency fault into restart churn or hide which revision is safe to restore.

How Should a Health Check Endpoint Separate Readiness and Liveness?

Start with invariants, not a list of technologies. A B2B notification service must not accept delivery work it cannot record durably. It must also avoid claiming success when the durable record is ambiguous. Those two rules make PostgreSQL part of readiness for request paths that create or transition delivery records. Redis is conditional: if it holds a replaceable cache and the application has a bounded fallback, cache failure can remain visible in metrics without removing the instance; if it supplies mandatory rate-limit state, deduplication state, or the only work queue used by the path, treating it as optional would make readiness dishonest. The probe contract follows the data invariant, not the dependency's brand or speed. Liveness has a narrower failure boundary: event-loop or process progress. It must not ask PostgreSQL or Redis for permission to keep the process alive, because dependency outages can last longer than a restart, and restarting healthy processes does not repair a remote system.

That boundary matters.

Keep the response boring. A status, contract version, build identifier, and check names are enough for routing and diagnosis; credentials, hostnames, customer identifiers, queue payloads, and raw exceptions do not belong in a public probe response.

Record the probe contract as an architecture decision

The decision is to separate three observers because they control different actions. The table is deliberately about failure effects and rollback, not feature counts.

Signal Question answered Dependencies Failure action Rollback value
Liveness Can this process still make progress? None over the network Restart only after process failure Avoids restart storms during shared outages
Readiness Can this revision safely accept new delivery work? Only dependencies required by the path Remove the instance from new traffic Lets old and new revisions be compared under one contract
External uptime Can a caller reach the deployed service path? DNS, network, edge, application Alert and investigate Detects failures that an in-process probe cannot see

The important limit is scope. Readiness is not proof that a notification reached a recipient, and external uptime is not proof that every tenant, channel, or background worker is healthy. Delivery outcomes need their own counters and latency distributions, partitioned cautiously so tenant or message identifiers do not create unbounded metric cardinality.

For metrics, use one unit per name, prefer base units, and include a suffix that identifies the unit, as detailed in the naming guidance linked below. A duration exported as seconds is easier to combine correctly than a mixture of milliseconds and seconds. Names and labels are an API too.

Put the critical path under one deadline

The following Python models the control flow rather than a framework-specific setup. Its key property is one readiness budget shared by concurrent checks. Sequential checks quietly multiply worst-case probe time, which makes a failing dependency consume more of the routing system's patience than the contract intended.

import asyncio
from dataclasses import dataclass
from typing import Awaitable, Callable


Check = Callable[[], Awaitable[None]]


@dataclass(frozen=True)
class ProbeResult:
    status_code: int
    body: dict


async def readiness(
    checks: dict[str, Check],
    required: set[str],
    timeout_seconds: float = 0.8,
) -> ProbeResult:
    async def run(name: str, check: Check) -> tuple[str, str]:
        try:
            await check()
            return name, "ok"
        except Exception:
            return name, "failed"

    try:
        async with asyncio.timeout(timeout_seconds):
            pairs = await asyncio.gather(
                *(run(name, check) for name, check in checks.items())
            )
    except TimeoutError:
        return ProbeResult(
            503,
            {"status": "not_ready", "contract": 1, "reason": "deadline"},
        )

    states = dict(pairs)
    unavailable = sorted(
        name for name in required if states.get(name) != "ok"
    )
    status = "ready" if not unavailable else "not_ready"
    return ProbeResult(
        200 if not unavailable else 503,
        {
            "status": status,
            "contract": 1,
            "checks": states,
            "unavailable_required": unavailable,
        },
    )


async def liveness() -> ProbeResult:
    return ProbeResult(200, {"status": "alive", "contract": 1})
Enter fullscreen mode Exit fullscreen mode

The 0.8 second budget is an example configuration, not a universal threshold. Set it below the caller's probe deadline, then validate it with failure injection: refused connections, stalled connections, exhausted pools, authentication rejection, and a dependency that accepts a connection but never completes the check. The useful test is not merely “does failure return 503?” It is “does every failure return within the promised budget without consuming the request pool needed for recovery?”

There is another trap. If the probe borrows from the same saturated connection pool as delivery requests, it may accurately report that the instance cannot take work, but repeated probes can worsen the saturation. Reserve tight acquisition limits, cancel timed-out work, and keep probe frequency low enough that observation does not become load.

This design has limitations. A shallow database query cannot prove that later writes will commit, an ok cache response cannot prove that a required key is fresh, and aggressive readiness gating can reduce available capacity during a partial outage; the trade-off is deliberate because refusing unsafe new work is preferable to accepting delivery state that cannot be recorded, but each team still has to define “unsafe” for its own path.

How does this make rollback safer?

During a rolling deployment, two revisions can coexist. A stable readiness contract lets the traffic layer stop sending new work to a revision that cannot satisfy the durable-write invariant while the previous revision remains eligible. The build identifier helps operators correlate state with deployment, but it must not alter the meaning of ready.

Test the transition, not just the endpoint. Bring up the new revision with its required dependency unavailable; it should remain not ready without failing liveness. Restore the dependency; readiness should recover. Then drain the instance and verify that accepted delivery work reaches a terminal, retryable, or explicitly failed state before shutdown. Rollback safety depends on admission and draining behavior together.

Metrics should distinguish probe executions from delivery outcomes. For example, a readiness result can be counted by a small bounded label such as result="ready" or result="not_ready"; notification delivery failures need similarly bounded dimensions such as channel and reason class, not a raw exception message. Alerting on both the external path and sustained delivery-failure ratios separates “the service cannot be reached” from “the service responds but cannot complete its job.”

Test both sides. A monitor that only calls readiness from inside the same network can miss an edge or DNS failure, while a monitor that only calls the public route cannot explain whether the database gate, cache policy, or deployment path caused the symptom.

No probe proves delivery.

The rejected single endpoint still has a valid use case

I reject one /health endpoint that checks every dependency for this notification service because one boolean would drive incompatible actions: restart, remove from traffic, and page an operator. It also makes a cache failure fatal even when the delivery path has a tested fallback, or makes a database failure look harmless when durable recording is mandatory.

A combined endpoint is still reasonable for a small, non-orchestrated internal service when no automated system interprets it as permission to restart or route traffic, and when a human uses its component results only for diagnosis. Even there, publish the semantics and impose a deadline. Once automation consumes the result, split the control signals before their meanings drift.

The final decision rule is compact: liveness covers local process progress, readiness covers safe admission for the current revision, and outside monitoring covers reachability. Delivery telemetry covers the business outcome. Preserve those boundaries across releases, and a failed notification deployment remains a routing and rollback decision rather than an argument about what “healthy” meant.

References

Top comments (0)