DEV Community

Jeff
Jeff

Posted on Originally published at powerduck.com

Health checks for APIs: /healthz, /readyz, /livez, version, and metrics in OpenAPI

A single /health endpoint that checks the database, the cache, and three downstream services, and returns 500 if any of them is briefly unreachable, is actively harmful. Kubernetes interprets a failed liveness probe as "this process is dead," kills the container, and starts a new one, which then hammers the same struggling dependency and makes the outage worse. The probe answered the wrong question. Operational health is at least three separate questions, and conflating them causes both cascading restarts and traffic sent to instances that cannot handle it.

Liveness, readiness, and startup are different

Probe Question it asks Failure action Should check dependencies?
Liveness (/livez) Is the process alive and not wedged? Restart the container No
Readiness (/readyz) Can this instance serve a request right now? Stop sending traffic, keep the pod Yes, the ones needed for the next request
Startup (/startz or startup probe) Has slow initialization finished? Restart only if it never starts Warm-up state

Liveness should be nearly impossible to fail except a true deadlock: if the HTTP server itself can answer, liveness is usually fine. Putting dependency checks in liveness turns a downstream blip into a restart storm. Readiness is where the database and cache belong; an instance that cannot reach its data store reports not-ready and is pulled from the load balancer without being killed, then returns to readiness when the dependency recovers. Startup probes protect slow-booting JVMs and migration runners so the other two probes do not fire before the app is ready.

The endpoint contract

Keep the responses simple, fast, and machine-readable. A widely used convention (popularized by Kubernetes itself) is the z suffix:

GET /livez   -> 200 OK
GET /readyz  -> 200 OK | 503 Service Unavailable
GET /startz  -> 200 OK | 503
Enter fullscreen mode Exit fullscreen mode

A healthy response can be a plain ok; a failed readiness response lists the failing components without leaking internals:

{
  "status": "unavailable",
  "checks": {
    "database": "ok",
    "cache": "unavailable",
    "objectStore": "ok"
  }
}
Enter fullscreen mode Exit fullscreen mode

Rules that keep probes safe:

  • Bounded and cheap. Every check has a short timeout; a dependency that hangs must not make the probe hang past the probe deadline. Probes are called constantly.
  • No side effects. A health check never mutates state, writes to the database, or counts against quotas.
  • Fail on hard dependency, ignore soft ones. Readiness includes only dependencies required for the next request. A non-critical analytics pipeline being down should not pull the API out of rotation.
  • Do not cascade. Protect downstream calls with circuit breakers so a health fan-out does not amplify an outage.
  • Return 503, not 200 with an error body, for not-ready; orchestrators and load balancers act on the status code.

What to expose, and to whom

Operational endpoints are a favorite reconnaissance target, so separate what is public from what is internal:

  • /livez and /readyz typically need to be reachable by the platform (kubelet, load balancer) and can return minimal detail publicly. Detailed per-component output is often restricted to an internal port, a management listener, or authenticated access.
  • /version (build SHA, version, deploy time) is useful for confirming a rollout but should not expose dependency versions or hostnames to anonymous callers; gate the detailed form.
  • /metrics (Prometheus) should be on an internal port or behind network policy; it reveals cardinality and internal names.
  • Never include secrets, connection strings, internal IPs, or stack traces in a health response.

A common pattern is two listeners: a public app port exposing coarse /livez and /readyz, and an internal management port exposing detailed health, version, and metrics, reachable only inside the cluster.

Documenting operational endpoints in OpenAPI

These endpoints are part of the contract for the platform team even if they are not business APIs, so document them, but tag and mark them clearly and keep them out of the public business namespace where appropriate:

paths:
  /readyz:
    get:
      operationId: getReadiness
      tags: [Operations]
      summary: Readiness probe; 503 when the instance cannot serve traffic
      security: []
      responses:
        '200':
          description: All required dependencies are reachable.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HealthStatus'
        '503':
          description: One or more required dependencies are unavailable.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HealthStatus'
  /livez:
    get:
      operationId: getLiveness
      tags: [Operations]
      summary: Liveness probe; fails only when the process is wedged
      security: []
      responses:
        '200': { description: Process is alive. }

components:
  schemas:
    HealthStatus:
      type: object
      required: [status, checks]
      properties:
        status:
          type: string
          enum: [ok, unavailable]
        checks:
          type: object
          additionalProperties:
            type: string
            enum: [ok, unavailable, timeout]
Enter fullscreen mode Exit fullscreen mode

security: [] on the operation explicitly makes it unauthenticated (an empty security requirement overrides any global security), which is correct for the coarse probes the platform must reach. If detailed health requires auth, model a second operation with the management security scheme rather than returning different data based on a hidden header.

Kubernetes wiring

The endpoint design only pays off if the probes are configured to match:

livenessProbe:
  httpGet: { path: /livez, port: 8080 }
  initialDelaySeconds: 5
  periodSeconds: 10
  failureThreshold: 3
readinessProbe:
  httpGet: { path: /readyz, port: 8080 }
  periodSeconds: 10
startupProbe:
  httpGet: { path: /startz, port: 8080 }
  failureThreshold: 30
  periodSeconds: 5
Enter fullscreen mode Exit fullscreen mode

The startup probe gives a slow instance up to 150 seconds to boot, after which liveness and readiness take over. Give each probe a generous failureThreshold so a single transient miss does not flap the instance; the platform should react to sustained failure, not one slow scrape.

What codegen, contract tests, and AI operators need

  • Documenting the probes means generated monitoring code and runbooks can rely on stable status codes instead of scraping logs.
  • Contract tests can assert that /readyz returns 503 with the structured body when a dependency is down and that /livez stays 200 in the same scenario, which proves the probes are correctly separated.
  • An AI operator or autoscaler reacting to incidents reads the readiness body to identify the failing component; a consistent checks map makes that automatable, while a free-text string does not.

Checklist

  1. Split liveness (process alive), readiness (can serve, includes hard dependencies), and startup (warm-up) into separate probes.
  2. Keep dependency checks out of liveness to avoid restart storms during downstream outages.
  3. Make probes cheap, bounded, side-effect free, and protected by circuit breakers.
  4. Return 200 versus 503 on the status code; include a structured per-component body for readiness.
  5. Expose coarse probes publicly with security: []; gate detailed health, version, and metrics on an internal port or auth.
  6. Never leak secrets, internal hosts, dependency versions, or stack traces.
  7. Wire Kubernetes startup, liveness, and readiness probes with sensible thresholds; test the "dependency down" case and prove liveness stays green.

Get these right and a database blip pulls traffic off the affected instances instead of restarting your whole fleet into the same outage.

You can document the operational endpoints, generate monitoring-friendly contracts, and test the dependency-down behavior in one local-first workspace, right in your browser. For the structured error body the 503 should reuse across the API, see RFC 9457 Problem Details for HTTP APIs.

Top comments (0)