DEV Community

Oleksandr Kuryzhev
Oleksandr Kuryzhev

Posted on Originally published at kuryzhev.cloud

Docker Healthchecks: Faster Deploys With Minimal Downtime

Originally published on kuryzhev.cloud


A deploy finishes, the new container shows "Up", and the first few requests still hit a 502 because the app is warming its caches. Or the opposite happens: the pipeline sits for two minutes waiting on a container that has been ready for ten seconds. In both cases, Docker healthchecks are usually misconfigured or missing. Small changes to them shorten rollouts and keep traffic away from containers that are not ready.

The tips below are independent, so apply whichever fits your setup. Behavior described here is the documented behavior for current Docker Engine and Compose releases. Verify defaults against the official HEALTHCHECK reference for your installed version.

1. Tune Docker healthchecks timing before anything else

The defaults are tuned for safety, not speed, so set every timing field explicitly.

The documented defaults are a 30-second interval, a 30-second timeout, 3 retries, and no start period. The first probe runs one interval after the container starts. So a container that becomes ready at second 2 can wait almost half a minute before it is marked healthy.

Use a short start_interval to probe quickly during startup, and a longer interval once the container is running. Note that start_interval only applies while the container is inside start_period, so set both.

services:
  api:
    image: registry.example.com/api:1.8.2
    healthcheck:
      test: ["CMD", "/app/api", "healthcheck"]  # exec form, no shell needed
      interval: 15s         # steady-state probe frequency
      timeout: 3s           # keep well below interval
      retries: 3
      start_period: 30s     # failures here do not count toward retries
      start_interval: 2s    # fast probing during start_period

Watch out for: start_interval needs a reasonably recent Engine (introduced with Docker Engine 25.0) and a recent Compose release. Older versions may reject or ignore it, so check your versions first.

2. Make the endpoint cheap and answer one question

A healthcheck should say "this process can serve requests", not "the whole platform is fine".

If the probe runs a database query and the database has a brief hiccup, every replica turns unhealthy at once. A proxy that honors health status then has nothing left to route to, so a small blip becomes a full outage. Check in-process state instead: the server is listening, migrations finished, and required caches are loaded.

Keep the handler fast and side-effect free. A probe that takes a lock, writes a row, or calls a third-party API adds load every few seconds on every replica. Reserve dependency checks for a separate diagnostic endpoint that monitoring calls, not the container runtime.

3. Do not assume curl exists in the image

A common broken healthcheck is one that calls a binary the image does not contain.

Slim and distroless images often ship without curl, and sometimes without a shell. The result is a container that is permanently "unhealthy" even though the app is fine.

The exec form avoids the shell entirely. In a Dockerfile, that is a JSON array after CMD; in Compose, it is ["CMD", ...]. A plain string after CMD runs through /bin/sh. A small subcommand built into your own binary is the most portable option.

# Use ONE of these per Dockerfile: if several HEALTHCHECK
# instructions are present, only the last one takes effect.

# Option A: BusyBox-based image, wget is available (shell form)
HEALTHCHECK --interval=15s --timeout=3s --start-period=30s --start-interval=2s \
  CMD wget -qO- http://127.0.0.1:8080/healthz || exit 1

# Option B: distroless, use the app's own binary (exec form, no shell)
HEALTHCHECK --interval=15s --timeout=3s \
  CMD ["/app/api", "healthcheck"]

Watch out for localhost resolving to IPv6 ::1 while the app listens only on IPv4. Using 127.0.0.1 explicitly removes that ambiguity.

4. Gate dependent services with service_healthy and --wait

Replace sleep loops with healthcheck-based ordering.

In Compose, depends_on with condition: service_healthy holds a service back until its dependency passes its check. This removes fixed delays like sleep 20, which are either too long or too short. The same healthchecks also let the CLI block until the stack is ready.

services:
  db:
    image: postgres:17
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U postgres"]  # ships with the image
      interval: 5s
      timeout: 3s
      retries: 10
  api:
    image: registry.example.com/api:1.8.2
    depends_on:
      db:
        condition: service_healthy
# Returns only when services are running/healthy, or fails on timeout
docker compose up -d --wait --wait-timeout 120

A CI job can then fail fast with a clear exit code instead of guessing when to start smoke tests. See the Compose healthcheck reference for the full field list.

5. Pair start-first updates with health in Swarm

Zero-downtime rollouts need the new task healthy before the old one stops.

In Docker Swarm, a service update can start the replacement task first. It removes the old task only after the new one reports healthy. Without a healthcheck, Swarm treats "process started" as success, which is exactly when the 502s happen. Add a check, then set the update order and a rollback policy.

docker service update \
  --image registry.example.com/api:1.8.3 \
  --update-order start-first \
  --update-failure-action rollback \
  --update-delay 5s \
  api

Here is what each update flag does:

  • --update-order start-first brings the new task up before the old task stops.
  • --update-failure-action rollback reverts the service if a new task fails or never becomes healthy within the update monitor window (--update-monitor).
  • --update-delay spaces out batches of replicas.

Start-first briefly runs old and new versions side by side, and briefly needs extra capacity. Confirm your database schema changes are backward compatible before relying on it.

6. Know who acts on an unhealthy status

Plain Docker Engine marks a container unhealthy and then does nothing about it.

The documented behavior is that the health status is informational on a standalone engine. It does not restart the container, and restart: unless-stopped only reacts to the process exiting. Swarm reschedules unhealthy tasks.

Some reverse proxies also act on health state. Traefik's Docker provider, for example, stops routing to unhealthy containers. Verify this for the proxy you run, because not all of them read Docker health state.

If you deploy with Compose on a single host, there are two practical patterns:

  • Put a health-aware proxy in front of the service.
  • Run a second replica and switch traffic only after --wait succeeds.

Treating the health flag as self-healing is a risky assumption that tends to surface only during an incident.

Watch out for restart loops in Swarm, or with an external watcher that restarts unhealthy containers. An overly strict check with few retries during slow starts can trigger them. Give start_period enough room for a cold boot on a busy host, not a fast laptop.

7. Read the health log when a rollout stalls

Docker keeps the last few probe results, so use them before changing any numbers.

When a container will not go healthy, the stored probe output usually names the cause. Look for connection refused, a missing binary, a timeout, or a non-200 response. Guessing at timings without reading it wastes deploy cycles. Inspect the state directly; with Compose, resolve the actual container ID first, because container names include the project prefix.

# Resolve the container for the Compose service "api"
CID=$(docker compose ps -q api)

# Current status: starting, healthy, or unhealthy
docker inspect --format '{{.State.Health.Status}}' "$CID"

# Last probe results with exit codes and output
docker inspect --format '{{json .State.Health.Log}}' "$CID"

# Watch health transitions as they happen
docker events --filter event=health_status

An exit code of 0 means healthy and 1 means unhealthy. Exit code 2 is documented as reserved, so make scripts return only 0 or 1. For more deployment patterns, browse the rest of the notes on kuryzhev.cloud.

Related

Top comments (0)