DEV Community

Cover image for Operating Containers: Healthchecks, Resource Limits, Restart Policies, and Env Config
Shubham Sharma
Shubham Sharma

Posted on Originally published at techdevmantra.com

Operating Containers: Healthchecks, Resource Limits, Restart Policies, and Env Config

A container that runs is not the same as a container that behaves. In development, "it started" is enough. In production you want more: the container should know when it is broken, stay within a memory and CPU budget, come back on its own after a crash, and get its config cleanly without secrets baked into the image. This post adds all four to a real stack, with commands and output captured on a live Docker Engine.

Tip

Key takeaways

  • A healthcheck lets Docker report a container as healthy or unhealthy, and Compose can hold a dependent back until its dependency is healthy.
  • Resource limits (--memory, --cpus, or deploy.resources.limits in Compose) cap what a container can use. Exceed memory and it is killed with exit code 137.
  • A restart policy (restart: unless-stopped or on-failure) brings a crashed container back automatically.
  • Load config from an env file so it stays out of the image, and know the precedence: a -e flag beats --env-file.

Prerequisites

Info

Get the code. The full stack for this post (compose file, health endpoint, env file) is in the docker-foundations repo, under 07-operating/. Clone it to follow along.

Healthchecks: does the container actually work?

A running container is not necessarily a working one. A web server can be up while its process is deadlocked. A healthcheck is a command Docker runs inside the container on an interval; if it passes, the container is healthy, if it fails enough times, unhealthy.

Here is the smallest demonstration. Run a container whose health depends on a file existing, then create and remove that file to flip its state:

docker run -d --name svc \
  --health-cmd="test -f /tmp/ok" --health-interval=2s --health-retries=2 \
  alpine sh -c "touch /tmp/ok; sleep 3600"
docker ps --format "{{.Names}}  {{.Status}}"
Enter fullscreen mode Exit fullscreen mode
svc  Up 5 seconds (healthy)
Enter fullscreen mode Exit fullscreen mode

Now break the check by deleting the file, wait for the interval to run, and look again:

docker exec svc rm /tmp/ok
docker ps --format "{{.Names}}  {{.Status}}"
Enter fullscreen mode Exit fullscreen mode
svc  Up 12 seconds (unhealthy)
Enter fullscreen mode Exit fullscreen mode

Recreate the file and it recovers:

docker exec svc touch /tmp/ok
docker ps --format "{{.Names}}  {{.Status}}"
Enter fullscreen mode Exit fullscreen mode
svc  Up 19 seconds (healthy)
Enter fullscreen mode Exit fullscreen mode

The status column tracks the health state live. In a real image you set this once with a HEALTHCHECK instruction in the Dockerfile, or a healthcheck: block in Compose.

Health-gated startup in Compose

The real payoff is dependency ordering. In the Compose post we saw that depends_on waits for a container to start, not to be ready. A healthcheck fixes that: depends_on with condition: service_healthy holds a service back until its dependency reports healthy.

Our stack has a web service that depends on a Postgres db with a pg_isready healthcheck. Bring it up:

docker compose up -d
Enter fullscreen mode Exit fullscreen mode
 Container op-demo-db-1   Started
 Container op-demo-db-1   Waiting
 Container op-demo-db-1   Healthy
 Container op-demo-web-1  Starting
 Container op-demo-web-1  Started
Enter fullscreen mode Exit fullscreen mode

Read that order: Compose started db, then waited until it was healthy, and only then started web. No more racing a database that has not finished booting.

docker compose ps --format "table {{.Service}}\t{{.Status}}"
Enter fullscreen mode Exit fullscreen mode
SERVICE   STATUS
db        Up 7 seconds (healthy)
web       Up 4 seconds (healthy)
Enter fullscreen mode Exit fullscreen mode

Resource limits: stay in budget

By default a container can use all the host's memory and CPU. One runaway process can starve everything else on the box. Limits fix that.

Memory

Cap memory with --memory (and --memory-swap to also cap swap). When a container tries to exceed its memory limit, the kernel kills it. Watch it happen: this Python container asks for 200MB with a 64MB cap.

docker run --name oom --memory=64m --memory-swap=64m \
  python:3-alpine python -c "bytearray(200*1024*1024)"
docker inspect oom --format 'OOMKilled={{.State.OOMKilled}}  ExitCode={{.State.ExitCode}}'
Enter fullscreen mode Exit fullscreen mode
OOMKilled=true  ExitCode=137
Enter fullscreen mode Exit fullscreen mode

OOMKilled=true and exit code 137 are the signature of a container killed for exceeding its memory limit. If you ever see a container mysteriously exit with 137, this is almost always why.

CPU

Cap CPU with --cpus. This container runs a busy loop that would otherwise peg a whole core, limited to half a CPU:

docker run -d --name cpuhog --cpus=0.5 alpine sh -c "while true; do :; done"
docker stats --no-stream --format "{{.Name}}  CPU={{.CPUPerc}}"
Enter fullscreen mode Exit fullscreen mode
cpuhog  CPU=49.71%
Enter fullscreen mode Exit fullscreen mode

The busy loop is held right at its half-a-CPU ceiling instead of consuming everything.

Limits in Compose

In a Compose file you set the same limits per service under deploy.resources.limits, which docker compose up applies (you do not need Swarm). After bringing the stack up, you can confirm the limit landed on the container:

docker inspect op-demo-web-1 --format 'memory={{.HostConfig.Memory}}  nanocpus={{.HostConfig.NanoCpus}}'
Enter fullscreen mode Exit fullscreen mode
memory=134217728  nanocpus=500000000
Enter fullscreen mode Exit fullscreen mode

That is the 128MB (134217728 bytes) and half-CPU (500000000 nanocpus) budget from the compose file, enforced on the running container.

Restart policies: self-healing

Processes crash. A restart policy tells Docker to bring the container back automatically. unless-stopped restarts it on any exit except a deliberate docker stop, and on-failure restarts only on a non-zero exit, optionally up to a retry cap.

Watch a crashing container recover. This one runs for a few seconds, then exits with an error, with a cap of three retries:

docker run -d --name crasher --restart on-failure:3 alpine sh -c "sleep 3; exit 1"
docker inspect crasher --format 'RestartCount={{.RestartCount}} Status={{.State.Status}}'
# ... a few crash-and-restart cycles later ...
docker inspect crasher --format 'RestartCount={{.RestartCount}} Status={{.State.Status}} ExitCode={{.State.ExitCode}}'
Enter fullscreen mode Exit fullscreen mode
RestartCount=0 Status=running
RestartCount=3 Status=exited ExitCode=1
Enter fullscreen mode Exit fullscreen mode

Docker restarted the crashing container three times, then stopped because it hit the :3 cap. On a real service you would use restart: unless-stopped (no cap) so it keeps recovering, which is what our compose file sets on both services.

Info

A restart policy is not a healthcheck. A restart policy reacts to the container process exiting. A healthcheck reacts to the app being unresponsive while the process is still up. Production services usually want both: restart: unless-stopped to recover from crashes, and a healthcheck so orchestrators and depends_on know when the app is actually ready.

Env config without secrets in the image

Hardcoding config into an image is a mistake: it bakes environment-specific values (and often secrets) into an artifact you push to a registry. Load them at run time instead. Docker pulls environment values from three places (the image's own ENV, an --env-file, and -e flags), and the precedence matters: a -e flag beats --env-file, and both beat a value baked into the image with ENV:

printf "GREETING=from_env_file\nONLY_IN_FILE=yes\n" > envfile
docker run --rm --env-file envfile -e GREETING=from_flag alpine env
Enter fullscreen mode Exit fullscreen mode
ONLY_IN_FILE=yes
GREETING=from_flag
# (PATH, HOSTNAME, HOME and other standard vars omitted)
Enter fullscreen mode Exit fullscreen mode

GREETING came out as from_flag: the explicit -e won over the file. ONLY_IN_FILE passed through untouched. The same order holds in Compose: a service's environment: block overrides its env_file:. Keep the file (.env) out of git, commit an .env.example template, and your secrets never enter the image.

The whole thing in one Compose file

All four behaviors live together in the stack's docker-compose.yml:

services:
  db:
    image: postgres:16-alpine
    environment:
      POSTGRES_USER: demo
      POSTGRES_PASSWORD: demo
      POSTGRES_DB: demo
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U demo"]
      interval: 3s
      timeout: 3s
      retries: 5
    restart: unless-stopped
    deploy:
      resources:
        limits:
          memory: 256M
          cpus: "0.50"

  web:
    image: node:22-alpine
    working_dir: /app
    command: node server.js
    volumes:
      - ./app:/app
    env_file: .env
    ports:
      - "8080:3000"
    depends_on:
      db:
        condition: service_healthy   # web starts only once db reports healthy
    healthcheck:
      test: ["CMD-SHELL", "wget -q -O- http://localhost:3000/health || exit 1"]
      interval: 3s
      timeout: 3s
      retries: 5
    restart: unless-stopped
    deploy:
      resources:
        limits:
          memory: 128M
          cpus: "0.50"
Enter fullscreen mode Exit fullscreen mode

docker compose up -d, and both services come up healthy, budgeted, self-healing, and configured from .env.

Common gotchas

Container exits with code 137

It was killed for exceeding its memory limit (or received a SIGKILL). Check docker inspect --format '{{.State.OOMKilled}}'. If true, raise the limit or fix the leak.

Healthcheck passes but the app is broken

Your check is too shallow. pg_isready proves Postgres accepts connections; a check that only pings the port proves less. Point the healthcheck at an endpoint that actually exercises the app, like a /health route that touches its dependencies.

depends_on: condition: service_healthy does nothing

The dependency has no healthcheck, so it can never report healthy. Add a healthcheck: block to the service you are waiting on.

The container keeps restarting forever

A restart policy plus a container that crashes instantly is a crash loop. Use on-failure:<n> to cap retries while you debug, check docker logs, and fix the underlying crash before switching back to unless-stopped.

My .env changes are ignored

Compose reads .env from the project directory for variable substitution, and env_file: for what a service sees. Make sure you edited the right one, and recreate the container (docker compose up -d again) so it picks up the new values.

Where to go next

Your containers now behave: they report health, respect limits, recover from crashes, and take config cleanly. The next step is making them safe.

  • Next in this series: Docker Security Basics, running as non-root, mounting the filesystem read-only, dropping capabilities, and keeping secrets out of images.

Verified on 2026-09-11 on a real Ubuntu 24.04.5 LTS system (arm64) with Docker Engine 29.8.0 and Compose v5.5.1. Captured: healthcheck transitions healthy/unhealthy/healthy; a Compose stack where db went Started to Waiting to Healthy before web started; a memory-limited container OOM-killed with ExitCode 137; a --cpus=0.5 busy loop held at 49.71% in docker stats; a crashing container restarted to RestartCount 3 under on-failure:3; -e overriding --env-file; and deploy.resources.limits applied as 134217728 bytes / 500000000 nanocpus on the web container.

Top comments (0)