A container that runs is not the same as a container that behaves. In development, "it started" is enough. In production you want more: the container should know when it is broken, stay within a memory and CPU budget, come back on its own after a crash, and get its config cleanly without secrets baked into the image. This post adds all four to a real stack, with commands and output captured on a live Docker Engine.
Tip
Key takeaways
- A healthcheck lets Docker report a container as healthy or unhealthy, and Compose can hold a dependent back until its dependency is healthy.
- Resource limits (
--memory,--cpus, ordeploy.resources.limitsin Compose) cap what a container can use. Exceed memory and it is killed with exit code 137.- A restart policy (
restart: unless-stoppedoron-failure) brings a crashed container back automatically.- Load config from an env file so it stays out of the image, and know the precedence: a
-eflag beats--env-file.
Prerequisites
- Docker installed and running. See Install Docker on macOS, Windows (WSL2), and Linux.
- Comfort with Compose, from Docker Compose: Run a Multi-Service Stack.
Info
Get the code. The full stack for this post (compose file, health endpoint, env file) is in the docker-foundations repo, under
07-operating/. Clone it to follow along.
Healthchecks: does the container actually work?
A running container is not necessarily a working one. A web server can be up while its process is deadlocked. A healthcheck is a command Docker runs inside the container on an interval; if it passes, the container is healthy, if it fails enough times, unhealthy.
Here is the smallest demonstration. Run a container whose health depends on a file existing, then create and remove that file to flip its state:
docker run -d --name svc \
--health-cmd="test -f /tmp/ok" --health-interval=2s --health-retries=2 \
alpine sh -c "touch /tmp/ok; sleep 3600"
docker ps --format "{{.Names}} {{.Status}}"
svc Up 5 seconds (healthy)
Now break the check by deleting the file, wait for the interval to run, and look again:
docker exec svc rm /tmp/ok
docker ps --format "{{.Names}} {{.Status}}"
svc Up 12 seconds (unhealthy)
Recreate the file and it recovers:
docker exec svc touch /tmp/ok
docker ps --format "{{.Names}} {{.Status}}"
svc Up 19 seconds (healthy)
The status column tracks the health state live. In a real image you set this once with a HEALTHCHECK instruction in the Dockerfile, or a healthcheck: block in Compose.
Health-gated startup in Compose
The real payoff is dependency ordering. In the Compose post we saw that depends_on waits for a container to start, not to be ready. A healthcheck fixes that: depends_on with condition: service_healthy holds a service back until its dependency reports healthy.
Our stack has a web service that depends on a Postgres db with a pg_isready healthcheck. Bring it up:
docker compose up -d
Container op-demo-db-1 Started
Container op-demo-db-1 Waiting
Container op-demo-db-1 Healthy
Container op-demo-web-1 Starting
Container op-demo-web-1 Started
Read that order: Compose started db, then waited until it was healthy, and only then started web. No more racing a database that has not finished booting.
docker compose ps --format "table {{.Service}}\t{{.Status}}"
SERVICE STATUS
db Up 7 seconds (healthy)
web Up 4 seconds (healthy)
Resource limits: stay in budget
By default a container can use all the host's memory and CPU. One runaway process can starve everything else on the box. Limits fix that.
Memory
Cap memory with --memory (and --memory-swap to also cap swap). When a container tries to exceed its memory limit, the kernel kills it. Watch it happen: this Python container asks for 200MB with a 64MB cap.
docker run --name oom --memory=64m --memory-swap=64m \
python:3-alpine python -c "bytearray(200*1024*1024)"
docker inspect oom --format 'OOMKilled={{.State.OOMKilled}} ExitCode={{.State.ExitCode}}'
OOMKilled=true ExitCode=137
OOMKilled=true and exit code 137 are the signature of a container killed for exceeding its memory limit. If you ever see a container mysteriously exit with 137, this is almost always why.
CPU
Cap CPU with --cpus. This container runs a busy loop that would otherwise peg a whole core, limited to half a CPU:
docker run -d --name cpuhog --cpus=0.5 alpine sh -c "while true; do :; done"
docker stats --no-stream --format "{{.Name}} CPU={{.CPUPerc}}"
cpuhog CPU=49.71%
The busy loop is held right at its half-a-CPU ceiling instead of consuming everything.
Limits in Compose
In a Compose file you set the same limits per service under deploy.resources.limits, which docker compose up applies (you do not need Swarm). After bringing the stack up, you can confirm the limit landed on the container:
docker inspect op-demo-web-1 --format 'memory={{.HostConfig.Memory}} nanocpus={{.HostConfig.NanoCpus}}'
memory=134217728 nanocpus=500000000
That is the 128MB (134217728 bytes) and half-CPU (500000000 nanocpus) budget from the compose file, enforced on the running container.
Restart policies: self-healing
Processes crash. A restart policy tells Docker to bring the container back automatically. unless-stopped restarts it on any exit except a deliberate docker stop, and on-failure restarts only on a non-zero exit, optionally up to a retry cap.
Watch a crashing container recover. This one runs for a few seconds, then exits with an error, with a cap of three retries:
docker run -d --name crasher --restart on-failure:3 alpine sh -c "sleep 3; exit 1"
docker inspect crasher --format 'RestartCount={{.RestartCount}} Status={{.State.Status}}'
# ... a few crash-and-restart cycles later ...
docker inspect crasher --format 'RestartCount={{.RestartCount}} Status={{.State.Status}} ExitCode={{.State.ExitCode}}'
RestartCount=0 Status=running
RestartCount=3 Status=exited ExitCode=1
Docker restarted the crashing container three times, then stopped because it hit the :3 cap. On a real service you would use restart: unless-stopped (no cap) so it keeps recovering, which is what our compose file sets on both services.
Info
A restart policy is not a healthcheck. A restart policy reacts to the container process exiting. A healthcheck reacts to the app being unresponsive while the process is still up. Production services usually want both:
restart: unless-stoppedto recover from crashes, and ahealthcheckso orchestrators anddepends_onknow when the app is actually ready.
Env config without secrets in the image
Hardcoding config into an image is a mistake: it bakes environment-specific values (and often secrets) into an artifact you push to a registry. Load them at run time instead. Docker pulls environment values from three places (the image's own ENV, an --env-file, and -e flags), and the precedence matters: a -e flag beats --env-file, and both beat a value baked into the image with ENV:
printf "GREETING=from_env_file\nONLY_IN_FILE=yes\n" > envfile
docker run --rm --env-file envfile -e GREETING=from_flag alpine env
ONLY_IN_FILE=yes
GREETING=from_flag
# (PATH, HOSTNAME, HOME and other standard vars omitted)
GREETING came out as from_flag: the explicit -e won over the file. ONLY_IN_FILE passed through untouched. The same order holds in Compose: a service's environment: block overrides its env_file:. Keep the file (.env) out of git, commit an .env.example template, and your secrets never enter the image.
The whole thing in one Compose file
All four behaviors live together in the stack's docker-compose.yml:
services:
db:
image: postgres:16-alpine
environment:
POSTGRES_USER: demo
POSTGRES_PASSWORD: demo
POSTGRES_DB: demo
healthcheck:
test: ["CMD-SHELL", "pg_isready -U demo"]
interval: 3s
timeout: 3s
retries: 5
restart: unless-stopped
deploy:
resources:
limits:
memory: 256M
cpus: "0.50"
web:
image: node:22-alpine
working_dir: /app
command: node server.js
volumes:
- ./app:/app
env_file: .env
ports:
- "8080:3000"
depends_on:
db:
condition: service_healthy # web starts only once db reports healthy
healthcheck:
test: ["CMD-SHELL", "wget -q -O- http://localhost:3000/health || exit 1"]
interval: 3s
timeout: 3s
retries: 5
restart: unless-stopped
deploy:
resources:
limits:
memory: 128M
cpus: "0.50"
docker compose up -d, and both services come up healthy, budgeted, self-healing, and configured from .env.
Common gotchas
Container exits with code 137
It was killed for exceeding its memory limit (or received a SIGKILL). Check docker inspect --format '{{.State.OOMKilled}}'. If true, raise the limit or fix the leak.
Healthcheck passes but the app is broken
Your check is too shallow. pg_isready proves Postgres accepts connections; a check that only pings the port proves less. Point the healthcheck at an endpoint that actually exercises the app, like a /health route that touches its dependencies.
depends_on: condition: service_healthy does nothing
The dependency has no healthcheck, so it can never report healthy. Add a healthcheck: block to the service you are waiting on.
The container keeps restarting forever
A restart policy plus a container that crashes instantly is a crash loop. Use on-failure:<n> to cap retries while you debug, check docker logs, and fix the underlying crash before switching back to unless-stopped.
My .env changes are ignored
Compose reads .env from the project directory for variable substitution, and env_file: for what a service sees. Make sure you edited the right one, and recreate the container (docker compose up -d again) so it picks up the new values.
Where to go next
Your containers now behave: they report health, respect limits, recover from crashes, and take config cleanly. The next step is making them safe.
- Next in this series: Docker Security Basics, running as non-root, mounting the filesystem read-only, dropping capabilities, and keeping secrets out of images.
Verified on 2026-09-11 on a real Ubuntu 24.04.5 LTS system (arm64) with Docker Engine 29.8.0 and Compose v5.5.1. Captured: healthcheck transitions healthy/unhealthy/healthy; a Compose stack where db went Started to Waiting to Healthy before web started; a memory-limited container OOM-killed with ExitCode 137; a --cpus=0.5 busy loop held at 49.71% in docker stats; a crashing container restarted to RestartCount 3 under on-failure:3; -e overriding --env-file; and deploy.resources.limits applied as 134217728 bytes / 500000000 nanocpus on the web container.
Top comments (0)