DEV Community

jidonglab
jidonglab

Posted on

Docker Compose depends_on Won't Wait for Postgres: 11/50 Runs Failed

My integration tests failed on a Tuesday with ECONNREFUSED 127.0.0.1:5432. I hit re-run and they passed. Then they failed again on Thursday. I had depends_on: [db] right there in my docker-compose.yml, so I assumed the database was up before my app started. That assumption was wrong, and the way Docker Compose depends_on actually works explained every flaky run I'd been blaming on "CI being CI."

So I stopped guessing. I wrote a loop that tore everything down with a fresh volume and booted the stack 50 times. 11 of 50 runs failed on the first database connection. That's 22%, which is far too high to call bad luck.

Then I fixed it the way every Stack Overflow answer says to, and it still failed. Twice. That second bug is the part almost nobody writes about.

TL;DR

  • depends_on: [db] (short form) only waits for the db container to start. It does not wait for Postgres to accept connections.
  • To wait for readiness, use the long form with condition: service_healthy and define a healthcheck on the database service.
  • The default healthcheck interval is 30 seconds, so a naive healthcheck can make your stack boot slowly. Set interval: 2s.
  • On a fresh volume, the official Postgres image runs a temporary init server that listens only on a Unix socket. Plain pg_isready can report "ready" during that phase. Use pg_isready -h 127.0.0.1 to force a TCP check.
  • Keep retry logic in your app anyway. Healthchecks gate startup only, not the rest of the container's life.

What does Docker Compose depends_on actually wait for?

Docker Compose depends_on, in its short form, waits for the dependency container to be created and started. That's it. "Started" means the process inside the container has launched, not that it's listening on a port or finished initializing.

Here's the setup I had:

services:
  db:
    image: postgres:16
    environment:
      POSTGRES_USER: app
      POSTGRES_PASSWORD: app
      POSTGRES_DB: app
  tests:
    build: .
    command: npm test
    depends_on:
      - db
Enter fullscreen mode Exit fullscreen mode

Compose starts db, sees the container running, and immediately starts tests. Postgres at that moment might be anywhere: still running initdb, still executing your init scripts, or restarting after initialization. My test runner opened a connection within about a second of boot. Sometimes Postgres was ready. Usually it was. 22% of the time it wasn't.

The reason it "usually works" is what makes this bug so annoying. On your laptop, the volume already exists, Postgres skips initialization, and it's ready in well under a second. In CI, every run gets a fresh volume, so every run pays the full init cost. The race only loses where you can't watch it.

How do I reproduce the depends_on race?

Run the stack in a loop with a fresh volume each time. This is the script I used:

fails=0
for i in $(seq 1 50); do
  docker compose down -v --remove-orphans >/dev/null 2>&1
  if ! docker compose up --abort-on-container-exit --exit-code-from tests >/dev/null 2>&1; then
    fails=$((fails+1))
  fi
done
echo "failed: $fails / 50"
Enter fullscreen mode Exit fullscreen mode

The -v matters. Without it, the named volume survives, Postgres skips init, and you'll convince yourself the bug doesn't exist. --exit-code-from tests makes docker compose up return the test container's exit code, so the loop can count failures.

How do I make Docker Compose wait until Postgres is ready?

Use the long form of depends_on with condition: service_healthy, and give the database a healthcheck. Compose will hold the dependent service until the healthcheck passes.

services:
  db:
    image: postgres:16
    environment:
      POSTGRES_USER: app
      POSTGRES_PASSWORD: app
      POSTGRES_DB: app
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 2s
      timeout: 3s
      retries: 30
  tests:
    build: .
    command: npm test
    depends_on:
      db:
        condition: service_healthy
Enter fullscreen mode Exit fullscreen mode

Three details here bite people:

  1. $$ is not a typo. Compose interpolates ${VAR} from your host environment when it parses the file. $$ escapes it so the literal ${POSTGRES_USER} reaches the container shell, where the variable actually exists.
  2. CMD-SHELL, not CMD. CMD runs the binary directly with no shell, so ${POSTGRES_USER} would never expand.
  3. Set interval yourself. The Docker default is 30 seconds. With the default, your tests can sit idle for half a minute waiting for the first check, even though Postgres was ready after three seconds. People "fix" the race this way and then complain their CI got slower.

I reran the 50-boot loop. Failures dropped from 11 to 2.

Two. Not zero.

Why does pg_isready say ready before Postgres accepts connections?

Because on a fresh volume, the official Postgres image starts Postgres twice, and the first one is lying to you, technically.

Here's the sequence the image's entrypoint runs on first boot:

  1. Run initdb to create the data directory.
  2. Start a temporary server with TCP disabled (it listens on the Unix socket only).
  3. Create your user and database, then run everything in /docker-entrypoint-initdb.d/.
  4. Stop the temporary server.
  5. Start the real server, listening on TCP.

pg_isready with no -h flag connects over the local Unix socket. During step 3, the temporary server is accepting socket connections, so pg_isready exits 0. Docker marks the container healthy. Compose releases your tests. Your tests connect over TCP to db:5432, where nothing is listening yet, because the real server hasn't started. ECONNREFUSED.

The window is small with an empty init directory. If you have a seed script with a few thousand rows in docker-entrypoint-initdb.d, the window grows with it, and so does your failure rate.

The fix is one flag. Force the healthcheck to use TCP, the same path your app uses:

    healthcheck:
      test: ["CMD-SHELL", "pg_isready -h 127.0.0.1 -U $${POSTGRES_USER} -d $${POSTGRES_DB}"]
      interval: 2s
      timeout: 3s
      retries: 30
      start_period: 10s
Enter fullscreen mode Exit fullscreen mode

The temporary server doesn't listen on TCP, so this check can't pass until the real server is up. start_period gives init some grace time: failed checks during that period don't count against retries.

Loop again: 0 of 50 failed. I ran it another 50 to make sure I wasn't celebrating noise. Still zero.

The general lesson: a healthcheck should test the same path your client uses. Socket vs TCP, localhost vs service name, the right database name. A healthcheck that checks something adjacent will eventually pass at the wrong moment.

What about migrations that must finish before the app starts?

Use condition: service_completed_successfully. It waits until the dependency container exits with code 0, which is exactly what a one-shot migration job does.

  migrate:
    build: .
    command: npm run migrate
    depends_on:
      db:
        condition: service_healthy
  app:
    build: .
    depends_on:
      db:
        condition: service_healthy
      migrate:
        condition: service_completed_successfully
Enter fullscreen mode Exit fullscreen mode

If migrate exits non-zero, Compose refuses to start app. That's the behavior you want: a failed migration should stop the deploy, not let the app boot against a half-migrated schema.

There are three conditions in total, and it helps to know all of them:

Condition Waits until
service_started Container is running (same as short form)
service_healthy Healthcheck reports healthy
service_completed_successfully Container exited with code 0

Is a healthcheck enough, or do I still need retries?

You still need retries in the app. Healthchecks only gate startup order inside Compose. They do nothing when Postgres restarts an hour later, when a network blip drops a connection, or when you deploy the same image to an orchestrator that has no depends_on at all.

My rule now: Compose healthchecks make local and CI boots deterministic. App-level connection retry with backoff makes production survivable. They solve different problems, and I stopped treating either as a replacement for the other.

One more CI tip: docker compose up -d --wait blocks until services are running and healthy, then returns. That's cleaner than a hand-rolled until pg_isready; do sleep 1; done loop in your workflow file.

The checklist I paste into every Compose file now

  • Long-form depends_on with an explicit condition, never the short list.
  • A healthcheck on every service someone depends on.
  • interval: 2s so readiness is noticed fast.
  • Healthcheck uses the same transport as the client (-h 127.0.0.1 for Postgres).
  • $$ for container-side variables, CMD-SHELL when you need a shell.
  • service_completed_successfully for migrations and seed jobs.
  • Retry logic in the app regardless.
  • Test with down -v, because a warm volume hides everything.

So does Docker Compose depends_on wait for the database?

No. Docker Compose depends_on, in its short form, only waits for the database container to start, not for Postgres to accept connections, which is why stacks with fresh volumes fail intermittently in CI. To wait for readiness, use depends_on with condition: service_healthy and add a healthcheck to the database service. For the official Postgres image, run pg_isready -h 127.0.0.1 so the check goes over TCP. Without -h, it can pass against the temporary socket-only server the image runs during first-boot initialization. In my case that one flag took a 50-run loop from 2 failures to 0.


Written by the developer behind Preterview, an interview prep platform.

Top comments (0)