DEV Community

Cover image for Docker homelab gotchas: docker ps called the image orphaned. It was live
Christian Anderson
Christian Anderson

Posted on

Docker homelab gotchas: docker ps called the image orphaned. It was live

My Docker host was 82% full, so I went looking for images to delete. docker system df said 6.67 GB was reclaimable. docker ps --format '{{.Image}}' printed bare SHA IDs instead of names for several running containers, which made one image, unclecode/crawl4ai, look like nothing was using it.

It was running. Healthy, up two weeks, serving traffic. I was one docker rmi away from taking out a live service on the strength of a formatting quirk.

What stopped me wasn't the image list. It was a listening port that nothing on my list owned.

for c in $(docker ps -q); do docker inspect -f '{{.Name}} {{.Config.Image}}' $c; done
ss -ltnp        # a port nobody claims is a container you mis-read
Enter fullscreen mode Exit fullscreen mode

That's the pattern for most of what follows: Docker's summary views are fine until you act on them.

How I actually run it

Nothing exotic. Docker runs inside unprivileged LXC containers on Proxmox (they need nesting=1 and keyctl=1), plus a small NAS running an appliance OS whose app manager is a layer over Docker. Single-purpose things, like my voice assistant's speech-to-text and text-to-speech servers, are one docker run line each. Anything with a database is a compose stack; the LLM tracing one is six containers.

Here's what bit.

Disk: the pressure was images, not data

Four guests were 82 to 86% full, and my instinct was to push data to the NAS. Profiling disagreed. The Docker LXC had 16 GB of images. Another guest had 11.4 GB of images and 11.3 GB of build cache. There was very little cold data at all: my weekly offload job's first run reclaimed 213 MB.

You can't move a running container's image layers onto NFS anyway, since overlayfs over NFS is slow and fragile. The fix was growing the guest disks, online, from a thin pool that was 7% used.

Two more lessons from the clean-up:

  • Per-tag sizes are a lie for planning. I removed 16 superseded tags of my own app, each listed at about 240 MB. They freed 60 MB, because the builds share nearly every layer. The whole clean-up, dangling images and build cache included, came to about 340 MB.
  • An image can be dead in ps and still be a dependency. Three NetBird images weren't running, but a compose file on disk still referenced them. What I did delete had no container ever and no references anywhere: 2.8 GB.

docker image prune -a is now on my "a human types it" list. On one host 4 of the 11 images were in use, and there's no fast path to re-pull if something happens to be stopped at the time.

Logs: nobody set a limit

Another Docker host's root filesystem reached 100%, 0 bytes free, with five failed systemd units. Nothing had alerted. The cause was that /etc/docker/daemon.json didn't exist, so Docker was using the default json-file log driver with no rotation and no size cap. The Home Assistant container's log was a single 1.9 GB file.

truncate -s 0 /var/lib/docker/containers/<id>/<id>-json.log   # 1.9G
journalctl --vacuum-size=100M                                  # 286M
apt-get clean                                                  # 271M
docker image prune -f        # dangling only, reclaimed 3.47G
Enter fullscreen mode Exit fullscreen mode

All four containers kept their uptimes. Then I wrote the daemon.json that should have been there all along: json-file, max-size: 50m, max-file: 3.

It didn't cap Home Assistant's log. Weeks later, the container still reports an empty log config:

docker inspect -f '{{.Name}} {{.HostConfig.LogConfig}}' $(docker ps -q)
# /homeassistant {json-file map[]}
Enter fullscreen mode Exit fullscreen mode

Daemon log defaults are copied into a container when it's created, not when Docker restarts. Every container that existed before the file stays uncapped until it's recreated. The only containers on that host with a cap are two whose compose file sets its own logging: options.

Restart policy: the file said one thing, the container another

A file browser on the NAS hard-exited one morning because it couldn't reach the SSO provider. It stayed dead for six days. RestartCount=0. Nothing retried and nothing alerted.

The compose file said restart: unless-stopped. The container had been created before that line was added and was never recreated, so it was still running restart=no. I'd "fixed" this once already with docker update, and it had reverted. The second time I checked the on-disk hostconfig.json rather than trusting docker inspect. Then I audited the NAS: all 18 running containers now have unless-stopped. Five didn't, including the SSO container that every login depends on. The reverse proxy had the same drift.

That appliance layer had its own go at me. My dashboard container was uninstalled by the NAS's app manager twice, with its config left on disk both times. The first time I reinstalled from the leftover compose file. The second time that file was gone too. I don't know why it happened, so I stopped letting the app manager own it:

docker run -d --name homepage --restart unless-stopped -p <port>:3000 \
  -v <config-dir>:/app/config \
  -v /var/run/docker.sock:/var/run/docker.sock:ro \
  -e HOMEPAGE_ALLOWED_HOSTS=<host>:<port>,localhost:<port> \
  ghcr.io/gethomepage/homepage:v1.8.0
Enter fullscreen mode Exit fullscreen mode

Leave out HOMEPAGE_ALLOWED_HOSTS on v1.x and every browser request gets a blank 400.

Configuration is read at create time

docker compose restart does not reload .env. env_file is read when the container is created. Use docker compose up -d --force-recreate.

Compose interpolates $ inside env_file values. My password manager's admin token is an Argon2 PHC string, $argon2id$v=19$m=19456,.... It reached the container as 51 mangled characters beginning =19=19456,, so the admin page returned 401 for every token, the right one included. Doubling every $ fixed it, and docker inspect showed the container getting 134 characters starting $argon2id$. The file being right proved nothing. The env file is named vw.env, not .env, because compose auto-loads .env as its own interpolation source.

environment: beats env_file. My app's compose file pinned three service URLs in environment:, so anyone setting them correctly in their own .env was silently overridden.

Stock compose defaults can contradict each other. The tracing stack's compose file sets DATABASE_URL to a hardcoded postgres:postgres instead of building it from POSTGRES_PASSWORD. Set only the password and Postgres changes while the app keeps using the literal default:

Error: P1000: Authentication failed against database server
Applying database migrations failed. ... Exiting...
Enter fullscreen mode Exit fullscreen mode

The web container restart-looped. The error mentions special characters that aren't URL-encoded, which was a red herring: the password was alphanumeric by construction. The fix is an explicit DATABASE_URL in .env, kept in step with the password.

Code baked into the image isn't live because the tests passed. My web app's code is copied in at build time, and compose mounts only the data volume. I committed three fixes, watched 661 tests pass, and both containers carried on running the previous version with the new code on disk beside them. Now I ask the container:

docker exec <container> python -c "from <pkg> import <module>; print(<module>.<CONSTANT_THE_FIX_ADDED>)"
Enter fullscreen mode Exit fullscreen mode

Networking: published ports and DOCKER-USER

My web app has a private instance and an internet-facing demo copy on the same Docker host. The private one is meant to be reached only through a reverse proxy on another box, but it was published on 0.0.0.0, so the demo container could reach it directly and skip the proxy.

Binding it to 127.0.0.1 fixed that and took the site down for a day, because the proxy is on a different machine. The real fix was to publish on the host's LAN address and restrict it in the DOCKER-USER chain. The shape:

iptables -N PRIVATE-IN
iptables -I DOCKER-USER -m conntrack --ctorigdst 192.0.2.10 --ctorigdstport 8080 -j PRIVATE-IN
iptables -A PRIVATE-IN -m conntrack --ctstate ESTABLISHED -j RETURN
iptables -A PRIVATE-IN -s 192.0.2.20 -j RETURN     # the proxy
iptables -A PRIVATE-IN -j DROP
Enter fullscreen mode Exit fullscreen mode

(Addresses and port are placeholders.) The proxy gets a 200 and everything else times out.

Before I added egress rules, the internet-facing demo could reach far more of my LAN than it had any business touching. Afterwards those internal services time out, while the handful it genuinely needs and the internet still answer. Getting there turned on three traps:

  • Matching -d never fires for published ports. By the time a packet reaches FORWARD, the destination has already been DNATed from the host address onto a 172.x container address. --ctorigdst matches what the client actually asked for.
  • DNS doesn't go through FORWARD. The container's resolver was the host's own mesh-VPN address, which is local, so queries land on INPUT. A mesh-range DROP in DOCKER-USER never touches DNS, which looks reassuring, but the separate INPUT -i br-... chain must allow port 53 or the container goes deaf.
  • The bridge name changes. It's br- plus the first 12 characters of the network ID, and recreating the network gives it a new ID. The script reads bridge and subnet from docker network inspect every run, and I re-run it after any compose down.

Docker rebuilds DOCKER-USER when it starts, so the script runs from a systemd unit with After=docker.service. It also had its own bug: the cleanup -D DOCKER-USER line left out the -s match it was inserted with, so it never deleted anything. I found 17 duplicate jumps. There's now one.

Healthy is not working

Two containers that docker ps would call fine:

  • My Cloudflare tunnel ran from a container that turned out to be a web UI for cloudflared, not a connector. After a recreate it logged No pre-existing config file found, showed healthy, and connected no tunnel at all. Every public hostname returned 530, SSO included. A plain cloudflare/cloudflared container with --restart unless-stopped fixed it.
  • The NAS has a Pascal-generation Quadro that, last I checked, falls off the PCIe bus hours after boot, so Immich's CUDA machine-learning container logs no CUDA-capable device and falls back to CPU. Immich v3 dropped Pascal support, so the moment the card is fixed that container will die with SIGILL, exit 132, and crash-loop. A working GPU is worse than a missing one until I change the image tag.

Rules I run by now

  • Never delete an image on the evidence of docker ps or docker system df. Inspect the containers and check the listening ports.
  • Check what the container received (docker inspect, docker exec ... env), not what the file says.
  • Changed .env? up -d --force-recreate, not restart.
  • Every container gets --restart unless-stopped, and I check it persisted.
  • If an app manager can uninstall it, anything I care about runs outside it.
  • Logs get a size cap on day one. A daemon.json default only reaches containers created after it, so recreate the old ones.
  • Firewall published ports by --ctorigdst in DOCKER-USER, derive the bridge name every run, and put the script behind After=docker.service.
  • "Healthy" means the process is up. It doesn't mean it's doing its job.

🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.

Top comments (0)