CI says the deploy succeeded, then docker ps -a shows Exited (1) or Exited (127). Don't restart yet: get the exit code, check what it usually means (and does not prove), then read the first log error. Read-only; authorized servers only. Replace YOUR_CONTAINER with the real name or ID.
1. Get the exit code
docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Image}}'
docker inspect --format '{{.State.ExitCode}} {{.State.OOMKilled}} {{.State.Error}}' YOUR_CONTAINER
With Compose, docker compose ps -a includes stopped ones; name is usually <project>-<service>-1.
Exited (1) 8 seconds ago is a single exit with code 1. Restarting (1) 3 seconds ago means a restart policy is looping; the bracketed number is from the last exit.
Inspect prints exit code, OOMKilled (true/false), and State.Error. Keep them apart. State.Error is set when Docker couldn't start the process; empty for most app crashes. Created means it never ran — usually a run-command or image problem.
docker inspect --format 'started={{.State.StartedAt}} finished={{.State.FinishedAt}} restarts={{.RestartCount}} policy={{.HostConfig.RestartPolicy.Name}}' YOUR_CONTAINER
A few seconds between started and finished is the "exits right after deploy" pattern. Timestamps are UTC (Z); convert before comparing with local logs.
2. Exit codes at a glance
| Code | Usual meaning |
|---|---|
| 0 | Main process finished successfully |
| 1 | Generic application error |
| 2 | Shell syntax error, wrong usage, or an app-defined error |
| 125 | Docker itself failed to run the container |
| 126 | Command found but not executable |
| 127 | Command not found |
| 132 | Illegal instruction, SIGILL (128 + 4) |
| 137 | Killed with SIGKILL (128 + 9) |
| 139 | Segmentation fault, SIGSEGV (128 + 11) |
| 143 | Stopped with SIGTERM (128 + 15) |
| 255 | Exit status −1, a failed ssh, or Docker lost the process |
Above 128 usually means "killed by signal (code − 128)". kill -l 137 prints KILL.
3. What each code means
0 — finished successfully
Expected for a migration or backup. A web service exiting 0 usually backgrounded itself or ran a one-shot. Check entrypoint/cmd:
docker inspect --format 'entrypoint={{json .Config.Entrypoint}} cmd={{json .Config.Cmd}}' YOUR_CONTAINER
0 only means the main process ended cleanly.
1 — generic application error
Most runtimes exit 1 for any uncaught exception or failed startup check. Read the logs (section 4) for the first error, then compare with this deploy's changes (image, env, config, migrations). The number alone does not name the error.
2 — syntax, usage, or app-defined error
No Docker-specific meaning — whatever the main process returned. Common sources:
-
Shell script that doesn't parse. Bash/dash exit 2 on syntax errors. An entrypoint with Windows (CRLF) endings often shows
syntax error: unexpected end of file(bash) orSyntax error: end of file unexpected(dash). Outsideif/case, CRLF can instead yieldnot found/ exit 127. -
Wrong builtin usage. Bash builtins return 2 for an invalid option or missing argument (e.g.
sourcewith no filename). -
Program rejecting arguments. Python
argparse→unrecognized arguments; Goflag→flag provided but not defined. Composecommand:typos show up here. -
Missing file.
python3 /app/main.pyexits 2 withcan't open file ... No such file or directoryif the path is wrong. -
Unrecovered Go panic. Prints
panic:plus a stack trace, then exits 2.
docker logs --timestamps YOUR_CONTAINER 2>&1 | head -n 30
bash -n entrypoint.sh # parse only
grep -c $'\r' entrypoint.sh # 0 = LF line endings
CRLF fix: sed -i 's/\r$//' entrypoint.sh, add *.sh text eol=lf to .gitattributes, rebuild. The log line decides the cause.
125 — Docker could not run the container
docker run or the daemon rejected the request; the process usually never started. Read the CI/deploy error and State.Error. Typical: unknown flag, pull denied, port already allocated (sudo ss -ltnp). Not an app bug. When Docker can't start it, status is often Created; when a shell inside hits 126/127, look in docker logs.
126 — found but not executable
Check the command and permissions inside the image. A bind-mounted script keeps the host execute bit. Compare docker image inspect --format '{{.Os}}/{{.Architecture}}' YOUR_IMAGE with uname -m. Mismatch → usually exec format error. 126 = can't execute; 127 = can't find.
127 — command not found
The program never started. Confirm the binary exists in this image and on PATH. Slimmer bases and Compose command: typos are common. Does not prove a crash after startup.
132 — illegal instruction (SIGILL)
132 is 128 + 4 (kill -l 132 → ILL). The CPU hit an instruction it couldn't run:
- Binary needs features this host lacks (AVX/AVX2/AVX-512, x86-64-v3); MongoDB 5.0+ on x86_64 documents AVX.
- VM hides features behind a basic hypervisor CPU model.
- QEMU emulating another arch hits unimplemented instructions.
- Deliberate trap (
ud2/__builtin_trap()) — a bug, not a CPU mismatch.
docker logs --timestamps YOUR_CONTAINER 2>&1 | tail -n 20
sudo dmesg -T | grep -i 'invalid opcode' | tail -n 10
grep -m1 '^flags' /proc/cpuinfo | grep -o -w -E 'avx|avx2|avx512f' | sort -u
A restart policy won't help. Rule out the trap case before buying a new server.
137 — SIGKILL
137 is 128 + 9. Check OOMKilled: true → memory lead; false → docker kill, timed-out docker stop, or another host process. 137 alone is not proof of OOM. Check kernel log, memory limits, and docker events before raising limits.
139 — SIGSEGV
Invalid memory access; kernel sent signal 11. Read logs around the exit, then sudo dmesg -T | grep -i segfault | tail -n 20. Note native-dependency or runtime upgrades. Kernel lines only count if their time matches finished.
143 — SIGTERM
A normal stop. Find who asked: redeploy, docker stop, Compose recreate, host shutdown, orchestrator. Docker keeps only recent events (last 256):
docker events --since 1h --until "$(date +%s)" --filter container=YOUR_CONTAINER
Not a crash — who sent SIGTERM, and did the replacement come up?
255 — exit(−1), failed ssh, or Docker lost the process
-
Program returned −1. Only the lowest 8 bits reach the parent, so
exit(-1)shows as 255 (bash/Python/Node/Go). -
Failed
ssh.sshexits 255 when ssh itself fails (refused, auth, host key), not when the remote command fails. Backup/deploy jobs that end insshexit 255 on connection failure. -
Docker lost the process. After a host crash, hard reboot or daemon crash, dockerd finds a container recorded as running with no saved exit status and marks it
Exited (255)— Docker's bookkeeping, not the app's verdict.
docker inspect --format 'exit={{.State.ExitCode}} started={{.State.StartedAt}} finished={{.State.FinishedAt}} policy={{.HostConfig.RestartPolicy.Name}}' YOUR_CONTAINER
docker logs --timestamps YOUR_CONTAINER 2>&1 | tail -n 30
uptime -s
sudo journalctl -u docker.service --since "24 hours ago" --no-pager | grep -i -E 'starting up|completed initialization'
ssh: errors in the last lines → fix key / known_hosts / network. finished matching boot (uptime -s) or a dockerd Starting up line, with a clean log stop → why did the host/daemon restart? finished is UTC; convert first.
4. Find the first error in the logs
docker logs --since 30m --timestamps YOUR_CONTAINER 2>&1 | head -n 80
docker logs --since 30m --timestamps YOUR_CONTAINER 2>&1 | grep -n -i -E 'error|fatal|panic|exception|denied|not found|refused' | head -n 20
Works on a stopped container until removed (--rm takes logs with it). Use 2>&1 — most errors are on stderr. Read the first error after deploy, not only the last line.
Common first lines (hypotheses until checked):
-
connection refused: dependency down or wrong address; inside a container,localhostis itself. -
permission denied: mount ownership or runtime user. -
no such file or directoryfor a script that exists: often CRLF on the shebang, or a missing interpreter. -
exec format error: wrong arch (see 126). - Missing env/config key: compare key names with the last good deploy; don't print secrets.
Redact passwords, tokens, connection strings and customer data before pasting logs.
Illustrative example
Illustrative example (not a real case): after a slimmer base image, docker ps -a shows Restarting (127) 4 seconds ago. Inspect: 127 false, empty error — process started, not OOM. started/finished ~1 s apart. First log line: /app/start.sh: 5: exec: gunicorn: not found. Entrypoint runs; new base lacks the app server. Restart policy or more memory won't help. Next: roll back the image tag, or rebuild with the dependency, then watch whether it stays up.
An exit code is a clue, not a diagnosis — combine the code, log lines and matching timestamps.
I work on OpsMate, which puts an SSH terminal and an AI assistant on one page, so you can run the commands above yourself or describe the exit in plain language and check against the raw output. Each time you ask the AI for help, recent terminal output is sent to the cloud AI along with your message, so remove sensitive information first. Every action is sorted into one of three risk tiers. Low-risk actions (such as a confident service restart or log rotation) run automatically. Medium-risk actions (such as changing config or recreating a container) are sent to you for approval first, and run automatically if no one responds within 30 minutes. High-risk actions (database schema, kernel parameter, network and firewall changes) only raise an alert and are never executed. When the system isn't sure, it treats the action as the next tier up; dangerous commands are always blocked. The full guide is on itops.sh.
Drafted with AI assistance and reviewed before publishing.
Top comments (0)