When a site starts returning 502 Bad Gateway, the instinct is to restart nginx, bump a timeout, or reboot the box. Most of the time none of that is needed, because nginx is the one part that's working. A 502 means nginx took the request, passed it to the upstream (your Node app, Gunicorn, PHP-FPM, a container), and got nothing usable back.
The fastest way out is to read the one log line nginx wrote at the moment of the failure. This post is the short version of how I do that on Ubuntu 22.04/24.04. Everything until the last section only reads state, and all of it is for servers you're authorized to access.
1. Make sure the 502 is really from your nginx
If you sit behind a CDN, its 502 page looks a lot like nginx's. Check from your laptop, then from the server itself:
curl -sS -o /dev/null -D - https://example.com/ | grep -i -E '^(HTTP|server|cf-ray)'
curl -sS -o /dev/null -w '%{http_code}\n' --resolve example.com:443:127.0.0.1 https://example.com/
A server: cloudflare or cf-ray header means the CDN answered. If the second command (run on the server, straight to local nginx) returns 200, the problem is between the CDN and your server, not behind nginx. If it's also 502, keep going.
Then confirm nginx itself is healthy:
systemctl status nginx --no-pager
sudo nginx -t
active (running) plus test is successful means stop suspecting nginx.
2. Find the line that matches the failing request
sudo tail -n 50 /var/log/nginx/error.log
sudo grep -E 'upstream|connect\(\)' /var/log/nginx/error.log | tail -n 20
If a server block has its own log, sudo nginx -T 2>/dev/null | grep -n -E 'error_log|proxy_pass|fastcgi_pass|upstream' shows where it goes and where each location proxies. (Don't paste nginx -T output anywhere public; it can contain credentials.)
A typical line:
[error] connect() failed (111: Connection refused) while connecting to upstream, client: 203.0.113.10, server: example.com, request: "GET /api/orders HTTP/1.1", upstream: "http://127.0.0.1:3000/api/orders"
The upstream: field is the exact address nginx tried. Comparing it with where the app actually listens solves a surprising number of 502s on the spot.
3. Map the message to a cause
| error.log says | What it usually means |
|---|---|
(111: Connection refused) |
Nothing listening at that address: app stopped, crashed, still starting, or on another port |
unix:...sock failed (2: No such file or directory) |
Socket file doesn't exist: service stopped, or nginx points at an old path (classic after a PHP upgrade) |
unix:...sock failed (13: Permission denied) |
www-data can't open the socket |
unix:...sock failed (11: Resource temporarily unavailable) |
All workers busy, socket backlog overflowed |
upstream prematurely closed connection |
App accepted the request then died: crash, killed worker, app-server timeout |
(104: Connection reset by peer) |
Process restarted or worker killed, or a reused keepalive connection was already closed |
upstream sent too big header |
Response headers (often a growing Set-Cookie) don't fit the buffer |
SSL_do_handshake() failed ... wrong version number |
proxy_pass https:// pointed at a plain-HTTP port |
no live upstreams |
Every server in the upstream {} group was marked failed earlier; scroll up for why |
upstream timed out (110: ...) |
That's a 504, not a 502: the app is slow, not gone |
Two things that look related but aren't 502s: host not found in upstream is logged at [emerg] when nginx loads config, so nginx won't start or reload at all. And (13: Permission denied) on a TCP address is the SELinux case on RHEL-family systems; stock Ubuntu nginx isn't confined by AppArmor, so there it's a plain file-permission problem.
If you saw "The proxy server received an invalid response from an upstream server", that page is Apache mod_proxy, not nginx.
4. Confirm it on the upstream side
Is the service up, and what did it log?
systemctl status YOUR_APP --no-pager
sudo journalctl -u YOUR_APP --since "30 min ago" --no-pager | tail -n 50
docker ps -a --format 'table {{.Names}}\t{{.Status}}\t{{.Ports}}'
Is something listening where nginx connects?
sudo ss -ltnp
sudo ss -lxp | grep -E 'php|gunicorn|uwsgi|\.sock'
ls -l /run/php/
Three mismatches I see a lot:
- App moved to 3001 after a deploy, nginx still sends to 3000.
-
proxy_pass http://localhost:3000while the app only listens on 127.0.0.1. nginx may resolvelocalhostto both 127.0.0.1 and[::1], and the IPv6 attempt is refused; the log showsupstream: "http://[::1]:3000/...". Writing127.0.0.1removes the ambiguity. -
php8.1-fpm.sockin the site config, but onlyphp8.3-fpm.sockexists after an Ubuntu 22.04 to 24.04 upgrade.
Does the app answer if you skip nginx?
curl -sS -o /dev/null -w '%{http_code} %{time_total}s\n' --max-time 10 -H 'Host: example.com' http://127.0.0.1:3000/
Refused here too means the app isn't serving. 200 here but 502 through nginx means look at the address mismatch, socket permissions, or the response headers. (PHP-FPM speaks FastCGI, so curl can't test it directly.)
Was a worker killed? For prematurely closed or reset by peer, check the same minute:
sudo journalctl -k --since "1 hour ago" --no-pager | grep -i -E 'out of memory|killed process'
sudo journalctl -u YOUR_APP --since "1 hour ago" --no-pager | grep -i -E 'worker timeout|killed|signal'
Gunicorn's [CRITICAL] WORKER TIMEOUT means Gunicorn killed the worker after its own --timeout (30 s default). Raising nginx timeouts won't help, because the app server cut the request first.
The Docker gotcha
Inside an nginx container, proxy_pass http://127.0.0.1:3000; points at the nginx container itself. Proxy to the Compose service name (http://app:3000), make the app listen on 0.0.0.0 inside its container, or, with nginx on the host, publish the port to loopback (127.0.0.1:3000:3000). Also remember nginx resolves hostnames at config load: if the app container is recreated with a new IP, nginx can keep using the old one until reload.
When the log points at buffers or keepalive
upstream sent too big header needs bigger buffers, and you have to raise them together:
location / {
proxy_pass http://127.0.0.1:3000;
proxy_buffer_size 16k;
proxy_buffers 8 16k;
}
Raising only proxy_buffer_size makes nginx -t fail on proxy_busy_buffers_size. For PHP-FPM the equivalents are fastcgi_buffer_size / fastcgi_buffers.
Occasional 502s with an upstream {} block that uses keepalive: if the app closes idle connections sooner than nginx expects, a request can land on a just-closed connection. Node.js closes idle keep-alive connections after 5 s by default (server.keepAliveTimeout), so make the app's idle timeout longer than nginx's upstream keepalive_timeout (60 s default), or lower the latter.
Changing things safely
Only now do we change anything, and only the one thing the log pointed at:
- Back up the file (
sudo cp /etc/nginx/sites-available/example.com /root/example.com.bak). - Make the single change.
-
sudo nginx -t, thensudo systemctl reload nginx(reload keeps connections; restart drops them). - Repeat the failing request with
sudo tail -f /var/log/nginx/error.logopen. - No better? Restore the backup, test, reload.
Don't raise every timeout "just in case", and don't chmod 0666 a socket. Fix what the log names.
I work on OpsMate, which puts an SSH terminal and an AI assistant on one page, so you can run the commands above yourself or describe the 502 in plain language and check the answer against the raw output. Each time you ask the AI for help, recent terminal output is sent to the cloud AI along with your message, so remove sensitive information first. Every action is sorted into one of three risk tiers. Low-risk actions (such as a confident service restart or log rotation) run automatically. Medium-risk actions (such as changing config or recreating a container) are sent to you for approval first, and run automatically if no one responds within 30 minutes. High-risk actions (database schema, kernel parameter, network and firewall changes) only raise an alert and are never executed. When the system isn't sure, it treats the action as the next tier up; dangerous commands are always blocked. The full guide, with every error message and the PHP-FPM and socket-permission details, is on itops.sh.
Drafted with AI assistance and reviewed before publishing.
Top comments (0)