<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: itops</title>
    <description>The latest articles on DEV Community by itops (@itops).</description>
    <link>https://dev.to/itops</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4170652%2F5338e322-d77e-4832-bae6-2dbc7835d217.png</url>
      <title>DEV Community: itops</title>
      <link>https://dev.to/itops</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/itops"/>
    <language>en</language>
    <item>
      <title>Nginx 502 Bad Gateway: read one log line before you touch anything</title>
      <dc:creator>itops</dc:creator>
      <pubDate>Fri, 09 Oct 2026 07:30:27 +0000</pubDate>
      <link>https://dev.to/itops/nginx-502-bad-gateway-read-one-log-line-before-you-touch-anything-16jo</link>
      <guid>https://dev.to/itops/nginx-502-bad-gateway-read-one-log-line-before-you-touch-anything-16jo</guid>
      <description>&lt;p&gt;When a site starts returning &lt;strong&gt;502 Bad Gateway&lt;/strong&gt;, the instinct is to restart nginx, bump a timeout, or reboot the box. Most of the time none of that is needed, because nginx is the one part that's working. A 502 means nginx took the request, passed it to the upstream (your Node app, Gunicorn, PHP-FPM, a container), and got nothing usable back.&lt;/p&gt;

&lt;p&gt;The fastest way out is to read the one log line nginx wrote at the moment of the failure. This post is the short version of how I do that on Ubuntu 22.04/24.04. Everything until the last section only reads state, and all of it is for servers you're authorized to access.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Make sure the 502 is really from your nginx
&lt;/h2&gt;

&lt;p&gt;If you sit behind a CDN, its 502 page looks a lot like nginx's. Check from your laptop, then from the server itself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-D&lt;/span&gt; - https://example.com/ | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'^(HTTP|server|cf-ray)'&lt;/span&gt;
curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code}\n'&lt;/span&gt; &lt;span class="nt"&gt;--resolve&lt;/span&gt; example.com:443:127.0.0.1 https://example.com/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A &lt;code&gt;server: cloudflare&lt;/code&gt; or &lt;code&gt;cf-ray&lt;/code&gt; header means the CDN answered. If the second command (run on the server, straight to local nginx) returns 200, the problem is between the CDN and your server, not behind nginx. If it's also 502, keep going.&lt;/p&gt;

&lt;p&gt;Then confirm nginx itself is healthy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;systemctl status nginx &lt;span class="nt"&gt;--no-pager&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;nginx &lt;span class="nt"&gt;-t&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;active (running)&lt;/code&gt; plus &lt;code&gt;test is successful&lt;/code&gt; means stop suspecting nginx.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Find the line that matches the failing request
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo tail&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 50 /var/log/nginx/error.log
&lt;span class="nb"&gt;sudo grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'upstream|connect\(\)'&lt;/span&gt; /var/log/nginx/error.log | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 20
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a server block has its own log, &lt;code&gt;sudo nginx -T 2&amp;gt;/dev/null | grep -n -E 'error_log|proxy_pass|fastcgi_pass|upstream'&lt;/code&gt; shows where it goes and where each location proxies. (Don't paste &lt;code&gt;nginx -T&lt;/code&gt; output anywhere public; it can contain credentials.)&lt;/p&gt;

&lt;p&gt;A typical line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[error] connect() failed (111: Connection refused) while connecting to upstream, client: 203.0.113.10, server: example.com, request: "GET /api/orders HTTP/1.1", upstream: "http://127.0.0.1:3000/api/orders"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;upstream:&lt;/code&gt; field is the exact address nginx tried. Comparing it with where the app actually listens solves a surprising number of 502s on the spot.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Map the message to a cause
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;error.log says&lt;/th&gt;
&lt;th&gt;What it usually means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;(111: Connection refused)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Nothing listening at that address: app stopped, crashed, still starting, or on another port&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;unix:...sock failed (2: No such file or directory)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Socket file doesn't exist: service stopped, or nginx points at an old path (classic after a PHP upgrade)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;unix:...sock failed (13: Permission denied)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;www-data&lt;/code&gt; can't open the socket&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;unix:...sock failed (11: Resource temporarily unavailable)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;All workers busy, socket backlog overflowed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;upstream prematurely closed connection&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;App accepted the request then died: crash, killed worker, app-server timeout&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;(104: Connection reset by peer)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Process restarted or worker killed, or a reused keepalive connection was already closed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;upstream sent too big header&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Response headers (often a growing &lt;code&gt;Set-Cookie&lt;/code&gt;) don't fit the buffer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;SSL_do_handshake() failed ... wrong version number&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;proxy_pass https://&lt;/code&gt; pointed at a plain-HTTP port&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;no live upstreams&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Every server in the &lt;code&gt;upstream {}&lt;/code&gt; group was marked failed earlier; scroll up for why&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;upstream timed out (110: ...)&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;That's a &lt;strong&gt;504&lt;/strong&gt;, not a 502: the app is slow, not gone&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two things that look related but aren't 502s: &lt;code&gt;host not found in upstream&lt;/code&gt; is logged at &lt;code&gt;[emerg]&lt;/code&gt; when nginx loads config, so nginx won't start or reload at all. And &lt;code&gt;(13: Permission denied)&lt;/code&gt; on a &lt;strong&gt;TCP&lt;/strong&gt; address is the SELinux case on RHEL-family systems; stock Ubuntu nginx isn't confined by AppArmor, so there it's a plain file-permission problem.&lt;/p&gt;

&lt;p&gt;If you saw "The proxy server received an invalid response from an upstream server", that page is Apache &lt;code&gt;mod_proxy&lt;/code&gt;, not nginx.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Confirm it on the upstream side
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is the service up, and what did it log?&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;systemctl status YOUR_APP &lt;span class="nt"&gt;--no-pager&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;journalctl &lt;span class="nt"&gt;-u&lt;/span&gt; YOUR_APP &lt;span class="nt"&gt;--since&lt;/span&gt; &lt;span class="s2"&gt;"30 min ago"&lt;/span&gt; &lt;span class="nt"&gt;--no-pager&lt;/span&gt; | &lt;span class="nb"&gt;tail&lt;/span&gt; &lt;span class="nt"&gt;-n&lt;/span&gt; 50
docker ps &lt;span class="nt"&gt;-a&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt; &lt;span class="s1"&gt;'table {{.Names}}\t{{.Status}}\t{{.Ports}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Is something listening where nginx connects?&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;ss &lt;span class="nt"&gt;-ltnp&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;ss &lt;span class="nt"&gt;-lxp&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'php|gunicorn|uwsgi|\.sock'&lt;/span&gt;
&lt;span class="nb"&gt;ls&lt;/span&gt; &lt;span class="nt"&gt;-l&lt;/span&gt; /run/php/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three mismatches I see a lot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;App moved to 3001 after a deploy, nginx still sends to 3000.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;proxy_pass http://localhost:3000&lt;/code&gt; while the app only listens on 127.0.0.1. nginx may resolve &lt;code&gt;localhost&lt;/code&gt; to both 127.0.0.1 and &lt;code&gt;[::1]&lt;/code&gt;, and the IPv6 attempt is refused; the log shows &lt;code&gt;upstream: "http://[::1]:3000/..."&lt;/code&gt;. Writing &lt;code&gt;127.0.0.1&lt;/code&gt; removes the ambiguity.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;php8.1-fpm.sock&lt;/code&gt; in the site config, but only &lt;code&gt;php8.3-fpm.sock&lt;/code&gt; exists after an Ubuntu 22.04 to 24.04 upgrade.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Does the app answer if you skip nginx?&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s1"&gt;'%{http_code} %{time_total}s\n'&lt;/span&gt; &lt;span class="nt"&gt;--max-time&lt;/span&gt; 10 &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s1"&gt;'Host: example.com'&lt;/span&gt; http://127.0.0.1:3000/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Refused here too means the app isn't serving. 200 here but 502 through nginx means look at the address mismatch, socket permissions, or the response headers. (PHP-FPM speaks FastCGI, so curl can't test it directly.)&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Was a worker killed?&lt;/strong&gt; For &lt;code&gt;prematurely closed&lt;/code&gt; or &lt;code&gt;reset by peer&lt;/code&gt;, check the same minute:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;journalctl &lt;span class="nt"&gt;-k&lt;/span&gt; &lt;span class="nt"&gt;--since&lt;/span&gt; &lt;span class="s2"&gt;"1 hour ago"&lt;/span&gt; &lt;span class="nt"&gt;--no-pager&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'out of memory|killed process'&lt;/span&gt;
&lt;span class="nb"&gt;sudo &lt;/span&gt;journalctl &lt;span class="nt"&gt;-u&lt;/span&gt; YOUR_APP &lt;span class="nt"&gt;--since&lt;/span&gt; &lt;span class="s2"&gt;"1 hour ago"&lt;/span&gt; &lt;span class="nt"&gt;--no-pager&lt;/span&gt; | &lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="nt"&gt;-E&lt;/span&gt; &lt;span class="s1"&gt;'worker timeout|killed|signal'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Gunicorn's &lt;code&gt;[CRITICAL] WORKER TIMEOUT&lt;/code&gt; means Gunicorn killed the worker after its own &lt;code&gt;--timeout&lt;/code&gt; (30 s default). Raising nginx timeouts won't help, because the app server cut the request first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Docker gotcha
&lt;/h2&gt;

&lt;p&gt;Inside an nginx container, &lt;code&gt;proxy_pass http://127.0.0.1:3000;&lt;/code&gt; points at the nginx container itself. Proxy to the Compose service name (&lt;code&gt;http://app:3000&lt;/code&gt;), make the app listen on &lt;code&gt;0.0.0.0&lt;/code&gt; inside its container, or, with nginx on the host, publish the port to loopback (&lt;code&gt;127.0.0.1:3000:3000&lt;/code&gt;). Also remember nginx resolves hostnames at config load: if the app container is recreated with a new IP, nginx can keep using the old one until reload.&lt;/p&gt;

&lt;h2&gt;
  
  
  When the log points at buffers or keepalive
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;upstream sent too big header&lt;/code&gt; needs bigger buffers, and you have to raise them together:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight nginx"&gt;&lt;code&gt;&lt;span class="k"&gt;location&lt;/span&gt; &lt;span class="n"&gt;/&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_pass&lt;/span&gt; &lt;span class="s"&gt;http://127.0.0.1:3000&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_buffer_size&lt;/span&gt; &lt;span class="mi"&gt;16k&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="kn"&gt;proxy_buffers&lt;/span&gt; &lt;span class="mi"&gt;8&lt;/span&gt; &lt;span class="mi"&gt;16k&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Raising only &lt;code&gt;proxy_buffer_size&lt;/code&gt; makes &lt;code&gt;nginx -t&lt;/code&gt; fail on &lt;code&gt;proxy_busy_buffers_size&lt;/code&gt;. For PHP-FPM the equivalents are &lt;code&gt;fastcgi_buffer_size&lt;/code&gt; / &lt;code&gt;fastcgi_buffers&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Occasional 502s with an &lt;code&gt;upstream {}&lt;/code&gt; block that uses &lt;code&gt;keepalive&lt;/code&gt;: if the app closes idle connections sooner than nginx expects, a request can land on a just-closed connection. Node.js closes idle keep-alive connections after 5 s by default (&lt;code&gt;server.keepAliveTimeout&lt;/code&gt;), so make the app's idle timeout longer than nginx's upstream &lt;code&gt;keepalive_timeout&lt;/code&gt; (60 s default), or lower the latter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Changing things safely
&lt;/h2&gt;

&lt;p&gt;Only now do we change anything, and only the one thing the log pointed at:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Back up the file (&lt;code&gt;sudo cp /etc/nginx/sites-available/example.com /root/example.com.bak&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Make the single change.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;sudo nginx -t&lt;/code&gt;, then &lt;code&gt;sudo systemctl reload nginx&lt;/code&gt; (reload keeps connections; restart drops them).&lt;/li&gt;
&lt;li&gt;Repeat the failing request with &lt;code&gt;sudo tail -f /var/log/nginx/error.log&lt;/code&gt; open.&lt;/li&gt;
&lt;li&gt;No better? Restore the backup, test, reload.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Don't raise every timeout "just in case", and don't &lt;code&gt;chmod 0666&lt;/code&gt; a socket. Fix what the log names.&lt;/p&gt;




&lt;p&gt;I work on OpsMate, which puts an SSH terminal and an AI assistant on one page, so you can run the commands above yourself or describe the 502 in plain language and check the answer against the raw output. Each time you ask the AI for help, recent terminal output is sent to the cloud AI along with your message, so remove sensitive information first. Every action is sorted into one of three risk tiers. Low-risk actions (such as a confident service restart or log rotation) run automatically. Medium-risk actions (such as changing config or recreating a container) are sent to you for approval first, and run automatically if no one responds within 30 minutes. High-risk actions (database schema, kernel parameter, network and firewall changes) only raise an alert and are never executed. When the system isn't sure, it treats the action as the next tier up; dangerous commands are always blocked. The full guide, with every error message and the PHP-FPM and socket-permission details, is &lt;a href="https://www.itops.sh/en/guides/nginx-502-bad-gateway/?utm_source=devto&amp;amp;utm_medium=post&amp;amp;utm_campaign=sprint_w2" rel="noopener noreferrer"&gt;on itops.sh&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Drafted with AI assistance and reviewed before publishing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>nginx</category>
      <category>devops</category>
      <category>linux</category>
      <category>ubuntu</category>
    </item>
  </channel>
</rss>
