DEV Community

Andrea Mancuso
Andrea Mancuso

Posted on

Cloudflare broke my website without expiring a single certificate

My site recently stopped working.

The frontend returned HTTP 500, while the API hostname behind it returned Cloudflare’s rather cheerful:

525: SSL handshake failed

My first suspicion was an expired certificate. Certificates are convenient suspects: they expire predictably, fail dramatically, and rarely retain legal counsel.

Unfortunately for the investigation, the certificate was fine.

The setup

The frontend at <domain>.com calls an API through:

https://<subdomain>.<domain>.com
Enter fullscreen mode Exit fullscreen mode

That hostname is proxied through Cloudflare. Behind Cloudflare is an Nginx server using a Let’s Encrypt ECDSA certificate and proxying requests to the actual API.

The request path looks roughly like this:

Browser
  → Cloudflare
    → Nginx origin
      → API
Enter fullscreen mode Exit fullscreen mode

The frontend returned 500 because its API requests were failing. The API returned 525 because Cloudflare could no longer complete TLS with Nginx.

Simple enough, except nothing relevant had knowingly changed on the origin.

The certificate was innocent

Certbot reported that the certificate was valid for another two months:

Certificate Name: <subdomain>.<domain>.com
Domains: <subdomain>.<domain>.com
Expiry Date: 2026-11-28
Enter fullscreen mode Exit fullscreen mode

Checking the certificate directly confirmed the same thing.

I then bypassed Cloudflare while preserving the hostname, SNI, and certificate verification:

curl \
  --resolve <subdomain>.<domain>.com:443:<origin-ip> \
  https://<subdomain>.<domain>.com/rpc
Enter fullscreen mode Exit fullscreen mode

The result was HTTP 401.

That was good news. The TLS handshake completed, the certificate verified, Nginx accepted the request, and the API rejected it because I had deliberately supplied no credentials.

Through Cloudflare, the same endpoint returned 525.

So:

Connection Result
Client → origin directly TLS succeeded
Cloudflare → origin TLS failed
Certificate validation Succeeded
Application reached through Cloudflare No

Cloudflare defines a 525 response as a failure while negotiating TLS with the origin. In this case, the origin logs were more specific:

SSL_do_handshake() failed
SSL routines::bad cipher
Enter fullscreen mode Exit fullscreen mode

Cloudflare was reaching the correct server. Nginx was simply rejecting its handshake before an HTTP request could exist.

The certificate, meanwhile, continued sitting there being valid and unhelpful.

Comparing a working virtual host

Another Cloudflare-proxied hostname on the same server still worked.

It used:

  • The same Nginx installation
  • The same OpenSSL installation
  • The same Cloudflare-to-origin path
  • The same ECDSA certificate type

The meaningful difference was its TLS configuration.

The working virtual host included Certbot’s current TLS policy:

include /etc/letsencrypt/options-ssl-nginx.conf;
ssl_dhparam /etc/letsencrypt/ssl-dhparams.pem;
Enter fullscreen mode Exit fullscreen mode

The broken virtual host did not. It inherited older global defaults:

ssl_protocols TLSv1 TLSv1.1 TLSv1.2 TLSv1.3;
ssl_prefer_server_ciphers on;
Enter fullscreen mode Exit fullscreen mode

The Certbot policy instead provided modern protocols, an explicit compatible cipher list, consistent session behavior, and:

ssl_protocols TLSv1.2 TLSv1.3;
ssl_prefer_server_ciphers off;
ssl_session_tickets off;
Enter fullscreen mode Exit fullscreen mode

Cloudflare documents the cipher suites it presents to origin servers and provides corresponding Nginx configuration guidance.

Apparently, “this configuration has worked for months” is no longer a supported TLS profile.

The fix

I added the existing Certbot policy to the affected virtual host:

server {
    listen 443 ssl http2;
    listen [::]:443 ssl http2;

    server_name <subdomain>.<domain>.com;

    ssl_certificate /etc/letsencrypt/live/<subdomain>.<domain>.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/<subdomain>.<domain>.com/privkey.pem;

    include /etc/letsencrypt/options-ssl-nginx.conf;
    ssl_dhparam /etc/letsencrypt/ssl-dhparams.pem;

    # Proxy configuration...
}
Enter fullscreen mode Exit fullscreen mode

Then I validated and reloaded Nginx:

sudo nginx -t
sudo systemctl reload nginx
Enter fullscreen mode Exit fullscreen mode

The results changed immediately:

API root:     404
API /rpc:     401
Frontend:     200
Enter fullscreen mode Exit fullscreen mode

The 404 and 401 were expected application responses. Their presence proved that requests were passing through Cloudflare, completing TLS, reaching Nginx, and arriving at the application again.

The frontend recovered without a redeployment.

Why did this fix it?

The older configuration allowed Nginx to impose its own cipher preference and otherwise relied on inherited OpenSSL defaults.

Cloudflare’s origin-facing handshake behavior had apparently changed—or traffic had moved to infrastructure using different handshake behavior. Either way, the previously tolerated configuration began hitting OpenSSL’s bad cipher path.

The replacement policy gave Nginx and Cloudflare an explicitly compatible TLS configuration. Because I applied the Certbot policy as a unit, I cannot identify the individual directive responsible. The explicit cipher configuration is an obvious candidate, while the change in cipher preference may also have affected negotiation.

That allowed Cloudflare’s preference to govern selection from the mutually supported cipher suites, avoiding the negotiation inconsistency.

I applied the known-working policy as a unit, so I cannot prove which individual directive cured the handshake without repeatedly breaking production to perform an admirably scientific series of outages. The Diff Isolator badge will have to wait.

The ssl_dhparam directive was probably not responsible because the certificate and likely negotiated suites use ECDSA and ECDHE. It accompanies Certbot’s standard configuration.

Did Cloudflare cause the outage?

With high confidence, yes—indirectly.

The origin had a latent compatibility problem: one virtual host was missing the current TLS policy. That configuration had nevertheless continued working until Cloudflare’s origin handshake behavior changed.

The evidence was:

  1. The certificate was valid.
  2. Direct TLS connections continued to work.
  3. Only Cloudflare-origin connections failed.
  4. Nginx logged bad cipher for those connections.
  5. A sibling Cloudflare-proxied virtual host with the current TLS policy worked.
  6. Applying that policy changed the public result from 525 to normal application responses immediately.
  7. No certificate renewal or application deployment was required.

Absolute proof would require Cloudflare’s internal rollout records. Sadly, those were not thoughtfully deposited in /var/log/nginx.

Cloudflare triggered the failure. The older Nginx configuration supplied the opportunity.

Lessons learned

Test the origin separately

Use curl --resolve to bypass a proxy while preserving the hostname:

curl \
  --resolve example.com:443:<origin-ip> \
  https://example.com/
Enter fullscreen mode Exit fullscreen mode

This distinguishes a certificate or origin problem from a proxy-to-origin compatibility problem.

A 525 occurs before your application

If Cloudflare returns 525, application logs may contain nothing because no HTTP request was created. Start with the TLS terminator’s logs.

Compare working virtual hosts

A healthy service on the same server provides a useful control group. Differences between two virtual hosts can be more informative than several pages of generic troubleshooting advice.

Keep TLS policy explicit and consistent

Do not let individual virtual hosts quietly inherit years-old defaults. Shared configuration is less exciting, which is generally a desirable property in cryptography.

A valid certificate does not guarantee a working TLS handshake

Certificates are only one part of negotiation. Protocol versions, cipher suites, key exchange, preference order, session handling, and proxy behavior can all ruin your afternoon independently.

The short version

Cloudflare began making an origin TLS handshake that one older Nginx virtual host could not handle. Nginx rejected it with bad cipher, Cloudflare returned 525, and the frontend converted its failed API calls into HTTP 500.

The certificate had not expired.

Adding the current Certbot TLS policy fixed the negotiation immediately.

Infrastructure is wonderful because a configuration can remain unchanged, pass direct tests, retain a valid certificate, and still stop working when somebody else improves something.


Enter fullscreen mode Exit fullscreen mode

Top comments (0)