DEV Community

Cover image for Cloudflare Tunnel 530 errors: the connector that was healthy and did nothing
Christian Anderson
Christian Anderson

Posted on

Cloudflare Tunnel 530 errors: the connector that was healthy and did nothing

It started as "Immich has gone down". Immich was fine. The server container was healthy and answered locally. What had gone was every public hostname I run through Cloudflare Tunnel, all at once, and it had been that way for about a day before anyone noticed.

This is what happened, the mistake that made me declare it fixed when it wasn't, and the checks I now use so the tunnel can't lie to me again.

A healthy container connecting nothing

On the evening of 3 September, around 22:00, two containers on my NAS were recreated. Both ran a community web UI wrapper for cloudflared, the Cloudflare Tunnel client. When they came back up, both logged this:

CONFIG: No pre-existing config file found
Enter fullscreen mode Exit fullscreen mode

That line is the whole incident. The wrapper is a web UI, not a connector. With no config it starts happily, shows as healthy in docker ps, and connects no tunnel at all. Both of its config directories were empty. From Cloudflare's side the tunnel had zero connectors, so every public hostname behind it returned 530.

That included my single sign-on portal. Anything behind SSO was unreachable from outside, not just the photo library. The photo library was just the first thing to be noticed.

The container's health check was checking the UI, and the UI was fine. Nothing in docker ps distinguishes "a web page is being served" from "traffic is flowing through a tunnel".

The fix: run an actual connector

Instead of trying to rebuild the wrapper's config, I ran the real thing, the official image with the tunnel token:

docker run -d --name cloudflared-connector --restart unless-stopped \
  cloudflare/cloudflared:latest tunnel --no-autoupdate run --token "$TOK"
Enter fullscreen mode Exit fullscreen mode

I pulled the token from the Cloudflare API (GET /accounts/{account}/cfd_tunnel/{tunnel}/token) using an API token already scoped to Tunnel edit and DNS edit. The tunnel then reported healthy with connections=4, and all five public hostnames were back.

A connector-only container has no UI and nothing to forget. The token is the whole configuration, because the routes live in Cloudflare's dashboard. That is exactly what I want from the piece that holds the front door open.

The two config-less wrapper containers kept running afterwards, doing nothing. They're harmless, but they're a good reminder that "running" and "working" are different claims.

The trap: split-horizon DNS makes local tests lie

Here is the part I'm least proud of. During the incident I tested a hostname from inside the network, got a clean redirect back and declared external access working. It was 530 the whole time.

My router's resolver is AdGuard Home, and it has rewrites for my own domain. Inside the house, a lookup for one of my hostnames returns the private address of my reverse proxy. Outside, public DNS returns Cloudflare's addresses. That is deliberate split-horizon DNS, and it means:

  • curl https://<my hostname>/ from any machine at home never touches Cloudflare.
  • It goes straight to the reverse proxy on the LAN, gets a perfect 302 to the login page, and proves nothing about the tunnel.

The tunnel could be completely dead and every local test would still pass.

How to test the real path

First, check whether you're in split-horizon territory at all. Compare what a public resolver says with what your own machine says:

dig +short @1.1.1.1 app.example.com   # what the internet sees
dig +short app.example.com            # what this machine sees
Enter fullscreen mode Exit fullscreen mode

If those answers differ, a plain curl from this machine is testing your LAN, not your tunnel.

Then force the request through the Cloudflare edge by pinning the hostname to the public answer, while keeping the right SNI and Host header:

EDGE=$(dig +short @1.1.1.1 app.example.com | head -1)
curl -sS -o /dev/null -w '%{http_code}\n' \
  --resolve app.example.com:443:"$EDGE" https://app.example.com/
Enter fullscreen mode Exit fullscreen mode

A 530 here means the edge has nowhere to send the request. A 200, 302 or 303 means the full path works: edge, tunnel, connector, origin.

--resolve also splits problems the other way. Point it at your reverse proxy's address instead, and it proves the proxy and certificate work independently of DNS. If that passes and the normal request fails, the fault is name resolution, not the service.

Ask the client's own resolver

The same mistake has caught me with DNS on the client side too. The week before, a household device had been handed a smart-DNS streaming box as its resolver by a DHCP tag, and that box answers nothing at all for my domain. I spent a session checking AdGuard, Cloudflare, the browser cache and the firewall, and they were all innocent.

The step that would have found it straight away: query the resolver the client is actually using, the one in its own DHCP lease, not the one you think it should be using. dig @1.1.1.1 and dig @router tell you about those resolvers, not about the device that's broken.

A public record pointing at a private address

One more split-horizon wrinkle turned up the same day. One of my hostnames had an A record in public DNS pointing at a private LAN address. It had only ever worked at home, because at home that address is reachable. From anywhere else it was a dead end. I moved it onto the tunnel, with the tunnel sending it to the reverse proxy and setting the origin server name so the proxy picks the right site block and still applies SSO. Then I verified it the only way that counts: GET and POST from the edge, both redirecting to the login portal.

Monitoring that was lying as well

On 13 September I found my Uptime Kuma "Cloudflare" monitor had been red for more than 2,700 heartbeats while the tunnel was fine. It had been polling the wrapper's web UI port, which no longer answered, so it got ECONNREFUSED. A monitor that's always red gets ignored, which is just as bad as one that's always green.

I replaced it with a push monitor that tests the path a real visitor takes:

  • a script on a systemd timer runs every five minutes;
  • it resolves two of my hostnames through 1.1.1.1, so it gets the public answer, not the split-horizon one;
  • it calls each with curl --resolve to the Cloudflare edge;
  • it only reports DOWN if both come back 530 or give no answer, so one flaky app can't page me about the tunnel;
  • Kuma expects a push every 600 seconds, so if the script itself dies, that goes red too.

It has a --dry-run flag so I can see what it would report without pushing anything.

The happy ending: a four-minute self-heal

This morning, 24 September, the NAS went down hard. Its journal simply stops mid-stream at 10:13:24 BST, with no shutdown sequence, which is what a power loss or hard reset looks like. It booted again at 10:14.

At 10:17 the connector container, set to restart unless-stopped, re-registered all four connections with Cloudflare. Total public outage: about four minutes. I didn't have to do anything.

Three weeks earlier, a container recreate had cost a day of silent outage. The difference wasn't a clever fix. It was running the real connector instead of a UI wrapper, and testing through the edge instead of through my own DNS.

What I'd tell myself three weeks earlier

  • A healthy container is not a working tunnel. Check the connector count on Cloudflare's side.
  • If your LAN resolves your own domain to a private address, every local curl is testing your LAN. Use curl --resolve against the public answer.
  • Compare dig @1.1.1.1 with your local answer before trusting any test.
  • When a client can't resolve something, query that client's own resolver first.
  • A monitor that probes the wrong thing is worse than no monitor. Probe the path a user takes.
  • Give the connector a restart policy and nothing else to remember. Then a power cut is a four-minute blip, not a day of "Immich is down".

The longer version of this, with the DNS, DHCP and tunnel traps side by side and a checklist, is a short field report: Home Network Traps.


🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.

Top comments (0)