Originally published on kuryzhev.cloud
The change window is open, the internal dashboard is supposed to be reachable at a public hostname with no inbound firewall rules, and the Cloudflare Tunnel you created an hour ago shows "Inactive" or returns a 502 to every browser. Nothing is listening on a public port, so there is no obvious place to look. This runbook covers the failure modes of Cloudflare Tunnel (formerly Argo Tunnel, still run by the cloudflared daemon) in the order they are usually diagnosed.
Symptoms
Match what you see to the list below before changing anything. Each symptom points at a different layer.
-
Error 1033 in the browser – Cloudflare resolved the hostname to a tunnel, but no
cloudflaredconnector is attached to it. - Error 1016 (origin DNS error) – the DNS record points at a tunnel UUID that was deleted or never existed.
-
502 Bad Gateway – the tunnel is connected, but
cloudflaredcould not get a valid response from the local service. -
404 with an empty body – the request reached
cloudflaredbut fell through to the catch-all ingress rule. - Tunnel flaps between Healthy and Degraded – connections to the edge drop and reconnect, often with timeout messages in the logs.
Start with the connector's own view of the world. The list and info subcommands need the account certificate (cert.pem) created by cloudflared tunnel login. On hosts that only have a tunnel token, check the tunnel's status and connectors in the dashboard instead.
# List tunnels and see which have active connections (requires cert.pem)
cloudflared tunnel list
# Show connector IDs, edge locations and origin IPs for one tunnel
cloudflared tunnel info my-internal-apps
# Read recent logs (systemd install)
journalctl -u cloudflared --since "30 min ago" --no-pager
If info shows no connectors, you have a connector problem. If it shows connectors but users still see errors, the problem is routing or the origin.
Root cause
A Cloudflare Tunnel has three independent links, and each can fail while the other two look fine.
-
Edge link. The outbound connection from
cloudflaredto Cloudflare's edge. -
Hostname mapping. The mapping from a public hostname to a service. It lives either in the dashboard (remotely managed tunnels) or in a local
config.yml(locally managed tunnels). -
Origin link. The connection from
cloudflaredto the actual origin service.
Many failures trace back to one of these documented behaviors:
-
cloudflaredonly makes outbound connections. If egress to the edge is blocked (the documented port is 7844, TCP and UDP), it never registers and you get 1033. -
localhostmeans "this process's own network namespace". Inside a container, it is not the host and not the neighboring container. - Ingress rules are evaluated top to bottom, and the last rule must be a catch-all. A hostname typo sends traffic to that fallback.
- A service URL with
https://and a self-signed or mismatched certificate fails TLS verification and surfaces as 502.
Watch out for: mixing management modes. If a tunnel was created in the dashboard (token-based), the dashboard configuration is authoritative and a local config.yml ingress section is ignored. Editing the wrong one is a common reason a "fix" changes nothing. Verify the current behavior in the official Cloudflare Tunnel documentation.
Fix #1: Restore the connector (1033 and Inactive status)
First confirm the process is running and can reach the edge. A connector that never registers is usually a firewall, proxy, or stale-credential issue.
# Is the service up?
systemctl status cloudflared
# Can this host reach the edge on the tunnel port over TCP?
# (hostnames per Cloudflare's firewall docs; verify the current list)
# Note: this does not prove that UDP 7844 (QUIC) is open.
nc -vz region1.v2.argotunnel.com 7844
nc -vz region2.v2.argotunnel.com 7844
# Force HTTP/2 over TCP to rule out filtered UDP (QUIC)
# Locally managed tunnel:
cloudflared tunnel --protocol http2 run my-internal-apps
# Token-based (remotely managed) tunnel:
cloudflared tunnel --protocol http2 run --token <TOKEN>
The default protocol setting is auto. It tries QUIC, which uses UDP, first and can fall back to HTTP/2 over TCP. Some corporate firewalls and cloud security groups allow outbound TCP only. In those networks the connector may fail or reconnect repeatedly, so forcing http2 is a quick way to confirm that diagnosis. If it works, choose one of two permanent fixes:
- Open outbound UDP 7844.
- Set the protocol in your configuration, using
protocol: http2inconfig.ymlor the--protocolflag in the service definition.
Other connector-level causes to rule out:
- An expired or rotated token in the service unit. Run
cloudflared service uninstall, then reinstall withcloudflared service install <TOKEN>using a fresh token from the dashboard. - A very old
cloudflaredbinary. Cloudflare supports only recent releases, so update to the current one and check the release notes for your version. - Clock drift on the host, which can break TLS to the edge. Check NTP.
If 1016 appears instead, the DNS CNAME points at a deleted tunnel. For a remotely managed tunnel, re-create the public hostname in the dashboard. For a locally managed tunnel, run cloudflared tunnel route dns --overwrite-dns my-internal-apps app.example.com so the record targets the live tunnel ID.
Fix #2: Correct the ingress and the origin address (502 and unexpected 404)
When connectors are healthy, the next suspect is the service URL. For a locally managed tunnel, validate the file before restarting anything. The example below shows the pieces that most often go wrong.
# /etc/cloudflared/config.yml (locally managed tunnel)
tunnel: 6ff42ae2-0000-0000-0000-example # tunnel UUID, placeholder
credentials-file: /etc/cloudflared/6ff42ae2-0000-0000-0000-example.json
ingress:
- hostname: app.example.com
service: http://localhost:8080 # plain HTTP to a local process
- hostname: grafana.example.com
service: https://localhost:3000
originRequest:
noTLSVerify: true # only for self-signed certs you accept; prefer caPool
- hostname: ssh.example.com
service: ssh://localhost:22
- service: http_status:404 # mandatory catch-all, must be last
Then test how a specific URL would be routed, without sending real traffic:
cloudflared tunnel ingress validate
cloudflared tunnel ingress rule https://app.example.com
The second command prints which rule matches. If it prints the catch-all, the hostname in the rule does not match what the user typed. Fix the spelling or add a wildcard rule.
For 502 specifically, test from the same host (or container) running cloudflared: curl -v http://localhost:8080. If that fails, the problem is the application, not the tunnel. If you use HTTPS to the origin, set originServerName to match the certificate's name, or supply a CA with caPool rather than disabling verification permanently.
Watch out for: noTLSVerify: true silences the error but removes protection on the hop between cloudflared and the origin. Treat it as a temporary diagnostic, not a final setting.
Fix #3: Repair container networking and unstable connections
In Docker or Kubernetes, localhost in an ingress rule is a very common 502 cause. Point the service at the container or Kubernetes Service name instead, and keep both on the same network.
# compose.yaml (no top-level "version" key needed in current Compose)
services:
cloudflared:
image: cloudflare/cloudflared:${CLOUDFLARED_VERSION:-latest} # pin CLOUDFLARED_VERSION in .env
command: tunnel --no-autoupdate run # token comes from the environment
environment:
TUNNEL_TOKEN: ${TUNNEL_TOKEN} # keep in .env or a secret store, never in git
restart: unless-stopped
networks: [internal]
app:
image: ghcr.io/example/app:1.4.2 # placeholder application image
networks: [internal]
networks:
internal: {}
With a token-based tunnel, set the public hostname's service in the dashboard to http://app:8080. In Kubernetes, use http://my-svc.my-namespace.svc.cluster.local:8080. Set CLOUDFLARED_VERSION in .env to a specific current release tag from the official registry. The latest fallback only exists so the file runs as written, so do not rely on it in production.
For flapping connections, check three things:
-
Replicas. Run at least two replicas of
cloudflaredso a single restart does not take the route down. - Resource limits. Look for limits that cause OOM kills on small containers.
-
Metrics and readiness. Inspect the connector metrics endpoint (the
--metricsflag sets the address) and its/readypath, which reports whether the connector currently has edge connections.
Prevention
Most of these failures are preventable with a few habits. Put them in your tunnel template once and reuse it across environments, the same way you would standardize other infrastructure patterns covered on kuryzhev.cloud.
- Pick one management mode per tunnel. Prefer remotely managed tunnels with Terraform or the API so configuration is reviewable, and avoid editing both places.
- Run two or more connectors on separate hosts or nodes for every production tunnel, and roll updates one at a time.
- Allow egress explicitly. Document outbound port 7844 (TCP and UDP) in your firewall or security-group baseline, and decide on protocol deliberately.
-
Pin and update
cloudflared. Use a pinned image tag, and review the release notes on a schedule so you never fall outside the supported window. - Alert on tunnel health. Use Cloudflare's tunnel health notifications, plus a probe of the public hostname, so you hear about 1033 before users do.
-
Protect the origin hop and the hostname. Put Cloudflare Access in front of internal apps, and keep TLS verification on between
cloudflaredand the origin. - Treat tokens as secrets. Store them in a secret manager, rotate them when staff change, and never bake them into images.
Before closing the ticket, run through a final checklist:
-
tunnel info(or the dashboard) shows connectors. -
ingress rulematches the intended service on locally managed tunnels. - A local
curlto the origin succeeds. - The public hostname returns the expected status.
Details can change between releases, so confirm flags and error-code meanings against the Cloudflare Tunnel documentation for your cloudflared version.
Top comments (0)