My site is small — five pages of static HTML served through a Cloudflare Worker. The thing that worried me most in the first month wasn't the content. It was the preview URLs.
The problem nobody warns you about
The moment you deploy to Cloudflare, you get extra hostnames for free: a workers.dev address, plus whatever preview or staging hostnames your setup generates. They serve exactly the same assets as production — same HTML, same CSS, same everything.
If any one of those hostnames gets crawled, you now have your entire site indexed under a domain you don't control, competing with your real domain for the same queries. And deleting the hostname later does not clean it up. The indexed URLs stay.
Three fixes that don't actually work
Ship a robots.txt in your build. It gets overwritten on the next build, and it's one file — it physically cannot say different things to different hostnames.
Put a noindex meta tag on every page. Preview and production serve identical HTML, so this noindexes production too. And it only takes effect after a crawl.
Password-protect the preview. Works, but it breaks your own testing and anything you want to hand to a collaborator.
The fix: let the Worker answer for robots.txt
The idea is simple. robots.txt is never a file in the build. The Worker intercepts that one path and returns a body based on the hostname that was actually requested.
Keep an explicit list of your production hostnames — apex, www, anything else you genuinely serve. If the request host is on that list, return the normal allow-everything file with the sitemap line pointing at production. If it is anything else, return a file that disallows everything.
Two details matter more than they look.
It's an allow-list, not a blocklist. New preview hostnames appear without you doing anything. A blocklist is always one deploy behind. An allow-list fails closed.
The sitemap line is hard-coded to production. Don't build it from the request host. Otherwise your preview's robots.txt advertises the preview's sitemap — which is exactly the URL you were trying to keep out.
Testing it takes about thirty seconds
Open the robots.txt of your production domain in a browser and read it. Then open the robots.txt of a preview hostname. The two files should say opposite things.
If they are identical, your request isn't reaching the Worker at all. The usual culprit is a robots.txt still sitting in your build output — a static file normally wins over a route, so the Worker never gets a chance to answer.
Gotchas I hit
Caching. The response varies by hostname, so any cache that doesn't key on the host will happily serve your preview the production file. Tell the response not to store — the file is a couple of hundred bytes, so there is nothing to gain by caching it and a whole class of bug to lose.
Hostname normalization. www versus apex, and case. Lowercase the hostname before comparing, and list every variant you actually serve.
The disallow form. It is the path form, disallow slash — not disallow star. Easy typo, silently wrong, and you will not notice until something is indexed.
Production rules now live in code. Once robots.txt is generated, you can't add crawl rules by editing a file. Keep production's rules in one readable block so nobody has to go hunting for them.
One thing worth knowing before you use this everywhere
Robots.txt blocks crawling, not indexing. A disallowed URL can still show up in results without a snippet if something links to it. That is fine here, because nothing links to your preview hostname — but it is the reason "just disallow everything" is not a general-purpose privacy tool.
For the same reason, disallow and noindex conflict: a disallowed URL never gets crawled, so the crawler never reads the noindex. Pick one as the primary mechanism. For a hostname that was never linked anywhere, disallow is enough. If a preview hostname may already have been crawled, add a noindex header on non-production hosts as well.
What I'd do differently
Set this up on day one. It is about twenty lines, and the cost of retrofitting is waiting out index removals on a domain you never wanted indexed.
This is the setup running on latenightraid.com — five static pages, one Worker, and no preview domain in the index.
The general version of the lesson: anything that should differ by hostname belongs in the Worker, not in a file that ships identically to every host.
If you're on Cloudflare Pages rather than Workers, the same pattern works in middleware, since the hostname is right there on the request.
Top comments (0)