<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: mgayy</title>
    <description>The latest articles on DEV Community by mgayy (@mgayy).</description>
    <link>https://dev.to/mgayy</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4165990%2F5678be6b-cef9-4903-81f2-4404ac03a469.png</url>
      <title>DEV Community: mgayy</title>
      <link>https://dev.to/mgayy</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mgayy"/>
    <language>en</language>
    <item>
      <title>A preview URL can outrank your real domain: keeping Cloudflare previews out of Google</title>
      <dc:creator>mgayy</dc:creator>
      <pubDate>Tue, 06 Oct 2026 10:23:02 +0000</pubDate>
      <link>https://dev.to/mgayy/a-preview-url-can-outrank-your-real-domain-keeping-cloudflare-previews-out-of-google-36o8</link>
      <guid>https://dev.to/mgayy/a-preview-url-can-outrank-your-real-domain-keeping-cloudflare-previews-out-of-google-36o8</guid>
      <description>&lt;p&gt;My site is small — five pages of static HTML served through a Cloudflare Worker. The thing that worried me most in the first month wasn't the content. It was the preview URLs.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem nobody warns you about
&lt;/h2&gt;

&lt;p&gt;The moment you deploy to Cloudflare, you get extra hostnames for free: a workers.dev address, plus whatever preview or staging hostnames your setup generates. They serve exactly the same assets as production — same HTML, same CSS, same everything.&lt;/p&gt;

&lt;p&gt;If any one of those hostnames gets crawled, you now have your entire site indexed under a domain you don't control, competing with your real domain for the same queries. And deleting the hostname later does not clean it up. The indexed URLs stay.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three fixes that don't actually work
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Ship a robots.txt in your build.&lt;/strong&gt; It gets overwritten on the next build, and it's one file — it physically cannot say different things to different hostnames.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Put a noindex meta tag on every page.&lt;/strong&gt; Preview and production serve identical HTML, so this noindexes production too. And it only takes effect after a crawl.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Password-protect the preview.&lt;/strong&gt; Works, but it breaks your own testing and anything you want to hand to a collaborator.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: let the Worker answer for robots.txt
&lt;/h2&gt;

&lt;p&gt;The idea is simple. robots.txt is never a file in the build. The Worker intercepts that one path and returns a body based on the hostname that was actually requested.&lt;/p&gt;

&lt;p&gt;Keep an explicit list of your production hostnames — apex, www, anything else you genuinely serve. If the request host is on that list, return the normal allow-everything file with the sitemap line pointing at production. If it is anything else, return a file that disallows everything.&lt;/p&gt;

&lt;p&gt;Two details matter more than they look.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It's an allow-list, not a blocklist.&lt;/strong&gt; New preview hostnames appear without you doing anything. A blocklist is always one deploy behind. An allow-list fails closed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The sitemap line is hard-coded to production.&lt;/strong&gt; Don't build it from the request host. Otherwise your preview's robots.txt advertises the preview's sitemap — which is exactly the URL you were trying to keep out.&lt;/p&gt;

&lt;h2&gt;
  
  
  Testing it takes about thirty seconds
&lt;/h2&gt;

&lt;p&gt;Open the robots.txt of your production domain in a browser and read it. Then open the robots.txt of a preview hostname. The two files should say opposite things.&lt;/p&gt;

&lt;p&gt;If they are identical, your request isn't reaching the Worker at all. The usual culprit is a robots.txt still sitting in your build output — a static file normally wins over a route, so the Worker never gets a chance to answer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotchas I hit
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Caching.&lt;/strong&gt; The response varies by hostname, so any cache that doesn't key on the host will happily serve your preview the production file. Tell the response not to store — the file is a couple of hundred bytes, so there is nothing to gain by caching it and a whole class of bug to lose.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hostname normalization.&lt;/strong&gt; www versus apex, and case. Lowercase the hostname before comparing, and list every variant you actually serve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The disallow form.&lt;/strong&gt; It is the path form, disallow slash — not disallow star. Easy typo, silently wrong, and you will not notice until something is indexed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Production rules now live in code.&lt;/strong&gt; Once robots.txt is generated, you can't add crawl rules by editing a file. Keep production's rules in one readable block so nobody has to go hunting for them.&lt;/p&gt;

&lt;h2&gt;
  
  
  One thing worth knowing before you use this everywhere
&lt;/h2&gt;

&lt;p&gt;Robots.txt blocks crawling, not indexing. A disallowed URL can still show up in results without a snippet if something links to it. That is fine here, because nothing links to your preview hostname — but it is the reason "just disallow everything" is not a general-purpose privacy tool.&lt;/p&gt;

&lt;p&gt;For the same reason, disallow and noindex conflict: a disallowed URL never gets crawled, so the crawler never reads the noindex. Pick one as the primary mechanism. For a hostname that was never linked anywhere, disallow is enough. If a preview hostname may already have been crawled, add a noindex header on non-production hosts as well.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;p&gt;Set this up on day one. It is about twenty lines, and the cost of retrofitting is waiting out index removals on a domain you never wanted indexed.&lt;/p&gt;

&lt;p&gt;This is the setup running on &lt;a href="https://latenightraid.com" rel="noopener noreferrer"&gt;latenightraid.com&lt;/a&gt; — five static pages, one Worker, and no preview domain in the index.&lt;/p&gt;

&lt;p&gt;The general version of the lesson: anything that should differ by hostname belongs in the Worker, not in a file that ships identically to every host.&lt;/p&gt;

&lt;p&gt;If you're on Cloudflare Pages rather than Workers, the same pattern works in middleware, since the hostname is right there on the request.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>seo</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
