<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: kacey noms</title>
    <description>The latest articles on DEV Community by kacey noms (@kacey_3785aafd260f9).</description>
    <link>https://dev.to/kacey_3785aafd260f9</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4171043%2F6b41e71c-d030-4973-a8e7-50a63cbaf16e.png</url>
      <title>DEV Community: kacey noms</title>
      <link>https://dev.to/kacey_3785aafd260f9</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/kacey_3785aafd260f9"/>
    <language>en</language>
    <item>
      <title>Keeping AI search bots in and AI training bots out on a small Next.js site</title>
      <dc:creator>kacey noms</dc:creator>
      <pubDate>Thu, 08 Oct 2026 10:37:40 +0000</pubDate>
      <link>https://dev.to/kacey_3785aafd260f9/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-nextjs-site-4e16</link>
      <guid>https://dev.to/kacey_3785aafd260f9/keeping-ai-search-bots-in-and-ai-training-bots-out-on-a-small-nextjs-site-4e16</guid>
      <description>&lt;p&gt;I look after the website for &lt;a href="https://www.capital-cs.com" rel="noopener noreferrer"&gt;Capital Complete Solutions&lt;/a&gt;, a property inventory company in Birmingham where I also work as an inventory clerk. It's a Next.js app on Vercel with Payload CMS, about 30 pages.&lt;/p&gt;

&lt;p&gt;This week I wanted to sort out one thing: let AI search tools find and cite our pages, but stop crawlers that only collect training data. I asked on the Cloudflare community and got more useful answers than I expected. Here's what I learned.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. AI companies run separate bots for training and for search
&lt;/h2&gt;

&lt;p&gt;The big providers split their crawlers by job, and each one is its own robots.txt token:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Training&lt;/th&gt;
&lt;th&gt;Search index&lt;/th&gt;
&lt;th&gt;Fetch for a user&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;GPTBot&lt;/td&gt;
&lt;td&gt;OAI-SearchBot&lt;/td&gt;
&lt;td&gt;ChatGPT-User&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic&lt;/td&gt;
&lt;td&gt;ClaudeBot&lt;/td&gt;
&lt;td&gt;Claude-SearchBot&lt;/td&gt;
&lt;td&gt;Claude-User&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Perplexity&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;PerplexityBot&lt;/td&gt;
&lt;td&gt;Perplexity-User&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common Crawl&lt;/td&gt;
&lt;td&gt;CCBot&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;td&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you also want to opt out of Gemini training, Google-Extended is the token for that, and it doesn't affect Google Search. Names do change, so check each provider's own docs before copying this.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. My robots.txt was letting everything in
&lt;/h2&gt;

&lt;p&gt;Someone on the thread looked at our live robots.txt and pointed out that everything was falling through the &lt;code&gt;User-agent: *&lt;/code&gt; group, so GPTBot and CCBot were allowed like anyone else.&lt;/p&gt;

&lt;p&gt;The fix is named groups. The catch I didn't know: a crawler follows the most specific group that matches its name and ignores &lt;code&gt;*&lt;/code&gt; completely. So once a bot has its own group, anything you put under &lt;code&gt;*&lt;/code&gt; no longer applies to it.&lt;/p&gt;

&lt;p&gt;In the App Router you can generate robots.txt from &lt;code&gt;app/robots.ts&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;MetadataRoute&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;next&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;robots&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nx"&gt;MetadataRoute&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Robots&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;userAgent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;GPTBot&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;ClaudeBot&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;CCBot&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt; &lt;span class="na"&gt;disallow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;userAgent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;*&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;allow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;disallow&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/admin&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="na"&gt;sitemap&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://www.capital-cs.com/sitemap.xml&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The search bots aren't named, so they fall into the &lt;code&gt;*&lt;/code&gt; group and stay allowed.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. robots.txt is a request, not a lock
&lt;/h2&gt;

&lt;p&gt;Well-behaved bots read it. Nothing forces anyone to. If you want enforcement, the traffic has to pass through something that can block it, like Cloudflare's AI Crawl Control. And that only works when Cloudflare is proxying your traffic. DNS only isn't enough, because the requests never touch Cloudflare.&lt;/p&gt;

&lt;p&gt;We're on Vercel, which advises against putting another proxy in front of it, so for now robots.txt is what we're relying on.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Be careful with the one-click switch
&lt;/h2&gt;

&lt;p&gt;One reply mentioned another thread where Cloudflare's one-click "Block AI bots" toggle on a Free plan returned a 403 to PerplexityBot as well as the training bots. If being cited matters to you, block bots one at a time instead.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Test it from outside
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; https://www.capital-cs.com/robots.txt
curl &lt;span class="nt"&gt;-I&lt;/span&gt; &lt;span class="nt"&gt;-A&lt;/span&gt; &lt;span class="s2"&gt;"GPTBot"&lt;/span&gt; https://www.capital-cs.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first shows exactly what bots see. The second only tests user-agent matching, not the bot's real IP. If you're only using robots.txt like us, it still returns a 200 for GPTBot, which is a good reminder of point 3.&lt;/p&gt;

&lt;p&gt;Small site, small change, but I'd been assuming "allow all" was the safe default, and for AI training it wasn't what we wanted. If you've split these differently, I'd like to hear how.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>nextjs</category>
      <category>seo</category>
      <category>webdev</category>
    </item>
  </channel>
</rss>
