DEV Community

kacey noms
kacey noms

Posted on

Keeping AI search bots in and AI training bots out on a small Next.js site

I look after the website for Capital Complete Solutions, a property inventory company in Birmingham where I also work as an inventory clerk. It's a Next.js app on Vercel with Payload CMS, about 30 pages.

This week I wanted to sort out one thing: let AI search tools find and cite our pages, but stop crawlers that only collect training data. I asked on the Cloudflare community and got more useful answers than I expected. Here's what I learned.

1. AI companies run separate bots for training and for search

The big providers split their crawlers by job, and each one is its own robots.txt token:

Provider Training Search index Fetch for a user
OpenAI GPTBot OAI-SearchBot ChatGPT-User
Anthropic ClaudeBot Claude-SearchBot Claude-User
Perplexity PerplexityBot Perplexity-User
Common Crawl CCBot

If you also want to opt out of Gemini training, Google-Extended is the token for that, and it doesn't affect Google Search. Names do change, so check each provider's own docs before copying this.

2. My robots.txt was letting everything in

Someone on the thread looked at our live robots.txt and pointed out that everything was falling through the User-agent: * group, so GPTBot and CCBot were allowed like anyone else.

The fix is named groups. The catch I didn't know: a crawler follows the most specific group that matches its name and ignores * completely. So once a bot has its own group, anything you put under * no longer applies to it.

In the App Router you can generate robots.txt from app/robots.ts:

import type { MetadataRoute } from 'next'

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      { userAgent: ['GPTBot', 'ClaudeBot', 'CCBot'], disallow: '/' },
      { userAgent: '*', allow: '/', disallow: ['/admin'] },
    ],
    sitemap: 'https://www.capital-cs.com/sitemap.xml',
  }
}
Enter fullscreen mode Exit fullscreen mode

The search bots aren't named, so they fall into the * group and stay allowed.

3. robots.txt is a request, not a lock

Well-behaved bots read it. Nothing forces anyone to. If you want enforcement, the traffic has to pass through something that can block it, like Cloudflare's AI Crawl Control. And that only works when Cloudflare is proxying your traffic. DNS only isn't enough, because the requests never touch Cloudflare.

We're on Vercel, which advises against putting another proxy in front of it, so for now robots.txt is what we're relying on.

4. Be careful with the one-click switch

One reply mentioned another thread where Cloudflare's one-click "Block AI bots" toggle on a Free plan returned a 403 to PerplexityBot as well as the training bots. If being cited matters to you, block bots one at a time instead.

5. Test it from outside

curl -s https://www.capital-cs.com/robots.txt
curl -I -A "GPTBot" https://www.capital-cs.com
Enter fullscreen mode Exit fullscreen mode

The first shows exactly what bots see. The second only tests user-agent matching, not the bot's real IP. If you're only using robots.txt like us, it still returns a 200 for GPTBot, which is a good reminder of point 3.

Small site, small change, but I'd been assuming "allow all" was the safe default, and for AI training it wasn't what we wanted. If you've split these differently, I'd like to hear how.

Top comments (0)