AI assistants find and cite websites in two different ways. Some crawlers collect pages to train future models. Others fetch pages so an assistant can search the web and link to your site in an answer. Your robots.txt file can treat these differently, but only if you know what each user-agent does.
One caveat first: robots.txt is a request, not a lock. Reputable crawlers honor it. Anything else needs server-level or firewall blocking.
The major AI crawlers
| User-agent | Owner | What it does |
|---|---|---|
| GPTBot | OpenAI | Collects content for training OpenAI models |
| OAI-SearchBot | OpenAI | Indexes pages for ChatGPT search |
| ChatGPT-User | OpenAI | Fetches a page when a ChatGPT user asks for it |
| ClaudeBot | Anthropic | Collects content for training Anthropic models |
| Claude-SearchBot | Anthropic | Indexes pages for Claude's search features |
| Claude-User | Anthropic | Fetches a page when a Claude user asks for it |
| PerplexityBot | Perplexity | Indexes pages for Perplexity answers and citations |
| Perplexity-User | Perplexity | Fetches a page when a Perplexity user asks for it |
| CCBot | Common Crawl | Builds a public web archive that many AI teams use for training |
| Google-Extended | A control token, not a crawler. Governs use of content for Gemini training and grounding | |
| Applebot-Extended | Apple | A control token. Governs use of Applebot-crawled content for Apple's AI training |
Google-Extended and Applebot-Extended are not separate crawlers. They are opt-out tokens that apply to content already crawled by Googlebot or Applebot. Blocking them does not remove you from Google Search, and AI Overviews follow Googlebot and standard directives such as nosnippet.
A recommended robots.txt block
If you have no AI rules today, your site is already open to all of these bots. The block below opts out of training while staying visible in AI search:
# Training crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /
# Search and answer crawlers: allow
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Keep your existing rules for Googlebot and * in their own groups.
Training vs. answer crawlers: the trade-off
Training crawlers feed future model versions. Blocking them signals that you do not want your content used for future training. It does not remove anything a model has already learned, and it is hard to measure its effect on your traffic.
Answer crawlers decide whether you appear in live answers. If you block PerplexityBot or OAI-SearchBot, your pages are less likely to be cited when people ask a relevant question. For most sites that want visibility, that is the bigger cost.
The block above is a reasonable default. Adjust it if you have licensing concerns, paywalled content or private material. Vendors also note that user-triggered fetchers such as ChatGPT-User may not apply robots.txt rules, so treat those rules as best effort.
How to verify your rules
- Open
https://yoursite.com/robots.txtin a browser. It should return 200 and plain text. A 404 means no rules, so everything is allowed. A 5xx error can stop crawlers entirely. - Test the rules with Python's standard library:
from urllib.robotparser import RobotFileParser
rp = RobotFileParser("https://yoursite.com/robots.txt")
rp.read()
for bot in ["GPTBot", "OAI-SearchBot", "PerplexityBot"]:
print(bot, rp.can_fetch(bot, "https://yoursite.com/blog/post"))
- Search your server logs or analytics for these user-agent strings. Confirm that the bots you allowed reach your pages, and that the bots you blocked request only
robots.txt. - Check your CDN or WAF. A bot-protection rule that returns 403 blocks crawlers regardless of robots.txt, and it can block the bots you meant to allow.
Python's parser does not handle every wildcard the way Google does, so test important paths with more than one tool.
Common mistakes
- Blocking Googlebot by accident. Google-Extended is not Googlebot. Blocking Googlebot removes you from Google Search.
-
Assuming a
*rule covers AI bots. A crawler follows the most specific group that matches its name and ignores*. If GPTBot has its own group, that group must contain every rule you want for GPTBot. - Using robots.txt for privacy. The file is public and it is only a request. Protect private pages with authentication.
-
Mixing up Disallow and Allow.
Disallow: /blocks the whole site, while an emptyDisallow:allows everything. A misplaced slash can hide your entire site from search. - Ignoring the CDN. Bot-protection rules can override what your robots.txt says.
- Expecting instant results. Crawlers cache robots.txt, and removal from AI answers takes time.
Test your robots.txt AI-crawler rules for free at https://citescore.vercel.app/free-check
Top comments (1)
i wonder if letting GPTBot in actually hurts my crawl budget or just adds extra load?