DEV Community

CiteScore
CiteScore

Posted on

Which AI crawlers should you allow in robots.txt? (GPTBot, PerplexityBot, Google-Extended and more)

AI assistants find and cite websites in two different ways. Some crawlers collect pages to train future models. Others fetch pages so an assistant can search the web and link to your site in an answer. Your robots.txt file can treat these differently, but only if you know what each user-agent does.

One caveat first: robots.txt is a request, not a lock. Reputable crawlers honor it. Anything else needs server-level or firewall blocking.

The major AI crawlers

User-agent Owner What it does
GPTBot OpenAI Collects content for training OpenAI models
OAI-SearchBot OpenAI Indexes pages for ChatGPT search
ChatGPT-User OpenAI Fetches a page when a ChatGPT user asks for it
ClaudeBot Anthropic Collects content for training Anthropic models
Claude-SearchBot Anthropic Indexes pages for Claude's search features
Claude-User Anthropic Fetches a page when a Claude user asks for it
PerplexityBot Perplexity Indexes pages for Perplexity answers and citations
Perplexity-User Perplexity Fetches a page when a Perplexity user asks for it
CCBot Common Crawl Builds a public web archive that many AI teams use for training
Google-Extended Google A control token, not a crawler. Governs use of content for Gemini training and grounding
Applebot-Extended Apple A control token. Governs use of Applebot-crawled content for Apple's AI training

Google-Extended and Applebot-Extended are not separate crawlers. They are opt-out tokens that apply to content already crawled by Googlebot or Applebot. Blocking them does not remove you from Google Search, and AI Overviews follow Googlebot and standard directives such as nosnippet.

A recommended robots.txt block

If you have no AI rules today, your site is already open to all of these bots. The block below opts out of training while staying visible in AI search:

# Training crawlers: opt out
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /

# Search and answer crawlers: allow
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
Allow: /
Enter fullscreen mode Exit fullscreen mode

Keep your existing rules for Googlebot and * in their own groups.

Training vs. answer crawlers: the trade-off

Training crawlers feed future model versions. Blocking them signals that you do not want your content used for future training. It does not remove anything a model has already learned, and it is hard to measure its effect on your traffic.

Answer crawlers decide whether you appear in live answers. If you block PerplexityBot or OAI-SearchBot, your pages are less likely to be cited when people ask a relevant question. For most sites that want visibility, that is the bigger cost.

The block above is a reasonable default. Adjust it if you have licensing concerns, paywalled content or private material. Vendors also note that user-triggered fetchers such as ChatGPT-User may not apply robots.txt rules, so treat those rules as best effort.

How to verify your rules

  1. Open https://yoursite.com/robots.txt in a browser. It should return 200 and plain text. A 404 means no rules, so everything is allowed. A 5xx error can stop crawlers entirely.
  2. Test the rules with Python's standard library:
from urllib.robotparser import RobotFileParser

rp = RobotFileParser("https://yoursite.com/robots.txt")
rp.read()
for bot in ["GPTBot", "OAI-SearchBot", "PerplexityBot"]:
    print(bot, rp.can_fetch(bot, "https://yoursite.com/blog/post"))
Enter fullscreen mode Exit fullscreen mode
  1. Search your server logs or analytics for these user-agent strings. Confirm that the bots you allowed reach your pages, and that the bots you blocked request only robots.txt.
  2. Check your CDN or WAF. A bot-protection rule that returns 403 blocks crawlers regardless of robots.txt, and it can block the bots you meant to allow.

Python's parser does not handle every wildcard the way Google does, so test important paths with more than one tool.

Common mistakes

  • Blocking Googlebot by accident. Google-Extended is not Googlebot. Blocking Googlebot removes you from Google Search.
  • Assuming a * rule covers AI bots. A crawler follows the most specific group that matches its name and ignores *. If GPTBot has its own group, that group must contain every rule you want for GPTBot.
  • Using robots.txt for privacy. The file is public and it is only a request. Protect private pages with authentication.
  • Mixing up Disallow and Allow. Disallow: / blocks the whole site, while an empty Disallow: allows everything. A misplaced slash can hide your entire site from search.
  • Ignoring the CDN. Bot-protection rules can override what your robots.txt says.
  • Expecting instant results. Crawlers cache robots.txt, and removal from AI answers takes time.

Test your robots.txt AI-crawler rules for free at https://citescore.vercel.app/free-check

Top comments (1)

Collapse
 
citedy profile image
Dmitry Sergeev •

i wonder if letting GPTBot in actually hurts my crawl budget or just adds extra load?