A lot of sites added this to robots.txt some time in the last two years:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: *
Disallow: /api/
That's a reasonable choice if you don't want your content used to train models. But many teams also have a CDN rule, a WAF setting or a blanket "block AI bots" switch that goes further. Without anyone deciding it, the site also disappears from AI search: ChatGPT search, Perplexity, Claude's web search.
The AI companies now run separate bots for training, for search indexing and for fetching a page when a user asks. You can treat each one differently. Most sites haven't, because the names aren't obvious.
Which bot does what
From each company's own documentation:
| Company | Bot | What it's for | If you block it |
|---|---|---|---|
| OpenAI | GPTBot |
Training data | Content shouldn't be used to train OpenAI's models |
| OpenAI | OAI-SearchBot |
ChatGPT search | Not shown in ChatGPT search answers, though navigational links can still appear |
| OpenAI | ChatGPT-User |
Fetches a page when a user asks | OpenAI says that because these are user-initiated, robots.txt rules may not apply |
| Anthropic | ClaudeBot |
Training data | Future content excluded from training |
| Anthropic | Claude-SearchBot |
Search indexing | May reduce your visibility in Claude's search results |
| Anthropic | Claude-User |
Fetches a page when a user asks | Claude can't retrieve your page for that user |
Google-Extended |
Gemini training and grounding | Google says it does not affect inclusion or ranking in Google Search | |
| Perplexity | PerplexityBot |
Perplexity's search index | Not surfaced in Perplexity search results |
Two things stand out.
Google's AI Overviews aren't on the list. They're built from the normal Search index, which Googlebot crawls. Google-Extended controls Gemini training and grounding, not whether you appear in AI Overviews. Blocking Googlebot to stay out of AI Overviews takes you out of Search entirely.
Training and search are separate switches. OpenAI's documentation says it plainly: a site can allow OAI-SearchBot to appear in search results while disallowing GPTBot to keep its content out of training.
A robots.txt that keeps you in AI search but out of training
If that's the trade-off you want:
# Search and user-requested fetches: allowed
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
Disallow: /api/
Disallow: /admin/
Allow: /
# Model training: blocked
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
Disallow: /
# Everyone else
User-agent: *
Disallow: /api/
Disallow: /admin/
Sitemap: https://example.com/sitemap.xml
One easy mistake: a crawler follows only the most specific group that matches its name and ignores User-agent: * completely. So any path you block for everyone (/api/, /admin/) has to be repeated in every named group, as above. Otherwise the named bots get access to it.
Whether to block training is a business decision, not a technical one. Some companies want their documentation in training data so that models describe their product correctly. The point is to make the choice on purpose.
Test it with a script, not by reading it
Reading robots.txt by eye is how the mistake above survives code review. Python's standard library can parse it:
# check_robots.py
from urllib.robotparser import RobotFileParser
SITE = "https://example.com"
BOTS = [
"GPTBot", "OAI-SearchBot", "ChatGPT-User",
"ClaudeBot", "Claude-SearchBot", "Claude-User",
"Google-Extended", "PerplexityBot", "Googlebot", "Bingbot",
]
PATHS = ["/", "/blog/", "/pricing", "/api/health", "/admin/"]
rp = RobotFileParser(f"{SITE}/robots.txt")
rp.read()
width = max(map(len, BOTS))
print(" " * width, *[p.ljust(12) for p in PATHS])
for bot in BOTS:
row = ["allow".ljust(12) if rp.can_fetch(bot, SITE + p) else "BLOCK".ljust(12) for p in PATHS]
print(bot.ljust(width), *row)
Here's what it printed for our own site, redliolabs.com:
/ /insights/ /services/ /contact
GPTBot allow allow allow allow
OAI-SearchBot allow allow allow allow
ChatGPT-User allow allow allow allow
ClaudeBot allow allow allow allow
Claude-SearchBot allow allow allow allow
Claude-User allow allow allow allow
Google-Extended allow allow allow allow
PerplexityBot allow allow allow allow
Googlebot allow allow allow allow
Bingbot allow allow allow allow
Every bot can reach every page, training crawlers included, because our file is a single User-agent: * / Allow: /. Seeing it as a grid turns that into a decision you can look at, not a default nobody remembers making. If your grid has a BLOCK you didn't expect, start there.
Put it in CI, with the expected grid committed alongside it, so a change to robots.txt that flips a cell fails the build.
(urllib.robotparser follows the original robots.txt rules. It applies the first rule that matches, while Google applies the most specific one, which is why the Disallow lines above come before Allow: /: the file then reads the same both ways. It also doesn't support * and $ wildcards in paths, so keep the file simple or test wildcard rules separately.)
Then check the layer robots.txt can't see
robots.txt is a request. Your CDN, WAF or bot-management settings are what actually answer. A request can be blocked there before robots.txt is ever consulted. Check two things.
1. What your server returns to each user agent:
for ua in "OAI-SearchBot/1.4" "Claude-SearchBot" "PerplexityBot" "GPTBot/1.4"; do
code=$(curl -s -o /dev/null -w "%{http_code}" -A "Mozilla/5.0 (compatible; $ua)" https://example.com/blog/)
echo "$ua -> $code"
done
A 403 or a challenge page means something other than robots.txt is deciding. This only tests rules based on the user-agent header. Some bot-management products also verify the bot's IP ranges, which a spoofed curl can't reproduce. OpenAI publishes its crawlers' IP ranges if you need to allow-list them.
2. Whether the bots actually reach you. Your access logs are the ground truth:
grep -Eo "OAI-SearchBot|ChatGPT-User|GPTBot|Claude-SearchBot|Claude-User|ClaudeBot|PerplexityBot" access.log \
| sort | uniq -c | sort -rn
If a search bot you allowed never appears, or only ever gets 403s, the block is upstream of robots.txt.
What this doesn't do
Allowing the search bots makes you eligible to be cited. It doesn't make anyone cite you. That depends on whether your pages answer the questions people ask, with specifics a summary can't replace. That's a content problem, not a configuration one. But it's worth fixing the configuration first, because no amount of good content helps if the crawler gets a 403.
If you're trying to work out which agencies or tools are worth paying for here, we wrote a guide to choosing a GEO agency, including the claims to be sceptical of.
Over to you: run the script on your own site and post the grid in the comments, especially if a cell surprised you. We're curious how many sites block AI search without meaning to.
From the team at Redlio Labs, where we build software and run SEO and AI search for B2B companies.
Top comments (2)
We need to generate a short comment, 1-2 sentences, casual, start with lowercase, specific reaction or question about this video. No promo, no URLs. Must follow style guidelines. Provide comment only. Something like: "i tried the quick test on my dev site and it turned out my robots.txt was actually blocking GPT, never noticed that before". Or ask: "does this also affect other LLMs like Claude or only ChatGPT?" Must be short. Check for any disallowed punctuation: no em-dash
Good question buried in there, so for anyone reading: it isn't only ChatGPT. Anthropic runs ClaudeBot for training, Claude-SearchBot for search and Claude-User for user requests, and Perplexity has PerplexityBot. The table in the post lists each one, and the script checks them all at once against your own robots.txt.