Three weeks ago I wrote about the difference between AI training crawlers and AI retrieval
agents:
GPTBot collects data to train models, but it's ChatGPT-User and OAI-SearchBot that fetch
your page live when ChatGPT answers a question — and only those can cite and link you.
Block them and you don't exist in AI answers, no matter how well you rank on Google.
A commenter said they hadn't realized these were different bots. That made me wonder how many
professional publishers haven't either. So I measured it.
Method
On August 1, 2026 I ran an automated check against the 50 biggest English-language news and
tech publishers. For each site:
- Fetch
robots.txtand evaluate it for 15 AI crawlers — 7 retrieval agents (OAI-SearchBot,ChatGPT-User,Claude-SearchBot,Claude-User,PerplexityBot,Perplexity-User,Amazonbot) and 8 training crawlers (GPTBot,ClaudeBot,CCBot,Google-Extended,Applebot-Extended,Bytespider,meta-externalagent,anthropic-ai) — using Google's documented robots semantics (most specific group wins, longest rule wins,Allowwins ties), for path/. - Fetch the homepage without executing JavaScript — because that's what retrieval agents do — and check whether the served HTML contains any readable text at all.
A site counts as citable only if both hold: no retrieval agent blocked, and actual text in
the served HTML. The tool is an open audit Actor I
built; four sites (NYT, Guardian, FT, Daily
Mail) bot-wall their homepage, so for those only the robots.txt half was evaluated — which
already settles their verdict.
Results
38 of 50 are not citable.
9 sites block all seven retrieval agents: CNN, NBC News, USA Today, HuffPost, The
Telegraph, Daily Mail, CNET, ZDNet, Mashable. Whatever these newsrooms publish, no AI answer
can ever quote it or link to it.
25 sites block in the wrong direction. They allow at least one training crawler while
blocking retrieval agents — meaning their content may train models, but the one thing that
sends readers back (a citation with a link) is off. The Verge, Wired, Ars Technica, The
Atlantic, Vox, The Guardian, The Washington Post and WSJ are all in this group. I doubt a
single one chose that trade on purpose.
5 sites fail by exactly one agent. ABC News and TechCrunch block only ChatGPT-User;
Axios, Tom's Hardware and VentureBeat block only Amazonbot. One robots.txt line away from
citable.
3 sites have the door open and the room empty. NPR, Politico and The Information allow
all seven retrieval agents — and serve a homepage whose HTML contains essentially no readable
text without JavaScript. Retrieval agents don't run JavaScript. A browser shows a normal page;
an AI agent gets nothing to quote. This failure is invisible in every browser-based audit.
Who gets blocked most tells its own story. PerplexityBot is blocked by 30 of 50 sites,
Amazonbot by 28, Anthropic's two retrieval agents by 24–25 — but OpenAI's OAI-SearchBot by
only 14. That's the licensing-deal era in one number: publishers with OpenAI deals let OpenAI's
citation bot in and block everyone else's.
The 12 citable sites: Fox News, CBS News, Business Insider, LA Times, Time, Slate, The
Independent, The Daily Beast, Semafor, Engadget, Gizmodo, PCMag.
Full table
Retrieval = citation agents allowed (of 7). Training = training crawlers allowed (of 8).
* = robots.txt only (homepage bot-walled).
| Site | Retrieval | Training | Citable | Why not |
|---|---|---|---|---|
| cnet.com | 0/7 | 1/8 | ❌ | robots.txt |
| cnn.com | 0/7 | 1/8 | ❌ | robots.txt |
| dailymail.co.uk * | 0/7 | 0/8 | ❌ | robots.txt |
| huffpost.com | 0/7 | 0/8 | ❌ | robots.txt |
| mashable.com | 0/7 | 1/8 | ❌ | robots.txt |
| nbcnews.com | 0/7 | 0/8 | ❌ | robots.txt |
| telegraph.co.uk | 0/7 | 0/8 | ❌ | robots.txt |
| usatoday.com | 0/7 | 0/8 | ❌ | robots.txt |
| zdnet.com | 0/7 | 1/8 | ❌ | robots.txt |
| bloomberg.com | 1/7 | 0/8 | ❌ | robots.txt |
| economist.com | 1/7 | 1/8 | ❌ | robots.txt |
| nytimes.com * | 1/7 | 0/8 | ❌ | robots.txt |
| arstechnica.com | 2/7 | 1/8 | ❌ | robots.txt |
| bbc.com | 2/7 | 0/8 | ❌ | robots.txt |
| cnbc.com | 2/7 | 0/8 | ❌ | robots.txt |
| marketwatch.com | 2/7 | 1/8 | ❌ | robots.txt |
| newyorker.com | 2/7 | 2/8 | ❌ | robots.txt |
| reuters.com | 2/7 | 0/8 | ❌ | robots.txt |
| theatlantic.com | 2/7 | 1/8 | ❌ | robots.txt |
| theverge.com | 2/7 | 1/8 | ❌ | robots.txt |
| vox.com | 2/7 | 1/8 | ❌ | robots.txt |
| wired.com | 2/7 | 2/8 | ❌ | robots.txt |
| wsj.com | 2/7 | 1/8 | ❌ | robots.txt |
| apnews.com | 3/7 | 3/8 | ❌ | robots.txt |
| nypost.com | 3/7 | 1/8 | ❌ | robots.txt |
| theguardian.com * | 3/7 | 2/8 | ❌ | robots.txt |
| newsweek.com | 4/7 | 2/8 | ❌ | robots.txt |
| forbes.com | 5/7 | 1/8 | ❌ | robots.txt |
| ft.com * | 5/7 | 1/8 | ❌ | robots.txt |
| washingtonpost.com | 5/7 | 2/8 | ❌ | robots.txt |
| abcnews.go.com | 6/7 | 3/8 | ❌ | blocks only ChatGPT-User |
| axios.com | 6/7 | 6/8 | ❌ | blocks only Amazonbot |
| techcrunch.com | 6/7 | 1/8 | ❌ | blocks only ChatGPT-User |
| tomshardware.com | 6/7 | 6/8 | ❌ | blocks only Amazonbot |
| venturebeat.com | 6/7 | 5/8 | ❌ | blocks only Amazonbot |
| npr.org | 7/7 | 8/8 | ❌ | no text without JavaScript |
| politico.com | 7/7 | 8/8 | ❌ | no text without JavaScript |
| theinformation.com | 7/7 | 8/8 | ❌ | no text without JavaScript |
| businessinsider.com | 7/7 | 4/8 | ✅ | |
| cbsnews.com | 7/7 | 7/8 | ✅ | |
| engadget.com | 7/7 | 8/8 | ✅ | |
| foxnews.com | 7/7 | 8/8 | ✅ | |
| gizmodo.com | 7/7 | 5/8 | ✅ | |
| independent.co.uk | 7/7 | 8/8 | ✅ | |
| latimes.com | 7/7 | 5/8 | ✅ | |
| pcmag.com | 7/7 | 8/8 | ✅ | |
| semafor.com | 7/7 | 7/8 | ✅ | |
| slate.com | 7/7 | 8/8 | ✅ | |
| thedailybeast.com | 7/7 | 8/8 | ✅ | |
| time.com | 7/7 | 8/8 | ✅ |
Caveats, honestly
- Some of this is deliberate. The NYT is suing OpenAI; several publishers are negotiating licenses. For them, blocking is leverage, not an accident. But "we allow model training and forbid citations" (25 sites) is a hard position to defend as strategy — and the sites that copied 2023's "block the AI bots" lists inherited these rules with none of the leverage.
- This is a snapshot (Aug 1, 2026) of path
/and the homepage. Sections may differ. - "Citable" means technically reachable for citation, not "gets cited".
Check your own site
Two failure modes, both invisible in a browser and in classic SEO tools:
-
robots.txt: look for your citation agents (
OAI-SearchBot,ChatGPT-User,Claude-SearchBot,Claude-User,PerplexityBot,Perplexity-User,Amazonbot). Blanket AI-blocklists and one-click CDN blockers usually hit these too. The 5-minute manual check is in my previous post. -
Server-rendered text:
curlyour page and look for your content in the HTML. If it only appears after JavaScript runs, retrieval agents see an empty shell.
For a single site you honestly don't need a tool — the manual check takes five minutes.
Where it stops being trivial:
- You manage a portfolio. An agency with 50 client sites doesn't read 50 robots.txt files quarterly. This entire study — 50 sites, both checks — ran automated in about 40 minutes.
- The answer changes behind your back. Robots.txt files drift: CDN one-click "block AI bots" switches, CMS updates, replatforming. Cloudflare's managed robots.txt added rules to a site of mine that I never wrote. A check that was green in March can be red in June with nobody having touched anything.
That's what I built SEO Health Auditor for:
it runs both checks (plus regular technical SEO) across any list of sites, from $0.05 per
page, and on a schedule it diffs against the previous run — so you learn about the drift
before your traffic does.
Raw data for all 50 sites available on request — happy to share the JSON.
Top comments (0)