Every big AI company runs a crawler. OpenAI has GPTBot, Anthropic has ClaudeBot, Google runs Google-Extended, Perplexity has PerplexityBot. A website can wave them through or turn them away with a few lines in a file called robots.txt. I wanted to know how many actually turn them away, so I pulled the data instead of guessing.
I took the top 1,000 sites from the Tranco ranking, fetched each one's robots.txt on October 2, 2026, and read what it said about the main AI crawlers. Here is what came back.
What "the top 1,000" actually is
One caveat up front, because it turned out to be interesting on its own. Of the 1,000 domains, only 485 served a robots.txt I could read. The rest are not websites in the normal sense. Around 320 were CDN and DNS infrastructure like akamaiedge.net and cloudfront.net. Another 68 were API endpoints such as googleapis.com with nothing to crawl. 63 returned a login page or a soft error instead of a real file, and 64 blocked the request outright. So every number below is out of the 485 sites that have a real, readable robots.txt.
About a quarter block at least one AI crawler
114 of the 485 sites, or 23.5%, block at least one of the eight major AI crawlers by name. Only 14 block all eight. So blocking happens, but blanket blocking is still rare. Most sites that block pick their targets.
Here is how often each crawler gets a full block:
| Crawler | Share of sites that block it |
|---|---|
| Common Crawl (CCBot) | 17.3% |
| Bytespider (ByteDance) | 16.3% |
| GPTBot (OpenAI) | 15.9% |
| ClaudeBot (Anthropic) | 15.1% |
| Google-Extended | 13.8% |
| Meta (meta-externalagent) | 12.6% |
| PerplexityBot | 12.0% |
| Applebot-Extended | 10.5% |
| Amazonbot | 9.1% |
Common Crawl sits at the top, and that makes sense. Its public archive feeds a lot of model training, so blocking CCBot is a way to block many AI companies at once. Bytespider, ByteDance's crawler, comes next and has a rough reputation for ignoring limits, which is probably why so many admins name it directly.
Sites also tend to treat AI crawlers as one group. Of the sites that block GPTBot, 79% also block ClaudeBot. When someone decides to close the door, they usually close it on everyone at once.
News sites are the real outlier
This is the sharp line in the data. Among the news and media sites in the set, 80% block at least one AI crawler. For everything else, it is 20%. News is four times more likely to say no.
And they are not shy about it. The New York Times, BBC, Bloomberg, CNBC, USA Today, NBC News and the Daily Mail block all eight crawlers I checked. CNN, Forbes, NPR and the Financial Times block seven.
The exceptions are the fun part. The Wall Street Journal blocks none of them by name. Neither does Fox News, Time, the Telegraph or the Independent. Reuters blocks one. Same industry, opposite bet. A lot of that comes down to licensing. A publisher that has signed a content deal with an AI company has a reason to leave the door open.
Marketplaces and platforms mostly stay open
Step outside news and the blocking drops off fast. Amazon blocks seven of the eight crawlers, but Etsy, Shopify and Booking.com block none of them by name. Wikipedia, Pinterest, IMDb, Substack and WordPress.com are wide open. Apple and Microsoft's main domains do not block them either.
So this is less about the whole web closing to AI and more about one corner of it. Newsrooms are pulling up the drawbridge. Most of the rest has not bothered yet.
64 sites would not even hand over the file
One more thing, and it matters if you pull data for a living. 64 of the 1,000 returned a 403 error to a plain request for robots.txt, a public file that exists specifically to be read. nih.gov, ScienceDirect and Stack Overflow were among them. That is a different kind of blocking, aimed at any automated request rather than any specific bot, and it is getting more common.
Why this is worth watching
If your business depends on being found or on pulling public data, this is a trend to keep an eye on. Blocking a training crawler like GPTBot is not the same as blocking the bot that fetches a page to answer a live question, and most sites that block are aiming at training, not answers. But the line keeps moving. The data a competitor publishes openly today may be closed to a crawler tomorrow, which is the whole reason a reliable way to collect and store the data you need is worth having before you need it.
I pulled all of this with a short script: read the Tranco list, fetch each robots.txt, parse the user-agent rules, count the blocks. If you want the same kind of web-scale snapshot for a question of your own, whether that is prices, listings, or who publishes what, that is the sort of thing we build at DataScrape Solutions.
Method, for anyone who wants to check it: Tranco top 1,000, fetched October 2, 2026. 485 of the 1,000 served a readable robots.txt; the rest were infrastructure domains, returned non-robots responses, or refused the request. "Block" means a full Disallow of the site root for that crawler's named user-agent.
Top comments (0)