A few posts back I measured 336 of 422 sites behind Cloudflare's managed bot rules refusing GPTBot and ClaudeBot at the edge with nothing in robots.txt to explain it (post here). Then I widened the test to 8,000 mid-ranked sites and found 178 that allow OAI-SearchBot or PerplexityBot in robots.txt and refuse the named request at the server anyway (post here).
Both of those were one-off samples. I turned the test into a running scan instead, and I am keeping every confirmed row in a public, continuously updated list.
Where it stands today
As of 2026-09-30 01:37 UTC, the list holds 750 sites. Each one was requested as OAI-SearchBot, as PerplexityBot, as Googlebot, as Bingbot (full user-agent strings) and as an ordinary browser, from the same place within a few minutes. A site is only listed when its robots.txt allows the AI crawler, the request naming that crawler was refused, and the other four requests were served.
CDN split, and this is the part I did not expect going in:
| CDN | Sites | Share |
|---|---|---|
| Cloudflare | 378 | 50.4% |
| Other | 362 | 48.3% |
| Vercel | 10 | 1.3% |
Just over half of the confirmed refusals sit behind Cloudflare. The rest run on every other CDN and host I have sampled. My earlier posts here were about Cloudflare specifically because that is where I started measuring, and it reads easily as a Cloudflare problem. On this sample it is not one CDN's setting; it is a pattern that shows up almost as often off Cloudflare as on it.
By Tranco rank band: 250k-500k 327, 100k-250k 174, 500k-1m 139, 50k-100k 79, 20k-50k 31. Fairly flat across the mid tail, no band standing out.
By country guess (top five, ccTLD first then hosting network registry country): Germany 47, United States 32, France 24, Russia 24, United Kingdom 22. 393 of the 750 have no usable guess (generic domain on a global host).
75 of the 750 have a contact page my prescreen could find on a static render. 16 have already had a note from me about the specific refusal; the other 734 have not.
How the test works, and its limits
- Same-session requests from one ordinary IP address. A firewall rule that checks the real crawler's published IP ranges would let it through even where this test sees a refusal, so a row is evidence of a user-agent rule, confirmed properly only in the site's own logs.
- Some owners block these crawlers on purpose. A
robots.txtthat says yes while the server says no is usually an accident, not always. - CMS comes from the homepage's generator tag and a few fingerprints; 507 of 750 came back unmatched. Of the rest, WordPress is the largest group at 176.
- A site drops off the list once its last test is more than 14 days old, until it is tested again. The scan runs in batches several times a day against the Tranco list, ranks 20,000 to 1,000,000. Adult, gambling and piracy hostnames are filtered out.
- No individual site from this mid-tail sample is named here or in any file I publish; the top-5,000 census from my earlier posts is the one dataset with named sites, and that one is already public under CC BY.
Where the list lives
The running list, the free 25-row sample by email, and the filtered CSV/API version are at /leads. Checking a single site stays free at the checker, and the 5,000-site census stays open data.
Top comments (0)