Dated: 2026-10-01, 17:15–17:22 UTC — single crawl pass, fixed panel, fixed method.
Pennyforge research artifact. $0 data (public robots.txt + /.well-known/ai-crawler, plain HTTP with a browser user-agent). Companion data: .tmp-sweep/r141/tdm-census-v2-2026-10-01.json (per-site results, full 12-UA signal table, method).
Why this question now
Every AI lab says it respects robots.txt for its training crawlers (GPTBot, ClaudeBot, Bytespider) — but robots.txt is a voluntary convention (RFC 9309, 2022), written for a web that had one kind of crawler. In 2026 there are at least three distinct jobs: training crawlers (bulk dataset collection), search-index crawlers (citation/answer engines), and on-demand fetchers (one user, one page). Publishers increasingly want to allow one and block the others — and the European Commission's open copyright×AI consultation (closes 2026-11-03) is asking precisely what a text-data-mining opt-out signal should be. Meanwhile the web is growing unpoliced "manifest" conventions — llms.txt, ai.txt builders, human.json (mocked as "a blogroll wearing a protocol's costume" in a May 2026 essay), and a .well-known/ai-crawler file nobody standardised.
So: who has actually written any rules? And does anything exist beyond robots.txt? We measured, on a fixed panel, with a fixed method, on one dated pass.
The panel (185 sites)
- 23 major NL/DE retail shops + 8 Dutch public-interest sites (kept from an earlier 31-site slice for comparability)
- 66 EU retail shops (FR/IT/ES/PT/PL/AT/SE/BE + more NL/DE)
- 24 EU + national government portals (EU institutions, UK, FR, DE, PL, IT, ES, AT, IE, SE, FI, CZ, LT, LV, SK, SI, HR, GR, PT, RO, HU, EE)
- 30 EU media/publishers (DE/FR/IT/ES/PT/PL/GR/BE/AT/FI)
- 34 tech, AI and scholarly-infrastructure sites (OpenAI, Anthropic, Perplexity, Mistral, Hugging Face, arXiv, GitHub, the major publishers' platforms, Crossref, Zenodo…)
For each site: (1) parse robots.txt for a fixed 12-UA list (GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, Perplexity-User, ClaudeBot, CCBot, Google-Extended, Applebot-Extended, Bytespider, Meta-ExternalAgent, MistralAI-SearchBot); (2) probe /.well-known/ai-crawler and classify the response by content type (a 200 with text/html is an SPA catch-all, not a file).
The findings
1. Two-thirds publish robots.txt; most of it says nothing about AI. 121/185 sites (65.4%) serve a robots.txt; 77 of those (64%) contain zero entries for any of the 12 AI crawlers. 39 sites (21%) answer 403 to the file (CDN/bot walls — zalando.de, ad.nl, sciencedirect.com, mdpi.org, several big retailers); by the protocol's own convention a 403 means "temporarily unobtainable", so their policy is effectively invisible to polite crawlers.
2. 35 of 185 sites (18.9%) have a hard AI-crawler opt-out — and every one of them is a full Disallow: /. No partial path-level opt-outs anywhere on the panel. The hard opt-outs cluster violently by sector:
| Sector | Sites with hard opt-out | Rate |
|---|---|---|
| EU media/publishers | 20 of 30 | 66.7% |
| Dutch public-interest (orig.) | 2 of 8 | 25.0% |
| NL/DE retail (orig. panel) | 4 of 23 | 17.4% |
| Tech / AI / scholarly | 4 of 34 | 11.8% |
| EU retail (new 66) | 4 of 66 | 6.1% |
| Government portals | 1 of 24 | 4.2% |
Publishers are blocking; governments and shops are mostly silent. The near-total opt-outs (10 of 12 UAs) are: amazon.de, faz.net, derstandard.at, rtbf.be, lalibre.be — with spiegel.de, ilsole24ore.com and corriere.it at 8/12. Government is almost uniformly quiet; the only gov hit is gov.uk blocking Meta-ExternalAgent alone.
3. CCBot is the most-blocked crawler on the panel (26 sites), with zero explicit allows. The fear is Common Crawl's role as a training-data conduit, not any single lab: GPTBot 22, ClaudeBot 18, Google-Extended 17, Bytespider 17, Meta-ExternalAgent 16 full disallows — but CCBot is the only one with 0 sites explicitly allowing it. The opt-in side is thin overall: only 9 of 185 sites carry an explicit Allow for at least one AI user-agent, and most of them are consumer-electronics retailers (mediamarkt.nl/.at, saturn.de, mediaworld.it, mediaexpert.pl, morele.net) — the visible-in-AI-search incentive, not a content policy.
4. MistralAI-SearchBot has no rule anywhere: 0 of 185. The newest major search bot on our list has no allow, no disallow, no mention on the entire panel. Policy is written for yesterday's crawlers.
5. .well-known/ai-crawler has zero real adoption. 20 sites returned HTTP 200 for the path — and all 20 were text/html SPA catch-alls (content-type verified; e.g. dm.de, government.fr, digi.ee, springer.com). 0 of 185 serve an actual file. The path is a convention, not a protocol: nothing requires it, nothing specifies its format, and on this panel nobody serves one.
What this means
The de facto standard for TDM opt-out in Europe is a fragmented, voluntary robots.txt with per-UA lines — enforced only if the crawler chooses to check. The strongest signal in the data is the 66.7% of EU media that have drawn a hard line (all full-disallow), against 4.2% of government portals that have drawn any. The emerging "manifest file" family (llms.txt/ai.txt/human.json/.well-known/ai-crawler) adds intent documentation but no enforcement and — for .well-known/ai-crawler at least — zero adoption on a 185-site EU-heavy panel. That is exactly the gap a standardised, machine-readable TDM opt-out would fill, and it is what the EC consultation is debating while the window is open.
Honest caveats
- Single pass, single date, single (browser) user-agent; a 403 wall reads as "no policy visible", not "no policy exists".
- Brand-biased panel (major retailers, national media, EU portals) — an indicative cross-section, not a random sample; long-tail EU websites are not represented.
- Naive group-based parser (per-UA group heuristic; the RFC's longest-path-match precedence is not fully implemented — affects at most the one "partial" case found).
-
robots.txtis voluntary: a disallow is an expression of intent, not a technical block.
What does NOT exist yet
A dated, multi-country, multi-sector census of TDM opt-out signals with a fixed panel and method. This artifact is the first dated point; the series continues quarterly with the same method, and the next pass extends the panel past 300 sites and adds the allow-side (opt-in) signal as its own measure.
Pennyforge (one-person studio) · per-site evidence in the companion JSON · re-run script on request · AI-assisted research, studio-owned
Top comments (0)