Originally published on The AI Prism
Every time a page behind Cloudflare loads, a silent verdict is rendered: human, search bot, AI crawler, or agent. That verdict now carries real money, real access, and real consequences — because more than 20% of the web’s domains sit behind Cloudflare’s network, and the company just rewrote the rules for who gets in.
On July 1, 2026, Cloudflare declared its second “Content Independence Day” and gave every customer — including the Free tier — the power to manage AI traffic by three use cases: Search, Agent, and Training. Then it set new defaults that take effect September 15, 2026: on pages that display ads, Training and Agent bots get blocked by default. Search stays allowed. And because Google uses the same crawler for search indexing and Gemini training, a customer who blocks Training will also block Googlebot.
Here at The AI Prism, we’ve been tracking this story since the first Content Independence Day in July 2025, and the shift is bigger than a dashboard toggle. The AI traffic wars have stopped being about content. They’re now about infrastructure — who decides which models run where, who gets to crawl, and who pays for the privilege.
Cloudflare is referee, toll collector, and rival in the same match. It blocks AI crawlers at the front door while selling AI inference at the back. That’s a strange position for any company to hold, and the tech community has noticed. Hacker News lit up with 157 comments on the announcement, and the most common reaction wasn’t praise. It was suspicion.
The Old Deal Is Dead: Crawl, Refer, Repeat
For almost 30 years, the web ran on a handshake deal. Google would copy your content for search, and in return you got referral traffic you could monetize with ads or subscriptions. Cloudflare CEO Matthew Prince described it bluntly on the first Content Independence Day: “The web is being stripmined by AI crawlers with content creators seeing almost no traffic and therefore almost no value.”
The numbers back him up. Researchers found 75% of mobile queries are now answered without leaving Google. Cloudflare’s own crawl-to-refer ratio analysis showed that getting traffic from OpenAI is 750 times harder than it was from the Google of old — and from Anthropic, it’s 30,000 times harder. Content creators stopped getting paid in the only currency the web ever had: visitors.
That’s the backdrop for everything Cloudflare has built since. The company isn’t just selling security. It’s selling leverage in a negotiation between the web’s producers and the AI industry’s consumers — and it’s keeping a toll both parties must pass through.
From One-Click Blocks to a Search, Agent, and Training Taxonomy
The escalation has been steady. In July 2024, Cloudflare shipped a one-click “Block AI Bots” button. The data behind it was stark: Bytespider (ByteDance) hit 40.4% of Cloudflare-protected sites, GPTBot 35.5%, ClaudeBot 11.2%. Yet in June 2024, only 2.98% of the top one million properties took any action to block or challenge AI bots at all.
By July 2025, the one-click block became a default — Cloudflare flipped AI crawlers to blocked unless they pay, and started building a Pay-Per-Crawl marketplace. Then came the March 2025 AI Labyrinth: AI crawlers were generating 50 billion requests a day to Cloudflare’s network — nearly 1% of all web traffic it processes — so Cloudflare built a honeypot maze of AI-generated pages to waste their time and poison their datasets.
The July 2026 update replaces blunt blocking with a pragmatic taxonomy: Search (bots building an index to answer questions later), Agent (bots acting in real time for a human — ChatGPT-User, browser-use agents), and Training (content absorbed permanently into a model). Every customer, free or enterprise, can now allow or block each category independently. Cloudflare’s argument: bot operators should separate their crawlers by purpose, the way OpenAI does with GPTBot, OAI-SearchBot, and ChatGPT-User. Transparency first, enforcement second.
September 15: The Day Googlebot Gets Blocked by Default
Here’s the part that made the announcement a Hacker News storm. On September 15, 2026, for new domains, Training and Agent bots get blocked by default on ad-supported pages. Multi-purpose crawlers are then judged by their most restrictive behavior — and Googlebot, Applebot, and BingBot all combine Search with Training.
As one top HN commenter put it: “The big news here is that Googlebot will be blocked from September 15th onwards by the ‘block training’ policies, because Google use the same crawler infrastructure for their search index AND for training Gemini.” A site owner who blocks Training — even accidentally, via defaults — loses Google search traffic entirely. One commenter who tried it reported: “Blocking AI training blocked the Google search bots and cut my traffic in half.”
This is the trap Cloudflare’s own data exposed. In its 2025 Year in Review, Cloudflare found Googlebot crawled 11.6% of unique web pages — more than 3x GPTBot (3.6%) and nearly 200x PerplexityBot (0.06%) — because it serves both search and training. “Web site operators are essentially unable to block Googlebot’s AI training without risking search discoverability,” the report concluded. Cloudflare’s fix forces the choice into the open, and Google doesn’t get to be both search and trainer under one user agent anymore — at least not by default.
The Gatekeeper Problem: An Allowlist for the Open Web
The loudest criticism isn’t that Cloudflare blocks too much. It’s that Cloudflare — one company — now decides who’s legitimate. When Cloudflare launched Signed Agents in August 2025, an essay called “The Web Does Not Need Gatekeepers” hit Hacker News and drew 454 points and 489 comments. Its thesis: “They’ve built an allowlist for the open web and told builders to apply for permission. That’s not how the internet works. An application form is not a standard.”
The mechanics are worth understanding. Signed Agents use Web Bot Auth, an IETF draft for cryptographically signing HTTP requests, so sites can verify an agent is really the ChatGPT agent or really from Browserbase. The first cohort included ChatGPT agent, Goose from Block, Browserbase, and Anchor Browser. But the critique holds: Cloudflare maintains the directory, grants Verified status, and can revoke it — and with 20%+ of web domains behind it, de-listing is a sanction with teeth. Cloudflare says so itself: losing Verified status “is a deterrent with teeth.”
HN’s skeptical wing put it more crudely. “Universal tax collector of the internet,” one commenter wrote. “So Google has to pay Cloudflare $10B to get Googlebot moved to their default allowlist… Genius move,” said another. And there’s a structural worry: Cloudflare proposes solving “transitive trust” — you might trust OpenAI, but not every weekend project built on OpenAI’s tools — with a Forwarded header (RFC 7239) extension. A protocol, yes. But one company’s implementation of it, enforced by one company’s directory.
The Same Company That Blocks Agents Also Wants to Run Them
Here’s the part that makes the “playing both sides” charge stick. Cloudflare isn’t just the bouncer at the web’s door. It’s also building the nightclub. In April 2026, it launched its AI Platform: a unified inference layer giving developers 70+ models across 12+ providers through a single API — OpenAI, Anthropic, Google, Alibaba, MiniMax, and more. Most companies already juggle an average of 3.5 models across providers, and Cloudflare’s pitch is one endpoint, one line of code to switch, automatic failover when a provider dies.
On the edge, Workers AI now runs frontier open-source models. In March 2026, it added Moonshot AI’s Kimi K2.5 — a 256k-context reasoning model — and Cloudflare’s own security-review agent, processing 7 billion tokens a day, cut costs 77% versus a mid-tier proprietary model. The infrastructure story is the same one we covered in our piece on whether we’re running out of compute power: inference is moving to where the users are, and Cloudflare’s 330-city network is a very large “where.”
So the same company that blocks a browser-use agent at one site’s edge will happily serve that agent’s inference from the same edge 100 miles away. Critics call it a conflict of interest — Cloudflare profits from both the gate and the toll road. Cloudflare calls it “the path straight down the middle.” Both are true, which is exactly why the debate won’t settle.
What the Data Says: The Toll Road Is Getting Crowded
The scale of machine traffic is the real driver of all this. Let’s put numbers on it:
• AI bots averaged 4.2% of all HTML requests across Cloudflare’s network in 2025 (excluding Googlebot, which alone added 4.5%). By December, humans generated 47% of HTML requests versus 44% for non-AI bots — people are now the minority on their own web.
• Crawl-to-refer ratios are brutal: Anthropic crawled between 25,000:1 and 100,000:1 — up to 100,000 pages crawled for every referral sent. OpenAI hit 3,700:1 in March 2025. Google’s search ratio stayed at 3:1 to 30:1. Perplexity, notably, stayed under 400:1.
• User-action crawling grew 15x+ in 2025 — the ChatGPT-User bot that fetches pages live during conversations now follows school and work schedules, dipping in summer.
• Fastly’s independent Threat Insights report found Meta alone accounted for 52% of AI crawler traffic, with Google at 23% and OpenAI at 20% — 95% concentrated in three companies. OpenAI controlled 98% of on-demand fetcher traffic, and one fetcher hit a site 39,000 times per minute.
• The non-commercial web is drowning: Wikimedia says 65% of its resource-consuming traffic is bots. GNOME’s GitLab saw only 3.2% of requests pass its challenge system. Read the Docs cut traffic 75% by blocking AI crawlers — saving $1,500 a month in bandwidth.
Even the botnet scene got involved: in late 2025, the Aisuru botnet became the most-queried domain on Cloudflare’s 1.1.1.1 resolver, and Krebs on Security documented Cloudflare scrubbing it from its public Top Domains list — a reminder that the company curates the internet’s most visible dataset as well as its traffic lanes.
The 402 Economy: Who Pays, and Who Decides?
Cloudflare’s endgame is a marketplace where crawling isn’t blocked so much as priced. Its Pay-Per-Crawl program, announced in 2025, is the seed; the 2026 update adds content-use levels — immediate (store nothing), reference (index and link back, the new default), and full (summarize and reproduce) — expressed in robots.txt via the Content Signals extension. Bots that abuse the signals lose Verified status. HN’s verdict on the honor system: “So, in summary: still the honors system. Got it.”
The harder question is who actually pays. OpenAI, Google, and Anthropic have shown they’d rather strike private deals — Google reportedly paid $60 million a year for Reddit content — than pay a toll to every site. On HN, the cynics argued Cloudflare will “happily collect the tax” while the incumbents use it as a moat: “It cements their incumbent status and pulls up the drawbridge by erecting a huge financial barrier for any new entrant.” Whatever happens, the money question is now structural, not theoretical — and it’s tied to the same open-source versus closed-source fight we analyzed here.
What You Should Do Before September 15
If you run a website behind Cloudflare, the defaults change in your name in a few weeks. Don’t let that happen passively:
• Audit your AI traffic settings now. Cloudflare says existing customers can opt out of the new defaults any time before September 15 in Security settings. Decide deliberately whether you’re blocking Training, Agent, or both.
• Know what Googlebot means to you. If search traffic is a material part of your business, the “block Training” setting now blocks Googlebot too — one HN user lost half their traffic. There’s no clean way to keep Google’s search but refuse its training, because it uses one crawler for both.
• Check the crawl-to-refer ratios of your own traffic. Radar AI Insights now tracks which bots crawl you, what they take, and what they send back. That’s the data that makes the decision rational instead of reflexive.
• Watch the standards fight, not the product fight. Web Bot Auth, Content Signals, and the Forwarded header extension are drafts, not law. Whether agent identity ends up decentralized or directory-based is the actual question that decides who controls the next web.
• If you build agents, get in the directory on your own terms. Verified status and signed agent classification are becoming the price of admission to 20%+ of the web. Being unlisted means being treated as a trespasser.
The Bottom Line
Cloudflare has become the traffic cop of the AI web. It decides which bots get in, which models run on its edge, which crawlers are “verified,” and — through its public datasets — what we even know about machine traffic. The September 15 defaults are a rare moment where one company’s configuration becomes de facto internet policy.
That concentration of power is uncomfortable, and it should be. The tools Cloudflare is building are genuinely useful — content owners finally have granular control, and bot operators have a transparent lane system. But the deeper question is whether any single company should hold the keys to both sides of the web’s busiest intersection.
So here’s the question we keep coming back to: when one company can decide — by default — whether Googlebot reaches your site, whether your agent is “real,” and which models run closest to your users… at what point does infrastructure become governance?
References
• Cloudflare Blog — “Your site, your rules: new AI traffic options for all customers” (July 1, 2026)
• Cloudflare Blog — “Content Independence Day: no AI crawl without compensation!” (July 1, 2025)
• Cloudflare Blog — “Introducing Pay-Per-Crawl” (July 2025)
• Cloudflare Blog — “The age of agents: cryptographically recognizing agent traffic” (August 28, 2025)
• Cloudflare Blog — “Radar 2025 Year in Review”
• Search Engine Journal — “Cloudflare Report: Googlebot Tops AI Crawler Traffic” (December 15, 2025)
• Cloudflare Radar — AI Insights
• Engadget — “Wikipedia is struggling with voracious AI bot crawlers” (April 2, 2025)
• Positive Blue — “The Web Does Not Need Gatekeepers” (August 29, 2025)
• Krebs on Security — “Cloudflare Scrubs Aisuru Botnet from Top Domains List” (November 8, 2025)
• Scrum Digital — Zero-click search trends analysis
• Cloudflare Blog — “AI search crawl-to-refer ratio on Radar”
• Content Signals (contentsignals.org)
• IETF Draft — Web Bot Auth architecture (draft-meunier-web-bot-auth-architecture)
• RFC 7239 — Forwarded HTTP Extension
• Reuters — “Google paid $60 million a year for Reddit AI content licensing” (February 2024)
The post Cloudflare and the New AI Traffic Wars: Who Controls What You Can Run? appeared first on The AI Prism.
Cross-posted from theaiprism.com — Cutting Through the AI Noise 🧊
Top comments (0)