AI agents are getting better at using the web.
They can search for information, read documentation, compare products, research topics, and collect data for users.
But there is a problem:
More websites are deciding which kinds of AI traffic they want to allow.
Cloudflare is one of the biggest pieces of infrastructure between websites and their visitors. So when Cloudflare changes how AI traffic is handled, it matters to anyone building software that fetches web pages.
In July 2026, Cloudflare introduced separate controls for three types of AI-related traffic:
- Search
- Agent
- Training
Then, in September 2026, Cloudflare changed the default settings for new domains.
Here is what developers need to know.
What changed?
Cloudflare's AI traffic controls were introduced in stages.
July 1, 2026
Cloudflare introduced three separate AI traffic categories:
Search
Crawlers that collect or index content so it can be used to answer questions later.
Agent
Automated software acting on behalf of a person, usually in real time.
Training
Crawlers collecting content to train or fine-tune AI models.
Website owners can decide how each type of traffic should be handled.
They can allow it, block it, or block it only on pages that display ads.
Existing Cloudflare customers keep their existing settings.
September 15, 2026
Cloudflare changed the default settings for new domains.
For new domains:
Traffic type New default
Search Allowed
Agent Blocked on pages with ads
Training Blocked on pages with ads
Website owners can change these settings.
The important part for AI developers is Agent traffic.
Why does this matter for AI agents?
Imagine a user asks an AI agent:
"Find the latest information about this product"
The agent may need to open several web pages, read them, and return the relevant information.
That is different from a traditional search crawler.
Cloudflare's definition of Agent is focused on automated software acting for a person in real time.
That is very close to how many AI agents work today.
So if more website owners use Cloudflare's new defaults, some AI agents may encounter more blocked pages.
But this won't happen to every website.
The new defaults apply to new domains onboarding to Cloudflare. Existing sites keep their current configuration unless the owner changes it.
Not every Cloudflare site will block AI agents
This is important.
Cloudflare isn't simply saying:
"AI crawlers are blocked"
The website owner decides what happens.
A site can:
- Allow Search traffic
- Allow Agent traffic
- Allow Training traffic
- Block one of them
- Block traffic only on pages with ads
- Block traffic everywhere
So a Cloudflare-protected website does not automatically mean an AI agent cannot access it.
You need to look at what the site actually allows.
What about Googlebot and other mixed crawlers?
Some crawlers do more than one job.
Cloudflare's documentation describes these as mixed-purpose crawlers.
For example, a crawler may be used for search indexing while also being associated with other AI-related activities.
Cloudflare has also added a Disallow AI Training option.
This allows a website owner to express a preference against AI training while still allowing accountable mixed-use crawlers to index the website for search.
So the exact setting chosen by the website owner matters.
What happens when your agent gets blocked?
A blocked request doesn't always mean the same thing.
You might receive:
- 403 Forbidden
- A Cloudflare challenge
- A rate-limit response
- Another error from the website And you cannot always tell from the response whether the website intentionally blocked AI traffic or whether something else went wrong.
For example, in Spicrawl, an upstream bot challenge can appear as:
ERR::UPSTREAM::CHALLENGE
The target website's HTTP status can also be returned through the X-Target-Status header.
This distinction matters.
Your scraper's HTTP response is not always the same as the target site's actual response.
How should developers handle this?
The answer isn't to keep trying until the block disappears.
If a website has clearly decided not to allow your automated traffic, treat that as a decision.
Here are some better practices.
1. Check the site's rules
Look at:
- robots.txt
- Terms of service
- API documentation
- Other access policies These can tell you what the website allows.
2. Use an official source when available
If the website provides:
- An API
- A feed
- A sitemap
- An official data source use it instead of repeatedly fetching the website.
3. Treat a deliberate block as a "no"
If a website intentionally blocks automated access, don't try to work around the decision.
Ask for permission or use another source.
4. Fetch only what you need
Don't download an entire website when you only need one piece of information.
For example, request the specific content you need and avoid unnecessary requests.
This is better for both your application and the website you're accessing.
5. Avoid repeated requests
Caching can help prevent fetching the same page again and again.
For example, Spicrawl's cache is enabled by default and can keep results for up to 48 hours.
6. Slow down when necessary
If a website starts returning errors or rate limits, don't immediately send more requests.
Reduce the request rate and back off.
7. Check the real target status
If you're using a scraping API, don't assume its HTTP status tells the whole story.
Check whether the response contains the target site's actual status.
What this means for AI developers
The web is becoming a more controlled environment for AI agents.
Search crawlers, AI training crawlers, and real-time AI agents are increasingly being treated as different types of traffic.
That means an AI agent can't simply assume:
"If a browser can open the page, my agent can open it too."
Access can depend on the website's policies, infrastructure, bot controls, rate limits, and Cloudflare configuration.
For developers, this means your web-access layer needs to handle failures properly and respect the decisions made by website owners.
What website owners can do
If you run a website behind Cloudflare, you now have more control over AI traffic.
You can separately manage:
- Search
- Agent
- Training You can also use Cloudflare's Disallow AI Training option to express that preference while allowing certain mixed-use crawlers to continue indexing for search. Cloudflare has also introduced Pay Per Use, currently in beta, for publishers who want to receive payment when AI products use their content.
The bigger picture
AI agents need access to the web to be useful.
At the same time, websites need control over how their content is accessed and used.
Cloudflare's new AI traffic controls are part of that shift.
For developers, the practical lesson is simple:
Don't treat every blocked page as a technical problem to bypass. First understand why access was denied and whether you're allowed to access the content.
As more websites define separate rules for search, AI agents, and AI training, responsible web access will become an important part of building reliable AI systems.
Originally published on Spicrawl
This article is adapted from the original: Cloudflare AI crawler blocking in 2026
About Spicrawl
Spicrawl is a web data platform for developers and AI agents, providing APIs for fetching web pages and returning clean data such as Markdown, JSON, HTML, and text.
Top comments (0)