DEV Community

Cover image for Cloudflare Separates AI Training From Search Indexing With New Crawler Controls
Ali Farhat
Ali Farhat Subscriber

Posted on Originally published at scalevise.com

Cloudflare Separates AI Training From Search Indexing With New Crawler Controls

Cloudflare has introduced new AI-crawler controls designed to solve a growing problem for website owners: allowing content to remain discoverable in search results without also permitting that content to be used for AI model training. The central addition, Disallow AI Training, publishes a no-training directive in a site's robots.txt while allowing eligible mixed-use crawlers to continue search indexing.

The change is part of Cloudflare's September 15, 2026 update on accountable mixed-use crawlers. As explained in Cloudflare's announcement on accountable-mixed-use AI crawlers, the company is moving away from its previous single Block AI Bots switch and toward separate policies for Search, Training, and Agent activity. For publishers, marketers, and businesses that depend on organic visibility, that distinction can be more useful than an all-or-nothing block.

What Cloudflare's new policy changes

A crawler can serve more than one purpose. It may index pages for a search product while also collecting material that could be used to train AI systems. Under a one-button AI block, site owners faced a blunt choice: permit that activity or risk limiting indexing by crawlers with mixed uses.

Cloudflare's updated approach separates the policy decision. The Disallow AI Training setting communicates that a site does not permit AI training, while search indexing can remain allowed. Cloudflare is also replacing Block AI Bots with distinct controls for three categories:

  • Search, for crawlers that index content for search.
  • Training, for crawlers that collect content for AI model training.
  • Agent, for AI agents that access websites to carry out tasks.

The company says new defaults effective September 15, 2026 can block Training and Agent activity while Search remains allowed. It also says mixed-use crawlers that combine Search and Training are blocked by configurations that block training. This policy is intended to prevent a crawler from using search-related access as a route to collect material for training.

Policy approach Search indexing AI training Agent activity
Previous Block AI Bots control Not separately controlled through the one-button setting Blocked when the setting was enabled Not separately controlled through the one-button setting
New granular AI-crawler policies Can remain allowed Can be disallowed through the Training policy and directive Can be controlled separately

Why robots.txt synchronization matters

Cloudflare is replacing Managed Robots.txt with Bot Preference Sync, which automatically keeps a zone's robots.txt aligned with its AI bot policies. The feature is enabled by default for new customers. Existing customers will be guided through a migration path.

That synchronization matters because a site's stated crawling preferences should match the policy selected in its Cloudflare configuration. Without it, a website owner could change a control in one place while its robots.txt gives a different signal to crawlers. Bot Preference Sync is intended to reduce that administrative gap.

Cloudflare also differentiates onboarding recommendations for ad-supported and non-ad-supported sites. Those presets affect the suggested Training, Search, and Agent settings, recognizing that sites with different business models may make different choices about discovery, content use, and automated access.

What the controls mean for content strategy

The practical value of this update is choice rather than automatic exclusion. A company that wants search engines to find product pages, documentation, articles, or local service content can keep search access open while expressing a separate preference against training use.

That can be relevant wherever content is a business asset. Original guides, product data, pricing explanations, support content, and research may help a company earn search traffic, yet the owner may not want that same material used for model training. Cloudflare's controls do not change the need to create useful, indexable pages, but they provide a more granular way to state access preferences.

The policy's effectiveness also depends on crawler operators honoring the directive. Cloudflare states that Apple, Google, and Microsoft have committed to honor the Disallow AI Training setting within a stated timeframe. That is significant because major crawler operators are needed for a robots.txt-based preference to have practical effect. The setting communicates a clear policy, rather than creating a technical guarantee that every crawler on the internet will comply.

How website owners should approach the rollout

The first decision is strategic: identify whether search visibility, AI training restrictions, and AI-agent access should be treated as separate permissions for the site. A business may reasonably choose different policies for each category based on how it uses content and how important search discovery is to its acquisition strategy.

Website owners using Cloudflare should then review the available AI Crawl Control and Bot Management policies, including whether their existing Block AI Bots configuration needs migration. They should also check how Bot Preference Sync affects the site's robots.txt and whether the onboarding preset reflects the site's model. The key is to treat these controls as part of a wider content-access policy, not as a substitute for monitoring search visibility or maintaining accurate site content.

For businesses that rely on organic discovery, the new separation reduces a previous trade-off. It makes it possible to signal opposition to training while preserving a path for search indexing, subject to crawler compliance and the specific policy selected.

AI search is changing how customers discover brands and content. If your company needs to protect its content choices without losing sight of where it appears in AI-generated answers, Scalevise's AI Visibility and GEO Checker can help identify how your brand is represented across relevant AI search experiences and turn the findings into practical visibility priorities. Start an AI Visibility scan to see where your business stands.

Frequently Asked Questions

What is Cloudflare's Disallow AI Training setting?

Disallow AI Training is a Cloudflare setting that publishes a no-training directive in robots.txt. It is designed to let site owners disallow AI training while allowing mixed-use crawlers to continue indexing content for search.

Does blocking AI training also block search indexing?

Not necessarily. Cloudflare's new granular controls separate Search from Training. Its September 15, 2026 defaults can block Training and Agent activity while Search remains allowed.

What happened to Cloudflare's Block AI Bots setting?

Cloudflare is deprecating the one-button Block AI Bots setting in favor of separate controls for Search, Training, and Agent crawler types. Existing customers will be guided through migration.

What is Bot Preference Sync?

Bot Preference Sync replaces Managed Robots.txt. It automatically synchronizes a zone's robots.txt with its AI bot policies and is enabled by default for new Cloudflare customers.

Will major crawler operators honor the no-training directive?

Cloudflare says Apple, Google, and Microsoft have committed to honor the Disallow AI Training setting within a stated timeframe. The directive remains a policy signal whose practical effect depends on crawler operators honoring it.


Conclusion

Cloudflare's new crawler policies give website owners a more precise alternative to a blanket AI-bot block. By separating search indexing, AI training, and agent access, the company enables sites to preserve search discoverability while publishing a clear restriction on training use. The rollout also makes robots.txt synchronization and policy review more important for organizations that want their stated content-access preferences to be consistent.

Top comments (0)