Cloudflare has introduced a “Disallow AI Training” setting intended to let sites remain accessible to search crawlers while expressing a refusal of AI training use. The 15 September change addresses crawlers that serve more than one purpose, where blocking access can affect both training and search.

The announcement also changes the meaning of the stronger “Block” options. These now apply to mixed-use crawlers, including Googlebot, Bingbot and Applebot, so selecting them can affect search access as well.

A preference and a network block do different jobs

Under Disallow AI Training, Cloudflare publishes the relevant preference through robots.txt using Bot Preference Sync. Mixed-use crawlers that it designates “Accountable” can continue visiting for search. Other training crawlers are blocked.

Cloudflare’s designation combines capabilities already available with commitments to deliver missing controls. It is therefore not a statement that every listed operator implements every part of the system today.

The distinction matters particularly for Bing. Cloudflare says Microsoft’s support for a domain-level no-training preference through robots.txt is targeted for early 2027. Until then, selecting the new setting does not automatically convey that preference to Bing through robots.txt. The release points to Microsoft’s existing controls for the interim.

That limitation belongs alongside the feature announcement. A site owner should not infer that one switch already produces identical treatment by every crawler.

Existing settings are being migrated

Cloudflare says customers’ existing preferences will generally carry over with their practical effect preserved. Previous Training selections of Block or Block on pages with ads migrate to Disallow AI Training under the described rules.

A site owner who wants mixed-use crawlers stopped entirely can choose Block, but should understand that it also stops their search crawling. The new controls separate Search, Training and Agent behaviour at the domain level.

Agents visiting a page on behalf of a user remain a separate category. Cloudflare says it is not introducing a comparable Disallow preference for agents at this stage, citing the lack of an established directive.

Training and summaries remain separate questions

The announcement distinguishes using content to train a model from displaying content in an AI summary. It describes more granular controls over summary use as future work, with a goal for early next year.

For publishers, the immediate task is to inspect the migrated configuration and the support offered by each relevant operator. Keeping a crawler reachable, expressing a use preference and enforcing a block are related but distinct actions.

This sits within the wider technology and security decisions facing publishers: access controls need to match the behaviour a site intends to permit. The release does not establish that a training preference preserves traffic levels or resolves every question about how content appears in AI products.

Questions

Does Disallow AI Training block search crawlers?

Accountable mixed-use crawlers remain allowed for search under the new setting; the stronger Block setting stops them entirely.

Does Bing already receive the new preference through robots.txt?

Cloudflare says that support is targeted for early 2027, so the setting does not automatically convey it that way at launch.

Are training and AI summaries controlled in the same way?

No. The announcement treats them separately and describes more granular summary controls as future work.

Sources