Can website owners stay visible in search without permitting their content to train AI models? Cloudflare mentions its new Disallow AI Training setting is designed to make that possible. The control separates search indexing from AI training for supported crawlers, addressing a long-standing issue for publishers whose search and training traffic came from the same bot.
Cloudflare Separates Search From AI Training
Cloudflare sorts crawler activity into 3 categories: Search, Training, and Agent. Search crawlers construct indexes, training crawlers gather materials for model development or fine-tuning, and agents visit sites on behalf of users.
Mixed-use crawlers formed a particular challenge because one crawler could support both search indexing and AI training. Blocking that crawler could therefore decrease a site’s search visibility while preventing training.
The new setting changes that model. When a website allows Disallow AI Training, Cloudflare publishes the suitable training preference through robots.Txt while persisting to permit qualifying combined-use crawlers to access of site for search. Other training crawlers can be blocked without impacting search traffic.
Cloudflare says fewer than 1% of sites on its network block search bots, while 17% use some mechanism to limit AI training. That gap illustrates why separate controls have to be increasingly crucial to publishers.
Apple, Google, and Microsoft Support the Model
Cloudflare has announced an Accountable designation for crawler operators that offer, or commit to providing, controls around training, AI summaries, transparency, and search independence.
Apple, Google, and Microsoft recently meet Cloudflare’s criteria. Under the new system, Applebot, Googlebot, and Bingbot can persist crawling for search when a publisher chooses Disallow AI Training, subject to the training-control mechanisms supported by each operator.
Cloudflare also determines crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. Those companies use separate search and training crawlers, permitting Cloudflare to limit training crawlers without blocking search crawlers.
AI Summaries Are the Next Challenge
Training controls solve only a part of the content management issue. Cloudflare is likewise working toward more granular controls over how publisher content seems in AI-generated search summaries.
The company claims that AI summaries can affect publishers differently relying on their business models. Sites funded by advertising might also prioritize traffic volume, while retailers might also value smaller numbers of visitors with stronger purchase intent. Cloudflare plans to offer publishers more control over how much of their content can appear in these experiences.











