Cloudflare separates AI search, training and agents. What brands need to know
Cloudflare has replaced its blunt AI bot controls with separate settings for search, training and agents. Here is what the change means for search visibility, AI citations and crawler strategy.
On 15 September 2026, Cloudflare changed the way website owners control AI crawlers. The important shift is not simply that there are more settings. It is that Cloudflare now treats search, model training and AI agents as different behaviours.
For brands, that distinction matters. A website may want to remain discoverable in search and in AI-generated answers without giving every crawler permission to use its content for model training. Until now, those objectives could be difficult to separate when a single crawler served more than one purpose.
Cloudflare's new controls are designed to make that choice more precise.
What Cloudflare changed
Cloudflare now groups AI-related crawler activity into three behaviours.
Search covers crawlers that build an index so content can be found later. Training covers crawlers used to train or fine-tune models. Agent covers user-directed systems that visit a page in real time on behalf of a person, such as chat fetch bots and browser-use agents.
The company is also deprecating its older Block AI Bots control in favour of these more granular settings. Cloudflare says existing customer preferences will be migrated automatically to the new Search, Training and Agent controls. Its documentation explains the updated policy options here.
The biggest addition is a setting called Disallow AI Training. It is intended for sites that want to express a no-training preference without unnecessarily removing themselves from search discovery.
That sounds like a subtle infrastructure change. In practice, it resolves an increasingly important problem for brands.
Why the old "block AI" approach was too blunt
Some crawlers are single-purpose. Others are mixed-purpose and can support both conventional search and AI-related uses.
Cloudflare specifically identifies Applebot, Bingbot and Googlebot as mixed-use crawlers. Under the updated controls, choosing Block for training can now block those crawlers completely, including their search function. That means a decision made to restrict AI training could also affect ordinary search discoverability.
Cloudflare's new Disallow AI Training option is designed to avoid that trade-off for crawler operators that support a separate no-training preference.
The demand for that distinction is already visible in Cloudflare's own data. The company says less than 1% of Cloudflare sites choose to block Search bots, while 17% use some mechanism to block Training. In other words, most site owners appear to see discovery and model training as two different decisions. Cloudflare published the figures in its 15 September announcement.
For marketing, SEO and content teams, that is the key takeaway: "Can this system find our content?" and "Can this company train on our content?" are no longer the same question.
Search, training and agents are three different decisions
The practical value of Cloudflare's model is that it forces website owners to think about crawler access by purpose rather than by a vague category called "AI bots".
A search crawler may index a product page so it can later surface in search results. A training crawler may collect the same page as part of a dataset used to improve a model. An agent may fetch the page at the moment a user asks it to compare products, check availability or summarise information.
Those activities can all touch the same URL, but the commercial implications are different.
For many brands, search discovery is desirable. Agent access may also be desirable where it helps a customer research or interact with the business. Model training is a separate policy question involving content rights, commercial value and organisational preference.
Cloudflare is effectively giving site owners a clearer control surface for those choices.
What "Disallow AI Training" actually does
Cloudflare says the new setting uses Bot Preference Sync to publish an appropriate no-training preference in robots.txt while allowing recognised mixed-use crawlers to continue operating for search.
For training-only crawlers operated by companies including Amazon, Anthropic, Meta and OpenAI, Cloudflare says it can block the training crawler without affecting a separate search crawler where the operator separates those functions.
Apple and Google already provide mechanisms that distinguish search from AI training. Apple uses Applebot-Extended, while Google uses Google-Extended. Both companies state that opting out of those extended uses does not affect conventional search ranking.
Microsoft is slightly different. Cloudflare says Microsoft is working towards a domain-level no-training preference in robots.txt, targeted for early 2027. Until that support is available, Cloudflare notes that selecting Disallow AI Training will not automatically communicate a no-training preference to Bing through robots.txt. Microsoft currently provides other controls, including its NOARCHIVE mechanism and Webmaster Tools.
That nuance is important. Disallow AI Training is not a universal protocol that every crawler interprets identically today. It is a control layer that combines Cloudflare's crawler classification and enforcement with the preferences supported by individual crawler operators.
robots.txt is a preference, not a security boundary
There is another distinction brands should understand.
A robots.txt file tells a crawler what a site owner wants it to do. It does not, by itself, technically prevent access. Cloudflare's own documentation is explicit that robots.txt compliance is voluntary.
Cloudflare can add an enforcement layer because it sits in front of the website. Its AI Crawl Control product can identify crawler traffic and allow, block or otherwise manage it according to the site's configuration. Cloudflare documents the difference between robots.txt preferences and enforced crawler controls here.
For teams reviewing their AI crawler strategy, this means there are two questions to ask: what preference does the website publish, and what access is actually being enforced?
Crawler access is not the same thing as AI visibility
This is where the conversation becomes particularly relevant to brands measuring AI search.
Allowing a crawler to access your website does not guarantee that ChatGPT, Gemini, Perplexity, Claude or another answer engine will mention, recommend or cite your brand. Crawlability is infrastructure. Visibility is an outcome.
Likewise, blocking a particular training crawler does not necessarily mean a brand disappears from AI-generated answers. An engine may still find information through a search index, retrieve a page using a separate crawler, cite a third-party source, or draw on information acquired through other permitted channels.
A useful way to think about the sequence is:
Accessible -> discoverable -> retrievable -> cited -> recommended.
Crawler controls influence the first stages. They do not determine the final answer.
That is why technical crawler configuration should be considered alongside AI visibility tracking, citation analysis and source influence. The practical question for a marketing team is not merely whether an AI bot can reach the site. It is whether the brand actually appears when prospective customers ask relevant questions.
What brands using Cloudflare should check now
- Review the new Search, Training and Agent settings. Do not assume the old Block AI Bots configuration maps to the policy your organisation wants today.
- Decide on each behaviour separately. Search discoverability, AI training and agent access have different commercial implications.
- Use Disallow AI Training carefully if your objective is "search yes, training no". A full Block setting can affect mixed-use crawlers such as Googlebot, Applebot and Bingbot.
- Review both
robots.txtand enforced crawler rules. A published preference and a network-level block are not the same thing. - Check the outcome, not just the configuration. Verify that important pages remain accessible, then monitor whether your brand continues to be mentioned and cited in the AI answers that matter to your customers.
Citations.io also has a free AI crawler checker that can help teams inspect how a site currently handles common AI crawler user agents.
What this means for AEO and AI search strategy
The technical foundations of search visibility are becoming more granular.
Traditional SEO teams are accustomed to thinking about Googlebot, indexability and robots.txt. AI search adds new layers: training crawlers, retrieval crawlers, user-triggered agents and source-level citation behaviour.
The result is not that every brand should simply allow every crawler. The better approach is to make an intentional decision about each use case, then measure the consequences.
For organisations working on answer engine optimisation or generative engine optimisation, crawler policy should now sit alongside content quality, technical accessibility, structured data, third-party authority and citation strategy.
A perfectly configured crawler policy cannot make an answer engine recommend a brand. But an accidentally restrictive policy can remove an important route through which content is discovered or retrieved.
The bigger shift: control and visibility are separating
Cloudflare's update reflects a broader change in how the web interacts with AI systems.
Website owners increasingly need controls that distinguish between being found, being used, and being acted on. Search, training and agents represent three different relationships between a website and an automated system.
For brands, the strategic opportunity is to stop treating "AI access" as a single on/off decision.
Decide where discovery creates value. Decide where training is acceptable. Decide how agents should interact with your pages. Then measure what actually happens in generated answers.
Citations.io tracks how ChatGPT, Gemini, Perplexity and Claude mention, rank and cite brands across AI-generated answers. Crawler configuration is one part of that picture. The outcome is whether your brand is visible when it matters.
Source note: This article is based on Cloudflare's 15 September 2026 announcement, its AI crawler policy documentation and its AI Crawl Control documentation. Crawler policies and operator support can change, so technical teams should verify current documentation before making production changes.

