For years, website owners have been given a fairly blunt choice when it comes to AI crawlers: allow them or block them.
That choice is becoming more complicated - and more important.
On 15 September 2026, Cloudflare rolled out changes that separate automated crawlers according to what they are actually doing with a website. Its controls now distinguish between Search, Training and Agent activity, while a new "Disallow AI Training" setting is designed to let publishers reject model training without unnecessarily sacrificing search discoverability.
For brands, publishers and marketing teams, this is more than a security setting. It reflects a wider change in how the web is being accessed by AI systems.
Being available for AI training, being indexed for search, being retrieved for an AI-generated answer and being visited by an autonomous agent are not the same thing. Treating them as one category can create unintended consequences for visibility.
What Cloudflare changed
Cloudflare now groups AI-related crawler behaviour into three broad categories:
Search - crawlers that collect or index content so it can be surfaced later through search and discovery experiences.
Training - crawlers that collect content for training or fine-tuning AI models.
Agent - automated systems acting in real time on behalf of a user, such as chat fetch bots or browser-use agents visiting a page to complete a task.
Cloudflare had already introduced these behavioural categories earlier in 2026. The significant change on 15 September is how its controls now handle mixed-use crawlers and the addition of its "Disallow AI Training" setting.
Cloudflare says the setting publishes an applicable no-training preference through robots.txt while allowing what it calls "Accountable" mixed-use crawlers to continue accessing the site for search. Training-only crawlers can be blocked without necessarily affecting search access.
The practical effect is that website owners have more granular choices than a single "block AI bots" switch.
Why the old "block AI" approach is too simplistic
A crawler's identity does not always tell you its purpose.
Some organisations use separate crawlers for separate functions. Others use a single crawler for more than one purpose. Cloudflare refers to these as mixed-use crawlers.
That creates a difficult trade-off. If a crawler is used both to build a search index and for AI training, blocking it completely can affect both activities.
Cloudflare specifically says that its full Block setting now applies to mixed-use crawlers including Applebot, Bingbot and Googlebot. If a website owner selects Block for training, those crawlers can be prevented from reaching the site altogether, including for search.
For brands that rely on organic discovery, that distinction matters.
A company might reasonably decide that it does not want its content used to train a model. That does not necessarily mean it wants the same content excluded from search indexes, retrieval systems or other discovery surfaces.
The important question is therefore no longer simply:
"Do we allow AI bots?"
It is:
"Which automated systems do we want to access our content, for which purposes, and what visibility do we expect in return?"
AI training is not the same as AI search
The terms are often grouped together, but training and retrieval are fundamentally different processes.
Model training typically happens before a user asks a question. Large amounts of information are used to develop or refine the model itself.
Search and retrieval happen closer to the point at which a user asks a question. A system may search an index, retrieve a current webpage or use another source to support the answer it generates.
This difference matters because a brand can have strong reasons to permit one form of access while restricting another.
A retailer may want current product information discoverable in search. A software company may want its documentation accessible to users asking technical questions. A professional services firm may want its research and expertise to be found and cited.
None of those objectives automatically require unrestricted use of the same content for model training.
Cloudflare's new controls are an acknowledgement that these use cases need to be treated separately.
AI agents add a third category
The Agent category makes the picture more interesting again.
An AI agent may access a website in real time because a user has instructed it to research something, compare products, collect information or complete a task.
That is different from both indexing and training.
For businesses, agent access could become commercially significant. A future customer might not visit a website directly before making a decision. Instead, their AI assistant could visit the site, read product information, check availability, compare options and return with a recommendation.
Blocking agents therefore has a different potential consequence from blocking training crawlers.
There are legitimate reasons to restrict automated agents, particularly where access involves sensitive areas, transactions, advertising models or resource-intensive activity. But brands should make that choice deliberately rather than assuming all AI-related traffic has the same value.
Mixed-use crawlers are where the trade-off becomes real
Cloudflare's announcement focuses heavily on mixed-use crawlers because they make publisher choice harder.
According to Cloudflare, Apple, Google and Microsoft meet, or have committed to meet, the requirements for its new "Accountable" designation. Those requirements include mechanisms for publishers to opt out of AI training and assurances that doing so will not affect traditional search results.
Cloudflare also says organisations including Anthropic and OpenAI operate separate search and training crawlers. In those cases, Cloudflare can block the training-specific crawler without necessarily blocking the crawler used for search-related discovery.
The implementation details will continue to evolve, and crawler policies should always be checked against the latest documentation from both Cloudflare and the crawler operator.
The broader principle is more durable: the purpose of a crawl matters as much as the name of the crawler.
What should brands check now?
Most Cloudflare customers do not need to make an immediate change. Cloudflare says existing preferences are being migrated automatically to the new controls.
However, this is a useful point for marketing, SEO and technical teams to review how AI access is configured.
1. Review your Cloudflare AI crawler settings
If your website uses Cloudflare, check the Search, Training and Agent controls in Security Settings.
Do not assume that an old "Block AI Bots" configuration maps perfectly to the outcome you want today. Cloudflare is deprecating that legacy control in favour of the more granular settings.
2. Decide separately how you feel about search, training and agents
Treat the three behaviours as separate policy decisions.
For example, a brand might choose to:
- allow search crawlers so its content remains discoverable;
- disallow AI training;
- allow agents on public information pages but restrict them elsewhere.
The right policy will depend on the site's commercial model, content and technical risk profile.
3. Review robots.txt as well as Cloudflare
Crawler control is not confined to one dashboard.
Review the site's robots.txt file and understand which directives are being published. Cloudflare's Bot Preference Sync can publish preferences on behalf of a site, so teams should know which system is controlling the final output. Our free AI crawler checker shows which AI crawlers a site's robots.txt currently allows or blocks.
It is also important to remember that robots.txt expresses a preference. Enforcement and crawler compliance can vary, which is one reason infrastructure-level controls are becoming more important.
4. Check your important pages remain accessible
Policies should be tested against the pages that matter commercially.
Product pages, documentation, service pages, research, help content and other high-value resources should not become inaccessible to important discovery systems by accident.
A technically correct blocking rule can still be commercially wrong if it removes content from the channels through which prospective customers are finding it.
5. Measure what happens after crawler access
Crawler access is only the first layer.
Allowing an AI search crawler to access a website does not guarantee that the brand will appear in generated answers. It does not guarantee a recommendation. And it certainly does not guarantee a citation.
That is where AI visibility measurement becomes necessary.
Brands need to understand whether they are actually being mentioned, which competitors appear alongside them, which sources influence generated answers and whether their own domains are being cited.
Crawler access is infrastructure. Visibility is the outcome.
The Cloudflare changes are important because they give website owners more control over the infrastructure of AI discovery.
But crawler access and AI visibility should not be confused.
A website can be fully crawlable and still have weak representation in generated answers. An AI engine may rely on third-party publishers, review sites, directories, forums, datasets or competitor content instead of the brand's own website.
Conversely, a brand can appear prominently in AI-generated answers because other authoritative sources discuss it, even when its own domain is not the primary citation.
This is why AI visibility requires a broader view than technical crawler configuration alone.
The useful questions are not just:
"Can an AI crawler access our site?"
They are also:
- Are AI engines mentioning our brand?
- Are we being recommended?
- Where do we appear relative to competitors?
- Which sources are being cited when our category is discussed?
- Is our own website influencing those answers?
- Are there important prompts where we are absent entirely?
Crawler policy determines what systems are allowed to access. AI visibility tells you what those systems actually do with the information available to them.
A new layer of web governance
Cloudflare's update is another sign that AI discovery is becoming part of normal web infrastructure rather than a separate technical curiosity.
For years, brands have managed search crawler access, indexing rules, canonicalisation and structured data as part of SEO. They will increasingly need similarly deliberate policies for AI training, retrieval and agent access, alongside answer engine optimisation.
The objective should not be to allow everything or block everything.
It should be to understand the role each type of automated access plays, decide which uses align with the organisation's objectives and then measure the resulting visibility.
As AI engines become a more important route through which people discover products, services and information, crawler configuration will increasingly sit alongside technical SEO, content strategy and digital governance.
And, as Cloudflare's latest changes demonstrate, "AI crawler" is no longer a sufficiently precise category on its own.
How Citations helps
Citations tracks how AI engines mention, rank and cite brands across generated answers.
It helps teams move beyond crawler access and understand the outcome: where a brand appears, which competitors are being surfaced, which sources influence answers and where there are opportunities to improve visibility.
Technical accessibility is part of the foundation. The next question is whether that accessibility is translating into meaningful presence across AI-generated discovery.
Sources
- Cloudflare, "Have it both ways: stay discoverable in search while disallowing AI training", 15 September 2026.
- Cloudflare Developers, "Block AI Bots".
- Cloudflare Developers, "AI Crawl Control".

