Web infrastructure company Cloudflare has introduced a new control that lets website publishers block their content from being used for AI model training while keeping it fully available in traditional search results — a decoupling that the crawler ecosystem has never cleanly supported before.
The feature, announced on the Cloudflare blog on September 15, arrives alongside a new "Accountable" designation for crawler operators. According to Cloudflare, Apple, Google and Microsoft are the first companies whose crawlers qualify. Anyone tracking breaking AI news this year knows the question of who gets to scrape published content has become one of the internet's most contested policy fights.
The Mixed-Use Crawler Problem
The core issue is what Cloudflare calls mixed-use crawlers: single bots, like Googlebot, Bingbot and Applebot, that fetch pages both to build search indexes and to collect training data for AI models. Historically, a site owner who wanted to opt out of AI training had one blunt instrument — the robots.txt disallow — and using it risked vanishing from search results entirely.
Cloudflare's new Disallow AI Training setting is named for the directive it publishes in a site's robots.txt file. When enabled, the site's no-training preference is synced into robots.txt, Accountable mixed-use crawlers remain allowed for search purposes, and every other training crawler is blocked. Cloudflare notes that blocking training-only crawlers — it names the training crawlers run by Amazon, Anthropic, Meta and OpenAI — never affected search in the first place.
What Makes a Crawler 'Accountable'
To earn the Accountable designation, a bot operator must meet or commit to meeting a set of requirements laid out in the blog post:
- A mechanism for site owners to opt out of AI training, via robots.txt or a similar standard
- A mechanism to opt out of AI summaries, set directly with the operator and, next year, through Cloudflare
- URL-level visibility into which pages were made available for training, plus metrics showing how content appeared in search
- Assurance that opting out of AI training will not affect traditional search rankings
Cloudflare says Apple, Google and Microsoft currently demonstrate they meet the qualifications, combining capabilities available today with time-bound commitments for features still in development.
New Settings, Deprecated Tools
The change reworks Cloudflare's bot management controls, which now classify crawler traffic into three behaviors: Search, Training and Agent. With Disallow AI Training added, the available settings are Allow, Disallow AI Training, Block on pages with ads, and Block.
Two long-standing tools are being deprecated. The broad "Block AI Bots" option gives way to the more granular Search, Training and Agent controls, and Managed Robots.txt is replaced by a system called Bot Preference Sync, with existing customers migrated automatically. Disallow AI Training will also become part of Cloudflare's recommended configuration for certain new domains.
The most consequential switch concerns the Block setting itself. Previously, Block and "Block on pages with ads" did not apply to mixed-use crawlers, because blocking them could damage search discoverability. As of September 15, those settings now apply to all training crawlers, including Applebot, Bingbot and Googlebot — meaning a site that chooses Block now accepts the search-ranking consequences of that choice.
No 'Disallow' for Agents — Yet
One gap remains. Cloudflare's Agent category covers user-directed agents, such as browser-use agents and chat fetch bots, visiting a page on a person's behalf. The company writes that agents do not create the same search-discoverability tradeoff as mixed-use crawlers, and that the internet lacks a well-established directive for expressing disallow preferences to agents. For now there is no Disallow setting for Agents, though Cloudflare says it will revisit the approach as standards such as ai-prefs mature.
Why It Matters for Publishers
For news organizations and content sites, the announcement is the closest thing yet to a workable middle path in the standoff between the publishing industry and AI developers. Publishers have argued that training on their work without payment or permission is extraction; AI companies have argued that crawling is standard practice and that robots.txt was never designed for AI-era distinctions.
By formalizing a designation for crawlers that separate search indexing from training use — and by enforcing it through default controls — Cloudflare is effectively asking AI labs and platform companies to compete on publisher-friendly terms. The three founding Accountable operators happen to be companies whose core businesses depend on search and operating systems, not model training, which is unlikely to be a coincidence.
The open question is enforcement and economics. Cloudflare can classify behavior and sync preferences, but the financial terms of training licenses remain a private-market matter between publishers and AI companies. What changes this week is more mundane and more meaningful: a site that says no to AI training can now do so without sacrificing its place in search results. For an industry watching referral traffic and training disputes unfold at the same time, that separation has been a long time coming.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →