Cloudflare's New Defaults Took Effect September 15: Training and Agent Crawlers Blocked by Default on Ad-Bearing Pages, With Apple, Google and Microsoft Agreeing to Separate Search From Training

The policy Cloudflare announced on July 1 took effect on September 15: on any page displaying advertising, "mixed-use" crawlers are blocked by default. Crawlers are classified by behavior into three categories — Search (crawling to build a search index, allowed by default), Training (crawling to train or fine-tune a model), and Agent (user-directed agents visiting a page on someone's behalf) — with mixed-use meaning a single crawler doing both Search and Training. The new defaults apply to newly onboarded domains, new sites added by existing customers, and all free-tier accounts that haven't opted out; existing paid customers keep their current settings. The squeeze is the mechanism: once a site blocks Training, any crawler that bundles training with search is blocked outright, losing the search half too — and Cloudflare names Googlebot, Applebot and Bingbot as falling into that category. The same day it shipped the escape valve: a Disallow AI Training setting available on every plan that lets a site stay discoverable in search while refusing to let the same crawler train on its content, which Apple, Google and Microsoft have implemented or committed to honor within stated timeframes. An "Accountable" designation also keeps accountable mixed-use crawlers allowed for search. The legacy "Block AI bots" option is deprecated as of the same date.

This Isn't Blocking AI — It's Forcing Crawler Operators to Separate Two Things

The core of this mechanism is not the default block, it is the squeeze: if you choose to block Training, an operator that bundles training and search into one crawler gets blocked entirely — search indexing included. The cost lands on whoever refuses to split. Site owners used to face a false dilemma: accept that your content feeds model training, or drop out of search, because the same bot did both. What Cloudflare did is hand that dilemma back, unchanged, to the crawler operators. Want to keep being indexed? Run and declare separate crawlers for separate purposes. So today's real news is not the changed default. It is that **Apple, Google and Microsoft have implemented or committed to honor** the Disallow AI Training setting that shipped the same day. This whole mechanism depends on the three major search entry points being willing to separate the two activities; with their assent, site owners have for the first time a genuine "keep search, refuse training" option. There is also the Accountable designation: accountable mixed-use crawlers stay allowed for search while every other training crawler is blocked — including the training-only crawlers from Amazon, Anthropic, Meta and OpenAI, since blocking those has no effect on search.

Check Your Dashboard Today Rather Than Next Time

How the rules are inherited creates two different risks pointing in opposite directions. If you are an **existing paid customer**, your settings were not touched. That sounds safe, but the practical meaning is that you may assume the new default protects you when it does not. If you are on a **free plan, a newly onboarded domain, or a site newly added under an existing account**, the default has changed — Training and Agent are now blocked on ad-bearing pages. That may be blocking access you actually wanted, such as certain user-directed agents. Both cases point to the same action: open the Cloudflare dashboard's security settings and review your AI bot policies. The defaults flipped yesterday, so this is a review rather than a precaution. Note too that the legacy "Block AI bots" option is deprecated as of the same date, so any configuration still relying on it needs migrating.

A Few Numbers Explain Why This Exists

Mixed-use crawlers now account for 36.6% of verified crawler traffic on Cloudflare's network, the single largest category, and its market report says 52% of crawler requests on the network now serve AI training, up from 22% in spring 2025. On the other side, fewer than 17% of site owners have placed any restriction on AI training, and fewer than 1% block search bots. Given that distribution, a default matters far more than any individual site's active choice — which is exactly why Cloudflare made the move in the defaults rather than shipping another toggle. The commercial layer changed in step: Pay Per Use replaces the per-fetch Pay Per Crawl, starting with Ceramic.ai and You.com as partners. Three open questions are worth keeping. There is no detailed rule for what counts as a "page that displays advertising" — whether it detects ad scripts, relies on a customer declaration, or something else. Which category a given crawler falls into is Cloudflare's call and can change without notifying publishers. And robots.txt remains only a statement of preference here: Cloudflare says plainly that compliance with it is voluntary, and the actual enforcement comes from AI Crawl Control.

via: Cloudflare blog on accountable mixed-use AI crawlers, TechCrunch, Help Net Security