
Cloudflare has announced a significant update to its AI bot management tools, moving beyond a simple block-or-allow toggle and introducing a three-category taxonomy that lets website owners control AI traffic based on how crawlers actually use their content.
From One Toggle to Three Categories
When Cloudflare launched its one-click “Block AI Bots” option a year ago, the primary concern was AI companies scraping content for model training without sending referral traffic in return. The conversation has since grown more complex. Not all automated AI traffic serves the same purpose, and site owners increasingly want finer control rather than an all-or-nothing switch.
The new system organizes automated AI traffic into three categories:
- Search: Bots that proactively index content to answer queries later. Cloudflare notes that site owners should expect referral traffic or comparable compensation in exchange for this access.
- Agent: Automated behavior acting on behalf of a real user in real time. This includes chat fetch bots and browser-driving agents that visit a site to complete a specific task, often with a person waiting on the result.
- Training: Crawlers that ingest content to train or fine-tune AI models, permanently incorporating the data into model architecture.
These three controls are available to all Cloudflare customers, including those on the free plan.
New Defaults Arriving in September
Cloudflare has announced that on September 15, 2026, new default settings will apply to all domains that onboard to the platform going forward. On pages that display advertising, Training and Agent categories will be blocked by default, while Search will remain allowed. The presence of ads is treated as a signal that the page is intended for human visitors and is part of the site owner’s monetization strategy.
A Push for Crawler Transparency
Alongside the new controls, Cloudflare is encouraging bot operators that perform multiple functions to split those functions into separate, clearly identified crawlers. If a company runs automation for indexing, for acting as an agent, and for training data collection, Cloudflare wants each behavior to run under a distinct crawler identity. The goal is to give site owners clearer visibility into why a crawler is visiting and what it will do with the content.
The updated classification system also accounts for bots that serve multiple purposes simultaneously, tracking all applicable categories rather than forcing a single label onto a crawler with mixed behavior.
Why It Matters for Site Owners
For smaller publishers, the stakes are practical. Blocking all crawlers risks losing discoverability in search results, but allowing unrestricted access means training data leaves the site with nothing in return. The new granular controls let owners permit indexing for search visibility while blocking training crawlers, a middle ground the previous single toggle could not offer.
Cloudflare frames the broader shift as a recognition that the term “AI bot” no longer describes a single behavior. As search engines themselves become answer engines that surface content directly rather than driving click-through traffic, the old assumptions about what crawlers do and what site owners get in return need a more precise framework to work from.