4 min read

Cloudflare splits AI traffic into three controllable types

Cloudflare is separating AI traffic into Search, Agent, and Training categories, with new controls, defaults, BotBase visibility, and content-use signals.

Image: Hacker News

Cloudflare is giving website owners finer control over automated traffic, replacing its broad “Block AI Bots” approach with separate controls for Search, Agent, and Training crawlers.

The company introduced its first Content Independence Day one year ago, arguing that the traditional web bargain—crawlers receive content and publishers receive referral traffic—no longer held when AI systems collected content without sending value back. It responded with a one-click blocking option and a Pay-Per-Crawl marketplace.

Cloudflare now says website owners need more flexibility than blocking all automation. For smaller sites, blocking every AI crawler can also mean disappearing from search, creating a trade-off between discoverability and content protection.

Cloudflare’s three AI traffic categories

The updated taxonomy classifies automated traffic by what it does on a site:

  • Search: Crawling or indexing content to answer questions later. Cloudflare says this behavior should generally generate referrals or other equitable compensation.
  • Agent: Acting in real time on behalf of a person, including chat-fetch bots such as ChatGPT-User and browser-use agents such as Gemini or Claude driving Chrome.
  • Training: Collecting content to train or fine-tune a model, permanently absorbing the data into the model’s underlying architecture.

A crawler can have more than one classification. Cloudflare is urging operators to separate automation by purpose, allowing site owners to understand why a bot is visiting and apply more precise access rules. Other classifications include Transact, Data Collection, Security Testing, SEO, Ads Verification, Social / Link Preview, Feed Fetching, and Monitoring & Operations.

Recommended reading

DeepSeek pauses fundraising after leaked compute comments

The new Search, Agent, and Training controls are available to all Cloudflare customers, including those on the Free tier. Cloudflare announced the updated options on July 1, 2026.

New defaults arrive September 15

On September 15, 2026, new domains joining Cloudflare will have Training and Agent crawlers blocked by default on pages displaying ads. Search will remain allowed by default, reflecting its role in directing visitors back to a site.

Multi-purpose crawlers will be governed by all of their applicable behaviors. As a result, crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who choose to block Training, even when those crawlers also perform Search functions.

Website owners can opt out of the new defaults in their Security settings before September 15. Cloudflare says it will continue notifying customers ahead of the change.

BotBase adds bot visibility for Enterprise

Cloudflare is also launching BotBase, a searchable database of known bots and agents for Enterprise Bot Management customers. The dashboard provides each bot’s classifications, lets customers filter traffic from a specific bot, and exposes its detection ID for use in Security rules.

Cloudflare plans to expand BotBase later this year from a visibility tool into a control center for automated content access.

The company is also adding rules based on how bots use content after retrieving it:

  • Immediate: Interact with content but store or reuse nothing.
  • Reference: Index, excerpt, and link back. This is the default.
  • Full: Summarize and reproduce content.

These preferences can be combined with classifications—for example, allowing Search, SEO, and Ads Verification bots only at the reference level.

Cloudflare is testing a new use field in its Content Signals extension for robots.txt: use=immediate, use=reference, or use=full. Managed robots.txt files that already include search=yes,ai-train=no will now also receive use=reference.

Cloudflare says it will track content use in BotBase and remove a bot’s Verified status if it abuses these signals. Bots that reproduce content in full cannot currently receive Verified status.

The meaning of Verified is changing as well. Verified bots will no longer be considered automatically allowed; their relevant category must also be permitted. Cloudflare says it is opening and making more transparent the verification process, which requires operators to represent themselves honestly and avoid abusing the access that status provides. The company is also experimenting with “transitive trust” for bots operated through third-party platforms, using the existing Forwarded header as part of the proposal.

Ava Chen

AI Editor

Ava covers the rapidly evolving world of artificial intelligence, from foundational models and research labs to the real-world economics of intelligence. With a background in computational linguistics, she cuts through the hype to find out what actually works. She firmly believes that benchmarks are just marketing until reproduced in the wild.

via Hacker News

/ Keep reading