Cloudflare is forcing AI companies to separate search bots from training bots
A new deadline for mixed-use crawlers
Cloudflare has been steadily tightening the rules for AI crawlers since it flipped from an opt-out to an opt-in blocking model in mid-2025 — the first major infrastructure provider to do so.
The next step lands September 15, 2026: Cloudflare will start blocking so-called 'mixed-use' AI crawlers by default on any ad-monetized page, unless the crawler's operator explicitly separates search-indexing traffic from AI-training and autonomous-agent traffic.
Why the split matters
Until now, a single crawler identity could quietly do both jobs — index a page for a company's AI-powered search product and also harvest it as training data — with no way for a publisher to allow one and block the other.
Forcing that split gives site owners a real lever: keep the crawler that gets your content surfaced in AI search results, block the one that's just scraping for model training, without losing visibility in the process.
Where this is headed
It comes bundled with a pay-per-crawl licensing option, which is the more consequential part long-term — it starts turning 'allow or block' into an actual commercial negotiation between publishers and AI companies, closer to how syndication deals already work.
Worth noting separately: robots.txt itself remains purely a voluntary signal with no enforcement behind it. Cloudflare's blocking happens at the CDN/server level regardless of what a crawler's operator claims to respect.
Based on reporting from TechCrunch
Is your own content set up for stories like these?
Run a free AI Readiness check — retrieval, extractability, schema.org signals, and a prioritized rewrite brief, scored the way an AI assistant actually reads your page.