Cloudflare's New AI Crawler Rules Will Reshape How You Get Cited

From 15 September 2026, Cloudflare blocks mixed-use AI crawlers by default on ad-supported pages, unless site owners opt back in. Here's what it means.

Cloudflare will block "mixed-use" AI crawlers by default from any page that carries adverts, starting 15 September 2026. In practice, bots that blend search, agent use and AI training will lose automatic access to monetised web pages unless the site owner opts back in, forcing AI companies to separate search crawling from training and pay publishers when their content actually gets used.

For anyone tracking their brand's visibility in ChatGPT, Gemini or AI Overviews, this is the biggest structural change to how AI systems access web content since crawling began at scale. It will not show up as a ranking drop or a citation league table, but it will quietly determine which sites AI models can even see.

What changed?

Cloudflare announced the policy on 1 July 2026, giving the AI industry a hard deadline to separate crawlers used for traditional search from those used for AI agents and model training. As TechCrunch reported, starting on 15 September 2026, Cloudflare's default settings will block mixed-use crawlers from any pages that host ads, unless the site owner adjusts the settings otherwise.

The new defaults will not touch every site overnight. According to TechCrunch, the changes will apply to new Cloudflare customers, new sites set up by existing customers, and all existing free customers. Existing paying customers who have already configured their own crawler settings will not be switched over automatically, according to AI Chat Daily.

Alongside the block, Cloudflare is relaunching its bot payment scheme. As TechCrunch explains, the company's earlier Pay Per Crawl marketplace, which let websites charge AI bots for scraping, is evolving into "Pay Per Use," aiming to pay publishers when their content creates value rather than simply when it's fetched. The first two partners testing this model are Ceramic.ai and You.com: when a publisher opts in, they're paid when their content appears in Ceramic's AI search results or when You.com accesses a piece of their premium content.

The move follows clear frustration inside Cloudflare about how little traffic AI companies send back in exchange for what they crawl. Data cited by Let's Data Science shows how lopsided the exchange has become: Cloudflare's 2025 crawl-to-referral ratios put Google at 14:1, OpenAI at 1,700:1 and Anthropic at 73,000:1, arguing the historic crawl-for-traffic bargain has broken down for AI crawlers.

Cloudflare CEO Matthew Prince framed the change as overdue given how the internet's traffic mix has shifted. Per Technology.org, Prince said: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge," referring to the recent milestone where bots surpassed human traffic online for the first time.

Google is the elephant in the room. Its crawler blends search indexing with AI training, and Cloudflare's messaging clearly targets that setup without naming it outright. Technology.org reports that Cloudflare calls out the world's largest search engine for holding roughly twice the information access of other AI companies, since staying discoverable often means accepting AI use too. Google does offer a workaround: Google-Extended lets site owners opt out of having content used for training and AI products such as Gemini Apps and Vertex AI, without affecting inclusion in Google Search, though its flagship Googlebot still crawls for Search, including AI features like AI Overviews and AI Mode.

Why does this matter for your business?

This is not a niche infrastructure story. Cloudflare sits in front of a substantial slice of the web, so its defaults effectively become policy for a huge number of sites without their owners lifting a finger. As The Next Web puts it, the company sits in front of a large share of the world's web traffic.

If your website runs display advertising and sits on Cloudflare, the crawlers that currently feed answers into ChatGPT, Perplexity or AI browsers may simply stop reaching your ad-supported pages from September, unless you or your web team actively re-enable access. That's a direct threat to AI visibility for exactly the kind of publisher-style content, guides, comparisons and explainers, that brands rely on to get cited in generative answers.

The flip side is leverage. Cloudflare is giving publishers a genuine negotiating position for the first time, rather than a binary choice between blocking everything or giving content away for free. Let's Data Science notes that major outlets are already backing the shift, with reporting naming organisations including The Associated Press, Time, The Atlantic and Reddit as participants in publisher advocacy.

There's also a data-quality upside buried in the announcement. Cloudflare says much of the current crawling activity is wasted anyway. Per TechCrunch, Cloudflare's data suggests over 50% of crawl traffic from AI crawlers is spent re-fetching unchanged pages, meaning a chunk of the current crawl-to-referral imbalance is inefficiency rather than genuine value extraction.

None of this happens in isolation. It lands amid a wider standoff between publishers, regulators and AI platforms over who gets to read the web for free. The Next Web reports that the UK is forcing Google to let publishers opt out of AI search without losing their ranking, and news publishers are suing OpenAI over training.

What should you do now?

First, find out whether you're on Cloudflare and what your current bot settings actually allow. Many site owners have never touched these settings and won't know they're about to be switched from "open by default" to "closed by default" on any page carrying ads.

Second, treat 15 September as a real planning date, not background noise. Decide deliberately which AI crawlers you want reaching your content, rather than letting a default setting make that call for you. If AI citations already drive meaningful referral traffic or brand mentions, blocking indiscriminately could cut off a channel you didn't realise you depended on.

Third, separate your thinking about "search visibility" from "AI training exposure." A crawler that helps you get cited in a live AI answer is doing something different from one that hoovers up your content to train a future model with no attribution at all. Cloudflare's whole policy is built around forcing that distinction into the open, and your content strategy should make the same distinction.

Finally, this is a good moment to check where your brand currently stands in AI-generated answers, since the crawler landscape you're being cited from is about to shift under your feet. Running a free check, such as the one available through Sited's audit at https://sited.online, gives you a baseline before the September changes take effect, so you can tell whether any drop in visibility later in the year is down to Cloudflare's new defaults or something else entirely.

Frequently asked questions

What is a "mixed-use" crawler?

It's Cloudflare's term for a bot that combines multiple jobs at once, typically search indexing, AI agent retrieval and model training, under a single crawler identity. As Proxycove explains, this is how Cloudflare refers to bots that collect data for two purposes simultaneously: search indexing and training AI models.

Will this block ChatGPT, Gemini or Perplexity from citing my site?

Not automatically for pure search-and-answer crawlers, but any crawler that Cloudflare classifies as mixed-use will lose default access to your ad-supported pages unless you opt back in. The practical effect depends heavily on how each AI company's crawlers are classified.

Does this affect Google Search rankings?

No. Cloudflare's plan keeps standard search indexing working as before; the change specifically targets crawlers that also harvest content for AI training or agent use. Google's own opt-out tool, Google-Extended, already lets sites block AI training use without losing Search inclusion.

What is "Pay Per Use" and how is it different from "Pay Per Crawl"?

Pay Per Crawl let publishers charge a fee every time a bot fetched a page. Pay Per Use goes further, aiming to pay publishers when their content creates value, not just when it's fetched, tying payment to actual use in an AI answer rather than a simple fetch.

Should small businesses worry about this if they're not big publishers?

Yes, if you run a content-heavy site with advertising and rely on AI citations for discovery. Even without ad revenue, it's worth checking your Cloudflare crawler settings now, since the classification of "mixed-use" bots could affect how future AI products access your content regardless of your monetisation model.

Sources