Cloudflare's Sept 15 AI Crawler Deadline: What It Means For You

From 15 September 2026, Cloudflare blocks AI Agent and Training crawlers by default on ad-funded sites. Here's what changed and what to check now.

From 15 September 2026, Cloudflare will start blocking AI "Agent" and "Training" crawlers by default on any page that carries adverts, unless the site owner opts out. That means the bots ChatGPT, Claude and Perplexity use to fetch pages in real time during a chat could be shut out of huge parts of the ad-funded web unless publishers check one setting before the deadline.

Here's what's actually changing, why it matters for your AI search visibility, and what to do about it right now.

What changed?

Cloudflare announced the update on 1 July 2026 as part of its second annual "Content Independence Day" initiative, replacing the old blunt "block AI bots" toggle with a three-way classification system, according to Cloudflare's own blog and its changelog.

The three new categories are Search, Agent and Training. Search crawlers proactively collect and index content to answer queries later, and these include traditional search bots as well as those used by AI models. Agent crawlers act in real time on a person's behalf, the kind of bot a live ChatGPT or Claude session uses to pull a page mid-conversation. Training crawlers collect content in bulk to build or fine-tune a model, with no immediate user-facing task attached, as explained by Help Net Security.

From the deadline, the defaults tip against two of those three categories on money-making pages. Starting 15 September 2026, new domains onboarding to Cloudflare, new sites created by existing customers, and all existing free-tier accounts get updated defaults: bots classified as Training or Agent are blocked on pages that display ads, while Search remains allowed, per Cloudflare's changelog and Hosting.com.

There's a sting for anyone relying on Google, Apple or Bing's crawlers too. Multi-purpose crawlers that combine Search with Training will be allowed or blocked according to all of their behaviours, meaning crawlers such as Googlebot, Applebot and Bingbot will be blocked by customers who have chosen to block Training, according to Arcalea and Technology.org. In practice, if a site blocks Training crawlers, Googlebot, Applebot and Bingbot get blocked too, even when Search crawlers stay allowed.

Cloudflare isn't leaving site owners in the dark. Website owners who don't want the new default settings can opt out through Security settings before 15 September, and Cloudflare says it will notify customers ahead of time so they have room to review and adjust, per KankaTech.

Separately, Google has confirmed it isn't following Cloudflare's related content-signals approach at all. Google spokesperson John Mueller reminded Reddit users that Google does not use llms.txt or llms-author.txt files, which suggests Google has no plans to follow Cloudflare's content signals, as reported by Loved By AI.

Why does this matter for your business?

This isn't a niche infrastructure tweak. Cloudflare sits in front of a vast share of the web, and its defaults now decide, by default, whether AI agents can fetch your pages the moment a user asks a question in a live chat session.

If your site runs ads and sits on Cloudflare, Agent crawlers, the exact mechanism a ChatGPT or Perplexity session uses to pull your page mid-conversation, are blocked unless you say otherwise. Get this wrong and you could quietly vanish from real-time AI answers even while still ranking normally in Google.

The Googlebot wrinkle is arguably worse. Because Googlebot crawls for both Search and AI training in a single bot, sites that block Training will also block Googlebot on those pages unless they explicitly opt out, as Arcalea notes. A well-intentioned move to keep AI models away from your content could accidentally strip you out of Google Search entirely.

This lands at a moment when AI answers already dominate a growing share of queries and citation supply is struggling to keep up. AI Overviews now show up on 43% of Google searches, up from 15% a year ago, while only 6.8% of US ChatGPT desktop queries carried a citation as of May 2026, even after fivefold growth over the year, according to Something Inc. and ROI Revolution. Losing agent access on top of that squeeze makes it harder to be the source an AI system pulls in.

There's also a broader signal here about who sets the rules for AI visibility. Cloudflare is trying to build a market-level bargaining system where a site gets something in return for being indexed, such as referral traffic. Meanwhile, the largest AI platform of all, Google, has publicly said it won't play by that rulebook, leaving publishers to navigate two different sets of expectations at once.

What should you do now?

First, if your site runs on Cloudflare, log in and check your bot settings before 15 September. Open Security, then Settings, then Configure AI bot policies, and make a deliberate choice before the deadline. Doing nothing is still a decision, just not one you got to make.

Second, decide deliberately, category by category. Do you want ChatGPT and Perplexity able to fetch your pages live when a user asks about you? If yes, make sure Agent crawlers are explicitly allowed, not left to the new default. Do you want your content used to train future models? That's a separate, genuinely optional choice under Training.

Third, if you rely on Googlebot for organic rankings, be extremely careful about blocking Training crawlers wholesale. As Arcalea points out, the update presents a real challenge for mixed-use crawlers such as Googlebot, which performs both search and training functions; blocking Training may inadvertently restrict search visibility. Test this on a staging environment or a single low-traffic page before rolling changes out site-wide.

Fourth, don't assume llms.txt or similar signal files will pick up the slack with Google specifically, given Mueller's comments that Google isn't using them. Structured access controls at the network level, like Cloudflare's, currently carry more real-world weight than file-based signals that crawlers can simply ignore.

Finally, this is a good moment to check where your brand actually stands across AI platforms right now, before any settings change; running a free check, such as Sited's free audit at https://sited.online, gives you a baseline to compare against once the September defaults take hold.

Frequently asked questions

Does this affect every website on Cloudflare?

No. The new defaults apply from 15 September 2026 to all newly onboarded domains, new sites created by existing customers, and all existing free-tier accounts. Existing paid customers with saved settings keep their current configuration unless they change it.

Will this block ChatGPT and Perplexity from citing my site in normal search results?

Not directly. Search crawlers, the type used for indexing content ahead of time, remain allowed by default. It's the real-time Agent crawlers, used when a chatbot fetches your page live during a conversation, that get blocked on ad-supported pages.

Could this accidentally hurt my Google rankings?

Yes, if you block Training crawlers without care. Blocking Training crawlers also blocks multi-purpose crawlers such as Googlebot, Applebot and Bingbot, even when Search crawlers are allowed. Review this setting specifically if organic search traffic matters to you.

Does Google follow Cloudflare's content signals?

No. Google spokesperson John Mueller has said Google does not use llms.txt or llms-author.txt files, which suggests Google has no plans to follow Cloudflare's content signals.

What's the single most important thing to do before 15 September?

Log into your Cloudflare dashboard, go to Security settings, and explicitly set your preferences for Search, Agent and Training crawlers rather than letting the new defaults apply automatically.

Sources