Dimly lit server room, rows of network racks with status lights, shallow depth of field

News: Cloudflare announced on 1 July 2026 that from 15 September 2026, new domains joining its platform will receive updated defaults: bots classified as "Training" and "Agent" will be blocked on pages that display advertising, while the "Search" category remains allowed (Cloudflare, "Your site, your rules", 1 July 2026).

The essentials in 20 seconds

  • The line going around, "Cloudflare will block Googlebot by default", does not match Cloudflare's text: the announced default leaves the Search category allowed.
  • The genuinely documented point sits elsewhere: the legacy "Block AI bots" option used to exclude mixed-purpose crawlers. It will stop excluding them.
  • Practical consequence: a site that ticked that box months ago changes behaviour without anyone touching it.
  • For AI visibility, the decisive category is not Training but Agent, which covers real-time page fetches.

An infrastructure announcement rarely reads as an editorial subject. This one is, because it shifts a decision most sites have never consciously made: who is allowed to read your pages, and for what purpose.

Since the 1 July announcement, one reading has taken hold across the trade press: on 15 September, Cloudflare will block Googlebot by default, and a fifth of the web will drop out of search results. Cloudflare's documentation describes something narrower, and more insidious.

The three categories Cloudflare created

The core of the change is not a block, it is a taxonomy. Cloudflare has stopped treating "AI bots" as one bloc and now sorts them by intent.

CategoryCloudflare's definitionDefault from 15 September
Search"Any behavior that collects or indexes your content, so it can answer questions about it later"Allowed
Agent"Automated behavior that is acting, usually in real time, on a person's behalf"Blocked on ad-bearing pages
Training"A crawler taking your content to train or fine-tune a model"Blocked on ad-bearing pages

This split carries a consequence Cloudflare owns: a crawler is allowed or blocked "according to all of their behaviors". Googlebot, Applebot and BingBot each do several things at once. They index for search and collect for training, under the same user agent. That is what a mixed-purpose crawler is, and that is where everything turns. We described the same single-agent problem in our analysis of Applebot and visibility inside Siri.

Why "Googlebot blocked by default" is a fragile reading

The announced default for new domains blocks Training and Agent on ad-bearing pages, and Cloudflare states in the same sentence that Search "will remain allowed by default". Googlebot sits in Search. On that basis, the shorthand "Googlebot blocked by default" does not hold.

But Cloudflare also writes, in its technical documentation, that mixed-purpose crawlers "will also be blocked by all configurations to block AI training". The new default is a configuration that blocks training. The two sentences meet, and the public documentation does not say explicitly which one prevails on a fresh domain.

What we are not claiming: from the public sources available on 10 August 2026, we do not know whether a new domain will see Googlebot blocked on its ad-bearing pages on 15 September. Cloudflare's two formulations point in different directions and the company has published no arbitration between them. We would rather flag that grey area than pick the most spectacular version, as part of the coverage has done.

The real trap: a box ticked six months ago

Here is the point that carries no ambiguity, and it concerns sites that already exist.

Many publishers enabled Cloudflare's legacy "Block AI bots" option at some stage. One click, a feeling of control, then on to something else. That setting, the documentation notes, excluded mixed-purpose crawlers: Googlebot kept coming through. From 15 September, that exclusion disappears, and the old checkbox starts covering the crawlers it used to spare.

The change requires no action on your part. That is exactly what makes it dangerous: there will be no notification in your analytics, no alert in Search Console, just an erosion of crawling on monetised pages, attributed a few weeks later to an imaginary algorithm update.

Not sure what is ticked inside your Cloudflare? That is true of most sites we audit. We look at your crawl configuration and your actual visibility in AI engines, and tell you what is blocking. .

The category nobody is watching: Agent

Coverage has concentrated on Training, because that is the political subject of the moment. For your visibility, the category that counts is Agent, and it is blocked by default on ad-bearing pages.

Cloudflare defines Agent as automated activity acting in real time on a person's behalf, and explicitly names chat fetch bots. That is the mechanism by which an assistant retrieves your page while a user is asking their question, in order to cite it in the answer. Blocking Agent is not refusing to feed a model: it is refusing to be cited at the exact moment somebody is looking for what you sell.

The distinction is structural for generative search. Training and real-time citation are two separate taps, and you can close the second while believing you closed the first. We covered the wider debate on blocking crawlers in our piece on publishers versus Common Crawl.

What to do before 15 September 2026

  1. Open the Security settings of your Cloudflare zone. Not to change anything, to observe. The question you need to be able to answer is simple: is AI bot blocking active on this domain, yes or no?
  2. Treat the legacy "Block AI bots" box as debt. If it is ticked, somebody ticked it, one day, for a reason. Recover that reason before 15 September, or untick it.
  3. Decide separately for Agent and for Training. These are two different judgements: one concerns your presence in AI answers, the other the training of models. Handling both in a single gesture is precisely the reflex Cloudflare has just made expensive.
  4. Identify your ad-bearing pages. The default applies only to them. If you monetise your content pages, then your content pages are in scope, which means the pages carrying your content strategy.
  5. Annotate the date in your reporting. Drop a marker on 15 September 2026. Without it, a crawl drop in October will be blamed on Google rather than on a checkbox in your own configuration.

Our take: the CDN becomes an editorial gatekeeper

What is at stake goes beyond a setting. Cloudflare, which writes no content, ends up deciding by default who may read everyone else's. For a publisher, that means part of your visibility policy is now defined inside an infrastructure console, by people who do not think of themselves as search professionals.

Our position fits in one sentence: the risk on 15 September is not that Cloudflare blocks Googlebot, it is that thousands of sites discover in October that they held an opinion on AI crawlers without knowing it. The right response is not to unblock everything in a panic, it is to open the settings page and decide once, knowingly. That is the same discipline we apply to on-page SEO: configuration you did not choose is configuration working against you.

What this article does not cover

We give no step-by-step procedure. The Cloudflare interface evolves and a stale screenshot does more damage than no instruction at all. Refer to Cloudflare's documentation at the moment you act.

We do not measure the true size of the affected estate. The default targets new domains, and several secondary reports extend it to existing free-tier accounts. We did not find that extension worded that way in Cloudflare's primary sources, so we do not repeat it as our own.

We do not cover pay-per-crawl monetisation. Cloudflare is separately developing mechanisms to charge crawlers for access. That is another subject, on another timeline.

Frequently asked questions

Will Cloudflare block Googlebot by default on September 15, 2026?
Not according to Cloudflare's own wording. The announced default for new domains blocks the Training and Agent categories on pages that display ads, and states that the Search category remains allowed by default. Googlebot belongs to the Search category. Cloudflare also writes that mixed-purpose crawlers are blocked by configurations that block Training, without publicly stating how the two rules interact on a new domain. We flag that grey area rather than settling it on Cloudflare's behalf.
What changes for the legacy "Block AI bots" option?
This is the clearest part of the announcement. Cloudflare's documentation states that mixed-purpose crawlers combining Search and Training will also be blocked by all configurations that block AI training, including the legacy Block AI bots option. It notes that this legacy setting previously excluded those mixed-purpose bots. In other words, a site that enabled that checkbox in the past will see its behaviour change without anyone touching anything.
What are the three bot categories Cloudflare defines?
Cloudflare defines Search as any behaviour that collects or indexes your content so it can answer questions about it later, Agent as automated activity acting in real time on a person's behalf, such as chat fetch bots, and Training as a crawler taking your content to train or fine-tune a model. A single crawler can fall into several categories at once.
How do I check or change my settings before September 15?
Cloudflare states that customers can mark their preference in their zone Security settings at any time leading up to September 15, 2026, confirming that they want no changes to Training crawlers that also crawl for Search purposes. The useful action is therefore to open those settings, observe the actual state of the configuration, and decide explicitly rather than inheriting a default.

Related reading

Editorial note. Disclosure: Cicéro is an SEO and GEO content agency. This analysis is editorial, is not sponsored, and carries no commercial relationship with Cloudflare or Google. The cicero.studio site uses Cloudflare Turnstile on its forms.

Sources

Alexis Dollé, founder of Cicéro
Alexis Dollé
CEO & Founder

Growth and SEO content strategist, I founded Cicéro to help businesses build lasting organic visibility, on Google and in AI-generated answers alike. Every piece of content we produce is designed to convert, not just to exist.

LinkedIn

Is your content ready for AI search?

Cicéro audits your visibility on Google, ChatGPT and Perplexity, then produces the content that makes you citable. Book your free diagnostic.