← Back to blog

Cloudflare's 15 September 2026 AI crawler defaults: what publishers need to decide

From 15 September 2026, Cloudflare blocks Training and Agent crawlers by default on pages that display ads, and judges multi-purpose crawlers by the strictest rule that applies to them. Any site blocking Training will also block Googlebot, Applebot and Bingbot. Two tools shipped in late August change how publishers should prepare, and Cloudflare's own executives describe the point of all this as negotiating leverage rather than revenue.


The most useful thing a publisher can do this week is not make a decision. It is to find out which decision was already made on their behalf. The risk in this deadline is rarely the new default; it is an old setting, chosen once in a dashboard nobody has opened since, that is about to mean something materially different from what it meant when it was chosen. Cloudflare has stopped asking whether a bot is "AI" and started asking what the bot does with the page after it fetches it: index it for a search result, act on it for a person waiting right now, or absorb it into a model. Those three behaviours are now separate switches for every customer including free accounts, and on 15 September 2026 the switches acquire new default positions.

What is actually changing on 15 September 2026

Two separate changes land on the same day, and they are often reported as one.

The first is a default. Cloudflare announced on 1 July 2026, in a post by Jin-Hee Lee and Bryan Becker, that "for all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default." Cloudflare's press release extends that further, stating that the changes "will also be made for all existing free customers that have not changed their settings by September 15, 2026".

The second is an enforcement rule, and it is the one with teeth. From the same date, multi-purpose crawlers will be "allowed/blocked according to all of their behaviors", with the most restrictive applicable rule winning. Cloudflare names the crawlers it means: Googlebot, Applebot and Bingbot. If you block Training, you block them, whether you did it through the new controls or through the older one-click "Block AI bots" toggle you may have flipped a year ago and forgotten about.

The controls themselves are already live. Only the defaults and the strictest-rule enforcement wait for September.

Search, Agent and Training: how Cloudflare now classifies crawlers

Cloudflare's three configurable categories are defined by behaviour rather than by operator.

Search is "any behavior that collects or indexes your content, so it can answer questions about it later". Cloudflare's position is that this behaviour should return something: "Site owners should expect to get referral traffic or other equitable compensation as a result."

Agent is "automated behavior that is acting, usually in real time, on a person's behalf, to get something done right now". Cloudflare's own examples are chat fetch bots such as ChatGPT-User and browser-use agents such as Gemini or Claude driving Chrome. The defining feature is that a human is waiting at the other end. This is the traffic blankspace calls a Live Search Agent, and it is the category that the new ad-page default switches off.

Training is "a crawler taking your content to train or fine-tune a model", where "your data is permanently absorbed into the underlying architecture of the AI".

These three sit inside a wider taxonomy of eleven classifications in BotBase, Cloudflare's bot directory, alongside Transact, Data Collection, Security Testing, SEO, Ads Verification, Social and Link Preview, Feed Fetching, and Monitoring and Operations. Only Search, Agent and Training are configurable by all customers today.

One further change is easy to miss and matters as much as the defaults: Verified status no longer means default-allowed. A verified bot is now allowable within its category, so allowing Search does not allow the same operator's training crawler. Bots that reproduce content in full cannot hold Verified status at all.

Why ad-supported pages are treated differently

Cloudflare's reasoning is stated plainly. "An ad is a signal that a website owner meant for a person to land there and see it", the company writes, and on those pages "we treat human attention as the end goal, and keep away the bots that may prevent this attention".

That is a coherent philosophy and it is worth naming what it assumes. It assumes that a bot arriving on an ad-supported page is a loss, because it consumes the content without generating the human impression the page was built to sell. On that logic the correct response is to keep the bot out.

It is also worth noting what Cloudflare has not published. The blog post says "the pages that display ads" and the press release says "pages with ads", but no technical detail has been given on how Cloudflare determines that a given page displays advertising. Publishers with mixed inventory, or with ad slots that only fill sometimes, should treat the boundary as unclear until Cloudflare documents it, and should test rather than assume.

Why blocking Training can also block Googlebot

This is the part most likely to cause an accident.

Most leading AI companies now run separate crawlers for discovery and for training, so a publisher can allow one and refuse the other. Google does not. In Cloudflare's own report, published the same day by Arielle Weiss, Zach Albertson and Emily Lanfear, the claim is direct: "Today, Google has access to about 2x more information than leading AI companies because Google leverages a mixed-use bot that makes it difficult for customers to participate in Google's search ecosystem without also participating in Google's AI ecosystem."

Google's counter-argument, which it has made consistently, is that Google-Extended already exists for exactly this purpose: it lets a site opt out of Gemini Apps and Vertex AI training without affecting inclusion in Google Search. The unresolved part is that Googlebot itself crawls for Search including AI Overviews and AI Mode, so the line between being indexed and feeding an AI answer is not one Google-Extended draws.

Whichever reading you prefer, the operational consequence from 15 September is the same. A Cloudflare-level block is a network block, not an advisory line in robots.txt, so it cannot be quietly ignored. If your Training setting is off and Googlebot is classified as mixed-use, Googlebot does not get in. Cloudflare's stated hope is that this pressure pushes mixed-use operators to split their crawlers; the company says more than a third of crawler activity on its network still comes from bots whose intent cannot be distinguished, and that it wants that number at zero within a year.

Cloudflare has also taken that argument to a regulator. Digiday has reported chief executive Matthew Prince meeting the UK Competition and Markets Authority in London to press the case that Google should have to separate its search and AI crawlers like everyone else. The September defaults and the regulatory lobbying are the same argument delivered through two different channels.

Bot Preference Sync, and why your robots.txt matters again

The most practically useful thing Cloudflare shipped in the run-up to the deadline arrived on 21 August 2026, and it addresses a problem most publishers did not know they had: the policy declared in their robots.txt and the policy enforced at their edge had drifted apart.

Bot Preference Sync generates a site's robots.txt directly from the Search, Agent and Training categories already configured in the Cloudflare dashboard. Cloudflare has made it available across plan levels from free through Enterprise, and it is on by default for new accounts. Existing Disallow directives are preserved, written between marked BEGIN and END blocks so that hand-authored rules and generated rules can coexist in the same file without one overwriting the other.

The reason this matters is narrower than it sounds and more important than it looks. Crawler operators have long justified inconsistent behaviour by pointing at the gap between what a site says publicly and what it enforces privately. A site that blocks a crawler at the network edge while its robots.txt stays silent has, on the operator's reading, published no preference at all. Bot Preference Sync closes that gap automatically, which strengthens a publisher's position in exactly the argument that follows any dispute about access: what did you actually ask for, and where did you ask for it.

A week later, on 28 August 2026, Cloudflare opened the operator side of BotBase, adding an automated review process for bot operators submitting their crawlers to the directory. Cloudflare cited roughly a sevenfold rise in annual new-bot submissions since 2023 as the reason for automating it. For publishers this is background rather than a task, but it is the supply side of the same system: the classifications that decide what your Search, Agent and Training switches actually do are only as good as the directory underneath them, and that directory is now growing faster than manual review could keep up with.

The use content signal

Alongside the blocking controls, Cloudflare has been testing a fourth field for Content Signals, the robots.txt extension it published at contentsignals.org in September 2025. The existing three fields cover search, ai-input and ai-train. The new field, use, describes what a bot may keep and reshare after crawling, at one of three levels: immediate, meaning interact but store and reuse nothing; reference, the default, meaning index, excerpt and link back; and full, meaning summarise and reproduce.

Sites already using Cloudflare's managed robots.txt have had use=reference appended automatically, so a managed file now reads Content-Signal: search=yes,ai-train=no,use=reference.

Two things to keep straight. The same three levels are being built as an enforceable setting for Bot Management customers, where they can be combined with classifications to express rules such as allowing Search, SEO and Ads Verification bots but only to the reference level. That is enforcement. The robots.txt field is not: Cloudflare states that content-use values "signal a website owner's preference, rather than issuing blocks directly".

The gap between those two is not academic. Google's John Mueller has said publicly that the content-signal directive has "no effects whatsoever for any crawler or LLM", and Search Engine Roundtable reported that Google added content-signal to the list of unsupported directives in its robots.txt repository. A preference expressed in a file is only as strong as the willingness of the reader to honour it. A rule applied at the network edge does not have that problem.

What Cloudflare says the blocking is actually for

It would be easy to read the September defaults as the front end of a payments system. They are not, and Cloudflare's own executives are unusually straightforward about it.

Speaking to Press Gazette, Cloudflare chief strategy officer Stephanie Cohen put the purpose in terms of negotiation rather than revenue: "We've seen lots of our customers use our tools so that they can create reliable scarcity for their content, and then negotiate better deals." Cloudflare named the Financial Times, The Atlantic, Ziff Davis, Conde Nast and the Associated Press among publishers working with it on that basis. The product being sold is leverage. The money, where it arrives, arrives through a bilateral licensing deal negotiated by humans, not through a toll collected at the edge.

The state of the actual rails supports that reading. Pay Per Crawl, launched in July 2025, was still in closed beta at the time of Press Gazette's reporting, and Cloudflare has described the pay-per-use model that succeeds it as at a very early stage. The Monetization Gateway, announced on 1 July 2026 to let customers charge for any resource behind Cloudflare using the x402 protocol, remains waitlist-only more than two months later, with no published pricing and no announced general availability date. On the buy side, the only AI companies Cloudflare has named as paying partners under the pay-per-use model are Ceramic.ai, which pays when publisher content appears in its AI search results, and You.com, which pays for on-demand access to premium content. No frontier model developer has publicly agreed to pay a publisher per crawl or per answer through any of these mechanisms.

That is not an argument against the tooling. It is an argument for reading the tooling accurately. If you are a publisher large enough to be in a licensing conversation, blocking is how you get a better number out of it, and Cloudflare's own customers say it works. If you are not large enough for anyone to negotiate with you, blocking removes you from the answer and returns nothing, and the September default will do that automatically unless you intervene.

Two further caveats belong here. Cloudflare's economics are not free: the Nieman Lab coverage of the "Same Gatekeepers, New Tollbooths" report puts Cloudflare's estimated take on its pay-per-crawl marketplace at around 30%, against TollBit and Sphere.ai charging publishers nothing and billing the AI companies instead, ScalePost at roughly 15% and ProRata.ai at a 50-50 split. Those are third-party estimates rather than disclosures and should be treated as such. And enforcement is imperfect: reporting collected under the heading that the pipes are leaky has documented AI scrapers bypassing publisher protections at scale, which is the standing limitation on every control described on this page. A network block works on traffic that identifies itself and does not actively evade. Not all of it does.

For scale on what the alternative rails have actually produced, TollBit's chief executive Toshit Panigrahi told Digiday that nearly 20% of the roughly 7,000 publisher sites on its network have earned money from its AI bot paywall, in amounts ranging from hundreds of dollars to tens of thousands per month. That is vendor-reported and unaudited, and it is still the most concrete public figure available for what charging machines for access currently earns.

What publishers should do in the next nine days

A short, ordered list of checks.

  1. Find out what your current setting actually is. Open zone Security settings and look at the Search, Agent and Training controls. If your organisation enabled the legacy "Block AI bots" toggle at any point, you are in scope for the strictest-rule change even if nobody has touched the dashboard since.
  2. Decide Search separately from Training. These were one decision until now and they are two decisions from September. Blocking Training with Search allowed is a coherent position; discovering on 16 September that Googlebot has been refused is not.
  3. Check whether Bot Preference Sync is writing your robots.txt. It is on by default for new accounts. If it is active, your public robots.txt now reflects your dashboard settings, which is usually what you want but is worth confirming rather than discovering. If it is not active, decide deliberately whether you want your declared policy and your enforced policy to match.
  4. Work out which of your pages Cloudflare will treat as ad-supported. Until the detection method is documented, verify behaviour on a representative sample rather than reasoning from your own ad configuration.
  5. Make the Agent decision deliberately. Agent traffic is the human-in-the-loop read: someone asked an assistant a question and it came to your page to answer it. Blocking it protects nothing that was going to render an ad impression anyway, because the agent will not render your ad stack. It does remove you from the answer.
  6. Be honest about whether you have a negotiation to strengthen. Blocking is leverage if someone is at the table. If nobody is, it is a decision to be absent from AI answers with no compensating revenue, and it should be taken knowingly rather than inherited from a default.
  7. Record the decision and the date. These defaults have now changed twice since July 2025. Whatever you choose, write down what you chose and why, so the next change is an amendment rather than an archaeology exercise.

Blocking is a position, not a business model

The September defaults are the clearest statement yet of a philosophy that has been implicit in the infrastructure layer for a year: if a bot cannot be made to pay, keep it out. For Training crawlers, that is a defensible position and often the right one, because a training fetch returns nothing and takes something permanent.

Agent traffic is a different case, and the ad-page default treats it as though it were not. An agent read happens because a person asked a question. It is the moment your content is being used to form an answer that a human will act on. Blocking it does protect the page from a bot that will never see an ad, but it also removes you from the answer, and it converts a real audience moment into nothing at all. The choice as currently framed is between giving that read away for free and refusing it. Neither of those is revenue.

blankspace exists to add the third option. It operates at the CDN edge, where these requests already arrive and where Cloudflare's own controls act, detects and verifies Live Search Agent traffic rather than trusting a user-agent string, and monetises the read in place by placing contextual brand mentions into the response the agent is assembling. The read still happens, the answer still gets written, and the publisher earns from it. That is complementary to the classification work Cloudflare has done, not opposed to it: the categories are what make a per-behaviour commercial decision possible in the first place. What blankspace disputes is the assumption underneath the ad-page default, which is that an agent on a monetised page can only ever be a cost. It is worth adding the same caveat that applies to everyone working at this layer, including blankspace: publisher-side advertising to AI agents is contested territory, no vendor can guarantee how a model treats what it is served, and the honest instruction is to measure your own traffic by agent type before deciding anything.

Frequently asked questions

What exactly changes on 15 September 2026?

Two things. Cloudflare sets new defaults so that Training and Agent crawlers are blocked on pages that display ads while Search remains allowed, applying to new domains, new sites on existing accounts, and existing free-tier customers who have not changed their settings. Separately, multi-purpose crawlers begin to be evaluated under all of their behaviours, with the most restrictive rule winning.

Will Cloudflare block Googlebot on my site?

Only if you block Training. From 15 September, crawlers that do both search and training, which Cloudflare identifies as Googlebot, Applebot and Bingbot, are judged by the stricter setting. If Training is blocked, they are blocked, including for sites that enabled the legacy "Block AI bots" option and have not revisited it. Sites that allow Training are unaffected.

Do the new defaults apply to my existing paid Cloudflare zone?

No. The ad-page defaults apply to new domains, new sites added by existing customers, and existing free customers who have not changed their settings. Configured paid zones keep their settings. The strictest-rule enforcement, however, applies to any customer who has chosen to block Training, on any plan.

What is Bot Preference Sync and do I need to do anything about it?

It is a Cloudflare feature, announced on 21 August 2026 and available from free through Enterprise plans, that writes your robots.txt from the Search, Agent and Training settings in your dashboard, preserving your existing Disallow rules between marked blocks. It is on by default for new accounts. You do not have to use it, but you should check whether it is running, because it changes what your site publicly declares about crawler access, and that declaration is what an operator will point to in any dispute.

Will blocking AI crawlers actually make anyone pay me?

It depends entirely on whether anyone is negotiating with you. Cloudflare's chief strategy officer has described the purpose of the blocking tools as creating scarcity so publishers can negotiate better deals, and Cloudflare names large publishers doing exactly that. The automated payment rails are much less mature: Pay Per Crawl has been in closed beta, the Monetization Gateway is still waitlist-only, and the only named AI companies paying under the pay-per-use model are Ceramic.ai and You.com. For a publisher without a licensing conversation to strengthen, blocking should be understood as a decision to withdraw rather than a route to revenue.