← Back to blogShould publishers block or monetise AI crawlers?

Should publishers block or monetise AI crawlers?

Blocking is no longer a free option. Academic research now puts the cost of a robots.txt block at around 7% of weekly traffic within six weeks, while citation data shows blocking does not reliably keep you out of AI answers. The workable position in 2026 is selective: block the crawlers that take and give nothing, monetise the ones that carry real audience, and treat your CDN default as the decision that actually matters.


For two years the block-or-monetise question was argued on principle. It can now be argued on evidence, and the evidence is uncomfortable for both camps. Publishers who blocked lost human traffic without gaining much protection. Publishers who allowed everything gave away a growing share of their audience for nothing. Between those two failures sits the position most publishers should actually hold: a per-crawler, per-page policy set at the edge, where the choice is not whether to open the door but what happens when something walks through it.

What the evidence now says about blocking

Three figures on blocking AI crawlers: news publishers who blocked lost around 7% of weekly visits within six weeks; 70.6% of the top 50 news sites blocking ChatGPT's retrieval bot still appeared in AI citation datasets; and roughly 79 to 82% of top news sites now block at least one AI training crawler. Earned yellow: 70.6% - the "unreliable protection" half of the thesis.

The strongest study to date is a Rutgers and Wharton working paper by Hangcheng Zhao and Ron Berman, "Strategic Response of News Publishers to Generative AI", last revised on SSRN in April 2026. Using a staggered difference-in-differences design across SimilarWeb, Semrush and Comscore data, it finds that news publishers who blocked large language model crawlers in robots.txt lost approximately 7% of weekly visits within six weeks of the block. The decline shows up in the Comscore household panel, which tracks real human browsing, so it cannot be dismissed as the mechanical removal of bot hits from a traffic count. The effect is concentrated in the largest publishers, and the authors attribute it mainly to lost brand exposure rather than lost referral clicks.

The second finding matters just as much. A BuzzStream analysis published in March 2026, covering four million citations across 3,600 prompts and drawing on Citation Labs' XOFU tracking data, found that among the top 50 news sites blocking ChatGPT's live retrieval bot, 70.6% still appeared in AI citation datasets. Blocking removed the publisher's visitors. It did not reliably remove the publisher's content from the answers.

Put together, the two results describe the worst possible outcome for a blanket blocker: measurable cost, unreliable protection. That does not make blocking wrong. It makes blanket blocking a poor default.

The three things an AI crawler can be doing

Log-scale bar chart of pages crawled per referral sent in June 2026: Anthropic at roughly 11,122 to 1, OpenAI at about 857 to 1, and Googlebot at around 5 to 1. Earned yellow: Anthropic at 11,122:1.

The purpose of the fetch determines the correct response, and the mix has shifted sharply.

Training crawlers gather content in bulk to train a model. There is no traffic back and no attribution, so the case for charging, licensing or blocking is strongest here. By June 2026, training crawlers accounted for 50.6% of AI bot traffic on Cloudflare's network, up from a small minority two years earlier.

Search and retrieval crawlers build an index an assistant queries to find and cite sources. Allowing these keeps content eligible for citation. On Cloudflare's network they had fallen to 10.7% of AI bot traffic by June 2026, which tells you how much of the crawl is no longer about being findable.

Live Search Agents fetch a specific page in real time to answer a live user question. This is the category closest to purchase intent and the one most worth monetising, because the read happens at the moment a person is making a decision.

The crawl-to-refer ratio is the quickest way to see which of these you are dealing with. Cloudflare Radar data for June 2026 put Anthropic's crawler at roughly 11,122 pages crawled per referral sent and OpenAI's at about 857 to 1, both improved on their spring figures but still orders of magnitude away from Googlebot at around 5 to 1. Cloudflare's Matthew Prince noted in June 2026 that automated requests had reached 57.5% of all HTML web traffic. A ratio in the thousands is a licensing or blocking conversation. A ratio in single or double digits is a monetisation conversation.

What changed in 2026: the default moved to your CDN

The most important development this year is that the decision is increasingly being made for you unless you make it yourself. From 15 September 2026, Cloudflare's default settings block mixed-use crawlers on any page that carries advertising, while continuing to permit search indexing. AI companies are being pushed to separate their search, agent and training crawlers or be blocked by default. The change applies to new customers, new sites created by existing customers and all existing free customers; existing paying customers can override it in the dashboard.

That is a policy decision about your inventory taken at the infrastructure layer. The first step in any 2026 block-or-monetise review is therefore not editing robots.txt. It is checking what your CDN's defaults already do on 15 September, and whether that matches what you actually want.

Cloudflare has paired the block with a Pay Per Use model that pays when AI uses content in an answer rather than when a bot fetches a page, and AWS shipped native per-request billing at CloudFront in June 2026. The practical effect is that "monetise" has stopped meaning "sign a vendor contract" and started meaning "configure a setting".

Why blocking Google is a different decision entirely

Comparison of blocking Google-Extended against blocking Googlebot: both stop training and Gemini grounding, only a Googlebot block removes you from AI Overviews, but a Googlebot block also removes you from Google Search results. Note: per the style guide, compare carries no yellow - the loss is shown by dimness, not colour.

Everything above assumes you can separate crawlers by purpose. With Google you largely cannot, which is why the July 2026 publisher revolt is a genuinely new situation rather than more of the same.

Google-Extended governs training and Gemini grounding. It does not remove your content from AI Overviews, which are assembled from the Googlebot index. Blocking Googlebot to escape AI Overviews therefore also removes you from search results. That is the trap, and it is why the question surfaced at board level in late July 2026, when Reddit, USA Today Co, Politico and Reuters were all reported to be weighing whether to cut off Google's crawlers. USA Today Co chief executive Mike Reed said it was "time to take a stand". Reported declines are severe: Politico down roughly 20% to 23% in US organic search traffic between mid-2025 and mid-2026, USA Today down anywhere from 18% to close to 50% depending on segment. Reddit, whose Google licensing agreement is reported at around $60m a year, has discussed withholding its corpus as leverage in a renewal negotiation.

The regulatory position is moving in publishers' favour, slowly. On 3 June 2026 the UK Competition and Markets Authority imposed a conduct requirement obliging Google to let publishers opt out of AI Overviews, AI Mode and Discover without affecting search rankings, plus a separate opt-out for fine-tuning and an attribution requirement. Google has nine months from that date to implement it. Until the toggle exists, blocking Google remains an all-or-nothing decision, and for most publishers the arithmetic still does not support taking it.

When blocking still makes sense

Block where the crawler takes value you cannot recover and offers nothing back. Unidentified scrapers and bots that ignore your rules are the clearest case. Block content you never want ingested for legal, contractual or editorial reasons. Block high-ratio training crawlers if you have no licensing conversation running and no intention of starting one. And block as a deliberate negotiating posture, accepting the traffic cost as the price of leverage, which is the logic several large publishers are now applying to Google.

News publishers already behave this way at scale. Studies through 2026 put the share of top news sites blocking at least one AI training crawler at roughly 79% to 82%, against under 10% among top retailers. A July 2026 survey of the top 1,000 websites found 40.9% unreadable to GPTBot and 18.4% closed to every AI crawler tested. Blocking is not a fringe position. It is simply not a free one.

When monetising makes sense

Monetise where the fetch represents a real, high-intent audience. Live Search Agent retrievals are the clearest case: the visit is happening anyway, the user is mid-decision, and content-layer advertising turns the read into revenue without altering what human readers see. Blocking that traffic forfeits the income and the presence in the answer at the same time.

This is the layer blankspace operates in, detecting Live Search Agent traffic at the CDN edge and injecting contextual brand mentions into what the agent retrieves, with revenue attributed the same day. blankspace reports a $15.02 CPM on this inventory from its own accounts; treat that as one operator's figure rather than a market benchmark, because there is no audited industry number for agent inventory yet.

The limitation both sides share: robots.txt is a request

Neither strategy works if the file is ignored. TollBit's "Pipes are Leaky" analysis found that roughly 30% of AI bot scrapes in Q4 2025 bypassed robots.txt entirely, and identified around 40 third-party scrapers reselling publisher content, including paywalled articles pulled in full from high-authority test sites. Compliance is voluntary by design.

This is the argument for moving policy from a text file to the network. A robots.txt directive asks a well-behaved crawler to stay away. A CDN rule decides what happens to the request regardless of who sent it. Enforcement, measurement and monetisation all live at the same layer, which is why the edge has become the natural control point rather than the site root. Some publishers are going further still: PPC Land reported in July 2026 that TIME had begun routing AI crawlers to a stripped-down markdown copy of its site, serving agents something different from what it serves people.

A practical decision framework

Start with what is already configured. Check your CDN's current and scheduled defaults, including the 15 September Cloudflare change if you sit behind it, before you change anything else.

Then get visibility. Classify AI traffic at the edge by bot and by purpose, because you cannot price or police traffic you cannot see, and most of it never reaches your analytics.

Then set policy by category rather than site-wide. Block unidentified scrapers and rule-breakers outright. Decide on training crawlers according to whether you want a licensing conversation, a negotiating position, or simply to keep content out of models. Allow retrieval crawlers to preserve citation eligibility, remembering that blocking them does not reliably remove you from answers anyway. Monetise Live Search retrievals on commercial, high-intent pages while holding sensitive editorial in observe-only.

Then review quarterly. Crawler identities, CDN defaults and the regulatory position have all changed materially within the past six months, and a policy set in 2025 is almost certainly wrong now.

Frequently asked questions

Does blocking AI crawlers reduce my traffic?

The Rutgers and Wharton working paper by Zhao and Berman, revised April 2026, found news publishers lost around 7% of weekly visits within six weeks of blocking, with the effect visible in human browsing panel data and concentrated among larger publishers. The authors attribute it primarily to reduced brand exposure in AI-mediated discovery rather than to lost referral clicks.

If I block a crawler, does my content stop appearing in AI answers?

Not reliably. BuzzStream's March 2026 study of four million citations found 70.6% of the top 50 news sites blocking ChatGPT's live retrieval bot still appeared in AI citation datasets. Content reaches models through syndication, aggregators, third-party scrapers and previously trained data, so a block removes your visitors more dependably than it removes your content.

Can I block AI Overviews without losing Google Search?

Not yet in most markets. Google-Extended covers training and Gemini grounding but not AI Overviews, which draw on the Googlebot index, so blocking Googlebot removes you from search as well. The UK CMA ordered Google on 3 June 2026 to provide a separate opt-out that does not affect rankings, and gave it nine months to comply, so the position should change during 2027.

What happens on 15 September 2026 if I use Cloudflare?

Default settings will block mixed-use AI crawlers on pages carrying advertising while still permitting search indexing. It applies to new customers, new sites and existing free customers; existing paying customers can override it in the dashboard. Check your configuration before the date rather than after it.

Is doing nothing a neutral choice?

No. Doing nothing means allowing everything for free while the volume grows, and automated requests already make up more than half of HTML web traffic according to Cloudflare's June 2026 figures. The safer version of inaction is observe mode: classify AI traffic at the edge without changing access, see the volume and mix, then set block or monetise policy per category once you know what you have.