← Back to blogDoes blocking AI crawlers hurt your SEO?

Does blocking AI crawlers hurt your SEO?

No. Blocking the bots that scrape your content to train models - GPTBot, Google-Extended, ClaudeBot - has no effect on your Google search rankings, because none of them is the crawler that builds Google's index. The real risk is subtler: block the wrong bot and you vanish from AI answers instead, and blocking rarely stops your content being cited anyway.


The confusion buried in this question is that "AI crawler" describes three different jobs, and only one of them has anything to do with how you rank. The crawler that decides your position in Google Search is Googlebot, and it is untouched by any of the AI opt-out tokens publishers argue about. The bots that train large language models are separate fetchers with separate names, and telling them to stay out changes nothing about your search visibility. So blocking GPTBot, ClaudeBot or Google-Extended does not cost you a single ranking position. What it can quietly cost you is something different - your presence inside AI answers - and that is the trade-off publishers should actually be weighing.

The three crawlers publishers keep confusing

Four figures - zero Google Search ranking positions lost by blocking AI training crawlers (Google); OAI-SearchBot had passed 55% web coverage by 2026 (Hostinger); 70.6% of top-50 sites blocking ChatGPT's retrieval bot are still cited (BuzzStream, Jan 2026); and sites blocking LLM crawlers lost roughly 7% of weekly traffic within six weeks (Strategic Response study).

Almost every mistake in this area comes from treating "bots" as one category. There are three, and they do different work. The first is the classic search crawler, Googlebot, which indexes your pages so they can rank in Google Search. Block it and your SEO collapses, because you are no longer in the index. The second is the AI training crawler - GPTBot from OpenAI, ClaudeBot from Anthropic, Google-Extended, Applebot-Extended - which collects content to train or ground foundation models. The third, newest and most often overlooked, is the AI search or retrieval crawler, such as OpenAI's OAI-SearchBot and PerplexityBot, which fetches pages in real time to build and cite a live answer.

The critical fact is that these are governed independently. A robots.txt rule aimed at GPTBot does not apply to OAI-SearchBot, and neither has any bearing on Googlebot. That separation is deliberate, and it is what makes the headline answer to this question so clear.

Why blocking AI training crawlers does not touch your Google rankings

Google states plainly that Google-Extended has no effect on a site's inclusion or ranking in Google Search. Google-Extended is not even a crawler in its own right - your pages are still fetched by Googlebot, and Google-Extended only governs whether that content may then be used to train Gemini and ground Google's generative products. Blocking it withdraws your work from AI training while leaving your search visibility completely intact. That is precisely why Google built a distinct token rather than folding the choice into an existing directive.

The same logic holds across vendors. OpenAI documents GPTBot and OAI-SearchBot as separate crawlers with independent rules, and none of them is the mechanism Google uses to rank you. This is why publisher-network analyses through 2026 find no measurable impact on Google Search rankings or indexation from blocking GPTBot. If the only bots you disallow are training crawlers, your SEO is not the thing at risk.

Where blocking really does cost you: AI search visibility

Comparison of blocking a training crawler (GPTBot) vs a retrieval crawler (OAI-SearchBot) - affects Google ranking (no / no), removes you from model training (yes / no), removes you from live AI answers (no / yes), reliably stops AI citing you (no / no), and what you forfeit (content for training / citations and referrals).

The genuine downside sits one layer over from SEO. If you block the retrieval crawlers - OAI-SearchBot for ChatGPT search, PerplexityBot for Perplexity - you can remove yourself from the live AI answers those products generate, because they can no longer fetch and cite your current pages. OpenAI's own guidance is to allow OAI-SearchBot if you want to appear in ChatGPT's search results, even where you block GPTBot for training. A Hostinger analysis of billions of bot requests found OAI-SearchBot had already passed 55 per cent web coverage by 2026, so this is now a mainstream discovery surface, not a fringe one.

This is the strategic framework the sophisticated publishers have settled on: block the training crawlers that take your content for nothing, allow the search crawlers that can send citations and the occasional click. Treating the two as one switch is the expensive mistake. Blocking everything with a blanket rule protects your training data and simultaneously deletes you from the answers where audiences increasingly begin their research.

Does blocking actually stop AI from citing you?

Less often than publishers assume. A January 2026 BuzzStream study, drawn from around four million citations across 3,600 prompts, found that 79 per cent of top news sites block at least one AI training bot and 71 per cent block retrieval bots - yet blocking did not reliably reduce how often they were cited. Among the top 50 news sites blocking ChatGPT's live retrieval bot, 70.6 per cent still appeared in AI citation datasets. Roughly 70 per cent of ChatGPT citations came from sites that block its retrieval bot, and around 95 per cent came from sites that block training bots.

The reason is that models draw on far more than a live fetch of your site: training data absorbed before you blocked anything, third-party summaries, syndication, aggregators, and other sites quoting you. Robots.txt is an honour-system instruction, not an access control, and it does nothing about the copies of your content that already exist elsewhere. So the mental model of "block the bot and disappear from AI" is largely wrong. You forgo the direct read and the chance of a first-party citation, but the model can still describe you from everyone else's version of your work.

The one thing you cannot block without hurting SEO

Flow - your page is crawled by Googlebot into one index that feeds both your ordinary Search ranking (unaffected) and the AI Overviews / AI Mode answer (takes the traffic); there is no opt-out token that removes the AI answer, so blocking it means leaving Search.

There is a single important exception, and it is the one that matters most in 2026. You cannot remove yourself from Google's AI Overviews or AI Mode without also removing yourself from Google Search, because Google builds those AI answers from the same index that Googlebot fills for ordinary search. There is no token that strips you out of the AI summary while leaving your blue-link ranking in place. Blocking Google-Extended does not do it - Overviews still draw on Googlebot-indexed content.

Regulators have noticed the trap. On 28 January 2026 Google acknowledged that its existing controls do not give publishers granular separation between AI features and ordinary listings, and said it was exploring new options, under pressure from the UK's Competition and Markets Authority, which has proposed letting publishers opt out of AI Overviews and AI Mode without a ranking penalty. Until such a control exists, the honest position is that the AI feature taking the most publisher traffic is the one you cannot block at all without sacrificing search.

What blocking costs in traffic

Blocking is not free even when it leaves rankings alone, because AI search crawlers can return referrals. A 2026 study of news publishers, "Strategic Response of News Publishers to Generative AI", found that sites which blocked LLM crawlers via robots.txt lost roughly 7 per cent of weekly traffic within six weeks, visible in human browsing-panel data. A separate analysis of the top 30 publishers that blocked AI crawlers put the total traffic decline at around 23 per cent, with human traffic measured by Comscore down about 14 per cent. These figures are contested and vary by site type - a small niche publisher and a global news brand face very different maths - but the direction is consistent. Every bot you block is also a route to citation and referral you are closing.

So should publishers block AI crawlers at all?

The decision is rarely the binary it is presented as. Blocking training crawlers is defensible and costs you no SEO, but it does not stop AI systems describing you and it forfeits nothing you were being paid for in the first place. Blocking retrieval crawlers protects nothing extra and can delete you from the answers audiences now start with. And the highest-traffic AI surface, Google's own, cannot be blocked without leaving Search. For most publishers the real question is not "block or allow" but "how do we get paid for the reads we cannot prevent".

That reframing is where the monetisation layer comes in. Blocking is a permission decision; it never generates revenue. blankspace approaches the same traffic from the opposite end: it detects AI and Live Search Agent reads at the CDN edge and places contextual brand mentions into the answer, so the retrieval that robots.txt cannot reliably stop becomes a monetisable event rather than a leak. Access control and monetisation are not the same lever, and treating blocking as a revenue strategy is the category error underneath most of these questions.

Frequently asked questions

Does blocking GPTBot hurt my Google rankings?

No. GPTBot is OpenAI's training crawler and has no role in how Google ranks or indexes your pages, which is Googlebot's job. Publisher-network analyses through 2026 find no measurable effect on Google Search rankings or indexation from blocking GPTBot. The only visibility you affect by blocking GPTBot is potential use of your content in OpenAI model training, not your search position.

What is the difference between an AI training crawler and an AI search crawler?

A training crawler, such as GPTBot or ClaudeBot, collects content to train or ground a model, and blocking it removes your work from that training set. A search or retrieval crawler, such as OAI-SearchBot or PerplexityBot, fetches pages in real time to build and cite a live answer, so blocking it can remove you from AI search results. They are governed by separate robots.txt rules, which is why a rule aimed at one does not apply to the other.

If I block AI crawlers, will AI stop mentioning my brand?

Usually not. A January 2026 BuzzStream study found most top news sites that block retrieval bots are still cited in AI answers, with more than 70 per cent of blocked sites appearing in citation datasets. Models draw on prior training data, syndication, aggregators and third-party summaries, so blocking your own site does not remove the many other copies of your content the model can rely on.

Can I remove my site from Google's AI Overviews without losing search rankings?

Not with today's tools. AI Overviews and AI Mode are built from the same index Googlebot fills for ordinary search, so there is no directive that removes you from the AI answer while keeping your blue-link ranking. Google acknowledged this gap on 28 January 2026 and, under pressure from the UK's Competition and Markets Authority, said it was exploring new controls, but until one ships the AI feature is not separately blockable.

What should a publisher do instead of blocking everything?

Separate the bots by purpose. Block the training crawlers that take content for nothing if that matches your rights strategy, allow the search crawlers that can return citations and referrals, and accept that Google's AI surface cannot be blocked without leaving Search. Then focus on measuring AI reads server-side and monetising the ones you cannot prevent, since blocking protects permission but never produces revenue.