← Back to blog

Which AI crawlers and bots visit your site? A 2026 guide to AI user agents

AI bots requesting your pages now split into four groups, not three: training crawlers, retrieval crawlers, live agents, and a growing set that will not identify itself at all. This updated reference names the current user agents, from Google-CloudVertexBot and Meta-ExternalFetcher to Bytespider and Grok, and sets out which ones actually honour the rules you set for them.


Pull a week of raw access logs and compare it with what the same site's logs looked like a year ago, and two things stand out. The first is how many new names have joined the list: PetalBot, GoogleOther, DuckAssistBot and MistralAI-User now request pages that, twelve months earlier, saw mostly GPTBot and ClaudeBot. The second is that naming the bot has stopped being the hard part. The harder question by late 2026 is which of those names can actually be trusted, because a handful of AI companies now cryptographically sign every request they send, while at least one large assistant routinely arrives looking like an ordinary Chrome browser rather than the crawler it really is.

The four things an AI bot might be doing on your site

Dozens of named AI bots now request publisher pages, but almost all of them are doing one of three declared jobs, and the job should drive your response. A training crawler gathers content in bulk to improve a model's general knowledge. A retrieval crawler builds and refreshes the index an assistant searches at the moment of a question. A live agent fetches a single page in real time because a user just asked something that needed current information, which makes it the category closest to purchase intent and the one most worth monetising.

The fourth thing a bot might be doing is not declaring any of this at all. A growing share of AI-related traffic arrives under a generic browser string, from a residential IP address, doing exactly what a training or retrieval crawler does while looking, to every standard tool a publisher owns, like an ordinary visitor. That category is covered in full elsewhere on this site; what matters here is that it changes how you should read the rest of this guide. A named user agent is a claim, not a guarantee, and two of the entries below, Bytespider and Grok, sit uncomfortably between the declared list and the undeclared one.

Training crawlers: gathering data to build a model

GPTBot is OpenAI's training crawler. ClaudeBot is Anthropic's equivalent. Google-CloudVertexBot is a newer addition to Google's declared fleet, used for training and grounding data gathered through Vertex AI rather than through Google's consumer search crawler. Meta-ExternalAgent is Meta's training and ingestion crawler, and by Cloudflare's own account it now sends the second-most requests of any bot on the open web, behind only Googlebot itself. Meta also runs a separate FacebookBot, which Cloudflare's bot directory now classifies alongside its AI training crawlers rather than purely as a link-preview fetcher. CCBot is operated by Common Crawl, whose open datasets remain widely used to train third-party models. Amazonbot, though it also supports retrieval for Amazon's own services, is documented by Cloudflare as carrying a training purpose too. Google-Extended and Applebot-Extended are not crawlers in this list at all; they are opt-out signals, covered in their own section below.

Disallowing these keeps your content out of model training, with little direct effect on whether you are cited in AI search answers, since the companies behind them mostly run separate crawlers for that job.

Retrieval and search crawlers: building the index an assistant queries

OAI-SearchBot is OpenAI's crawler for its search index, distinct from GPTBot, so a publisher can allow ChatGPT search presence while keeping content out of training. PerplexityBot indexes content for Perplexity's answer engine. Claude-SearchBot is Anthropic's retrieval crawler, controllable separately from ClaudeBot. Applebot is Apple's single crawler for Spotlight, Siri and Safari search, distinct from the training opt-out token that shares part of its name. Cloudflare's own guidance also names PetalBot, run by Huawei to power Petal search and related AI services, and GoogleOther, a Google crawler kept separate from Googlebot proper and used across several parts of Google's ecosystem. Allowing these crawlers tends to preserve your eligibility to appear and be cited in AI search answers.

Live agents and assistants: fetching one page to answer one question

ChatGPT-User is the agent OpenAI dispatches when a user's prompt requires fetching a live page. Perplexity-User is Perplexity's equivalent, and Claude-User is Anthropic's. Meta-ExternalFetcher has joined this group as Meta's user-triggered fetcher, kept distinct from the bulk training work done by Meta-ExternalAgent. DuckDuckGo runs DuckAssistBot, which crawls pages in real time to compose its AI-assisted answers, and Mistral has added MistralAI-User for the same purpose behind its own assistant. These bots activate per user question and retrieve the specific page needed to compose an up-to-date answer. Blocking them removes you from the moment a user is actively researching; monetising the read, rather than blocking it, is the alternative that captures value from that high-intent moment instead of forfeiting it.

Opt-out tokens are not crawlers: Google-Extended and Applebot-Extended

Two of the most-discussed entries on any bot list never fetch a single page. Google-Extended is a control token used in robots.txt to opt out of content being used to improve Google's generative models; ordinary crawling is still done by Googlebot. Applebot-Extended works the same way for Apple: its own documentation states plainly that it "does not crawl webpages" and exists solely to govern how content that Applebot has already collected may be used to train Apple's foundation models. That distinction became more consequential on 8 June 2026, when Apple confirmed at WWDC26 that Siri now runs partly on Google's Gemini models and rewrote its Applebot documentation to state that crawled data may train the models behind Apple Intelligence, Siri and related services. A publisher can disallow Applebot-Extended and Google-Extended without losing search visibility on either platform, because the underlying crawl and the training permission are governed separately.

The bots that do not play by their own rules: Bytespider and Grok

Two names deserve separate treatment because a declared user agent does not reliably describe what actually happens in the logs.

Bytespider is ByteDance's crawler, built to gather training data for Doubao and the wider model family behind TikTok. Its declared user agent carries the contact address spider-feedback@bytedance.com, which remains the most reliable signal that a request is genuinely ByteDance's rather than a spoofed copy, and it has a long-documented record of crawling pages that robots.txt asks it to leave alone. Its scale has been volatile rather than steadily rising: a monthly crawler report tracking Cloudflare Radar's AI Insights data put Bytespider at roughly 10.1 per cent of AI crawler traffic in May 2026, a level some analysts read at the time as a sign it would soon overtake GPTBot, only for the same tracking to show it falling back to 7.3 per cent in June and 4.7 per cent in July. The lesson is not that Bytespider has become harmless; it is that any single month's ranking of AI crawlers is a snapshot, not a trend line, and policies should be reviewed on that basis rather than set once and forgotten.

Grok, xAI's assistant, is the more difficult case, because its declared crawlers, including user agents such as GrokBot, xAI-Grok and Grok-DeepSearch, are not what most of its retrieval traffic actually presents. Behavioural reporting through 2026 has repeatedly found Grok's page reads arriving under outdated Chrome and mobile Safari strings, generic HTTP client signatures, and rotating residential IP addresses, which makes a robots.txt rule aimed at its declared name largely ineffective against the traffic that matters. A disallow directive only works if the bot identifies itself and chooses to obey; Grok frequently does neither.

Why the user-agent string alone is no longer the whole story

Every name in this guide is a string a bot declares about itself, and a string can be spoofed or simply ignored, as Bytespider and Grok both show in different ways. Reliable identification has traditionally meant checking a request against the bot owner's published IP ranges and reverse DNS records, and that check still matters. What has changed in 2026 is that a real technical alternative now exists for the crawlers willing to use it. Web Bot Auth, an application of the IETF's RFC 9421 HTTP Message Signatures standard, lets a bot operator cryptographically sign every outbound request, so a publisher's server or CDN can verify with genuine certainty whether a request claiming to be a given company's agent was produced by that company's private key. Cloudflare integrated the mechanism into its Verified Bots Programme in July 2025, and by 2026 Google's AI-browsing agent and OpenAI's ChatGPT agent were both signing requests this way, with AWS WAF, Vercel, Shopify and Akamai adding support on the receiving end. A chartered IETF working group has been aiming for standards-track publication, but publishers do not need to wait for the paperwork: the signed share of total claimed AI traffic is still a minority, so the practical approach is to run cryptographic verification alongside the older reverse-DNS and IP-range checks rather than replacing one with the other.

The commercial stakes of getting this right are rising. A September 2026 analysis of Cloudflare's own robots.txt data, drawn from a network snapshot taken on 31 August 2026, found that publishers are not simply blocking "AI" as a category; they are blocking training crawlers far more often than they block the retrieval and user-action bots run by the same companies. Bytespider carried a disallow-to-allow ratio of roughly 5.8 to 1 across the sites sampled, ahead of Meta-ExternalAgent at about 4.3 to 1 and ClaudeBot at about 2.4 to 1, while ChatGPT-User sat close to parity at roughly 1.1 disallows for every allow. That pattern only holds up, and only stays honest, if the bot on the other end of the request is reliably who it claims to be.

How to see which bots are actually visiting you

Because most of the crawlers and agents named above do not execute JavaScript, none of them appear in client-side analytics tools such as Google Analytics, and the ones that deliberately disguise themselves, like Grok's retrieval traffic, will not appear as a distinct entry anywhere in a standard bot report either. Seeing the full picture means measuring below the browser, at the server or CDN edge, where every request is visible and can be checked against IP ranges, reverse DNS and, increasingly, a cryptographic signature rather than a trusted header. blankspace provides this as its analytics layer, classifying each AI request as a live search agent, training crawler, search crawler or unclassified traffic, attributing it to its owner where that is verifiable, and showing volume and page-level breakdown before any JavaScript would have run. That visibility is the basis for deciding, per bot and per page, whether to block, allow or monetise what is reading your content.

Frequently asked questions

What is the difference between GPTBot and ChatGPT-User?

GPTBot is OpenAI's training crawler, which gathers content in bulk to improve model knowledge over time. ChatGPT-User is the live agent that fetches a specific page in real time to answer a particular user's question. The two are declared separately in OpenAI's documentation and can be allowed or blocked independently in robots.txt.

Is Google-Extended a crawler?

No. Google-Extended is a control token used in robots.txt to opt a site out of having its content used to improve Google's generative models; it never fetches a page itself. Ordinary crawling is done by Googlebot, and the newer Applebot-Extended works the same way for Apple, governing training use of content that Applebot has already collected rather than crawling anything independently.

Can I allow AI search but block AI training?

Yes, for any company that runs separate crawlers for the two jobs. OpenAI splits GPTBot from OAI-SearchBot, Anthropic splits ClaudeBot from Claude-SearchBot, and Meta now splits its bulk training crawler, Meta-ExternalAgent, from its user-triggered fetcher, Meta-ExternalFetcher. Each pair can be allowed or disallowed independently in robots.txt.

Why doesn't a robots.txt block actually stop Bytespider or Grok?

For Bytespider, a robots.txt block generally does work against requests carrying its declared user agent, but the crawler has a documented history of also fetching disallowed pages regardless, which is why many publishers pair a robots.txt rule with server or CDN-level enforcement. Grok is the harder case: its retrieval traffic frequently arrives under ordinary browser user agents and residential IP addresses rather than its declared crawler names, so a rule written against GrokBot or xAI-Grok never sees the request it was meant to stop. A disallow directive only works if the bot identifies itself honestly and chooses to comply, and neither is guaranteed here.

How do I know a bot is really who it claims to be?

Traditionally, by checking the request against the operator's published IP ranges and reverse DNS records rather than trusting the user-agent header alone. Since 2025, a growing number of AI crawlers, including Google's AI-browsing agent and OpenAI's ChatGPT agent, also cryptographically sign each request under the Web Bot Auth standard, which a server or CDN can verify against the operator's published public key. Signed traffic is still a minority of the total, so the two verification methods, legacy and cryptographic, should currently run side by side.