← Back to blog

What is ClaudeBot, and what does it mean for publishers?

ClaudeBot is Anthropic's web crawler, the bot that collects public web content to help train the Claude models. Publishers can allow or block it independently in robots.txt, but the sharper question is economic: ClaudeBot currently reads more pages and returns fewer visitors than any other major AI crawler.


A line ending in "ClaudeBot/1.0" in your server logs is Anthropic's training crawler at work, fetching pages so their text can help build the next generation of Claude models. It is a training bot, not a search engine and not the thing that fetches a page when someone asks Claude a live question - those are separate Anthropic user agents with separate jobs. ClaudeBot exists to gather public web content at scale, it identifies itself openly, and it honours robots.txt, so any publisher can exclude it. The decision that actually matters is not whether you can block it, but whether blocking, allowing, or monetising the read is the right call for your business - because ClaudeBot takes a great deal from the open web and, on the current numbers, sends almost nothing back.

What is ClaudeBot?

ClaudeBot is the automated web crawler operated by Anthropic, the company behind the Claude family of AI models. Its purpose is to collect publicly available web content that may be used to train and improve those models. It behaves like a conventional search-engine crawler in mechanical terms - it requests pages, reads the HTML, and moves on via links - but the content it gathers is destined for a model training corpus rather than a searchable index that returns clicks.

Anthropic runs more than one bot, and this is the single most important thing for publishers to understand. Alongside ClaudeBot, Anthropic has historically operated legacy user agents including anthropic-ai and Claude-Web, and it now documents a three-bot framework covering distinct, current use cases. Each bot carries a different user-agent string, does a different job, and triggers a different consequence when you block it. Treating "the Claude bot" as one thing is the most common mistake in an AI-crawler robots.txt policy.

ClaudeBot vs Claude-User vs Claude-SearchBot: which Anthropic bot is which?

Anthropic updated its crawler documentation in 2026 to formally separate three named agents, and the distinction changes what blocking each one actually does.

ClaudeBot collects public web content that may be used to train Anthropic's generative models. If you block ClaudeBot in robots.txt, Anthropic says it will exclude your site's future content from its training datasets. This is the bot to target if your objection is specifically to being used as training data.

Claude-User handles user-initiated fetches. When a person asks Claude a question that requires reading a live page, Claude may retrieve that page using the Claude-User agent. This is not bulk training crawling - it is a single fetch on behalf of one user in one conversation. Blocking Claude-User stops Claude from pulling your pages into individual user answers.

Claude-SearchBot indexes content to improve the quality and relevance of Claude's search features. It is the closest analogue to a traditional search crawler in Anthropic's fleet, because the content it gathers feeds a retrieval layer that can surface and cite your pages. Blocking Claude-SearchBot removes you from that search index.

The controls are independent. Blocking ClaudeBot does not block Claude-SearchBot or Claude-User, and vice versa. That independence is the whole point: a publisher can refuse to be training data while still allowing Claude to find and cite them in a live answer, or make the opposite choice. Anthropic also runs a separate agent for its Claude Code developer tool. The practical implication is that a blanket "Disallow" aimed at one string leaves the others free to keep reading.

Does ClaudeBot obey robots.txt, and how do you block it?

Yes. Anthropic states that ClaudeBot, Claude-User and Claude-SearchBot all honour robots.txt directives, including the non-standard Crawl-delay directive that lets you slow a bot down rather than ban it outright. Blocking is done per user agent. To exclude the training crawler you add a block targeting ClaudeBot; to exclude search indexing you target Claude-SearchBot; to exclude live user fetches you target Claude-User. Because you may still want to appear in Claude's cited search results while refusing to be training data, most publishers should decide bot by bot rather than issuing one catch-all rule.

A worked example: a "Disallow: /" under a "User-agent: ClaudeBot" block tells Anthropic to keep your future content out of training, but leaves Claude-SearchBot and Claude-User untouched unless you add blocks for them too. There is no single switch that turns off "Claude" in one line, which is exactly why the three-bot split matters.

Why blocking ClaudeBot by IP address does not work reliably

Some publishers try to stop AI crawlers at the network level by blocking IP ranges. With ClaudeBot this is fragile. Anthropic does not publish dedicated crawler IP ranges, and its bots operate from public cloud-provider addresses rather than a fixed, documented block. Anthropic's own guidance warns that blocking the IP addresses its bots use may not work correctly, because doing so can also prevent the bot from reaching your robots.txt file in the first place - and if the crawler cannot read the file, it cannot honour the rules in it. The reliable, supported control is robots.txt by user agent, not an IP firewall. Verifying the crawler by IP is likewise harder than with Googlebot, which publishes verifiable ranges; ClaudeBot's honesty rests on its user-agent string and Anthropic's stated behaviour rather than a published address list.

How much does ClaudeBot take, and how much does it give back?

This is where ClaudeBot becomes a monetisation question rather than a permissions one. Cloudflare tracks a crawl-to-refer ratio - how many pages a company's crawlers request for every one human visitor that company sends back to the web. In the week of 25 May to 1 June 2026, Anthropic's crawl-to-refer ratio stood at roughly 11,122 to 1 on Cloudflare Radar's data. That had improved from about 13,528 to 1 in April, but it remained the worst ratio of any major AI operator. For comparison, OpenAI sat around 857 to 1 and traditional Googlebot around 5 to 1 over comparable periods.

The gap is structural, not incidental. Google's crawler feeds a search product that has historically returned clicks to publishers, so its ratio is close to balanced. Anthropic, for most of ClaudeBot's life, has run a training-heavy crawler without a mass-market search product that routes traffic back, so the read is close to one-directional. Claude does now cite sources in its answers, and search features are growing, but the volume of referral traffic those citations generate is a fraction of the volume of content being read. On raw crawl share, ClaudeBot is also enormous: Anthropic's combined AI crawling footprint was roughly level with OpenAI's in mid-2026, and in May 2026 GPTBot narrowly overtook ClaudeBot on share of AI crawl requests after ClaudeBot had led earlier in the year.

That asymmetry is the reason a growing share of large publishers now block Anthropic at least partially. On one 2026 sample of 107 leading sites, 39 of them (around 36 per cent) blocked at least one Anthropic crawler by mid-June, with ClaudeBot blocked by 38 and Claude-SearchBot by only 19. The gap between those two numbers is telling: publishers are more willing to refuse training ingestion than to refuse search indexing, because search retrieval at least carries the possibility of a citation and a click.

Should publishers block ClaudeBot, or monetise the read?

Blocking is a blunt instrument. It stops the training use you may object to, but it also removes your content from a product millions of people now use to answer questions, and it does nothing to capture value from the reads that continue while your policy propagates. The more useful framing is a decision matrix rather than a yes or no. If your priority is keeping your work out of model training, block ClaudeBot specifically and leave Claude-SearchBot open so you remain eligible to be cited. If your priority is brand presence in AI answers, keep Claude-SearchBot and Claude-User open and accept the training read as a cost of visibility. If your priority is revenue, blocking captures nothing - a blocked crawler pays you the same as an unblocked one, which is zero.

This is the point blankspace was built around. The crawl-to-refer numbers describe a market failure: content is read at industrial scale and returns almost no traffic, so the traditional monetisation model - a human landing on a page and seeing an ad - never fires. blankspace operates at the CDN edge, where the read actually happens, to detect Live Search Agent and LLM traffic and turn that read into revenue rather than an uncompensated extraction. Blocking ClaudeBot is a way to opt out. Monetising the read is a way to be paid for the traffic you are already receiving but cannot currently see or bill for. For most publishers the honest answer is a blend: block the uses you object to on principle, and monetise the ones you cannot realistically stop.

Frequently asked questions

Is ClaudeBot the same as the bot that reads a page when I ask Claude a question?

No. ClaudeBot is Anthropic's training crawler, which collects content in bulk to help train the Claude models. When a person asks Claude a question that requires reading a live webpage, Anthropic uses a different agent, Claude-User, for that single user-initiated fetch. A third agent, Claude-SearchBot, indexes content for Claude's search features. They are controlled separately in robots.txt, so blocking one does not block the others.

Does ClaudeBot respect robots.txt?

Yes. Anthropic states that ClaudeBot, Claude-User and Claude-SearchBot all honour robots.txt directives, including the Crawl-delay directive for slowing a bot rather than banning it. If you block ClaudeBot specifically, Anthropic says it will exclude your site's future content from its model-training datasets. Compliance depends on the bot being able to read your robots.txt file, which is one reason IP-level blocking can backfire.

Can I block ClaudeBot by IP address?

It is not recommended. Anthropic does not publish fixed crawler IP ranges and its bots use public cloud-provider addresses, so IP blocking is unreliable and can even stop the crawler from reaching your robots.txt file, which prevents it from honouring your rules at all. The supported method is to block by user-agent string in robots.txt.

Why do people say ClaudeBot takes more than it gives?

Because of its crawl-to-refer ratio, the number of pages Anthropic's crawlers request for every human visitor Anthropic sends back. On Cloudflare Radar's 2026 data that ratio was roughly 11,000 to 1, far worse than OpenAI's or Google's, because ClaudeBot has historically been a training crawler without a mass search product that returns clicks. Publishers reading those figures increasingly conclude that allowing the read for free is a poor trade.

Should I block ClaudeBot or find a way to be paid for it?

It depends on your objection. If you specifically do not want your content used for model training, block ClaudeBot and consider leaving Claude-SearchBot open so you can still be cited. If your concern is that valuable content is being read at scale for no return, blocking captures nothing, because a blocked crawler pays exactly the same as an unblocked one. Monetising the read at the edge, which is what blankspace does, aims to turn that traffic into revenue rather than simply switching it off.