← Back to blog

What does AI crawler traffic actually cost publishers to serve?

The bandwidth bill is almost always trivial. A mid-sized publisher serving ten million crawler requests a month is looking at roughly fifty dollars of egress, and on a flat-rate CDN it is zero. The cost that actually shows up is bot management, and at Politico it now takes a quarter of hosting spend.


Two entirely different numbers get filed under the same heading when publishers talk about what AI crawlers cost them. The first is the marginal cost of delivering bytes to a bot, which is a function of page weight, cache hit rate and whatever your CDN charges per gigabyte. The second is the cost of working out which bots are on your site, deciding what to do about them, and paying somebody to enforce that decision. The first is usually small enough to round away. The second is the one turning up in hosting budgets, and almost no publisher has separated them before deciding whether to block, allow or charge.

Why the bandwidth bill is smaller than publishers expect

Work the arithmetic on a realistic case. Take a mid-sized publisher serving ten million HTML page requests a month to declared AI crawlers. Crawlers generally request the document and not the image, video and script payload that surrounds it, so the relevant unit is a compressed HTML response, call it sixty kilobytes. Ten million of those is roughly 600 gigabytes of egress.

Amazon CloudFront lists the first 10TB of monthly data transfer out to the United States at 8.5 US cents per gigabyte. That puts the bill at about fifty dollars. Fastly, the most expensive of the major CDNs at list price, is around twelve cents per gigabyte, so about seventy-two dollars. Push the same publisher to a hundred million crawler requests a month and you are still inside the low hundreds. And if the site sits behind Cloudflare, which does not meter bandwidth at all and charges a fixed monthly plan fee instead, the marginal cost of the next crawler request is exactly nothing.

This is why the reporting does not find what people assume it will find. A Media Operator's Bron Maher, surveying publishers in April 2026, concluded that AI crawler surges "don't seem to be translating into big increases in hosting costs" and that the damage was landing elsewhere. That finding sits uncomfortably alongside the volume data from the same piece, which cites TollBit's State of the Bots report finding one AI crawler visit for every 31 human visits in Q4, up 25 per cent on the previous quarter and 60 per cent over six months. Volume went up sharply. The egress line did not.

The reason is caching. Normal traffic gets served from CDN edge nodes and never touches origin. AI crawler traffic mostly hits the same popular URLs as everybody else, so it mostly gets cached too, and cached bytes are the cheap kind.

When the bandwidth bill is not small

The arithmetic inverts in three specific situations, and they are worth naming because they are the cases where a publisher genuinely does have a five-figure problem.

The first is media. Wikimedia Foundation reported a 50 per cent increase in bandwidth for multimedia downloads between January 2024 and early 2025, driven by crawlers scraping openly licensed images and video. Its site reliability team found that almost two-thirds of the most resource-intensive traffic on Wikimedia Commons came from bots, despite bots accounting for only 35 per cent of pageviews. Crawlers pull from the long tail of the archive, which is precisely the content no cache holds.

The second is uncached endpoints and bulk artefacts. Read the Docs documented a single crawler downloading 73TB of zipped HTML in May 2024, nearly 10TB of it in one day, at a cost of over 5,000 dollars in bandwidth charges. When the team blocked AI crawlers, daily traffic fell from 800GB to 200GB and saved roughly 1,500 dollars a month.

The third is burst. Fastly's threat research team, analysing traffic between April and July 2025, recorded fetcher request volumes exceeding 39,000 requests per minute against individual sites. That is not an egress problem, it is an origin capacity problem, and it produces the effect Tampa Bay Times chief executive Conan Gallaty described to AdExchanger: "It taxes our ability to serve customers." Real readers get a slow page because a bot is saturating the origin.

The one publisher number on the record

The most useful figure any publisher has put into the public domain is not a bandwidth number at all. Amelia Binder, SVP of global government affairs at Axel Springer, told AdExchanger in August 2026 that 25 per cent of Politico's hosting costs now go toward bot management, and that it was something the company "didn't budget for".

That is worth sitting with. A quarter of a major publisher's hosting spend is going not on delivering content to bots but on identifying, classifying and filtering them. Detection vendors, WAF rules, challenge infrastructure, log analysis and the engineering time to keep all of it current are a recurring operating cost that grows with the sophistication of the adversary rather than with the volume of bytes. It is also, unlike egress, a cost that blocking does not eliminate. Blocking is the thing you are paying for.

It should be treated as one publisher's self-reported allocation rather than an industry benchmark, and it comes from a source with an interest in the number being large, since Axel Springer is a News/Media Alliance member and the alliance authored the Stealth Bot Prohibition Act now before the House, which would fine undisclosed crawlers 53,000 dollars per violation. It is still the only figure of its kind on the record, and it points at the right line in the budget.

How to work out your own figure

IAB Tech Lab's Content Monetization Protocols Working Group published draft bot management guidance in May 2026 built on exactly this premise: that most content owners have, in the document's words, "little to no understanding" of what it costs to deliver their content to bots. The guidance is deliberately not prescriptive about what to allow or block. Hillary Slattery, senior director of product management and programmatic at the Tech Lab, was blunt about that when asked whether the document tells publishers what to permit: "That's between the content owners and their gods."

What it does say is that there is a measurement floor, and it is three questions: which bot is crawling your content, how often, and how much it is costing you.

Ask your CDN, because it will not volunteer

All three answers live with your CDN, and Slattery's practical point is that CDNs are not generating these reports unless a client asks for them. Fastly and Akamai can both produce a breakdown of bot traffic by user agent, request count and associated cost. That request is an email, not a project.

Once you have request counts by user agent, the cost model is arithmetic. Multiply requests by average response size to get gigabytes. Multiply gigabytes by your contracted rate, not list price, since anyone doing meaningful volume is on a committed rate well below the published one. Separate cached from uncached requests, because the uncached ones are the expensive ones and they are usually a small minority doing most of the damage. Then add the line nobody puts in the model: the fraction of your bot management tooling and engineering time attributable to AI crawlers specifically.

Serve cheaper bytes

The cost-reduction tactic the Tech Lab guidance recommends is to stop sending bots the full human page. A crawler ingesting an article does not need the header, the footer, the navigation, the consent banner, the ad stack or the imagery. Serving a plain-text or markdown representation to declared bots is, in Slattery's framing, equally effective and cheaper, and it is a straightforward CDN-level content negotiation rule. It also improves how cleanly the content parses, which is a separate benefit worth having.

Why cost is the wrong number to optimise

Here is the problem with building a crawler strategy around cost. Cost avoidance has a ceiling, and the ceiling is your hosting bill. If a publisher blocks every AI crawler perfectly and pays nothing to do it, the best possible outcome is that it stops spending whatever it was spending, which for most publishers is a three-figure monthly egress number plus some share of a bot management budget that blocking does not actually remove. That is the entire prize.

The value exchange, which is the term the Tech Lab guidance uses, is the number that is not capped. The relevant comparison is not cost against zero, it is cost against what the crawler returns. Crawl-to-refer ratios are the crude version of that metric and they are unflattering: Slattery cited Claude at roughly 8,800 crawls per referral as of April 2026, against a 65,000 to 1 ratio reported by Business Insider in January, and DuckDuckGo at 1.5 to 1. Cloudflare's own Radar data, broken out by industry by David Belson, showed News and Publications sites getting materially better ratios than the web average, with Anthropic at 2,500 to 1, OpenAI at 152 to 1 and Perplexity at 32.7 to 1 in an early snapshot. All of these have moved and should be pulled live rather than quoted from memory.

Referrals are also not the only return, and for a growing number of publishers they are not the return that matters. Citation, licensing payment and paid retrieval are all forms of value exchange that a crawl-to-refer ratio cannot see. A crawler that sends no clicks but pays per fetch, or that places the publisher's content in front of a buyer inside an AI answer, is not the same proposition as one that takes and returns nothing, even though the ratio looks identical. This is the ground blankspace works on, treating the retrieval itself as the monetisable event at the CDN edge rather than waiting for a click that may never come.

Slattery's closing advice is the right posture for the whole question, and it applies as much to cost as to access: "Don't turn everything off. It's not a light switch; it's a mixing board."

What changes on 15 September 2026

The decision is about to get made for a large share of the web by default. Cloudflare's new mixed-crawler rules take effect on 15 September 2026, setting defaults on new domains to allow search but block training and agent use on pages carrying ads. Chris Dicker, chief executive of Candr Media, made the point to Digiday that the power of the change lies in inertia rather than in the toggle: if even a fraction of the roughly 20 per cent of the web behind Cloudflare leaves the defaults alone, the supply of freely scrapable content shrinks and the marginal cost of crawling rises for the crawler rather than the publisher.

That is the direction of travel worth planning against. The cost of serving a crawler was never the interesting number. The price of access is.

Frequently asked questions

How much does AI crawler traffic cost a typical publisher in bandwidth?

For a mid-sized publisher serving around ten million crawler requests a month, the egress cost is roughly fifty to seventy-five dollars at major CDN list rates, and lower on a contracted rate. On Cloudflare, which does not meter bandwidth, the marginal cost is zero. The bill only becomes material when crawlers pull uncached media, archives or bulk downloads, as Wikimedia and Read the Docs both found.

Does blocking AI crawlers save money?

It saves the egress, which for most publishers is a small number, and it does not save the bot management cost, which is the larger one. Detection, classification and enforcement are what blocking consists of, so choosing to block increases that line rather than reducing it. Blocking is a content and leverage decision rather than a cost decision.

How do I find out what AI crawlers are costing my site specifically?

Ask your CDN for a bot traffic report broken down by user agent, request count and cost. IAB Tech Lab's May 2026 bot management guidance is explicit that CDNs including Fastly and Akamai can produce this but generally will not unless a client requests it. Then split cached from uncached requests, apply your contracted per-gigabyte rate, and add the attributable share of your bot management tooling.

Why do crawlers cost so little to serve when they generate so much traffic?

Because CDN caching absorbs almost all of it. AI crawlers largely request the same popular HTML documents that human readers do, so those responses are served from edge cache and never reach origin. Crawlers also skip most of the page weight, requesting the document rather than the images, video and scripts that make a human pageview expensive.

Is the cost of AI crawling a good argument for blocking?

It is a weak one on its own. Cost avoidance is capped at what you currently spend, which is usually a three-figure monthly egress number, while the value a crawler does or does not return is uncapped in both directions. The stronger framing is the value exchange: what each crawler takes, what it gives back in referrals, citation or payment, and whether that trade is one you would sign.