← Back to blog

What is a stealth crawler, and what can publishers actually do about one?

A stealth crawler is an automated tool that scrapes a site without declaring what it is or why it is there, typically by presenting as an ordinary browser on a residential IP address. On large news sites, roughly a quarter of all traffic is estimated to be stealth crawling that publishers never register as bot traffic at all. New York and Congress are both moving to outlaw it, and the objections to those bills are more serious than most publishers realise.


Open your access logs and the crawlers you can name are the easy part. GPTBot, ClaudeBot and PerplexityBot arrive with a user-agent string, honour or ignore your robots.txt in public, and can be blocked with one line of configuration. The harder problem is the request sitting underneath them: a Chrome user agent, a residential IP address in a suburb somewhere, a believable referrer, and a fetch pattern that only looks wrong when you aggregate thousands of sessions. That request is a stealth crawler, and it is defined by what it withholds rather than by what it does. It retrieves your content without identifying the software making the request, the company operating it, or the purpose the content will be put to.

What is a stealth crawler?

A stealth crawler is a web crawler that accesses and collects content from a site without disclosing its identity or its purpose. Digiday's August 2026 explainer on the term defines it as a crawler that will not self-identify and instead mirrors the browsing patterns that make it look like a human visitor. In practice, stealth crawlers either ignore robots.txt outright, or route around it by fetching through a third-party service that has no relationship with the publisher at all.

Two things follow from that definition and both matter commercially.

The first is that "stealth" is a statement about disclosure, not about intent. Not every undeclared crawler is malicious. Academic researchers, security teams, price-comparison services and investigative journalists all use unidentified automated tools for legitimate work, and they do so precisely because declared crawlers get blocked. The Electronic Frontier Foundation makes this argument at length, and it is the reason the legislation described below is contested rather than uncontroversial.

The second is that a stealth crawler is invisible to the tooling most publishers rely on. If a request looks like a browser, it lands in your analytics as a session, in your ad stack as an impression opportunity, and in your bot reports as nothing at all. You cannot block what you cannot see, and you cannot licence what you cannot count.

How do stealth crawlers hide?

Three techniques do most of the work.

Generic user agents. The crawler presents itself as an ordinary Chrome, Safari or Firefox client rather than as a named bot. Nothing in the request declares an operator.

Residential and ISP proxy networks. Rather than crawling from identifiable cloud infrastructure, traffic is routed through pools of consumer IP addresses, so it originates from the same address space as your real audience. Chris Dicker, CEO of Candr Media, put the limitation plainly to Digiday: Cloudflare only governs what is behind Cloudflare, and well-funded scrapers can use residential proxies or simply buy data from second-hand brokers who scraped it elsewhere.

Scraping as a service. The AI company never touches your servers. An intermediary crawls, aggregates and resells, which severs the evidential link between the crawl and the eventual commercial use. Lindsay Van Kirk, svp of innovation at People Inc., described the symptom at an IAB Tech Lab event in May 2026: the company found its content appearing in applications it had no relationship with, and traced access back to unauthorised usage that defaulted through to home proxy networks as a last-resort route in.

Frederick Jahn of Centinel Analytica, who reverse-engineered scraping tools before moving to the defensive side, told Digiday his team has "pretty much reverse-engineered any protection that is out there". His point is that the shared proxy infrastructure makes it nearly impossible to tie specific AI companies to specific crawls with any certainty. That opacity is not a side effect. It is the product.

How much publisher traffic is stealth crawling?

The honest answer is that nobody knows precisely, because the measurement problem is the whole problem. The credible estimates converge on a large number.

Cloudflare Radar data cited by Digiday in August 2026 puts more than half of all web traffic on a bot basis. The bipartisan bill introduced in Congress in July 2026 opens with the same claim, that automated bots now account for more than half of all global web traffic.

HUMAN Security's 2026 State of AI Traffic and Cyberthreat Benchmark Report, published in April 2026 and drawn from more than one quadrillion digital interactions across 2025, found AI scraper traffic grew 597% from January to December 2025, with AI-driven traffic overall up 187% across the year. Media and streaming sites absorbed 41% of AI scraper traffic, the largest single category. Scraping attacks now affect close to 20% of site traffic for the median organisation, roughly double 2022 levels.

On the specific question of the stealth share, Jahn's estimate for major news brands is the most useful figure in circulation: 20% to 30% of traffic comes from identified crawlers, and roughly a quarter of total traffic is stealth crawling that mimics human users. That second number is traffic most publishers do not currently see as bot traffic at all.

The commercial scale is contested in the same direction. Media analyst Matthew Scott Goldstein's work on what he calls the scraper economy put it at around $1bn. Danielle Coffey, president and CEO of the News/Media Alliance, told Digiday she has seen data suggesting a multi-billion dollar market.

Stealth crawlers, declared bots and mixed-use crawlers: what is the difference?

The access landscape has fragmented into three tiers, and conflating them leads publishers to the wrong policy.

Declared crawlers identify themselves accurately and can be allowed, blocked or metered per publisher. This is the layer that infrastructure gatekeepers like Cloudflare, Fastly and Akamai now govern in a meaningful way.

Mixed-use crawlers declare themselves but serve more than one purpose, typically search indexing plus AI training or agentic retrieval. Cloudflare's chief strategy officer Stephanie Cohen told Digiday these crawlers put content creators in an impossible position, and the company has said that from 15 September 2026 its defaults will allow search while blocking training and agent use on pages carrying ads. Google is the obvious case, and Cloudflare has been careful not to promise it can force a separation.

Stealth crawlers sit outside that machinery entirely. They are the reason the mixed-use debate, however consequential, does not resolve the leakage problem. Cloudflare governs what runs through Cloudflare. Everything else is unpriced.

Publishers should be clear that a declared retrieval agent is not the enemy in this taxonomy. A named Live Search Agent fetching your page in order to answer a live user query is a countable, addressable, monetisable event. A stealth crawler is the same fetch with the commercial metadata stripped out.

Why does undisclosed crawling matter commercially?

The obvious harm is infrastructure cost. Publishing executives have told Digiday their servers have been overloaded by millions of scrapes over time, and both the New York and federal bills lead on server burden.

The more consequential harm is that stealth crawling breaks the precondition for every commercial arrangement publishers are currently trying to build. Goldstein has been the clearest on this, writing that transparency is not the endgame but the precondition: every licensing conversation stalls at the same place, because publishers have no reliable way to prove what was taken, by whom, and at what volume. Once a crawler has to say its name, he argues, scraping stops being free and starts being a line item.

That argument generalises beyond licensing. A publisher cannot run a pay-per-crawl arrangement against traffic it cannot attribute. It cannot report AI-era audience to advertisers on a basis it cannot audit. It cannot decide rationally whether to block or monetise a given operator when a quarter of the relevant traffic is filed under "humans". Undisclosed crawling is, in the end, a measurement failure with a revenue consequence attached.

What would the New York Stealth Crawler Prohibition Act require?

New York's legislature passed the Stealth Crawler Prohibition Act, S.9934A and its Assembly companion A.11292, in early June 2026. It is the first state law of its kind in the US. The mechanics matter more than the headline.

Who is covered. A "covered news source" is a publication or service that performs a public-information function comparable to that traditionally served by journalism organisations, makes a substantial expenditure of labour, skill and money to create and distribute content, publishes or updates at least monthly with a process for error correction, and has at least 1,000 monthly active users in New York. Broadcast, cable, satellite and digital services are all in scope.

What a crawler must do. Disclose identity and purpose at or before the point of access, by presenting a valid and accurate user-agent string naming the software product, its version and the company behind it, and by disclosing the specific nature and purpose of the crawler, including all uses the content could be put to, in a format the publisher can access.

What is prohibited. Deploying a stealth crawler, meaning one that does not comply with those disclosure requirements, in a manner that would damage, impair or burden the operation of a covered news source or otherwise cause it economic harm.

Enforcement. The New York Attorney General can seek injunctions and civil penalties of up to $15,000 per day for each violation. Separately, and this is the provision the bill's opponents concentrate on, a journalism provider can obtain a subpoena before filing any action, ordering a service provider to disclose information sufficient to identify an alleged violator.

Timing. The bill awaits Governor Hochul's signature. Clark Hill's analysis notes that action is unlikely before the 7 November 2026 gubernatorial election, that the bill may well be modified before signature, and that the provisions take effect 180 days after signature, putting the earliest realistic effective date at around 7 May 2027. Sponsors were Assembly Member Steven Otis and State Senator Mike Gianaris.

What is the federal Stealth Bot Prohibition Act?

On 23 July 2026, Representatives Valerie Foushee, Laurel Lee and Gus Bilirakis introduced the bipartisan Stealth Bot Prohibition Act in the US House. It requires automated web crawlers to accurately identify themselves and disclose their purpose when accessing websites, prohibits bots that intentionally misrepresent their identity or impersonate human users in connection with generative AI services, and authorises the Federal Trade Commission to enforce through civil penalties, with state attorneys general also able to act.

The supporter list reads as a who's who of the publishing side: the News/Media Alliance, News Corp, The New York Times, Condé Nast, Hearst Magazines, Advance Local, Axel Springer US, McClatchy, USA TODAY Co, Vox Media, Newsmax, Tampa Bay Times and Reddit, alongside the rights bodies ASCRL and the Society of Composers and Lyricists, and, notably, the AI developer Bria.

Bria's position is the interesting one. Chief AI strategy officer Vered Horesh said the company is an AI developer asking for the rule to apply to itself, on the basis that companies scraping in the dark have something to hide. Condé Nast CEO Roger Lynch and News Corp chief executive Robert Thomson both put their names to considerably more colourful statements. Coffey framed the enforcement logic simply: bad bots must identify themselves, and once they do, publishers can stop them or sue them.

The bill's immediate task is attracting co-sponsors. Coffey has predicted the process could take a while.

What is the case against these bills?

Publishers should understand the opposing argument, because it is not frivolous and it will shape whatever version of these laws eventually passes.

The Electronic Frontier Foundation's Tori Noble argued on 20 July 2026 that stealth crawlers are simply automated tools for accessing public web data without disclosing the user's identity, and that anonymous crawling enables some of the most publicly beneficial uses of the open web. The examples are concrete. The Markup used crawlers identifying as ordinary Firefox browsers to analyse whether Amazon prioritised its own brands over better-rated competitors. ProPublica used a tool simulating an ordinary customer to show Amazon steering shoppers to more expensive products. Security researchers and privacy tools crawl anonymously by necessity. In each case, a declared crawler would have been blocked, and the reporting would not exist.

The EFF's sharper objection is to the New York bill's subpoena mechanism, which it characterises as giving websites the power to unmask anyone using an unidentified crawler without any evidence that they broke the law. On 3 August 2026, the EFF joined 18 civil society organisations, with the Center for Democracy and Technology also publishing its own call, urging Hochul to veto.

The substantive technical criticism is worth sitting with: the EFF's position is that the real harm is over-aggressive crawling rather than anonymity, that server strain can be addressed with technical measures targeting conduct rather than identity, and that unmasking therefore does not solve the problem it is sold as solving. A publisher can accept that critique and still want disclosure, because the publisher's interest is in pricing access rather than in stopping load. But the two goals are genuinely different, and the bills conflate them.

What can publishers do about stealth crawlers before any of this becomes law?

Legislation is, on the most optimistic reading, a 2027 instrument in one state. The practical work is available now.

Measure the gap before you argue about policy. Reconcile server-side request logs against your client-side analytics. The difference between what your origin serves and what your JavaScript records is where undeclared automated traffic lives. Until you have that number for your own domain, industry estimates are a substitute for evidence.

Raise the cost rather than chasing perfection. Van Kirk's framing is the most useful one in circulation: friction is the goal. People Inc. went from blocking roughly 2,100 user agents to over 30,000 with a block-all approach, and her argument is that adding two full seconds of latency to the majority of scrapers is a good outcome even when they get through, because every scraper forced to pay a residential proxy network is margin coming out of the scraping business.

Put the challenge early. Jahn's practical advice is client-side "are you human" challenges on first page load rather than deeper in the session, systematic blocking of non-essential crawlers, and rigorous benchmarking of any vendor claiming more than 90% detection.

Separate the stealth question from the retrieval question. A block-all posture aimed at undeclared scraping should not sweep up the declared retrieval agents that generate citations, referrals and, increasingly, addressable inventory. Decide your policy per operator and per purpose, not per bot.

Reconcile your posture with the 15 September Cloudflare defaults. If you sit behind Cloudflare, the mixed-use crawler change lands in weeks and will alter what reaches your origin. Auditing stealth traffic after that change without knowing what the change itself did to your logs will produce a misleading baseline. Take the measurement now.

Where does blankspace sit in this?

blankspace operates at the CDN edge, which is where the distinction between declared and undeclared retrieval is observable in the first place. The relevant capability here is detection rather than enforcement: identifying when a request is a Live Search Agent fetching a page to answer a live user query, separating that from ordinary human sessions and from undeclared automated traffic, and making the declared portion countable and monetisable rather than merely visible.

That is deliberately a narrower claim than the legislation makes. Edge detection does not unmask a residential-proxy scraper that has decided to look exactly like a reader, and no vendor should tell a publisher otherwise. What it does is ensure that the traffic which is willing to declare itself gets treated as an audience with commercial value rather than as an unclassified line in a log file. The argument holds regardless of which vendor a publisher uses to do it.

Frequently asked questions

Is using a stealth crawler illegal?

Not currently, in most circumstances, in the US. New York's Stealth Crawler Prohibition Act would make it unlawful to deploy an undisclosed crawler against a covered news source in a way that burdens its operation or causes economic harm, but the bill is unsigned and would not take effect until roughly 180 days after any signature. The federal Stealth Bot Prohibition Act, introduced in July 2026, is at the co-sponsor stage. Separately, undisclosed scraping may already give rise to copyright, contract or computer-misuse claims depending on the facts, which is the basis of several existing publisher lawsuits against AI companies.

How can I tell whether traffic is a stealth crawler or a real reader?

Not from a single request, which is the point. Detection relies on aggregate signals: mismatches between server-side requests and client-side analytics, fetch patterns that traverse archives faster than a person could read, absence of the secondary asset requests a real browser would make, and behavioural fingerprinting at the edge. Any vendor claiming near-perfect identification of a well-run residential-proxy scraper should be benchmarked rather than believed.

Is a stealth crawler the same thing as an AI crawler?

No. The categories overlap but are not the same. Most named AI crawlers, including GPTBot, ClaudeBot and PerplexityBot, declare themselves and are therefore not stealth crawlers by definition, whatever a publisher thinks of their behaviour. Conversely, plenty of undeclared crawling has nothing to do with AI, including price monitoring, security research and academic work. Stealth describes the disclosure posture, not the end use.

Would the New York law apply to my publication if I am not based in New York?

Potentially, yes. The threshold in the bill as passed is tied to the audience rather than the publisher's location: a covered news source needs at least 1,000 monthly active viewers, listeners, users or subscribers in New York, along with the editorial and investment criteria. A UK or European publisher with a modest New York readership could therefore fall in scope. The bill may be modified before signature, so the final thresholds are not settled.

Does blocking stealth crawlers cost me AI visibility?

It should not, if the blocking is properly targeted. Visibility in AI answers comes overwhelmingly from declared retrieval agents and from indexed content that named crawlers have already read. The risk is a blunt block-all posture that catches those alongside the undeclared traffic. The safe sequence is to decide allow, block or meter per named operator and per purpose first, then apply aggressive friction to everything undeclared.