Robots.txt was designed for crawlers that turn up on their own schedule. Google-Agent turns up only when a person asks an AI system running on Google infrastructure to go and do something on the web: research a product, compare options, fill in a form. Because a human triggered the visit, Google files it under "user-triggered fetchers", a category its documentation says generally ignores robots.txt rules. A publisher who adds a Disallow rule for Google-Agent has therefore recorded a preference, not set a control.
What is Google-Agent?
Google-Agent is the user agent Google lists for agents hosted on its infrastructure that navigate the web and take actions when a user asks them to. It appears on Google's official list of user-triggered fetchers, which Google last updated on 19 August 2026. Search Engine Journal reported that the entry was added quietly on 20 March 2026.
Google publishes two user-agent variants, one mobile and one desktop, and both contain the token compatible; Google-Agent;. It also publishes the IP ranges these requests come from in a file called user-triggered-agents.json. Google's page does not say which consumer products currently send the traffic, so publishers should treat it as the identity for a class of Google agent activity, not for a single named product.
Does Google-Agent obey robots.txt?
Not as a rule. Google's documentation says user-triggered fetchers generally ignore robots.txt, on the logic that the request was initiated by a person rather than by an automated crawl. Google's overview of its crawlers makes the same distinction: common crawlers always respect robots.txt for automatic crawls, while fetchers run at a user's request are described separately.
The practical consequence is that robots.txt cannot be the only line of defence for anything you want kept away from AI agents. Search Engine Journal's coverage and TollBit's analysis both reached the same conclusion: restricting this kind of visitor means server-side authentication or access rules. Robots.txt still has a role for declaring intent, and Known Agents notes that it can be used to request a block, but compliance is up to the operator.
How is Google-Agent different from Googlebot and Google-Extended?
The three names do different jobs, and conflating them is the most common mistake in publisher robots.txt files.
Googlebot is Google's search crawler. It visits continuously to build the index and respects robots.txt. Google-Extended is a control token in robots.txt that lets a publisher state whether content may be used for Gemini model purposes; it is not a bot that fetches pages. Google-Agent is a fetcher that arrives in real time on behalf of a specific person and reads a page, or acts on it, at that moment.
The difference matters for measurement as well as control. Googlebot visits are about indexing, and Google-Agent visits are about a live task, so a Google-Agent request is closer to an AI retrieval than to a crawl. Publishers who want to understand that category in more depth can read our explainer on what a Live Search Agent is.
What happened to Project Mariner?
Project Mariner was Google's experimental AI browsing tool, and it is widely reported as the first product to use Google-Agent. Android Headlines reported that Google shut Mariner down on 4 May 2026 after a 17-month run in Google Labs, with its capabilities being folded into Gemini Agent and Chrome's auto browse features.
That does not make Google-Agent obsolete. Google's fetcher documentation was still being updated in August 2026, and the successors are more likely to widen agent traffic than to shrink it. Chrome's auto browse, announced in January 2026 for Google AI Pro and AI Ultra subscribers in the United States, is a good example of why: when an agent drives a normal browser on a user's behalf, there may be no distinctive user agent to spot.
Known Agents still lists a separate GoogleAgent-Mariner user agent that runs a real Chrome build, executes JavaScript and sets cookies. As of 8 October 2026 it reports that about 8 per cent of the top sites it tracks block that agent in robots.txt. It also states that the agent publishes no verification method, so a matching string is only a clue.
How does Google-Agent compare with other user-triggered AI fetchers?
Google is not alone in taking this position, but the major operators do not behave identically, and the differences affect how much weight a robots.txt rule carries.
| Fetcher | Operator | Stated robots.txt position |
|---|---|---|
| Google-Agent | Generally ignores robots.txt (user-triggered fetcher) | |
| ChatGPT-User | OpenAI | Documentation says robots.txt rules may not apply to user-initiated actions |
| Perplexity-User | Perplexity | Says it generally ignores robots.txt for user-initiated requests |
| Claude-User | Anthropic | Anthropic says all three of its bots respect robots.txt |
The behavioural evidence points the same way. TollBit's State of the Bots report for the first half of 2026, as covered by Search Engine Journal on 14 August 2026, found that ChatGPT-User reached disallowed pages on more sites than any other bot it tracked. In an earlier TollBit report covering the second half of 2025, about 30 per cent of AI bot scrapes violated explicit robots.txt restrictions. Both are TollBit measurements from its own network, so they describe the sites TollBit monitors rather than the whole web.
How can publishers identify Google-Agent traffic?
Start with what Google provides, and then assume it will not be enough.
Google lists the user-agent token, the published IP ranges, and a reverse DNS pattern. It warns that user-agent strings can be spoofed, so a string match alone should never grant or deny access. Google is also experimenting with Web Bot Auth, in which the agent signs its requests and identifies itself as https://agent.bot.goog. Signed requests are the strongest signal available because a signature cannot be copied from a log file the way a string can.
Agents that drive an ordinary browser are harder. HUMAN Security, a vendor that sells bot-management products, describes Mariner traffic as looking like a standard Chromium client with no signed identity, and advises relying on behaviour, request paths and attempted actions rather than identifiers. That is a vendor view, but it matches the underlying logic: if the identity cannot be trusted, the behaviour is what is left to classify.
What can publishers actually do about Google-Agent?
Because robots.txt will not do the work, the controls sit elsewhere. Four steps are worth taking, in order.
First, measure. Log which agents visit, which pages they read and which actions they attempt. Server-side logs and CDN logs see this traffic, while JavaScript-based analytics usually do not, which is why it tends to be missing from standard reports.
Second, decide a policy by action rather than by name. HUMAN suggests allowing content browsing, treating low-risk conversions with care, and denying login, account changes and checkout by default. For a publisher the equivalent split is usually between reading public articles, reading subscriber-only content, and submitting forms.
Third, enforce at the edge. Verification against Google's published IP ranges, or against Web Bot Auth signatures where present, is far more reliable than a user-agent rule. Check that your firewall rules do not block Google's ranges by accident, because a blanket bot rule can stop legitimate agent visits before they reach the page.
Fourth, decide whether access should be monetised. An agent that reads an article to answer a question is consuming the page without loading the ads that normally pay for it. Systems that sit at the CDN layer, including the one blankspace operates, can classify this kind of visit and decide what to serve it, which is a different question from whether to block it.
Should publishers block Google-Agent?
Blocking is a business decision, not a technical default. A person asked an agent to visit your page, so a hard block may simply send that person to a competitor that is readable. On the other hand, the visit generates no ad impression and may deliver the content to the user without a click.
Two cautions apply. Google lists Google-Agent separately from Googlebot, but its documentation does not spell out how a block would interact with Google Search, so test any rule on a narrow set of paths before rolling it out. And because the user-agent string can be spoofed, a block based on the string alone will stop honest traffic and miss dishonest traffic, which is the wrong way round.
Frequently asked questions
Does adding Google-Agent to robots.txt block it?
Not reliably. Google's documentation says user-triggered fetchers generally ignore robots.txt, so a rule for Google-Agent records your preference without enforcing it. Enforcement needs a server, firewall or CDN rule that verifies the request first.
Is Google-Agent the same as Googlebot?
No. Googlebot is the search crawler that builds Google's index and respects robots.txt. Google-Agent is a separate fetcher that visits in real time when a person asks an AI agent to do something, and it is listed under a different category in Google's documentation.
How do I verify that a request really comes from Google-Agent?
Check the source IP against the ranges in Google's user-triggered-agents.json file and use the documented reverse DNS checks, rather than trusting the user-agent string, which Google warns can be spoofed. Where requests carry Web Bot Auth signatures from https://agent.bot.goog, verifying the signature is stronger still.
What replaced Project Mariner?
Reporting says Google shut Project Mariner down on 4 May 2026 and is folding its capabilities into Gemini Agent and Chrome's auto browse features. Google's fetcher page does not say which products send Google-Agent traffic today, so monitor your logs rather than assuming.
Do ChatGPT-User, Claude-User and Perplexity-User follow robots.txt?
They differ. OpenAI's documentation says robots.txt rules may not apply to ChatGPT-User, Perplexity says it generally ignores robots.txt for user-initiated requests, and Anthropic says all three of its bots respect robots.txt. Independent measurements from TollBit suggest compliance is uneven in practice, so treat robots.txt as a signal rather than a guarantee for every operator.
