Project Mariner, Google's browser agent prototype, was shut down on May 4, 2026. Google-Agent, the user agent Google added to its crawler documentation in March, is still there. Those two facts get mixed up a lot, and most early coverage got one detail wrong: Google-Agent generally ignores robots.txt.
Correction, September 17, 2026
The first version of this article said Google-Agent respects robots.txt and recommended controlling it there. Google's documentation says the opposite. It also treated Project Mariner as an ongoing product and claimed Google has three crawlers. I've rewritten the article against Google's current documentation.
What happened to Project Mariner
Google DeepMind showed Project Mariner in December 2024 as a research prototype: an agent that takes a task in plain language and works through websites to finish it. In May 2025 it opened to Google AI Ultra subscribers in the US. On May 4, 2026, Google discontinued it, according to the notice on Google Labs as cited by Wikipedia.
The idea didn't die with it. Chrome's auto browse, launched in the US in January 2026 for Google AI Pro and Ultra subscribers, lets Gemini click through sites inside the browser, and the Gemini app has its own agent mode.
What Google-Agent is
Google lists Google-Agent among its user-triggered fetchers, the category for requests a person started rather than a crawl schedule. Google says it's used "by agents hosted on Google infrastructure" that browse the web and take actions when a user asks them to. It was added on March 20, 2026.
Note what that description doesn't say. It doesn't name a product, and it's limited to agents running on Google's servers. An agent working inside someone's own Chrome is a different setup, and Google hasn't documented which user agent those visits carry. Don't assume every Gemini-assisted visit shows up as Google-Agent in your logs.
The user agent string itself looks like a normal Chrome browser with compatible; Google-Agent added, in a mobile and a desktop variant. That's worth knowing for log filters: search for the token, not for a full string.
Does Google-Agent respect robots.txt?
Generally, no. Google says it plainly for the whole category: "Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules." A Disallow for Google-Agent is a request it isn't built to follow.
That changes the practical advice. If an area of your site shouldn't be reachable by an agent, robots.txt is the wrong tool, because the agent is standing in for a person who could open that page anyway. Protect it the way you'd protect it from a person: with authentication, server-side checks and rate limits. robots.txt still matters for Googlebot and for the training opt-out, just not here.
Google-Agent vs. Googlebot vs. Google-Extended
These three get lumped together, but only one of them is a crawler in the classic sense, and Google runs far more than three. Its common crawlers alone include Googlebot, its image, video and news variants, GoogleOther and more.
| What it is | Follows robots.txt? | Effect of blocking | |
|---|---|---|---|
| Googlebot | Crawler for Google Search | Yes | Pages drop out of Search |
| Google-Extended | A robots.txt token, not a crawler with its own user agent | It is the robots.txt rule | Content isn't used to train Gemini or for grounding in Gemini Apps and Vertex AI. No effect on Search. |
| Google-Agent | User-triggered fetcher for Google-hosted agents | Generally no | A robots.txt block isn't reliable. Use server-side rules. |
Google-Extended is the one people misjudge most. It has no user agent of its own, because the crawling happens with Google's existing user agents. It only tells Google what it may do with the content afterwards.
How OpenAI, Anthropic and Perplexity handle the same question
Every major AI company now runs a bot like this. They don't agree on robots.txt.
- OpenAI says for ChatGPT-User that "robots.txt rules may not apply", a change it made to its bot documentation in December 2025.
- Perplexity says Perplexity-User generally ignores robots.txt.
- Anthropic says all three of its bots, including Claude-User, honor robots.txt, according to its help center.
So Anthropic is the exception, not the rule. Three of the four big operators treat user-triggered fetches like a browser visit. I think that's defensible for a single page a user asked for, and less so for an agent that clicks through twenty pages and submits a form. Either way, it's the documented behavior, and your setup should assume it.
How to identify and verify Google-Agent
The user agent string can be faked by anyone, so a log line that says Google-Agent proves nothing on its own. Google gives three ways to check:
- Match the IP address against Google's published
user-triggered-agents.jsonrange file. - Run a reverse DNS lookup. Genuine requests resolve to a
google-proxy-…google.comhostname, which should resolve back to the same IP. - Look for a signature. Google says it's experimenting with the Web Bot Auth protocol for Google-Agent, using the identity
https://agent.bot.goog, so requests can be verified cryptographically instead of by IP.
Many WAFs and CDNs with a verified-bots list do the first two for you. Check whether yours already knows Google-Agent before writing custom rules.
What to do on your site
Nothing dramatic. In Cloudflare's July 2025 breakdown, user-triggered and undeclared AI fetches together made up under 5% of AI bot traffic, so most sites don't need a special policy. Four things are worth doing:
- Check that your bot protection isn't blocking verified Google-Agent requests by accident, if you want agents to be able to complete tasks like a booking or a quote request. A strict "block everything that isn't Googlebot" rule catches it.
- Stop relying on robots.txt for anything sensitive. Admin areas, internal search and APIs need authentication or server-side limits, whether the visitor is an agent or not.
- Make forms work without guesswork. Clear labels, standard input types and visible error messages help assistive technology, and they help agents in exactly the same way.
- Start logging the token. Filter for
Google-Agent, verify a sample against the IP file, and you'll have a baseline once agent traffic grows.
The Crawler Access Checker shows how your robots.txt treats Google-Agent and what your server returns to its user agent. The second part is the one that counts here, because a 403 from your firewall stops the agent and a robots.txt rule doesn't. For the other AI bots, the AI crawlers guide has the full list.
FAQ
Check what AI agents see on your site
The free Crawler Access Checker tests 37 bots, including Google-Agent, against your robots.txt and your server's real response.
Run the Crawler Access Checker