Corrected on September 17, 2026
The first version of this guide got RFC 9309's group merging wrong, claimed disallow rules free crawl budget for other pages, credited Yandex with Crawl-delay and Host support, and recommended Search Console's robots.txt tester, retired in 2023. Its AI crawler table listed undocumented tokens (DeepSeekBot, xAI-Web-Crawler, cohere-ai), called the retired anthropic-ai token active, described Google-Extended as a retrieval crawler and claimed, without a findable source, that the major AI labs honor robots.txt for all their bots. A "live audit of 10 guides" contradicted itself and had no saved data, so I removed it. Lumina's Crawler Access Checker misread stacked User-agent lines; that's fixed too.
In October 2025 Heise published an obituary for robots.txt, arguing that AI crawlers have made the protocol meaningless. That's half right. Robots.txt was never enforceable, and the number of bots requesting a typical site has grown from a few search engines to dozens of AI crawlers and fetchers, each with its own policy. Some of them, notably the fetchers that act on a user's request, say openly that robots.txt generally doesn't apply to them.
The other half: the crawlers most sites actually care about, Googlebot, Bingbot, GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot among them, document that they follow robots.txt. So the file still controls a lot. It just isn't a lock.
What Robots.txt Actually Is
Robots.txt is a plain-text file at the root of your host (yoursite.com/robots.txt) that tells crawlers which paths they may request. The format is standardized in RFC 9309, published by the IETF in September 2022, 28 years after the protocol first appeared in 1994.
Crawlers that follow the protocol fetch robots.txt before crawling a host, parse it and cache it. Google caches it for up to 24 hours. The protocol is voluntary: nothing technically forces a bot to obey. For the bots that commit to it, though, robots.txt is the simplest way to control what they crawl.
One detail many guides skip: robots.txt controls crawling, not indexing. A URL that's disallowed won't be fetched, but Google can still index it without a snippet if other pages link to it. Mistake number one below covers what to do instead.
Why Robots.txt Still Matters
Robots.txt does two jobs today. It keeps crawlers away from URL spaces that are useless to crawl, and it's the one standard place where you can tell AI training crawlers and AI search crawlers different things.
The first job is often oversold. Google's crawl budget documentation is written for very large sites, roughly a million pages or more, or medium sites whose content changes daily. It also says Google won't move the crawl budget you free up by blocking URLs to other pages, unless Google is already hitting your server's capacity limit. For a typical site, blocking internal search results or endless filter combinations is about keeping junk out of the crawl, not about getting important pages crawled faster.
The second job is newer. Robots.txt is where you can say "no" to AI training and "yes" to AI search in one file, with the caveat that user-triggered fetchers play by different rules (see which AI bots actually follow robots.txt).
The File Format (RFC 9309)
RFC 9309 defines groups made of one or more User-agent lines followed by Allow and Disallow rules, plus comments starting with #. Other records, like Sitemap and Crawl-delay, aren't part of the standard, but the RFC allows crawlers to support them.
Here's a small valid file:
User-agent: *
Disallow: /admin/
Disallow: /search?
Allow: /admin/help/
User-agent: Googlebot
Disallow: /staging/
Sitemap: https://example.com/sitemap.xml
Three rules decide how a crawler reads it. First, several User-agent lines in a row share the rules below them. Second, a crawler uses the group that names it, and only falls back to User-agent: * if no group matches its name. If a file has several groups for the same user-agent, the crawler merges them. Third, within the group, the most specific (longest) matching path wins, and when Allow and Disallow match equally, Allow wins. Order doesn't matter.
That second rule hides a trap in the example above. Googlebot has its own group, so it ignores the * group entirely. It may crawl /admin/ and /search? because its own group only blocks /staging/ (mistake 4 below).
Paths support two wildcards: * matches any sequence of characters and $ marks the end of the URL. Disallow: /*.pdf$ blocks every URL ending in .pdf. User-agent names don't support wildcards, so User-agent: GPT* doesn't match GPTBot.
Three things look like part of the spec but aren't. Crawl-delay: Bing honors it, Google ignores it, and Yandex stopped using it in February 2018. Host: an old Yandex record that Yandex no longer lists. And a noindex line in robots.txt: Google ended support for that in September 2019.
The AI Crawler Layer: Which User-Agents Exist
Operators split their bots by purpose, and the purpose decides what blocking does. Training crawlers collect content for future models. Search crawlers build an index for answers with citations. User-triggered fetchers load a page because someone asked about it. Control tokens aren't bots at all: you use them in robots.txt to opt out of a use, and the normal crawler reads the rule.
| Operator | User-agent | Purpose |
|---|---|---|
| OpenAI | GPTBot | Training crawler for OpenAI's models |
| OpenAI | OAI-SearchBot | Search crawler for ChatGPT search results |
| OpenAI | ChatGPT-User | Fetches a page when a ChatGPT user's request needs it; OpenAI says robots.txt rules may not apply |
| Anthropic | ClaudeBot | Training crawler for Claude models |
| Anthropic | Claude-SearchBot | Search crawler for Claude's search results |
| Anthropic | Claude-User | Fetches a page when a Claude user's request needs it; Anthropic says it honors robots.txt |
| Perplexity | PerplexityBot | Search crawler for Perplexity's index |
| Perplexity | Perplexity-User | Fetches a page for a user's question; Perplexity says it generally ignores robots.txt |
Google-Extended | Control token, not a crawler: opts content out of Gemini model training and grounding in Gemini Apps and Vertex AI; no effect on Search or AI Overviews | |
Google-Agent | Google-hosted agents acting on a user's request; generally ignores robots.txt | |
| Apple | Applebot-Extended | Control token, not a crawler: opts content out of training Apple's AI models |
| Common Crawl | CCBot | Crawler for the public Common Crawl dataset, which many AI models are trained on |
| Meta | Meta-ExternalAgent | Crawler for AI training and Meta products |
| Mistral | MistralAI-User | Fetches pages for user requests in Le Chat |
| ByteDance | Bytespider | Crawler attributed to ByteDance; no official documentation, and reports say it doesn't reliably follow robots.txt |
You'll find more names in many block lists. anthropic-ai and Claude-Web are retired Anthropic tokens that ClaudeBot replaced. cohere-ai, DeepSeekBot and xAI-Web-Crawler aren't documented by Cohere, DeepSeek or xAI; Cohere even says it doesn't crawl the web to train its models. A rule for them does no harm, but you can't count on it doing anything either. Traffic numbers and what large publishers block are in the AI crawlers guide.
The useful split is training versus search. If you want out of training data but still want to be cited in AI search, block the training crawlers and allow the search crawlers. For OpenAI, that means a group that disallows GPTBot while OAI-SearchBot stays allowed. For Anthropic, disallow ClaudeBot and keep Claude-SearchBot allowed. For Google, a disallow for Google-Extended covers Gemini training without touching Search.
The 5 Patterns That Work
Most sites need one of these, or a mix of two.
1. The minimal everything-allowed file
For a site with nothing to block:
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
All crawlers welcome, whole site allowed, sitemap listed. An empty Disallow or no robots.txt at all has the same effect, but this version is explicit and carries the Sitemap line.
2. Keep junk URLs out of the crawl
For sites with internal search, session parameters or tracking parameters:
User-agent: *
Disallow: /search?
Disallow: /*?session=
Disallow: /*?utm_
Sitemap: https://example.com/sitemap.xml
Block URL patterns that generate endless variations, never the real product, article or service pages.
3. AI training opt-out, AI search allowed
For publishers who don't want to be training data but do want citations:
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: Meta-ExternalAgent
User-agent: Bytespider
Disallow: /
# Search crawlers such as OAI-SearchBot, Claude-SearchBot
# and PerplexityBot fall through to this group
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Two limits. Google-Extended also covers grounding in Gemini Apps and Vertex AI, so this blocks more than training at Google. And Bytespider has no documented robots.txt behavior, so treat that line as a request, not a block.
4. Block AI crawlers as far as robots.txt can
For paywalled, licensed or sensitive content:
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: Perplexity-User
User-agent: Google-Extended
User-agent: Google-Agent
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
User-agent: MistralAI-User
User-agent: Bytespider
Disallow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Be clear about what this does. The documented crawlers, like GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot, will stay out. ChatGPT-User, Perplexity-User and Google-Agent generally don't apply robots.txt to user-requested fetches, so listing them states your wish without enforcing it. If you need those fetches stopped, block them at the server or CDN, for example with a WAF rule based on their published IP ranges or Cloudflare's bot controls.
5. Section-specific rules
For sites where AI training bots may read the blog but not the docs:
User-agent: GPTBot
User-agent: ClaudeBot
Disallow: /docs/
Allow: /
User-agent: *
Allow: /
Sitemap: https://example.com/sitemap.xml
Longest match wins, so Disallow: /docs/ beats Allow: / for URLs under /docs/, and the rest of the site stays open for both bots.
The 6 Most Common Mistakes
1. Disallowing a page you want deindexed
To remove a page from Google, the noindex tag is the right tool. If the page is also disallowed in robots.txt, Googlebot never fetches it, never sees the noindex, and can keep the URL in the index based on links. The right order: leave the page crawlable, add noindex, wait until Google has recrawled and dropped it, and only then add a disallow if you want to stop future crawling.
2. Blocking CSS or JavaScript
Older files often block /wp-content/, /assets/ or /js/. Google renders pages like a browser and needs those files to see the page as users do. Without them, content and layout can be misread. Allow CSS and JavaScript unless you have a very specific reason not to.
3. Using robots.txt as a security layer
Anyone can read your robots.txt. Listing /admin/, /staging/ or /backup/ shows exactly which paths exist. Protect private areas with authentication, IP allowlists or a VPN. Use robots.txt for paths you don't want crawled, not paths you want hidden.
4. Giving a bot its own group and forgetting the shared rules
Once a crawler finds a group with its name, it ignores User-agent: *. A file that blocks /admin/ for everyone and then adds User-agent: Googlebot with a single rule has just unblocked /admin/ for Googlebot. Repeat the shared rules in every named group, or list the bot next to * in the same group.
5. Robots.txt from two sources
Each host serves exactly one robots.txt. If a CMS plugin generates a virtual one and your repository also contains a static file, only one of them is served, depending on how the server routes the request, and it may not be the one you edited. Fetch yoursite.com/robots.txt yourself after every change and check that what's live is what you intended. Remember too that each subdomain and protocol needs its own file.
6. Missing Sitemap line
The Sitemap line tells every crawler where your sitemap is, not just the search engines where you submitted it. WordPress (since 5.5), Shopify and Wix add it automatically; many custom setups don't. One line is enough: Sitemap: https://yoursite.com/sitemap.xml. Use several lines for several sitemaps.
Robots.txt vs. Noindex vs. WAF: Which Tool When
The question is whether the bot may fetch the URL, may index it, or shouldn't get a response at all.
| Tool | What it does | When to use |
|---|---|---|
| robots.txt Disallow | Asks compliant bots not to fetch the URL. Doesn't prevent indexing if the URL is linked elsewhere. | Endless URL variations, internal search, AI training opt-out. |
| noindex meta tag | Lets the page be fetched but keeps it out of search results. Per page. | Thank-you pages, thin templates, pages you want removed from the index. |
| X-Robots-Tag header | Same as noindex, sent as an HTTP header, so it works for PDFs and other files. | Keeping PDFs, downloads or other non-HTML files out of the index. |
| WAF / firewall block | Refuses the request at the server or CDN. Actual enforcement. | Bots that ignore robots.txt, user-triggered fetchers you don't want, undeclared scrapers, paywalled content. |
| HTTP authentication | Requires a login before any content is served. | Genuinely private content, staging environments, internal tools. |
How to Audit Your Robots.txt (3 Methods)
Method 1: Lumina Crawler Access Checker
Enter your domain in Lumina's Crawler Access Checker. It fetches your robots.txt, parses the groups and shows a site-level verdict for 37 search, AI and social bots: allowed, partially blocked or blocked. It also sends live requests with each bot's user agent, so you can see when the server or CDN answers a bot differently from what robots.txt says. It doesn't evaluate a single path against the longest-match rule, so for questions about one specific URL, use the Search Console tools below.
Method 2: Search Console robots.txt report and URL Inspection
In Search Console under Settings, the robots.txt report shows which robots.txt files Google found for the top hosts of your property, when it last fetched them, and any warnings or errors. You can request a recrawl there after an urgent change. The old robots.txt tester that let you check individual URLs was retired in late 2023. To check whether Google may crawl a specific URL, use URL Inspection instead.
Method 3: Direct fetch and manual review
Open yoursite.com/robots.txt in a browser. Check that it returns HTTP 200, has a Sitemap line, lists the bots you intend, and contains no surprises: plugin-added paths, leftover staging rules or duplicate groups. A 4xx response means crawlers treat the site as unrestricted. A 5xx response makes Google treat the whole site as disallowed for a while, which is worse than no file at all. Repeat the check after deploys and plugin updates.
Which AI Bots Actually Follow Robots.txt
Robots.txt is voluntary for every bot. What matters is what each operator documents, and that differs by bot type.
Most training and search crawlers state that they follow robots.txt. OpenAI documents this for GPTBot and OAI-SearchBot, Anthropic for ClaudeBot, Claude-SearchBot and Claude-User, Perplexity for PerplexityBot, and Common Crawl for CCBot.
User-triggered fetchers are the exception, as the table above shows for ChatGPT-User, Perplexity-User and Google-Agent. Anthropic's Claude-User is the documented exception among the big four. For these bots, robots.txt expresses a preference, and a firewall rule is the only real control.
Then there's behavior outside the documentation: Wired reported in June 2024 that Perplexity accessed content despite robots.txt blocks, Cloudflare published evidence in August 2025 of undeclared Perplexity crawlers getting around blocks, which Perplexity disputed, and Bytespider has no official documentation at all. That's why sites with valuable content combine robots.txt with bot rules at the CDN or server.
One more limit: blocking CCBot keeps your pages out of future Common Crawl snapshots, but it doesn't remove what's already in older snapshots, and it doesn't stop the many models that were trained on them.
A 5-Step Robots.txt Workflow
Fetch yoursite.com/robots.txt and run your domain through Lumina's Crawler Access Checker. Note which AI bots you block, which you allow, whether the Sitemap line is there and whether any named group has lost the shared rules.
Run Crawler Access Checker →Three realistic choices: allow everything, opt out of training while allowing AI search, or block AI crawlers as far as robots.txt allows plus a firewall rule for user-triggered fetchers. Write the decision down so the next person knows the file's intent.
See the 5 patterns →Open the Crawl stats report in Search Console and look for many requests to internal search, parameter URLs or calendar pages. Those are candidates for a Disallow. Don't expect it to speed up crawling of other pages unless your server is at its limit.
Check your sitemap →Build it from the patterns above. Keep it short, check every named group for the rules it should share, and avoid copying a long file from another site whose needs differ from yours.
Test before deploy →Deploy, fetch the live file yourself, and run the Crawler Access Checker again. Check the robots.txt report in Search Console over the next day, since Google caches the file for up to 24 hours.
Verify after deploy →FAQ
Where to Start
If you only do one thing this week, run your domain through Lumina's Crawler Access Checker and read your live robots.txt next to the result. Look for three things: a named group that lost the shared rules, an AI policy that doesn't match what you want, and a missing Sitemap line.
If you have more time, go through the five-step workflow. For most publishers the most useful change is pattern 3, opting out of training while allowing AI search, together with a clear note on which user-triggered fetchers robots.txt can't stop.
Audit your robots.txt now
Lumina's Crawler Access Checker parses your robots.txt for 37 search, AI and social bots and checks how your server answers each of them. Free, no signup.
Run Crawler Access Checker →