Home About Support Blog Ask AI
Dashboards +
On-Page SEO +
Technical SEO +
SERP & Content +
Local SEO +
Get the Chrome Extension
Free SEO & GEO Tool

AI Bot & Robots.txt Checker

Check which AI bots and crawlers are allowed or blocked by robots.txt rules, then verify with live server checks using each bot's real User-Agent.

Last updated: September 2026

How this AI bot and robots.txt checker works

Most crawler checkers only test Googlebot. But in 2026, GPTBot, ClaudeBot, PerplexityBot, and Google-Extended matter just as much for AI search visibility. This AI bot and robots.txt checker tests 37 crawlers — including Googlebot, OpenAI's new OAI-AdsBot (ChatGPT ads validator), and every major AI bot — then fires a live request at your homepage with each one's User-Agent string. Google-Extended and Applebot-Extended are robots.txt tokens without a user agent of their own, so for those two the server column only shows how your server treats that string.

That second check is the one people miss. A bot can be "allowed" in robots.txt but still get a 403 from a Cloudflare rule or a CDN bot-protection layer. This tool shows both layers: what your robots.txt says, and what actually gets through. It works the other way round too. Google-Agent, ChatGPT-User and Perplexity-User fetch pages when a person asks for them, and their operators say robots.txt generally doesn't apply there. A Disallow line won't stop them, only a server rule will. And watch the search bots: GPTBot collects training data, but ChatGPT search fetches pages with OAI-SearchBot. Block that one at the server and no amount of content work gets you cited there.

Common robots.txt errors this checker finds

Most mistakes are dumb, not subtle. The classic: a Disallow: / left over from a staging environment that nobody removed before going live. The checker shows it as BLOCKED for every bot it applies to. A missing robots.txt shows up too, which means every URL is crawlable by default.

Path rules need a human eye. Disallow: /folder also blocks /folder-archive and /folder.html, because rules match by prefix. The checker lists each bot's restricted paths under "rules" (hover the badge), and the raw file sits at the bottom of the results, but it doesn't test a single URL against them.

Is my robots.txt blocking search engines?

Run this checker against your live URL. A site-wide block on Googlebot shows up as BLOCKED in the first result row. Partial blocks show as "rules", and those are where most accidents hide: a wildcard pattern (Disallow: /*.php) meant for one path that matches thousands. Google caches robots.txt for up to 24 hours, so a fix can take a day to register. Search Console's robots.txt report lets you request a recrawl of the file.

robots.txt tester vs. robots.txt checker

Google retired the robots.txt Tester inside Search Console in late 2023. The replacement is a short report under Settings → robots.txt that only shows the currently fetched file. It doesn't let you test custom user agents or preview rule changes. This checker fills part of that gap: it runs your live file against 37 named crawlers and checks what the server actually returns to each. It doesn't take a draft file or a custom User-Agent, so test rule changes on staging first.

Explore more tools

FAQ

What is a robots.txt file?+
A text file at your site's root (example.com/robots.txt) that tells crawlers which pages they can or cannot access. It uses User-Agent, Allow, and Disallow directives to control crawling behavior per bot.
What does "rules" mean vs "BLOCKED"?+
"Rules" means the bot can crawl your site but certain paths are restricted (e.g. /admin/, /api/). This is normal and healthy. "BLOCKED" means Disallow: /, so a bot that honors robots.txt won't crawl any page. User-triggered fetchers like Google-Agent, ChatGPT-User and Perplexity-User generally ignore robots.txt, so for them the live server check decides. Claude-User is the exception and does honor it.
Why does the server check show 403 for some bots?+
A 403 means the server actively rejects that bot regardless of robots.txt rules. This is typically done via firewall rules, CDN settings, or server-side bot detection. The bot cannot access your site even if robots.txt allows it.
What happens without a robots.txt?+
All crawlers assume they have permission to access every page. This tool shows all 37 bots as "allowed" in that case. You can still control access via server-side rules (which the live check reveals).
What is Crawl-Delay?+
A robots.txt directive that tells bots to wait a certain number of seconds between requests. It limits crawling speed to reduce server load. Google ignores it, and Search Console's crawl rate setting is gone too. To slow Googlebot down in an emergency, Google says to return 500, 503 or 429 for a day or two.
AI crawler check on every page

Lumina shows robots.txt rules and the X-Robots-Tag for the page you're on, plus, once you connect GA4, a report of visits referred by ChatGPT, Perplexity and other AI assistants. Free.

Add Lumina to Chrome — Free