AI Bot & Robots.txt Checker
Check which AI bots and crawlers are allowed or blocked by robots.txt rules, then verify with live server checks using each bot's real User-Agent.
How this AI bot and robots.txt checker works
Most crawler checkers only test Googlebot. But in 2026, GPTBot, ClaudeBot, PerplexityBot, and Google-Extended matter just as much for AI search visibility. This AI bot and robots.txt checker tests 37 crawlers — including Googlebot, OpenAI's new OAI-AdsBot (ChatGPT ads validator), and every major AI bot — then fires a live request at your homepage with each one's User-Agent string. Google-Extended and Applebot-Extended are robots.txt tokens without a user agent of their own, so for those two the server column only shows how your server treats that string.
That second check is the one people miss. A bot can be "allowed" in robots.txt but still get a 403 from a Cloudflare rule or a CDN bot-protection layer. This tool shows both layers: what your robots.txt says, and what actually gets through. It works the other way round too. Google-Agent, ChatGPT-User and Perplexity-User fetch pages when a person asks for them, and their operators say robots.txt generally doesn't apply there. A Disallow line won't stop them, only a server rule will. And watch the search bots: GPTBot collects training data, but ChatGPT search fetches pages with OAI-SearchBot. Block that one at the server and no amount of content work gets you cited there.
Common robots.txt errors this checker finds
Most mistakes are dumb, not subtle. The classic: a Disallow: / left over from a staging environment that nobody removed before going live. The checker shows it as BLOCKED for every bot it applies to. A missing robots.txt shows up too, which means every URL is crawlable by default.
Path rules need a human eye. Disallow: /folder also blocks /folder-archive and /folder.html, because rules match by prefix. The checker lists each bot's restricted paths under "rules" (hover the badge), and the raw file sits at the bottom of the results, but it doesn't test a single URL against them.
Is my robots.txt blocking search engines?
Run this checker against your live URL. A site-wide block on Googlebot shows up as BLOCKED in the first result row. Partial blocks show as "rules", and those are where most accidents hide: a wildcard pattern (Disallow: /*.php) meant for one path that matches thousands. Google caches robots.txt for up to 24 hours, so a fix can take a day to register. Search Console's robots.txt report lets you request a recrawl of the file.
robots.txt tester vs. robots.txt checker
Google retired the robots.txt Tester inside Search Console in late 2023. The replacement is a short report under Settings → robots.txt that only shows the currently fetched file. It doesn't let you test custom user agents or preview rule changes. This checker fills part of that gap: it runs your live file against 37 named crawlers and checks what the server actually returns to each. It doesn't take a draft file or a custom User-Agent, so test rule changes on staging first.
Explore more tools
Meta Tag Analyzer
Full meta tag audit for any URL.
llms.txt Generator
Create AI crawler guides.
Heading Checker
Analyze H1-H6 hierarchy.
Security Headers
Check HTTP security headers and server config.
JS vs No-JS
See what content disappears without JavaScript.
FAQ
Lumina shows robots.txt rules and the X-Robots-Tag for the page you're on, plus, once you connect GA4, a report of visits referred by ChatGPT, Perplexity and other AI assistants. Free.
Add Lumina to Chrome — Free