Is GPTBot blocked on my site?
A site can publish welcoming robots.txt rules while its CDN or firewall refuses the same crawlers. This checker asks for your homepage as a browser and as each of four AI crawlers, compares the answers, and shows what robots.txt says alongside, so you can see whether the two agree.
A robots.txt rule is not an edge block
robots.txt is a text file of requests, read by crawlers that choose to follow the Robots Exclusion Protocol. An edge block is different: your CDN or firewall refuses the request before your site ever sees it. The site owner usually cannot see that happening. Your pages load in a browser, your robots.txt says every crawler is welcome, and the refused requests are answered above your application, so they often do not show up in its own logs.
What this tool does
It requests your homepage as a normal desktop browser, then again with the user-agent strings OpenAI and Perplexity publish for GPTBot and OAI-SearchBot and PerplexityBot, and the string ClaudeBot is seen sending (Anthropic documents the token, not the full string), and fetches your /robots.txt. That is six requests to start with, one for the browser, one for each of the four agents and one for robots.txt. Your own site's redirects are followed and another site's are not, and the check stops at 18 requests in total or after 50 seconds, whichever comes first. For each agent it compares the answer with the browser's (status, response size, and whether your own page title came back) and reports whether the agent was served, refused, or given a bot-check page. Beside that it shows what robots.txt allows or disallows for the same agent, and flags the two cases where they disagree: robots.txt allows an agent that your server refuses, and robots.txt disallows an agent that your server would serve.
Where edge blocks come from
They live in CDN, firewall or hosting settings: a bot-management option, a custom rule that matches user-agent text, a rule that refuses traffic from hosting-provider networks, or a security plugin. A setting meant to stop scrapers can catch AI crawlers as well. This tool cannot tell you which rule refused a request, only that something did.
What “challenged” means
A challenged request got an interstitial, a short “checking your browser” page, instead of your content, often with a 200 or 503 status. Crawlers do not solve challenges, so to them your site is that blank page. We recognise it from the response itself: a different status, a much smaller body, or a missing page title.
What this cannot tell you
We send each user-agent string from our own server's address, not from the vendor's network. A site that verifies crawlers by address will refuse us and still serve the real crawler, so “blocked” here means a request with that user-agent from an unknown address was refused, not proof about the vendor's own requests. ClaudeBot's string is the form seen in server logs, not one Anthropic publishes. It is one homepage request per agent: rate limits and rules for particular paths are not tested. Google-Extended is not checked here: Google documents it as a robots.txt token with no user-agent string of its own, so no request can be made as it. Its rule is covered by the robots.txt AI tester. If even the browser request is refused or challenged, we say we cannot judge the site and show no per-agent result at all.
Checking your own access logs
Logs show what actually arrived. Search your CDN's logs and your server's for the crawler names and read the status code beside each: a run of 403s next to GPTBot is an edge block. A user-agent line can be written by anyone, so treat it as a claim; the vendor pages linked above say how each company identifies its crawler. The full scan has a log check: on a report you can paste a slice of an access log and it counts which AI crawlers appear and what they got. See the methodology for how the crawler test is scored.
Questions
How do I tell whether GPTBot is blocked on my site?
Two things can block it. Your robots.txt can disallow it, and your CDN or firewall can refuse it without your robots.txt saying anything. This tool checks both and shows them side by side. If the robots.txt column says allowed but the server column says blocked or bot-check, the refusal is at the edge and changing robots.txt will not fix it.
My robots.txt allows GPTBot. Why does the tool say it is refused?
Because robots.txt is only a request that a crawler reads before it fetches. A CDN or firewall rule decides separately, and it can answer a request with an error or a bot-check page whatever robots.txt says. Look at the bot settings and custom rules in your CDN or firewall, and at any rule that refuses hosting-provider networks.
Does a blocked result prove OpenAI, Anthropic or Perplexity cannot reach my site?
No. We send the user-agent strings OpenAI and Perplexity publish, and the string ClaudeBot is seen sending (Anthropic documents the token, not the full string), from our own server, not from their networks. A rule that checks the crawler's address could refuse us and still let the real crawler in. Treat a blocked result as strong evidence that a request with that user-agent from an unknown address is refused, and check your access logs to see what the real crawlers get.
Why does the tool say it cannot judge my site?
The comparison only means something when a normal browser request succeeds. If that request was refused, challenged or got no answer, every crawler would appear to match the refusal, and reporting all of them as blocked or all as served would be a guess. So the tool says it cannot judge the site and shows no per-agent result.
To test a robots.txt file on its own, use the robots.txt AI tester.
Run the full check
This is one of the checks the full scan runs. The scan also reads your llms.txt and sitemap, looks at up to twenty-five pages, and writes the files you are missing. Read how it is scored.
Last reviewed . Maintained by Cite Files — corrections to [email protected].