Skip to content
Cite Files

All free tools

robots.txt tester for AI crawlers

Whether a crawler is allowed depends on which user-agent group matches it and which rule is most specific, which is easy to get wrong by eye. Paste a robots.txt, or enter your address and we fetch it, and the tester shows for each AI and answer-engine crawler whether it is allowed or disallowed, and which rule decided.

Up to 250,000 characters. Pasted text is read in your request and never fetched from anywhere.

Compared with the rules only; this path is not requested. Start it with /.

Snippet builder

The robots.txt group that would allow or disallow one agent for the whole site. It is a snippet to copy, not a recommendation. A group that names an agent replaces the * group for that agent; it does not add to it.

Rule
User-agent: GPTBot
Disallow: /

How a crawler picks its group and its rule

A robots.txt is made of groups: one or more User-agent lines followed by Allow and Disallow rules. Under RFC 9309, a crawler finds the group whose user-agent line most specifically matches its own name and uses User-agent: * only when no group does. Inside that one group, the longest matching path rule decides, and Allow wins when two rules are the same length. This tester follows that order. Like the full scan, it matches a user-agent line as a case-insensitive part of the crawler’s name, and its rules may use * and a closing $.

Why the * rules do not apply to an agent with its own group

Groups are not layered. Once a crawler has a group of its own, the * group is not consulted for it at all:

User-agent: *
Disallow: /private

User-agent: GPTBot
Allow: /

Here GPTBot is allowed at /private, because the only group that applies to it has no rule against that path. Every other agent in the list is disallowed there. This is the most common way a file does the opposite of what its author meant.

One company, several tokens

Vendors publish more than one user-agent token, each for a different job: gathering content for model training, building a search index, or fetching a page because a person asked an assistant about it. A rule for one token does not govern another unless the file says so. Which token does which job is stated by the vendor, not by this page: see OpenAI, Anthropic, Perplexity and Google. The tester evaluates 27 tokens and lists the 11 the scanner treats as critical first.

What this tool does and does not do

Pasted text is evaluated and never fetched. With an address, it requests /robots.txt on that site as CiteFilesBot, and nothing else: one request, or up to three if your own site redirects it, for example from http to https or to the www address. A redirect to another site is reported and not followed, and the check gives up after 12 seconds. The path you enter is compared with the rules and is never requested. It reads User-agent, Allow, Disallow, Sitemap and the Content-Signal line, quotes the last two, and does not validate other directives.

It does not read your robots.txt before fetching it, which would be circular. It asks for the one named file directly, because robots.txt is a file published for crawlers to read. The llms.txt generator and the full scan do honour robots.txt before reading pages.

It does not test whether a crawler is actually served your pages. robots.txt is a request, not an enforcement, and an edge rule can refuse a crawler the file allows; the AI crawler access checker compares the two. When the same agent is named in two groups, RFC 9309 combines them but this tester reads only the first, so it flags the file for a manual check.

Reading the result

“No rule matched” means nothing in the group covers that path, so the agent is allowed. A file that answers 200 with an HTML page is not a robots.txt, so no verdicts are shown; the same goes for a server error or a refused request. A file with no robots.txt at all is shown as every agent allowed.

Questions

Does a Disallow rule under User-agent: * block GPTBot?

Only if the file has no group that names GPTBot. If it does, GPTBot follows that group on its own and the rules under User-agent: * are ignored for it, so a rule meant for every crawler has to be repeated inside each named group.

What happens when a site has no robots.txt?

Nothing is disallowed, so every crawler is allowed by default. This tester shows every agent as allowed and says the file is absent. If the server answers with an error instead of a missing-file status, it shows no per-agent result.

Which rule wins when an Allow and a Disallow both match?

The longer matching rule wins, whatever order the lines are in. If the two are the same length, Allow wins.

If robots.txt allows a crawler, will it be able to read the page?

Not necessarily. robots.txt is a request that well-behaved crawlers follow, not something the server enforces. A firewall, bot-protection rule or CDN setting can refuse a crawler that robots.txt allows, which this tester does not check.

Run the full check

This is one of the checks the full scan runs. The scan also tests whether those crawlers are actually served your pages, reads your sitemap and llms.txt, and writes the files you are missing. Read how it is scored.

3 free checks a day, no account needed. Takes about a minute — a free account raises it to 25 and keeps your reports.

Last reviewed . Maintained by Cite Files — corrections to [email protected].