Skip to content
Cite Files

All free tools

ai.txt generator

ai.txt is where a site states its terms for AI training use, which is a separate question from whether assistants may fetch and cite it. This tool writes one that matches the position your robots.txt already takes, rather than choosing a licensing stance for you.

Used only to label the file. Domain name only, without https:// or a path.

Nothing is fetched. Leave it empty to see what the file says when there are no rules.

What ai.txt is for

An ai.txt file states a site's terms for AI training use of its content. robots.txt (RFC 9309) answers a different question: which crawlers may fetch which paths. Retrieval, where an assistant fetches your page now to answer a question and cite you, depends on access. Training, where your content becomes part of a model, is a question of stated terms. A site can want the first and not the second, and neither file decides that for the other.

What this tool does

It reads the robots.txt you paste or fetch and checks the main AI user-agents the scan tracks, such as GPTBot, ClaudeBot and PerplexityBot. If robots.txt disallows any of them from the whole site, the file reserves training rights to match. If it disallows none, the file records that retrieval and citation are already permitted and takes the training position from your Content-Signal line if it has an ai-train entry. With no such entry it leaves training as an explicit choice, with both alternatives written out and the reservation active as a template default. If no robots.txt rules are found at all (the file is missing, served as a web page, empty or unreadable), the file says so, says that is not the same as the site having stated a position, and presents its training stance as a template default for you to confirm. It copies your existing position; it does not recommend one, and Cite Files has no view on how you should license your work.

If your robots.txt carries a Content-Signal line with ai-train=yes, the file permits training to match; with ai-train=no it reserves. Either way the file repeats the line in a note and the explanation says which stance is active and why. A Content-Signal line does not override a robots.txt that disallows an AI user-agent: then the file reserves, and if that robots.txt also says ai-train=yes the file states plainly that the two disagree, so you can resolve it in robots.txt and generate again.

What it does not do

An ai.txt is a declaration, not an enforcement mechanism. Compliant crawlers follow robots.txt; refusing one that ignores it takes a rule at your edge. This page does not know which crawlers read an ai.txt. In fetch mode the tool requests /robots.txt and nothing else (at most 3 requests if your own site redirects, for example from http to https to www; a redirect to another site is not followed), so it does not read an ai.txt you already publish. A pasted robots.txt is never fetched. The full check does read an existing ai.txt, and it flags the combination that matters: granting training use while robots.txt disallows AI crawlers from reading the site at all.

If you want a different position

Change robots.txt first, then generate again, so the two files never disagree. You can also edit the generated file by hand; its comments mark which lines to delete or uncomment.

Questions

Is ai.txt a standard?

No. It is a convention rather than a ratified standard, and adoption is partial. It records what you intend; it does not enforce anything.

Will the generated ai.txt stop AI companies training on my site?

No. It states your terms, and whether a crawler reads it is up to the crawler. Compliant crawlers follow robots.txt. Refusing one that ignores it takes a rule at your edge, such as a CDN, firewall or server rule.

My robots.txt allows every crawler. Why does the file have a reservation switched on?

If your robots.txt has no Content-Signal line with an ai-train entry, it says nothing about training, so there is no position of yours to copy. The file then ships with the reservation (Stance A) active as a template default and the permit (Stance B) commented out, and its own comments say robots.txt gave no training signal and that you should confirm it. If your robots.txt declares ai-train=yes, the permit (Stance B) is active instead; with ai-train=no the reservation is active because you said so. Delete or uncomment whichever lines you want before you publish it.

Is my robots.txt or the generated file stored?

We keep a count of how many times each tool was run each day. To enforce the hourly limit, your IP address is held in a rate-limit counter that is deleted within two days. The address you test, the text you paste and the result are not stored. In fetch mode the tool requests one file, /robots.txt, as CiteFilesBot: at most 3 requests if your own site redirects, and a redirect to another site is not followed. A pasted file is never fetched. Each connection is limited to 30 runs an hour.

Run the full check

This is one of the checks the full scan runs. The scan also tests whether crawlers are actually served your pages, grades the ai.txt you already have, and writes the other files you are missing. Read how it is scored.

3 free checks a day, no account needed. Takes about a minute — a free account raises it to 25 and keeps your reports.

Last reviewed . Maintained by Cite Files — corrections to [email protected].