Skip to content
Cite Files

All free tools

llms.txt validator

An llms.txt file is meant to give an assistant a short, readable map of a site, but a file with structural errors is harder for a tool to use, and a site that answers /llms.txt with an HTML page has no file at all. Paste yours, or enter your address and we fetch it, and the validator lists what is wrong.

We request /llms.txt on that site as CiteFilesBot, and nothing else: one request, or up to three if your own site redirects it. A redirect to another site is not followed.

What an llms.txt file is

It is a markdown file at the root of a site, set out in the llms.txt proposal. It opens with one H1 title, then a blockquote that says in a sentence or two what the site is. After that come H2 sections, each a list of links written like - [Page name](https://example.com/page): what it covers. A section called Optional marks links an assistant may skip when it is short of room. The file is a curated index of the pages worth reading, not a sitemap.

What this tool does and does not do

Pasting a file grades the text and makes no request at all. Entering an address requests /llms.txt on that site as CiteFilesBot (the bot page lists what it fetches). That is one request, or up to three if your own site redirects it, for example from http to https or to the www address. A redirect to another site is reported and not followed. The check gives up after 12 seconds. It does not fetch your home page, your llms-full.txt or any link inside the file.

It does not read your robots.txt first. It asks for the one named file directly, because /llms.txt is a file published for crawlers to read. The llms.txt generator and the full scan do honour robots.txt before reading pages.

It checks structure: the title, the summary, the sections, whether links are absolute and described. It does not check that the links load, that the descriptions are true, or that any assistant reads the file. If you paste without entering your domain, the count of links pointing elsewhere is not shown, because a pasted file has no address to compare against. Every other finding is unaffected.

How to read the score

The quality score runs from 0 to 100 and grades the file itself. It starts at 100 and loses 34 for each critical problem, 12 for each major one and 4 for each minor one, never going below 0. It is the same grading the full scan applies to your llms.txt, where presence and quality together carry 10 of the scan's 100 points; the methodology lists the rest. When there is no file there is no score, only a reason. A 404 means there is no file. A 403 means the server refused us, so we cannot tell, and that is reported as inconclusive rather than as a failure.

The HTML-shell trap

Many single-page apps and catch-all routes answer every unknown path with 200 OK and the site's HTML page. A check that only looks at the status reports a file that does not exist, and a crawler that asks for /llms.txt gets your app shell. The validator reads the body. If it is an HTML document, it says that the server answers 200 at /llms.txt but is serving a page, and it does not grade that page as a badly written llms.txt. An HTML error page behind a 404 is simply a missing file and is reported that way.

llms-full.txt

A companion file, called llms-full.txt by convention, carries the full text of the pages the index points to in one document, so an assistant can read the content without crawling each link. This tool does not fetch or grade it. The full scan does, as a separate check.

Questions

Does an llms.txt change how AI assistants cite my site?

Not on its own, and this tool cannot show that it does. llms.txt is a proposal, and nothing obliges a crawler or an assistant to fetch it. What a good file can do is give a tool that does read it a short, accurate map of your site. If your server turns AI crawlers away, a perfect file is never read, which is why the full scan weighs crawler access more heavily than this file.

Why does the validator call my file an HTML page when I can see it in my browser?

A browser shows whatever the server sends. If the address returns your site's page layout instead of plain text, the body is HTML and the validator reports it as such. Fetch the address with curl -i, or view the source: a real file starts with a hash and a title, not an angle bracket.

Does a high score mean my links work?

No. The score grades structure only. The tool does not request any link in the file and does not check that a description is true, so a file can score 100 and still point at pages that no longer exist.

Is my pasted text or my address stored?

We keep a count of how many times each tool was run each day. To enforce the hourly limit, your IP address is held in a rate-limit counter that is deleted within two days. The address you test, the text you paste and the result are not stored. Pasted text is graded and never fetched. Each connection is limited to 30 runs an hour.

Run the full check

This is one of the checks the full scan runs. The scan also checks whether crawlers are actually served your pages, reads your robots.txt and sitemap, and writes the files you are missing. Read how it is scored.

3 free checks a day, no account needed. Takes about a minute — a free account raises it to 25 and keeps your reports.

Last reviewed . Maintained by Cite Files — corrections to [email protected].