Skip to content
Cite Files

All free tools

llms.txt generator

Generate an llms.txt for a site from its own words. The tool reads each page title and meta description and arranges them into the llms.txt format. It does not summarise, rewrite or add anything, so every line traces to text you already publish.

Reads your robots.txt, homepage, one sitemap and up to ten of its pages, then stops. Your own site's redirects are followed, another site's are not, and it stops at 20 requests or after 45 seconds. Six runs an hour from one connection.

What this tool does

It reads your homepage, one sitemap and up to ten of the pages that sitemap lists, then writes a file in the llms.txt format: a title, a one-line summary and a list of links. The title is your site name, taken from og:site_name or the homepage title. The summary is your homepage meta description, copied word for word. Each link is a page title, followed by that page's own meta description. No language model is involved, and the same site gives the same file twice.

How it chooses the ten pages

It reads your robots.txt first and never lists a page that it closes to CiteFilesBot. From the sitemap it drops duplicates, other hosts, addresses with query strings and files such as PDFs, then takes the ten with the shortest paths, keeping sitemap order between pages of equal depth. If the sitemap is an index of other sitemaps, it does not open them: the pages come from links on your homepage instead, and the result says so. With no readable sitemap, only the homepage is listed.

A page that does not return 200 HTML, is marked noindex, redirects, or names a different canonical is left out and shown under Left out with its reason.

What goes into each entry

If more than half of your page titles end with the same tail, such as a pipe and your brand name, the tail is removed from the link labels so they read as page names. Otherwise titles are copied as published. A page with no meta description is listed with no description. That gap is real and the file shows it, because the description is what lets an assistant choose a link without opening it. Fix it on the page and run the tool again. If your homepage has no meta description, the summary line is a visible TODO for you to write. Two or more pages under the same first folder, such as /docs/, get a section named after that folder.

How the full report differs

The full scan reads up to twenty-five pages. Its generated llms.txt, available with a free account, uses a language model to compress your page text into descriptions. That can read better where your meta descriptions are thin, and it also means a model chose the wording, so each line needs checking against the page. This tool makes the other trade: plainer output, with every word one you already publish. The model is never asked what your business does.

Limits

Ten pages, one sitemap, public pages only. The tool does not log in or run JavaScript, so a title or description that a script writes after the page loads is invisible to it. A run asks for your robots.txt, your homepage, one sitemap and up to ten pages, three pages at a time, which is 13 requests when nothing redirects. Your own site's redirects are followed and another site's are not. A hard cap of 20 requests in total, redirects included, stops the run and says so. It also stops after three failures in a row, and returns what it has after 45 seconds. Each connection gets six runs an hour. Treat the output as a draft: read it, fix the TODO if there is one, and put it at the root of your site as /llms.txt.

Questions

Why does the tool not write descriptions for me?

A description it wrote would be a claim about your site that nobody has checked. Using only your own titles and meta descriptions means every line traces to a page you published, so the only thing left to decide is whether you still stand by that wording. If a page has no description, its entry has none.

Why is one of my pages listed without a description?

Its served HTML has no meta description tag that we could read. A page that fills in its head with JavaScript after loading looks the same to this tool, because it does not run scripts. Add a description to the page itself and generate the file again.

Why were some of my pages left out?

A page is left out when robots.txt blocks CiteFilesBot from it, when it does not return a 200 HTML page, when it is marked noindex, when it redirects, or when its canonical points to a different page. Only ten pages are read in any case, shallowest paths first. The Left out list under the result names the reason for each page.

Will publishing this file make AI assistants cite my site?

This tool cannot show that it will. llms.txt is a proposal, and nothing obliges an assistant to fetch it. The generator builds the file and the validator grades its structure. Neither measures whether anyone reads it.

Run the full check

This is one of the checks the full scan runs. The scan also checks whether crawlers are actually served your pages, grades any llms.txt you already have, and writes the other files you are missing. Read how it is scored.

3 free checks a day, no account needed. Takes about a minute — a free account raises it to 25 and keeps your reports.

Last reviewed . Maintained by Cite Files — corrections to [email protected].