Questions and answers
One page for both products: whether an AI assistant can reach, read and cite a site, and whether an AI agent could transact with it. Every answer here cites the line of code or copy it describes, so it can never say a different thing than what actually runs.
Citation Readiness
What does the Citation Readiness score measure?
One hundred points across five categories: machine access (30), discovery files (25), content structure (20), structured data (15) and trust signals (10). Every point is earned on something we measured — a request we sent, a file we found, a heading we counted — never on a model's opinion. Grades run from A at 85 and above down to F below 40.
Why does machine access carry the most points?
Fourteen of access's thirty points come from whether AI crawlers are actually served at the edge, because nothing else in the rubric matters if the crawler never receives the page. A flawless llms.txt behind a WAF that answers 403 to ClaudeBot is worth nothing at all, so the precondition is weighted heaviest, not the polish on top of it.
Sources: The /llms.txt proposal · Anthropic: how to block the crawler
What does 'turned away at the edge' mean, and how do you detect it?
We request your homepage once as a current desktop browser, then again as GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot, and compare each crawler's answer to the browser's. We send the user-agent strings OpenAI and Perplexity publish for GPTBot, OAI-SearchBot and PerplexityBot, and the string ClaudeBot is seen sending, because Anthropic documents the token but not a full string. Google-Extended is not requested: Google documents it as a robots.txt token with no user agent of its own, so we read its rule from robots.txt instead. Blocked means an error status where the browser got a page. Challenged means an interstitial — a bot-check page crawlers do not solve — decided structurally: status class, response size relative to the browser response, and whether your own title tag survived, corroborated by known challenge-page phrases rather than keyword-matched alone.
Sources: OpenAI: crawlers and user agents · Anthropic: how to block the crawler · Perplexity crawlers · Google common crawlers
What does the free summary show, and what does a free account add?
Without an account: a Citation Readiness score out of 100, the category breakdown, and the top three problems. With a free account: every finding with its evidence, a prioritised fix list, both model tests in full, and generated llms.txt, llms-full.txt, ai.txt, robots.txt, JSON-LD, an AGENTS.md template and an MCP starting point.
Sources: The /llms.txt proposal · ai.txt (Spawning) · JSON-LD 1.1
How are llms.txt and llms-full.txt graded, and what do you generate?
llms.txt is worth ten points: four just for being present and genuine — not a 200 OK HTML shell mistaken for a file — and up to six more scaled by a quality score out of 100 and its link count. llms-full.txt is worth five: three for existing, five if it also carries 300 or more words and at least one heading. With a free account we generate both from your site's own published text, alongside ai.txt, robots.txt, JSON-LD, an AGENTS.md template and an MCP starting point.
Sources: The /llms.txt proposal · ai.txt (Spawning)
What are the answerability test and the bait test, and how are quotes verified?
The answerability test hands a model only the text a crawler receives from your pages and asks seven fixed questions a customer would ask — what you do, who it's for, where you operate, how to reach you, what it costs, who runs it, how current it is. An answer only counts if the model returns a verbatim quote, and we check server-side that the quote actually occurs in the text we sent; if not, the model was filling a gap, and the question is marked unanswerable. The bait test asks the opposite: with no grounding instruction, one model describes your organisation the way it would to someone who asked, a second model must find a verbatim quote supporting each statement, and we verify that quote the same way.
What does CiteFilesBot fetch, and how do I block it?
It fetches a site only when somebody — usually its owner — asks us to audit it: robots.txt, the sitemap, a handful of conventional file paths, and up to twenty-five pages, identifying itself as CiteFilesBot/1.0 (+https://citefiles.com/bot; audit requested by a visitor). The one exception is the crawler-access test, which deliberately sends the user-agent strings of GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot against your homepage, so we can see how your edge treats them. Three are strings their vendors publish, and ClaudeBot's is the form seen in server logs. To block us, add 'User-agent: CiteFilesBot' and 'Disallow: /' to your robots.txt — we honour it, and a scan of a site that disallows us reports that it could not read the site.
Sources: RFC 9309: Robots Exclusion Protocol · OpenAI: crawlers and user agents · Anthropic: how to block the crawler · Perplexity crawlers · Google common crawlers
What does running a scan do to my server?
Up to twenty-five read-only GET requests, usually well under a minute of traffic. We identify ourselves honestly, request only public pages, do not solve challenges, do not rotate addresses, and do not touch anything behind a login. No forms are submitted and nothing is written to your site — the whole audit is arithmetic on things we fetched, never an action taken against it.
How do weekly watches work, and when do you email?
You can ask us to re-check a site once a week. We email you only when something material moved: the score shifted by two or more, a crawler started being turned away, a file appeared or vanished, or a new serious problem exists — never as a routine heartbeat, because a monitoring message that usually says nothing gets filtered, and then the one that matters gets filtered with it. Every scan records which scanner version produced it, so a change caused by us improving how we measure is labelled as ours rather than shown as your site's movement. A watch that keeps failing backs off and stops after five consecutive failures.
How long do you keep my data, and how do I have it deleted?
Scans not attached to an account are deleted after 90 days, enforced by a scheduled job. Scans in an account stay until you delete them or close the account. You can delete your own account immediately from your account page — that removes your scans, reports, watches, batches, badges and feedback and cannot be undone — or write to [email protected] and we act within 30 days. We do not sell or share what we hold, and nothing you scan is sent to a third-party model provider.
How do I remove my site from the public crawler-access record?
Write to [email protected] from an address at the domain in question, and we remove it — we do not ask why. Site names are published there only when that has been deliberately enabled in the first place; an unnamed scan already contributes only to the aggregate, never to a named row.
Why are your own scores and tiers published on your own site?
Because a versioned score is only meaningful if you can check what changed and why, and 'trust us' is not something you should accept from the party doing the scoring. Several rubric fixes were found by running our own tools on our own site, and each happened to raise our own score — one point here, two there, four elsewhere — so instead of asserting each was defensible we publish the changelog and let you read the reasoning yourself.
The full rubric — every check, its weight, and how it is measured — is at how the score works.
Agent Commerce Readiness
What is Agent Commerce Readiness?
A second audit on the same scan. The Citation Readiness score asks whether an AI assistant can reach, read and cite a site; this asks whether an AI agent could transact with it. Eight read-only checks are bundled into a tier, T0 up to T5, by published predicates. There is no score, on purpose.
Does it change my Citation Readiness score?
No. It has its own rubric version, its own table and its own predicates, and nothing in it feeds the score. Running it moves no number anywhere.
Who can run it, and does it run on its own?
Only the signed-in owner of a scan, from the report page, and only when they ask. It never runs automatically on a site.
What does it do to my server?
Never more than 40 read-only requests, usually well under that. It honours your robots.txt for every path, sends exactly one POST (an MCP server/discover, and only after a GET has answered 405), and never submits a form or a payment.
Sources: Model Context Protocol specification (2026-07-28) · RFC 9309: Robots Exclusion Protocol · Model Context Protocol specification
What do the tiers mean?
T0 Not agent-readable, T1 Agent-readable, T2 Agent-operable, T3 Agent-integratable, T4 Agent-transactable, T5 Agent-payable — each rung a bundle of reproducible facts decided by a published predicate, checkable with curl; the full predicate for each is on the methodology page. Ungraded is not a tier: it means the browser baseline itself was refused or a first-rung cell could not be read, so nothing could be compared.
Is Ungraded a bad grade?
No. Ungraded is not a tier. It means the browser baseline itself was refused or a first-rung cell could not be read, so nothing could be compared. A site we could not measure has produced no evidence, and T0 is a claim we only make when we looked and found agents refused or prices illegible.
What does citefiles.com grade on its own rubric?
T0, Not agent-readable. Our pricing page is /pro, and it carries no price yet — it describes a product that is not for sale — so the first rung fails honestly: the page was found, read as a browser and as each agent, and no price appears in it. Our MCP endpoint answers 405 to a GET and describes itself to a server/discover request, so T3's interface predicate would be met, but the ladder is climbed rung by rung and T1 is not. Until this build the same site read Ungraded, because the scanner only recognised a pricing page by the words pricing, plans or price in its path; the changelog records the change.
Sources: Model Context Protocol specification (2026-07-28) · Model Context Protocol specification · MCP specification: transports
What is DIVERGENT?
A site can serve an agent a real commerce page whose price or title differs from what a browser sees at the same URL. That site is still served, so it can still climb the ladder; the flag is raised beside the tier with both sets of prices shown, and the reader judges. It never moves the tier.
Does it test whether an agent can actually pay?
No — we never settle a challenge. T5 means the site ANSWERS with a machine-payable challenge (x402 or L402) that a credentialed agent could settle; that is an offer we can see with one request, not a transaction. Cloudflare's pay-per-crawl and a bare 402 with no machine-readable challenge stay recorded as watch items, never graded, and completing a trade leaves nothing a GET can see.
Sources: x402 payment protocol · L402: Lightning HTTP 402 protocol
What are UCP, MCP, OpenAPI and x402, in one line each?
UCP is a public manifest at /.well-known/ucp describing a merchant's agent-commerce services. MCP is the Model Context Protocol, an endpoint that answers 405 to GET and describes itself to one server/discover request. OpenAPI is a JSON description of an HTTP API with an openapi or swagger key. x402 (v1 or v2) is an HTTP 402 carrying a machine-readable challenge in a header that a credentialed agent could settle without a human; L402 is the same idea over a Lightning macaroon. We record the challenge shape and grade it at T5; we never settle one.
Sources: Model Context Protocol specification (2026-07-28) · Universal Commerce Protocol · Model Context Protocol specification · MCP specification: transports · OpenAPI Specification 3.1.0 · x402 payment protocol · L402: Lightning HTTP 402 protocol
How do I check a predicate myself?
Every predicate is a curl. Request your pricing page with a browser user-agent and again with an agent's published user-agent string and compare status, size and title. Fetch /.well-known/oauth-authorization-server and read grant_types_supported: client_credentials must appear explicitly, because the RFC 8414 default excludes it. GET your MCP path and expect 405; anything else is not an MCP endpoint by the spec.
Sources: RFC 8414: OAuth 2.0 Authorization Server Metadata · MCP specification: transports
Why is the T2 rung about robots.txt at all?
Because robots.txt is what a site says about agents and the edge is what it does, and the two are allowed to disagree. A robots.txt that allows agents while the edge challenges them is a contradiction, and only the enforced one is what an agent meets. Coherent-closed is honest and fails T2 honestly.
Sources: RFC 9309: Robots Exclusion Protocol
How do I remove my site, or stop the check running?
Write to [email protected] from an address at that domain and we remove it without asking why. To stop the check reading your site, disallow CiteFilesBot in robots.txt; the audit honours it for every path.
Sources: RFC 9309: Robots Exclusion Protocol
Why is the public record empty?
It publishes nothing below eight graded sites, and naming a host on it needs two things at once: its tier reproduced by hand from a different network than the one that measured it, recorded against that exact tier and rubric version, and this page's publishing switch turned on. A later audit that moves the tier, or a rubric-version bump, silently un-names the host. Every audit in it ran because an owner asked, so it is a fact about that sample and not about the web.
The eight cells, the tiers, and how to check every predicate yourself with curl are at Agent Commerce Readiness.
Not answered here? What our crawler fetches and how to block it is at CiteFilesBot.
Last reviewed . Maintained by Cite Files — corrections to [email protected].