Skip to content
Cite Files

Rubric v1.2.0 · scanner v1.11.0 · what has changed

How the score works

A score whose arithmetic is secret is just an opinion with a number on it. Here is all of it: every check, the weight it carries, and how each one is measured.

The rubric

One hundred points across five categories. Access dominates because nothing else matters if the crawler never receives the page — a perfect llms.txt behind a WAF that returns 403 to ClaudeBot is worth nothing at all.

Machine access

30 points
  • 14AI crawlers are served at the edge
  • 10robots.txt permits AI crawlers
  • 6Content is present without JavaScript

Discovery files

25 points
  • 10Sitemap (present, declared in robots.txt, carries lastmod, not stale)
  • 10llms.txt (present, and its own quality)
  • 5llms-full.txt

Content structure

20 points
  • 5Every page has a meaningful title
  • 4Pages carry a meta description
  • 3Exactly one H1 per page
  • 4Content is broken up by headings
  • 4Pages have enough text to answer a question

Structured data

15 points
  • 6Pages publish JSON-LD
  • 4The site declares who publishes it
  • 2WebSite / WebPage entities present
  • 3Content-specific schema types (the ones we recommend for your sector)

Trust signals

10 points
  • 4About and contact pages exist
  • 2Content is attributed to someone
  • 2Content is dated
  • 2Machine-addressed extras (agent docs, security.txt)

Grades: A is 85 and above, B is 70, C is 55, D is 40, F is below that. Broken JSON-LD costs two points on top of scoring nothing, because structured data that fails to parse signals care without delivering it.

The crawler access test

We request your homepage as a current desktop browser, then again as GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. For GPTBot, OAI-SearchBot and PerplexityBot we send the user-agent string the vendor publishes. Anthropic documents the ClaudeBot token but not a full string, so for ClaudeBot we send the form seen in server logs. Then we compare the four answers against the browser one.

Google-Extended is not requested. Google documents it as a robots.txt token with no user agent of its own, so no request can be made as it; its rule is read from robots.txt with the other tokens, below.

A verdict of blocked means an error status where a browser got a page. Challenged means an interstitial: a bot-check page, which crawlers do not solve. We decide that structurally — status class, response size relative to the browser response, and whether your own <title> survived — rather than by looking for vendor phrases, because block-page wording changes constantly and a keyword test produces confident nonsense in both directions.

This is a single request per crawler from one address at one moment. It cannot see rules that depend on request volume, geography or reputation, and a site that was fine when we asked can be blocking an hour later. Treat a clean result as evidence, not a guarantee.

robots.txt

Read the way RFC 9309 says a crawler reads it: the most specific matching user-agent group wins, * is the fallback, the longest matching rule decides, and Allow beats Disallow on equal length. A file containing the word “Disallow” is not necessarily blocking anyone, and a file that looks welcoming can still block everything.

We evaluate the root path for 27 crawler tokens. We also detect Cloudflare’s managed robots.txt, which is injected above your own file and disallows the AI crawlers — so the rules you wrote are never reached.

File detection, and the phantom-file problem

Many sites answer 200 OK with their HTML shell for any path they do not recognise. A checker that treats 200 as “found” will report an llms.txt on a site that has never had one.

So before anything else we request a deliberately impossible path. If that returns 200, every subsequent 200 is compared against the shell it returned. A file only counts as present when the body is a file rather than a page.

Paths we check:

  • /llms.txt
  • /llms-full.txt
  • /.well-known/llms.txt
  • /soul.md
  • /memory.md
  • /AGENTS.md
  • /ai.txt
  • /.well-known/ai-plugin.json
  • /.well-known/mcp.json
  • /openapi.json
  • /.well-known/security.txt
  • /humans.txt

ai.txt, and the distinction it exists for

Almost everyone conflates two separate questions, so we keep them apart. Retrieval is whether an assistant may fetch your page to answer a question now, and cite you — governed by robots.txt and by whether your edge actually serves the crawler. That is what most of this report is about. Training is whether your content may become weights in a model. That is a licensing question, and ai.txt is where it is declared.

Wanting the first and not the second is a perfectly coherent position, and a common one. Without a declaration nobody can tell which you meant, and silence tends to be read as whatever suits the reader.

Two caveats we print every time: ai.txt is a convention rather than a ratified standard, and adoption is partial. It records intent, it does not enforce anything. It matters because a stated reservation is evidence of intent and an unstated one is not.

We read it with the same RFC 9309 grammar as robots.txt, and flag the contradiction that actually matters: granting training use while disallowing the crawlers from reading the site at all. The reverse — allowing retrieval and reserving training — is coherent and we leave it alone. The ai.txt we generate is derived from the position your robots.txt already expresses; we do not hold a view on how you should licence your work.

Agent Commerce Readiness — a second question on the same scan

A plainer explanation, with questions and answers, is at /agent-commerce-readiness.

ACR rubric v1.1.0 · scanner v1.11.0

The Citation Readiness rubric above answers “can an AI assistant reach, read and cite this site?” This is a different question run on the same scan: “can an AI agent transact with it?” — authenticate without borrowing a human’s password, discover an API, read a price, obtain a credential of its own.

There is no score here, on purpose. The capability matrix is the ground truth — eight cells, each a reproducible fact about what your server said to us — and the tier is a name for a bundle of those facts, decided by published predicates. Every predicate can be checked with curl. It never feeds the Citation Readiness score, the same firewall we keep around technical health, and no language model is consulted anywhere in it. It runs only when you ask for it, sending never more than 40 read-only requests against your site, usually well under — never a form, never a payment — and it honours your robots.txt for every path it touches.

The five states

Every cell we measure lands in exactly one of five states: value (asked, answered, and the answer carries a verdict), absent (asked, answered authoritatively, and there is nothing there), unreadable (asked, and we could not get an answer we trust — timeout, network error, an unparseable body), blocked (asked, and the site refused us), and inconclusive (the control this cell depends on failed, so the measurement has no reference point). Absent is not unreadable is not blocked; a site we could not read has not earned a bad grade, it has produced no evidence.

The eight cells

Sourced from the same table the matrix and the public index read, so this list can never describe a check differently than the one that actually ran.

  • Agent access to the commerce page — The same commerce page requested as a browser and again as each AI agent, compared for whether the price or title an agent sees matches what a browser sees.
  • Price legibility — Whether a price is somewhere a machine can read it without running the page. Prices are read as text because none of the nine real pricing pages we surveyed marked them up with JSON-LD; a rubric keyed on schema would score everybody zero and tell nobody anything.
  • Rate-limit and error shape — When an agent hits your API unauthenticated, is it told how to behave: standard rate-limit headers, a machine-readable error body, or neither.
  • robots.txt vs the edge — What your robots.txt says about agents against what your edge actually does to them, because the two are allowed to disagree and frequently do.
  • Machine credentials (OAuth) — Whether an agent can obtain its own credentials, read from OAuth authorization-server and OpenID discovery documents. client_credentials must be declared explicitly, because RFC 8414 says an absent grant-types array defaults to authorization_code and implicit — so silence on this one is a negative, not an unknown.
  • UCP manifest — /.well-known/ucp, the one agent-commerce protocol with a mandatory public discovery document, validated against the shape two live merchants actually serve rather than an idealised reading of the spec.
  • OpenAPI document — A discoverable, parseable API description. A 200 only counts if the body is JSON with an openapi or swagger key, because a single-page app will happily serve its app shell at any path we guess.
  • MCP endpoint — The spec mandates a 405 to a GET on an MCP endpoint, which is how we fingerprint one for free; only where that fires do we send the single server/discover request, a parameterless, cacheable metadata read the spec itself marks public.

The tiers

The tier is the highest rung whose predicates are all proven true. An unknown cell stops the climb there and is not counted against the site — you cannot reach a tier through a cell we could not read, and it never denies one either. T0 is a positive claim — we looked, and agents are refused or cannot read a price — never a default.

TierPredicate
T0 · Not agent-readableAgents are refused the commerce pages, or no price is readable without running the page.
T1 · Agent-readableAn agent gets the commerce page, and the prices are text the server sent.
T2 · Agent-operableAn unauthenticated API request is answered with something a machine can act on (rate-limit headers, a machine-readable error, or a structured JSON answer), and what robots.txt says about agents agrees with what the edge does.
T3 · Agent-integratableA real machine interface exists — a parseable OpenAPI document, a live MCP endpoint, or a valid UCP manifest. Any one.
T4 · Agent-transactableAn agent can obtain its own credentials — client_credentials declared explicitly, or dynamic client registration.
T5 · Agent-payableThe site answers with a machine-payable challenge — x402 (v1 or v2) or L402 — so an agent that holds credentials could settle without a human. We observe the challenge; we never settle it.
UngradedNot a tier: the browser baseline itself was refused, so nothing could be compared. A site we could not measure has not earned a grade; it has produced no evidence.

DIVERGENT

A site can serve an agent a real commerce page whose price or title differs from what a browser sees on the same URL — silent wrongness, because nothing signals it, and an agent quoting the page quotes a number the human next to it never saw. That site is still served, so it can still climb the ladder: DIVERGENT is raised beside the tier, first-class, with both sets of prices shown, and the reader judges. Folding it into the tier would be an editorial weight, and the first site it capped would have grounds to argue with it. If that ever changes, it changes here, with a version bump and an entry in the changelog.

Watch items — recorded, not scored

Observed on none of the 12 real sites we measured when this shipped. They are recorded so that the day adoption arrives, the rubric changes with a changelog entry, not a retraction.

  • Payment challenge (402 family) — The 402 family, recorded as an observed challenge: x402 v1, x402 v2 and L402 are graded at T5 — a machine-payable challenge a credentialed agent could settle. Cloudflare's pay-per-crawl is a crawler toll, not a commerce offer, and a bare 402 with no machine-readable challenge is payment-signalling, not agent-payable — both stay recorded only.
  • RFC 9728 protected resource metadata — A protected-resource metadata document, when one is published, recorded alongside its authorization servers.
  • agents.json — An agents.json manifest, when one is published, recorded with its declared schema and capabilities.
  • ai-plugin.json — Recorded for staleness only — the OpenAI plugin programme it belonged to ended in April 2024, not as a capability.

What we do not measure, and why

  • Structured receipts: they sit behind a completed transaction, and we never transact.
  • The OpenAI/Stripe Agentic Commerce Protocol and Google's AP2: both describe exchanges between parties who have already agreed to trade, and leave no server-side artefact a GET request can see.
  • Anything that needs a login we have no credential for. When you give a scan a staging basic-auth username and password, it is reused for ACR requests to that scan's own origin only — never sent to any other host a probe discovers — and we still cannot reach a login only a human browser session establishes.

Sources: RFC 8414: OAuth 2.0 Authorization Server Metadata · RFC 7591: OAuth 2.0 Dynamic Client Registration · RFC 9728: OAuth 2.0 Protected Resource Metadata · RFC 9457: Problem Details for HTTP APIs · IETF draft: RateLimit header fields for HTTP · OpenAPI Specification 3.1.0 · Model Context Protocol specification · Universal Commerce Protocol · x402 payment protocol · L402: Lightning HTTP 402 protocol

Where our tools disagree

We recommend declaring Content-Signal in robots.txt, and Lighthouse marks any robots.txt containing it as invalid. Both of those are true at once, so rather than pick a side we measured it: removing that single line takes this site’s Lighthouse SEO score from 92 to 100 and changes nothing about what a crawler does.

Lighthouse’s parser predates the directive. We keep the line and accept the eight points, because the declaration is the point and an SEO sub-score is a proxy. If you follow the advice and your number moves, that is why — it is flagged in your technical report rather than left for you to discover.

The answerability test

We give a language model only the text a crawler receives from your pages — no search, no browsing, no outside knowledge — and ask it seven questions a customer would ask an assistant about you.

An answer only counts if the model returns a verbatim quote from your text, and we then check that quote really appears in what we sent. If it does not, the model was filling a gap rather than reading one, and the question is marked unanswerable. Those discarded answers are counted and shown to you, because that is precisely what happens when a real assistant is asked about you and your pages come up short.

The bait test

The answerability test asks fixed questions. This asks the opposite: when a model is asked about you and your pages come up short, what does it say instead?

One model is given your page text and asked to describe your organisation the way it would to someone who asked — with no instruction to stay grounded, because that is how a real assistant behaves. A second, independent model then has to find a verbatim quote supporting each statement, and we check that quote actually occurs in the text we supplied. Statements that survive neither gate are the shape of the thing an assistant will tell a customer that your site never said.

This is a measure of your content’s ambiguity, not of the model’s honesty. Where your pages are specific, models stay accurate.

Sector, and how it is decided

Sector drives two things: which schema.org types we recommend, and which peer group your score is compared against. It is decided by rules, not by a model, and the report shows the signals that produced it — a classification you cannot audit is not worth acting on.

Two kinds of signal, and the difference matters. Presence signals are settled by a single page: a checkout page means a shop however small, a LegalService entity means a practice. Share signals only mean something in proportion — nearly every company has a blog, so having one says nothing, while two thirds of the site being articles says publisher. Counting both the same way is how a payments company gets classified as a newspaper, which is exactly what happened before this distinction existed.

When the evidence points two ways at once we say “general” rather than pick. The recommended schema list for a sector is fixed, so two competitors get identical advice and neither gets a type invented for them.

Peer comparison

Your score against the sites we have scanned in your sector: the median, the middle half, and where you sit. One scan per site — the most recent — so a site scanned forty times during a fix session cannot dominate its own sector.

Below eight comparable sites we widen to every sector; below eight again we show nothing, because a median of three is not a benchmark.

The caveat is the important part. Our sample is people who came here because they suspected a problem, so it skews low. Beating this median means you are ahead of sites whose owners were already worried — not ahead of the web. Treat it as a floor.

Watching for changes

You can ask us to re-check a site once a week. We email you only when something material moved: the score shifted by two or more, a crawler started being turned away, a file appeared or vanished, or a new serious problem exists. Cosmetic churn does not earn an email, because a monitoring message that usually says nothing gets filtered — and then the one that matters gets filtered with it.

Every scan records which build of the scanner produced it. When we improve how we measure, comparisons across that change are labelled as ours rather than presented as your site’s movement. This exists because of a real false alarm: a sampling improvement moved a site by six points and the watch told its owner their site had improved. It had not — we had.

A watch that fails backs off, and gives up after five consecutive failures rather than re-scanning an abandoned domain every week forever.

What the model is not allowed to do

Every measured check on your report is arithmetic on things we fetched. No score is ever set by a model.

The model writes the commentary, and it compresses your own page text into descriptions for your llms.txt. It is never asked what your business does, because it would answer. Where a generated file needs a fact we have no source for — your social profiles, your licence, your correction policy — you get a visible TODO instead of a plausible invention.

If the model is unavailable, the report still arrives: you get the measured half in full and a note saying the commentary was skipped.

Sampling

We read your homepage plus up to twenty-four more pages, chosen from your sitemap where you have one and from your homepage links where you do not. Selection favours the pages that describe the business — about, services, pricing, docs, contact — and caps any one section at three pages, so a large blog cannot crowd out everything that matters. A site’s nine-hundredth post tells us nothing the third did not.

Pages your robots.txt disallows are excluded from the sample. Partly courtesy, but mainly because including them measures the wrong thing: judging your content on a sign-in form you explicitly told crawlers to skip scores you on a page no crawler will ever read. We found this on our own site, where /login and /register were quietly dragging the content score down.

This means structure and schema figures are a sample, not a census. Where a category is scored on a proportion, the report says how many pages that proportion is out of.

How we fetch

Requests identify themselves as CiteFilesBot except in the crawler-access test, where the point is to send a specific agent string. We read one page at a time per site with a small amount of parallelism, cap every response, and stop after a fixed budget. We only ever request public pages over HTTP and HTTPS on the standard ports. More about our crawler.

Found a check that is wrong, or a site we score unfairly? Tell us at [email protected] — the rubric is versioned so corrections are traceable.

Last reviewed . Maintained by Cite Files — corrections to [email protected].