SEOscanner.app

Free check · llms.txt

llms.txt checker

Check whether a site serves /llms.txt, what the server returns, and whether the file follows the proposal's format. Google says Search ignores llms.txt, and this tool says so too.

Reads robots.txt and /llms.txt on the site's preferred origin. No pages are crawled. Free, no account needed.

What the checker reads

It requests /llms.txt from your site's preferred origin after reading robots.txt, which it obeys for its own user agent. The results give the status, content type and size, say whether the server sent an HTML page instead of a text file, and set the text beside the format the proposal describes: an H1 name, a blockquote summary and H2 sections of links. Link targets that do not parse as URLs are listed.

Nothing here is scored, and nothing here changes the full audit's score. The scanner records llms.txt for information only, whether the file exists or not, and never recommends adding or removing one. The AI crawler methodology explains why AI-facing signals sit outside the score.

What is established, and what is not

The evidence on llms.txt sorts into three tiers: what Google has documented, what the proposal itself says, and what no primary source supports.

Documented by Google

Google addresses the file directly in its guide to generative AI features in Search:

“Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”

[G-AI-OPT] Mythbusting generative AI search: what you don't need to do

The same passage leaves the door open for other uses:

“It's completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files.”

[G-AI-OPT] Mythbusting generative AI search: what you don't need to do

Both sentences sit in the part of Google's guide headed “Mythbusting generative AI search: what you don't need to do”. The scanner's myth guard, the list of beliefs it will never score, records two related claims:

From the myth guard, which lists every claim the scanner refuses to score and why.
ClaimVerdictWhat the source says
llms.txt is required for AI visibilityContradicted (Google)Google Search ignores llms.txt; it "will neither harm nor help" [G-AI-OPT]
Special markup or files are needed for AI searchContradicted (Google)"there's no special schema.org markup you need to add" [G-AI-OPT]

Described by the proposal

llms.txt is a format proposed by Jeremy Howard. It is not a web standard and not a search engine feature, and its author is the primary source for what the format is, never for how a search engine treats it. The proposal defines the structure this checker compares against:

  • “An H1 with the name of the project or site. This is the only required section”[LLMSTXT]
  • “A blockquote with a short summary of the project, containing key information necessary for understanding the rest of the file”[LLMSTXT]
  • “Zero or more markdown sections delimited by H2 headers, containing “file lists” of URLs where further detail is available”[LLMSTXT]
  • “The “Optional” section is used, by convention, for secondary information: links an agent can skip when a shorter context is needed.”[LLMSTXT]

Its intended reader is an assistant at work: “llms.txt information is instead used on demand, when an agent needs information about a topic while assisting a user.”[LLMSTXT] The proposal is also candid about where the format has taken hold: “llms.txt files are used most heavily for software documentation, where coding agents follow them to find API references and tutorials.”[LLMSTXT]

Not established

No source in the scanner's registry establishes llms.txt as a ranking factor for any search engine. The registry holds no Bing statement about the file, and its records of the AI crawler operators say nothing about reading it. This page therefore makes no claim that llms.txt affects visibility in search results or in AI answers, in either direction.

If you keep one

The proposal's own test is the most practical one:

“Test your file by asking an agent questions about your content, giving it only your llms.txt as a starting point.”

[LLMSTXT] Example

Beyond that, a few problems make a file unusable whatever its purpose:

  • The server returns an HTML page at /llms.txt, such as the home page or a styled error page served with status 200. An agent expecting Markdown receives markup.
  • The first line is not an H1, the one section the proposal requires.
  • Links are malformed or point at pages that have moved. The checker flags targets that do not parse; it does not fetch them.
  • The summary describes a site that has since changed. This checker cannot judge that; the review described below reads the file against your pages.

The full scanner's advisory AI review reads llms.txt alongside your pages and can draft one when it would help. It is labeled as model output and never enters the score. To see which AI crawlers your robots.txt admits, use the AI crawler checker.

The rules this checker runs

Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.

Primary sources

The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.

This page checks one thing. The full audit runs all 70 scored rules.

It covers 16 categories on up to 200 pages, orders the fixes by the strength of the evidence behind them and prints its score with the arithmetic. Every finding cites the primary source it rests on. No account is needed.

Run the full 70-rule audit