SEOscanner.app

Free check · robots.txt

robots.txt checker and validator

Fetch a site's robots.txt, parse it the way Google documents, and see which URLs Googlebot and bingbot may crawl, which lines are ignored and which sitemaps it declares.

Reads /robots.txt, the sitemaps it declares (or /sitemap.xml) and the home page. Sitemap URLs are tested against the rules without being fetched. Free, no account needed.

What the checker reads

The checker requests /robots.txt from the host your address finally lands on, following up to five redirects, the minimum Google documents. It parses the file with the algorithm Google describes and RFC 9309 standardizes. Each crawler obeys the group that names it most specifically; inside that group the longest matching path wins, and an Allow beats a Disallow of equal length.

It then tests the home page and every URL in your sitemaps against the Googlebot and bingbot groups. Those URLs are not fetched, because the question is narrower: does the file permit fetching them? A sitemap URL that robots.txt blocks is a page you have asked search engines to find and then told them not to read.

The results show:

  • whether robots.txt answers, and with which status;
  • whether it is valid UTF-8 and within Google's 500 KiB parsing limit;
  • every group, rule and Sitemap line as parsed, with the lines Google ignores marked;
  • which intended URLs Googlebot and bingbot may not fetch, with the rule and line responsible;
  • CSS, JavaScript and images your home page loads that the file blocks;
  • noindex lines, which have no effect in robots.txt.

What robots.txt controls, and what it does not

robots.txt governs crawling. RFC 9309 defines its two rule types in a sentence:

“These lines indicate whether accessing a URI that matches the corresponding path is allowed or disallowed.”

[RFC9309] 2.2.2 The "Allow" and "Disallow" Lines

Indexing is a different matter. Google says robots.txt “is not a mechanism for keeping a web page out of Google”[G-ROBOTS-INTRO]. A disallowed URL “can still be indexed if linked to from other sites”[G-ROBOTS-INTRO], and in that case “the search result won't have a description”[G-ROBOTS-INTRO]. For keeping a page out of results, the same documentation points elsewhere:

“To keep a web page out of Google, block indexing with noindex or password-protect the page.”

[G-ROBOTS-INTRO] Introduction to robots.txt

Two common errors follow from treating robots.txt as an indexing control. A noindex line inside the file is simply dropped: “Specifying the noindex rule in the robots.txt file is not supported by Google.”[G-NOINDEX] And a noindex on a page the file blocks is never read at all:

“For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”

[G-NOINDEX] Block Search indexing with noindex

Common robots.txt mistakes

A robots.txt that errors or times out

A 404 is harmless; Google reads it as though no file existed. A 5xx status, a 429 or a timeout is another story. Google describes what happens next:

“For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can't fetch a new version, for the next 30 days Google will use the last good version, while still trying to fetch a new version.”

[G-ROBOTS-SPEC] Handling of errors and HTTP status codes

RFC 9309 goes further and tells every crawler that “If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow.”[RFC9309] Serving the file from static storage or the edge keeps it answering during an outage.

A named group that silently drops the shared rules

Crawlers do not merge a named group with the * group. Google uses “the group with the most specific user agent that matches”[G-ROBOTS-SPEC], and Bing's guidance is that with a bingbot group present, “BingBot will ignore all the other directives”[B-ROBOTS-2012]. Adding User-agent: Googlebot with a single rule therefore frees Googlebot from every rule written under User-agent: *. The checker lists the rules each crawler loses this way.

Blocking the files a page needs to render

Google's mobile guidance asks site owners to “Let Google crawl your resources.”[G-MOBILE] Its robots.txt introduction is explicit for the case where missing CSS, JavaScript or images make a page harder for its crawler to understand: “don't block them”[G-ROBOTS-INTRO].

Relative Sitemap lines

A line such as Sitemap: /sitemap.xml is common and cannot be used. Of the sitemap field, Google writes:

“It must be a fully qualified URL, including the protocol and host, and doesn't have to be URL-encoded.”

[G-ROBOTS-SPEC] sitemap

Directives Google never reads

Google supports four fields: user-agent, allow, disallow and sitemap. Its specification names the most familiar casualty, noting that “other fields such as crawl-delay aren't supported”[G-ROBOTS-SPEC]. Bing said in 2012 that bingbot honors crawl-delay, so the checker marks such lines as ignored by Google rather than calling them errors. Paths that do not start with / or * are ignored as well.

Size and encoding

“Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored.”[G-ROBOTS-SPEC] Rules past that point are not merely late; they do not exist for Google. The format itself is fixed too: “The robots.txt file must be a UTF-8 encoded plain text file and the lines must be separated by CR, CR/LF, or LF.”[G-ROBOTS-SPEC]

What this checker cannot tell you

It reads the file; it does not pretend to be Googlebot. Every request identifies the scanner honestly, so a firewall, CDN or bot-protection layer that answers search crawlers differently stays invisible to it. The full audit lists that as a manual check. The scanner also obeys robots.txt for its own user agent, and any URL the file closes to it is reported as not evaluated instead of being fetched anyway.

Search engines are not the only readers. For OpenAI, Anthropic, Perplexity, Apple, Common Crawl and Meta, the AI crawler checker applies the same file to each documented token and quotes the operator on what that token controls.

The rules this checker runs

Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.

  • SEO-ROBOTS-04robots.txt is reachableC1 Mechanism· w8
  • SEO-ROBOTS-02robots.txt is UTF-8 textC1 Mechanism· w8
  • SEO-ROBOTS-03robots.txt is smaller than 500 KiBC1 Mechanism· w8
  • SEO-ROBOTS-05ASearch crawlers may fetch every intended URLC1 Mechanism· w8
  • SEO-ROBOTS-05BResources needed for rendering are crawlableC2 Recommended· w4
  • SEO-ROBOTS-08Sitemap lines are fully qualified URLsC1 Mechanism· w8
  • SEO-IDX-03noindex is not placed in robots.txtC1 Mechanism· w8
  • SEO-IDX-02noindex is not hidden behind a robots.txt disallowC1 Mechanism· w8
  • SEO-ROBOTS-01robots.txt summaryC4 Informational· w0
  • SEO-ROBOTS-07Named crawler groups replace the * groupC4 Informational· w0
  • SEO-ROBOTS-10Lines Google ignoresC4 Informational· w0

Primary sources

The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.

This page checks one thing. The full audit runs all 70 scored rules.

It covers 16 categories on up to 200 pages, orders the fixes by the strength of the evidence behind them and prints its score with the arithmetic. Every finding cites the primary source it rests on. No account is needed.

Run the full 70-rule audit