What the checker reads
The checker requests /robots.txt from the host your address finally lands on, following up to five redirects, the minimum Google documents. It parses the file with the algorithm Google describes and RFC 9309 standardizes. Each crawler obeys the group that names it most specifically; inside that group the longest matching path wins, and an Allow beats a Disallow of equal length.
It then tests the home page and every URL in your sitemaps against the Googlebot and bingbot groups. Those URLs are not fetched, because the question is narrower: does the file permit fetching them? A sitemap URL that robots.txt blocks is a page you have asked search engines to find and then told them not to read.
The results show:
- whether robots.txt answers, and with which status;
- whether it is valid UTF-8 and within Google's 500 KiB parsing limit;
- every group, rule and
Sitemapline as parsed, with the lines Google ignores marked; - which intended URLs Googlebot and bingbot may not fetch, with the rule and line responsible;
- CSS, JavaScript and images your home page loads that the file blocks;
noindexlines, which have no effect in robots.txt.
What robots.txt controls, and what it does not
robots.txt governs crawling. RFC 9309 defines its two rule types in a sentence:
“These lines indicate whether accessing a URI that matches the corresponding path is allowed or disallowed.”
Indexing is a different matter. Google says robots.txt “is not a mechanism for keeping a web page out of Google”[G-ROBOTS-INTRO]. A disallowed URL “can still be indexed if linked to from other sites”[G-ROBOTS-INTRO], and in that case “the search result won't have a description”[G-ROBOTS-INTRO]. For keeping a page out of results, the same documentation points elsewhere:
“To keep a web page out of Google, block indexing with noindex or password-protect the page.”
Two common errors follow from treating robots.txt as an indexing control. A noindex line inside the file is simply dropped: “Specifying the noindex rule in the robots.txt file is not supported by Google.”[G-NOINDEX] And a noindex on a page the file blocks is never read at all:
“For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”
Common robots.txt mistakes
A robots.txt that errors or times out
A 404 is harmless; Google reads it as though no file existed. A 5xx status, a 429 or a timeout is another story. Google describes what happens next:
“For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can't fetch a new version, for the next 30 days Google will use the last good version, while still trying to fetch a new version.”
RFC 9309 goes further and tells every crawler that “If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow.”[RFC9309] Serving the file from static storage or the edge keeps it answering during an outage.
A named group that silently drops the shared rules
Crawlers do not merge a named group with the * group. Google uses “the group with the most specific user agent that matches”[G-ROBOTS-SPEC], and Bing's guidance is that with a bingbot group present, “BingBot will ignore all the other directives”[B-ROBOTS-2012]. Adding User-agent: Googlebot with a single rule therefore frees Googlebot from every rule written under User-agent: *. The checker lists the rules each crawler loses this way.
Blocking the files a page needs to render
Google's mobile guidance asks site owners to “Let Google crawl your resources.”[G-MOBILE] Its robots.txt introduction is explicit for the case where missing CSS, JavaScript or images make a page harder for its crawler to understand: “don't block them”[G-ROBOTS-INTRO].
Relative Sitemap lines
A line such as Sitemap: /sitemap.xml is common and cannot be used. Of the sitemap field, Google writes:
“It must be a fully qualified URL, including the protocol and host, and doesn't have to be URL-encoded.”
Directives Google never reads
Google supports four fields: user-agent, allow, disallow and sitemap. Its specification names the most familiar casualty, noting that “other fields such as crawl-delay aren't supported”[G-ROBOTS-SPEC]. Bing said in 2012 that bingbot honors crawl-delay, so the checker marks such lines as ignored by Google rather than calling them errors. Paths that do not start with / or * are ignored as well.
Size and encoding
“Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored.”[G-ROBOTS-SPEC] Rules past that point are not merely late; they do not exist for Google. The format itself is fixed too: “The robots.txt file must be a UTF-8 encoded plain text file and the lines must be separated by CR, CR/LF, or LF.”[G-ROBOTS-SPEC]
What this checker cannot tell you
It reads the file; it does not pretend to be Googlebot. Every request identifies the scanner honestly, so a firewall, CDN or bot-protection layer that answers search crawlers differently stays invisible to it. The full audit lists that as a manual check. The scanner also obeys robots.txt for its own user agent, and any URL the file closes to it is reported as not evaluated instead of being fetched anyway.
Search engines are not the only readers. For OpenAI, Anthropic, Perplexity, Apple, Common Crawl and Meta, the AI crawler checker applies the same file to each documented token and quotes the operator on what that token controls.
The rules this checker runs
Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.
- SEO-ROBOTS-04robots.txt is reachableC1 Mechanism· w8
- SEO-ROBOTS-02robots.txt is UTF-8 textC1 Mechanism· w8
- SEO-ROBOTS-03robots.txt is smaller than 500 KiBC1 Mechanism· w8
- SEO-ROBOTS-05ASearch crawlers may fetch every intended URLC1 Mechanism· w8
- SEO-ROBOTS-05BResources needed for rendering are crawlableC2 Recommended· w4
- SEO-ROBOTS-08Sitemap lines are fully qualified URLsC1 Mechanism· w8
- SEO-IDX-03noindex is not placed in robots.txtC1 Mechanism· w8
- SEO-IDX-02noindex is not hidden behind a robots.txt disallowC1 Mechanism· w8
- SEO-ROBOTS-01robots.txt summaryC4 Informational· w0
- SEO-ROBOTS-07Named crawler groups replace the * groupC4 Informational· w0
- SEO-ROBOTS-10Lines Google ignoresC4 Informational· w0
Primary sources
The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.
- Google Search technical requirementsG-TECH · last updated 2025-12-18 · verified 2026-10-03
- How Google interprets the robots.txt specificationG-ROBOTS-SPEC · last updated 2026-08-31 · verified 2026-10-03
- Introduction to robots.txtG-ROBOTS-INTRO · last updated 2025-12-10 · verified 2026-10-03
- Block Search indexing with noindexG-NOINDEX · last updated 2025-12-10 · verified 2026-10-03
- Robots meta tag, data-nosnippet, and X-Robots-Tag specificationsG-ROBOTS-META · last updated 2026-03-24 · verified 2026-10-03
- Mobile site and mobile-first indexing best practicesG-MOBILE · last updated 2025-12-10 · verified 2026-10-03
- To crawl or not to crawl, that is BingBot's questionB-ROBOTS-2012 · last updated posted 2012-05-03 · verified 2026-10-03
- RFC 9309: Robots Exclusion ProtocolRFC9309 · last updated September 2022 · verified 2026-10-03