SEOscanner.app

Rule · robots.txt

robots.txt is reachable

SEO-ROBOTS-04C1 Mechanism· w8

Class
C1 Mechanism
Weight
8
Severity on fail
critical
Applies to
Google

Test this rule on a live site with the robots.txt checker or the AI crawler checker. Free, no account needed.

Scope and evidence

Evaluated on each in-scope origin. The book marks this requirement SEARCH-ENGINE-SPECIFIC, STANDARD.

A single scan observes one moment, but the documented consequence applies whenever the condition occurs, so it is scored. A 404 is not a failure: Google treats it "as if a valid robots.txt file didn't exist."

Fail conditions

Each condition has a stable code and a reason template. The evaluator fills the placeholders from observed values only; no sentence in a report is generated.

  • server_error

    robots.txt on {origin} returned {outcome}. For the first 12 hours, Google stops crawling the site but keeps trying to fetch robots.txt; after that it uses the last good version for up to 30 days. RFC 9309 tells crawlers to assume complete disallow when robots.txt is unreachable.

  • unreachable

    robots.txt on {origin} returned {outcome}. For the first 12 hours, Google stops crawling the site but keeps trying to fetch robots.txt; after that it uses the last good version for up to 30 days. RFC 9309 tells crawlers to assume complete disallow when robots.txt is unreachable.

The documented fix

Serve robots.txt from static storage or the edge so it answers 200 (or a deliberate 404) even during maintenance or overload.

Sources, quoted

These are the passages the rule rests on, stored once in the knowledge base and emitted exactly as written. A citation marked for specific conditions appears only on findings with those codes.

  • “For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can't fetch a new version, for the next 30 days Google will use the last good version, while still trying to fetch a new version.”

    [G-ROBOTS-SPEC] How Google interprets the robots.txt specification · Handling of errors and HTTP status codes · last updated 2026-08-31

    https://developers.google.com/crawling/docs/robots-txt/robots-txt-spec
  • “If the robots.txt file is unreachable due to server or network errors, this means the robots.txt file is undefined and the crawler MUST assume complete disallow.”

    [RFC9309] RFC 9309: Robots Exclusion Protocol · 2.3.1.4 Unreachable Status · last updated September 2022

    https://www.rfc-editor.org/rfc/rfc9309.html