SEOscanner.app

Rule · robots.txt

Search crawlers may fetch every intended URL

SEO-ROBOTS-05AC1 Mechanism· w8gate: indexing

Class
C1 Mechanism
Weight
8
Severity on fail
critical
Applies to
Google and Bing

Test this rule on a live site with the robots.txt checker, the XML sitemap checker, the noindex checker or the AI crawler checker. Free, no account needed.

Scope and evidence

Evaluated on each member of I. The book marks this requirement ELIGIBILITY REQUIREMENT, SEARCH-ENGINE-SPECIFIC, and the scanner treats it as a gate: indexing rule: a failure on an intended page removes that page from the eligible set instead of entering the achieved share.

Fail conditions

Each condition has a stable code and a reason template. The evaluator fills the placeholders from observed values only; no sentence in a report is generated.

  • disallowed_googlebot

    {url} is disallowed for Googlebot by {rule} (robots.txt line {line}, group {group}). That Googlebot isn't blocked is one of Google's technical requirements for indexing. Google may still index the URL from links elsewhere, but without reading the page, and the result won't have a description.

  • disallowed_bingbot

    {url} is disallowed for bingbot by {rule} (robots.txt line {line}, group {group}). Under the Robots Exclusion Protocol, bingbot should not crawl it.

The documented fix

Remove or narrow the rule. If the URL should not be in search results, take it out of sitemaps and use noindex on a crawlable URL instead, because robots.txt does not keep pages out of Google.

Sources, quoted

These are the passages the rule rests on, stored once in the knowledge base and emitted exactly as written. A citation marked for specific conditions appears only on findings with those codes.