- Class
- C1 Mechanism
- Weight
- 8
- Severity on fail
- critical
- Applies to
Test this rule on a live site with the robots.txt checker or the noindex checker. Free, no account needed.
Scope and evidence
Evaluated on fetched URLs with an effective Google noindex. The book marks this requirement REQUIRED, SEARCH-ENGINE-SPECIFIC.
Fail conditions
Each condition has a stable code and a reason template. The evaluator fills the placeholders from observed values only; no sentence in a report is generated.
noindex_behind_disallow{url} carries noindex ({directive_source}), but robots.txt disallows it for Googlebot ({rule}, line {line}). Google cannot read rules on pages it may not crawl, so this noindex is ignored and the URL can still be indexed from links.
The documented fix
Remove the disallow for URLs that rely on noindex.
Sources, quoted
These are the passages the rule rests on, stored once in the knowledge base and emitted exactly as written. A citation marked for specific conditions appears only on findings with those codes.
“For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”
[G-NOINDEX] Block Search indexing with noindex · Block Search indexing with noindex · last updated 2025-12-10
https://developers.google.com/search/docs/crawling-indexing/block-indexing“If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.”
[G-ROBOTS-META] Robots meta tag, data-nosnippet, and X-Robots-Tag specifications · Combining robots.txt rules with indexing and serving rules · last updated 2026-03-24
https://developers.google.com/search/docs/crawling-indexing/robots-meta-tag