What the checker reads
Enter a page and the checker requests robots.txt and the sitemaps it declares, then the page itself, the home page and the pages they link to, twelve in all. For each page it records:
- the final HTTP status, after any redirects;
- whether robots.txt lets Googlebot and bingbot fetch it, and the rule and line that decide;
- every
robotsandgooglebotmeta tag and everyX-Robots-Tagheader, as served, and the rules Google applies once they are combined; - the canonical URL when it names another page, and whether that target is blocked or carries noindex;
- whether your sitemaps list the page as one you want indexed.
Every URL in your sitemaps is also tested against robots.txt without being fetched, so a listed page that Googlebot may not crawl shows up even when the checker never loads it.
What a page needs before it can be indexed
Google's technical requirements are short. The page has to answer with a success code: “Google only indexes pages that are served with an HTTP 200 (success) status code.”[G-TECH] It has to be public: “If a page is made private, such as requiring a log-in to view it, Googlebot will not crawl it.”[G-TECH] Googlebot has to be allowed to crawl it. And it must not ask to be left out, which is what noindex does:
“Do not show this page, media, or resource in search results.”
Meeting all four makes a page eligible and nothing more. Google says so on its technical requirements page, and the checker repeats it:
“Just because a page meets these requirements doesn't mean that a page will be indexed; indexing isn't guaranteed.”
| Claim | Verdict | What the source says |
|---|---|---|
| Submitting or listing a URL guarantees indexing | Contradicted | Google does not guarantee it will crawl, index or serve a page, even one that follows Search Essentials [G-HOW-SEARCH-WORKS][G-URL-INSPECTION] |
Where a noindex can live
Google reads the rule “with either a <meta> tag or HTTP response header”[G-NOINDEX]. A robots meta tag addresses every crawler that reads it, and a googlebotmeta tag addresses Google alone. The header also reaches files that have no HTML head to hold a tag, and a header value scoped to another crawler, such as bingbot: noindex, does not count for Google.
When several rules apply they are added together rather than ranked: “In the case of conflicting robots rules, the more restrictive rule applies.”[G-ROBOTS-META] In Google's words, “the search engine will use the sum of the negative rules”[G-ROBOTS-META]. A related rule, unavailable_after, takes effect on a date: “Do not show this page in search results after the specified date/time.”[G-ROBOTS-META] The checker treats a date that has passed exactly like a noindex.
Noindex rules Google never reads
A noindex on a page robots.txt blocks
Blocking a page in robots.txt can look like the stronger way to keep it out. Google is clear that it works the other way round:
“For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”
The blocked page is never fetched, so its rules are never seen: “If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.”[G-ROBOTS-META] Meanwhile a blocked URL “can still be indexed if linked to from other sites”[G-ROBOTS-INTRO], and when that happens “the search result won't have a description”[G-ROBOTS-INTRO]. To keep a page out of search results, let crawlers fetch it and let the noindex do its job. The robots.txt checker shows which rule blocks a URL.
A noindex line in robots.txt
Some sites still add Noindex: lines to robots.txt. “Specifying the noindex rule in the robots.txt file is not supported by Google.”[G-NOINDEX] The checker reports every such line with its line number.
A noindex that JavaScript removes
Changing the robots meta tag in the browser is unreliable, for a reason Google spells out:
“When Google encounters the noindex tag, it may skip rendering and JavaScript execution, which means using JavaScript to change or remove the robots meta tag from noindex may not work as expected.”
This checker reads the HTML as the server sends it. The command-line scanner's rendered mode also compares it with the page after JavaScript has run.
Noindex and canonical tags together
A page that carries noindex while naming another URL as its canonical sends two different messages. Google advises against using noindex to pick a canonical, because “it will completely block the page from Search.”[G-CANON] The same applies to a canonical that points at a page carrying noindex. For robots.txt the advice is a plain instruction: “Don't use the robots.txt file for canonicalization purposes”[G-CANON]. The canonical tag checker covers the rest of canonicalization.
Which pages count as meant to be indexed
An issue needs intent. A noindex on a thank-you page is a choice; the same tag on a page your sitemap lists is a contradiction. The checker takes the home page and every URL in your sitemaps as the pages you want indexed, and judges status, crawl access and noindex for those. Without a sitemap it infers the set from links and treats a noindex or a disallow as deliberate, listing it as a note rather than an issue.
The full audit applies the same rule, and publishing a sitemap is what makes the check strict. The sitemap checker shows what yours lists.
What this checker cannot tell you
Whether Google has actually indexed a page. That answer lives in Search Console's URL Inspection tool and Page indexing report, which only the site's owner can open. The checker also cannot see a firewall or bot-protection layer that answers Googlebot differently from its own honest user agent, and it obeys robots.txt itself, so a page closed to every crawler is reported as not fetched instead of being loaded anyway.
The rules this checker runs
Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.
- SEO-HTTP-01Intended URLs return HTTP 200C1 Mechanism· w8
- SEO-HTTP-02Intended URLs are publicly accessibleC1 Mechanism· w8
- SEO-ROBOTS-05ASearch crawlers may fetch every intended URLC1 Mechanism· w8
- SEO-IDX-06Intended URLs carry no noindexC1 Mechanism· w8
- SEO-IDX-02noindex is not hidden behind a robots.txt disallowC1 Mechanism· w8
- SEO-IDX-03noindex is not placed in robots.txtC1 Mechanism· w8
- SEO-CANON-07noindex and robots.txt are not used to choose canonicalsC2 Recommended· w4
- SEO-IDX-01noindex inventoryC4 Informational· w0
- SEO-IDX-05How robots directives combineC4 Informational· w0
Primary sources
The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.
- Google Search technical requirementsG-TECH · last updated 2025-12-18 · verified 2026-10-03
- In-depth guide to how Google Search worksG-HOW-SEARCH-WORKS · last updated 2025-12-18 · verified 2026-09-28
- How HTTP status codes, and network and DNS errors affect Google SearchG-HTTP · last updated 2026-02-04 · verified 2026-10-03
- Introduction to robots.txtG-ROBOTS-INTRO · last updated 2025-12-10 · verified 2026-10-03
- Block Search indexing with noindexG-NOINDEX · last updated 2025-12-10 · verified 2026-10-03
- Robots meta tag, data-nosnippet, and X-Robots-Tag specificationsG-ROBOTS-META · last updated 2026-03-24 · verified 2026-10-03
- How to specify a canonical URL with rel="canonical" and other methodsG-CANON · last updated 2026-07-10 · verified 2026-10-03
- Understand the JavaScript SEO basicsG-JS · last updated 2026-03-04 · verified 2026-10-03
- Page indexing reportG-PIR · last updated not stated · verified 2026-09-28
- URL Inspection toolG-URL-INSPECTION · last updated not stated · verified 2026-09-28
- To crawl or not to crawl, that is BingBot's questionB-ROBOTS-2012 · last updated posted 2012-05-03 · verified 2026-10-03
- RFC 9309: Robots Exclusion ProtocolRFC9309 · last updated September 2022 · verified 2026-10-03