SEOscanner.app

Free check · Indexing

Noindex and indexability checker

Check what decides whether a page can be indexed: its status code, whether robots.txt lets Googlebot and bingbot fetch it, every noindex in its robots meta tags and X-Robots-Tag header, and noindex rules Google can never read.

Reads robots.txt and the sitemaps it declares, then loads the URL you enter, the home page and the pages they link to, up to 12 pages, and up to 6 canonical targets they name. Sitemap URLs are tested against robots.txt without being fetched. Free, no account needed.

What the checker reads

Enter a page and the checker requests robots.txt and the sitemaps it declares, then the page itself, the home page and the pages they link to, twelve in all. For each page it records:

  • the final HTTP status, after any redirects;
  • whether robots.txt lets Googlebot and bingbot fetch it, and the rule and line that decide;
  • every robots and googlebot meta tag and every X-Robots-Tag header, as served, and the rules Google applies once they are combined;
  • the canonical URL when it names another page, and whether that target is blocked or carries noindex;
  • whether your sitemaps list the page as one you want indexed.

Every URL in your sitemaps is also tested against robots.txt without being fetched, so a listed page that Googlebot may not crawl shows up even when the checker never loads it.

What a page needs before it can be indexed

Google's technical requirements are short. The page has to answer with a success code: “Google only indexes pages that are served with an HTTP 200 (success) status code.”[G-TECH] It has to be public: “If a page is made private, such as requiring a log-in to view it, Googlebot will not crawl it.”[G-TECH] Googlebot has to be allowed to crawl it. And it must not ask to be left out, which is what noindex does:

“Do not show this page, media, or resource in search results.”

[G-ROBOTS-META] Valid indexing and serving rules: noindex

Meeting all four makes a page eligible and nothing more. Google says so on its technical requirements page, and the checker repeats it:

“Just because a page meets these requirements doesn't mean that a page will be indexed; indexing isn't guaranteed.”

[G-TECH] Google Search technical requirements
From the myth guard, which lists every claim the scanner refuses to score and why.
ClaimVerdictWhat the source says
Submitting or listing a URL guarantees indexingContradictedGoogle does not guarantee it will crawl, index or serve a page, even one that follows Search Essentials [G-HOW-SEARCH-WORKS][G-URL-INSPECTION]

Where a noindex can live

Google reads the rule “with either a <meta> tag or HTTP response header”[G-NOINDEX]. A robots meta tag addresses every crawler that reads it, and a googlebotmeta tag addresses Google alone. The header also reaches files that have no HTML head to hold a tag, and a header value scoped to another crawler, such as bingbot: noindex, does not count for Google.

When several rules apply they are added together rather than ranked: “In the case of conflicting robots rules, the more restrictive rule applies.”[G-ROBOTS-META] In Google's words, “the search engine will use the sum of the negative rules”[G-ROBOTS-META]. A related rule, unavailable_after, takes effect on a date: “Do not show this page in search results after the specified date/time.”[G-ROBOTS-META] The checker treats a date that has passed exactly like a noindex.

Noindex rules Google never reads

A noindex on a page robots.txt blocks

Blocking a page in robots.txt can look like the stronger way to keep it out. Google is clear that it works the other way round:

“For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”

[G-NOINDEX] Block Search indexing with noindex

The blocked page is never fetched, so its rules are never seen: “If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.”[G-ROBOTS-META] Meanwhile a blocked URL “can still be indexed if linked to from other sites”[G-ROBOTS-INTRO], and when that happens “the search result won't have a description”[G-ROBOTS-INTRO]. To keep a page out of search results, let crawlers fetch it and let the noindex do its job. The robots.txt checker shows which rule blocks a URL.

A noindex line in robots.txt

Some sites still add Noindex: lines to robots.txt. “Specifying the noindex rule in the robots.txt file is not supported by Google.”[G-NOINDEX] The checker reports every such line with its line number.

A noindex that JavaScript removes

Changing the robots meta tag in the browser is unreliable, for a reason Google spells out:

“When Google encounters the noindex tag, it may skip rendering and JavaScript execution, which means using JavaScript to change or remove the robots meta tag from noindex may not work as expected.”

[G-JS] Use meta robots tags carefully

This checker reads the HTML as the server sends it. The command-line scanner's rendered mode also compares it with the page after JavaScript has run.

Noindex and canonical tags together

A page that carries noindex while naming another URL as its canonical sends two different messages. Google advises against using noindex to pick a canonical, because “it will completely block the page from Search.”[G-CANON] The same applies to a canonical that points at a page carrying noindex. For robots.txt the advice is a plain instruction: “Don't use the robots.txt file for canonicalization purposes”[G-CANON]. The canonical tag checker covers the rest of canonicalization.

Which pages count as meant to be indexed

An issue needs intent. A noindex on a thank-you page is a choice; the same tag on a page your sitemap lists is a contradiction. The checker takes the home page and every URL in your sitemaps as the pages you want indexed, and judges status, crawl access and noindex for those. Without a sitemap it infers the set from links and treats a noindex or a disallow as deliberate, listing it as a note rather than an issue.

The full audit applies the same rule, and publishing a sitemap is what makes the check strict. The sitemap checker shows what yours lists.

What this checker cannot tell you

Whether Google has actually indexed a page. That answer lives in Search Console's URL Inspection tool and Page indexing report, which only the site's owner can open. The checker also cannot see a firewall or bot-protection layer that answers Googlebot differently from its own honest user agent, and it obeys robots.txt itself, so a page closed to every crawler is reported as not fetched instead of being loaded anyway.

The rules this checker runs

Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.

  • SEO-HTTP-01Intended URLs return HTTP 200C1 Mechanism· w8
  • SEO-HTTP-02Intended URLs are publicly accessibleC1 Mechanism· w8
  • SEO-ROBOTS-05ASearch crawlers may fetch every intended URLC1 Mechanism· w8
  • SEO-IDX-06Intended URLs carry no noindexC1 Mechanism· w8
  • SEO-IDX-02noindex is not hidden behind a robots.txt disallowC1 Mechanism· w8
  • SEO-IDX-03noindex is not placed in robots.txtC1 Mechanism· w8
  • SEO-CANON-07noindex and robots.txt are not used to choose canonicalsC2 Recommended· w4
  • SEO-IDX-01noindex inventoryC4 Informational· w0
  • SEO-IDX-05How robots directives combineC4 Informational· w0

Primary sources

The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.

This page checks one thing. The full audit runs all 70 scored rules.

It covers 16 categories on up to 200 pages, orders the fixes by the strength of the evidence behind them and prints its score with the arithmetic. Every finding cites the primary source it rests on. No account is needed.

Run the full 70-rule audit