SEOscanner.app

Free check · Sitemaps

XML sitemap checker and validator

Find a site's sitemaps through robots.txt or /sitemap.xml, confirm each returns 200 and parses as a sitemap, count the URLs it lists and test a sample of them.

Reads robots.txt, every declared sitemap and the index files it names (or /sitemap.xml), then loads the home page and up to 25 listed URLs. Free, no account needed.

What the checker reads

Sitemaps are found exactly as the full audit finds them. Sitemap lines in robots.txt come first, and only when there are none does the checker try /sitemap.xml. A sitemap index is followed one level down, so each child file appears in the results beside the index that named it. Every file is decompressed when gzipped, parsed with DTDs and external entities disabled, and measured against the protocol's limits.

With the files read, the checker loads the home page and up to 25 listed URLs in byte order. That sample shows whether the addresses you submit answer 200, redirect, fail, or declare some other page canonical. The full audit goes on to load up to 200.

What a sitemap does

A sitemap lists the URLs you want found, and Google reads it as a canonical signal too: “All pages listed in a sitemap are suggested as canonicals”[G-CANON]. That is why a listed URL that redirects, or that names another page as canonical, counts against the sitemap rather than being ignored: the two signals disagree.

For Google it is optional. The starter guide says of a sitemap that “this isn't required”[G-STARTER]. The checker reports a site with no sitemap as a practice to add, never as a failure, and the full audit scores that for Bing only, because of the lastmod signal described below. Several popular beliefs about sitemaps fail against the primary record:

From the myth guard, which lists every claim the scanner refuses to score and why.
ClaimVerdictWhat the source says
A sitemap improves rankingsNot supportedSitemaps help discovery and do not guarantee crawling or indexing [G-SITEMAP-OVERVIEW]
Submitting or listing a URL guarantees indexingContradictedGoogle does not guarantee it will crawl, index or serve a page, even one that follows Search Essentials [G-HOW-SEARCH-WORKS][G-URL-INSPECTION]
Sitemap priority and changefreq matterContradictedGoogle ignores both; Bing says they "do not influence how your content is crawled or ranked" [G-SITEMAP-BUILD][B-SITEMAP-2025]

lastmod is the one optional field both engines describe using. Google relies on it only “if it's consistently and verifiably (for example by comparing to the last modification of the page) accurate.”[G-SITEMAP-BUILD] Bing describes a broader role for it:

“The lastmod field in your sitemap remains a key signal, helping Bing prioritize URLs for recrawling and reindexing, or skip them entirely if the content hasn't changed since the last crawl.”

[B-SITEMAP-2025] Why lastmod Still Matters for AI Powered Indexing

The protocol settles which date belongs there: “the date must be set to the date the linked page was last modified, not when the sitemap is generated.”[SITEMAPS-ORG]

Common sitemap problems

A file that does not load or does not parse

A sitemap that answers 404, serves an HTML error page or breaks off mid-element yields no URLs at all. String concatenation is the usual culprit; one unescaped ampersand is enough. The protocol is strict on this point: “All data values in a Sitemap must be entity-escaped. The file itself must be UTF-8 encoded.”[SITEMAPS-ORG]

URLs outside the sitemap's scope

Scope covers the host and the directory. “all URLs listed in the Sitemap must use the same protocol (http, in this example) and reside on the same host as the Sitemap.”[SITEMAPS-ORG] Google adds that “a sitemap can only contain descendant URLs of the directory where the sitemap is hosted from.”[G-HREFLANG] A sitemap at /blog/sitemap.xml may therefore list only URLs under /blog/. Out-of-scope entries are not half-honored, either: “URLs that are not considered valid are dropped from further consideration.”[SITEMAPS-ORG] Declaring the sitemap in the robots.txt of the host whose URLs it lists lifts the restriction.

Relative URLs

A loc of /pricing is not a URL a crawler can use on its own. “Use fully-qualified, absolute URLs in your sitemaps.”[G-SITEMAP-BUILD]

Files past the size limits

“All formats limit a single sitemap to 50MB (uncompressed) or 50,000 URLs.”[G-SITEMAP-BUILD] Large sites split their URLs across several files and list those in a sitemap index, which has the same limits.

Redirects and non-canonical URLs in the list

Sitemaps generated from a database can list old slugs, tracking variants or the http address of an https page. Each such entry sends a signal that the page itself contradicts. List only final, self-canonical URLs, generated from the same function the templates use for their canonical tags.

lastmod written at build time

A lastmod that changes on every deploy, or sits in the future, says nothing about the page. The checker flags future dates, values outside the W3C Datetime format, and, for Bing, values without a time. Bing asks for ISO 8601 dates “use standard ISO 8601 date formatting for lastmod values, including both the date and time”[B-SITEMAP-2025].

What this checker cannot see

Sitemaps submitted directly in Google Search Console or Bing Webmaster Tools may be accepted for a verified property beyond the protocol's scope rules. The checker cannot see submissions, so it applies the protocol as written. Whether your sitemaps are submitted at all is a manual check in the full audit.

Only a sample of listed URLs is loaded, so a clean sample is encouraging rather than conclusive. robots.txt is checked here only for its Sitemap lines; the robots.txt checker tests every listed URL against its rules.

The rules this checker runs

Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.

Primary sources

The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.

This page checks one thing. The full audit runs all 70 scored rules.

It covers 16 categories on up to 200 pages, orders the fixes by the strength of the evidence behind them and prints its score with the arithmetic. Every finding cites the primary source it rests on. No account is needed.

Run the full 70-rule audit