SEOscanner.app

Free check · Canonicalization

Canonical tag checker

Read the rel=canonical tags on a page and the pages it links to, then check placement, absolute URLs, conflicting canonicals and whether each target loads without a redirect.

Loads the URL you enter and the home page, then pages they link to, up to 12 pages in all. It also requests the canonical targets they name and the http/https and www variants of the host. Free, no account needed.

What the checker reads

Enter a page and the checker loads it, the home page and the pages those two link to, twelve pages in all. On each it reads every rel="canonical" link element, records whether the HTML parser kept it in the <head> or pushed it into the body, and collects any Link: rel="canonical" response header. Each canonical that names a different URL is then requested, so the results show what the target answers first, where it finally lands and which canonical it declares for itself.

The http, https, www and bare versions of your host are requested as well. A variant that serves the same pages at a second address is a canonicalization problem in its own right, and a permanent redirect to the preferred address is the remedy both Google and Bing describe.

How canonicalization works

A canonical link names the URL you want treated as the main version of a page. Google describes its canonicalization methods as signals of different strength. A redirect is “A strong signal that the target of the redirect should become canonical.”[G-CANON], while “All pages listed in a sitemap are suggested as canonicals”[G-CANON]. None is mandatory; Google says of the methods that “none of them are required.”[G-CANON] The checker therefore never counts a missing canonical as an issue.

What matters is that the signals you do send agree:

“Don't specify different URLs as canonical for the same page using different canonicalization techniques”

[G-CANON] General canonicalization best practices

RFC 6596, which defines the link relation, adds a condition only a person can judge: the canonical target has to carry the same content as the page, or a superset of it. The checker cannot see that, and the full audit lists it on its manual review checklist. Ordinary duplication does not break Google's spam policies either:

From the myth guard, which lists every claim the scanner refuses to score and why.
ClaimVerdictWhat the source says
Duplicate content triggers a penaltyContradictedOrdinary duplication is "not a violation of Google's spam policies" [G-CANONICALIZATION]

Common canonical mistakes

A canonical the parser moves into the body

Google accepts the element in one place: “The rel="canonical" link element is only accepted if it appears in the <head> section of the HTML”[G-CANON]. A tag-manager <iframe>, a tracking pixel or stray text early in the head makes the parser close the <head> before it reaches the canonical, which then lands in the body where Google will not accept it. The results name the element that closed the head. Google's advice is to “make sure at least the <head> section is valid HTML”[G-CANON].

Relative and protocol-relative URLs

“Use absolute paths rather than relative paths with the rel="canonical" link element.”[G-CANON] RFC 6596 is more permissive and lists among the things a target may do: “Specify a relative IRI (see [RFC3986], Section 4.2).”[RFC6596] The checker follows Google's stricter guidance, and treats //www.example.com/page as relative too.

Two canonicals that disagree

A theme and an SEO plugin that each print one produce exactly this, as does an HTML tag that has drifted away from a Link header. The RFC leaves no room for it: “Specify only one canonical link relation for a resource.”[RFC6596]

A target that redirects, fails or points elsewhere

RFC 6596 names the targets to stay away from:

“Avoid designating the target (canonical) as:”

[RFC6596] 3. The Canonical Link Relation
  • “The source IRI of a permanent redirect (for HTTP, this refers to 300 and 301 response codes, defined in Sections 10.3.1 and 10.3.2 of [RFC2616]).”[RFC6596]
  • “An IRI that also specifies a canonical link relation to an IRI other than itself.”[RFC6596]
  • “An IRI that returns an error code, such as a 4xx response in HTTP (Section 10.4 of [RFC2616]).”[RFC6596]

The checker follows each target and reports which of the three applies. A 308 is treated like a 301. Temporary redirects (302, 303, 307) are allowed by the RFC and never reported.

Canonicals combined with noindex or a robots.txt block

“Don't use the robots.txt file for canonicalization purposes”[G-CANON], Google says, and it advises against noindex for the same job because “it will completely block the page from Search.”[G-CANON] A canonical that points at a URL robots.txt blocks, or at a page carrying noindex, is the combination both statements warn against.

http canonicals on an https site

“Google prefers HTTPS pages over equivalent HTTP pages as canonical”[G-CANON], and certificate errors or HTTPS-to-HTTP redirects can “cause Google to prefer HTTP very strongly.”[G-CANON] The results mark any canonical whose scheme differs from its page. When that http address then redirects permanently to https, the target is the source of a permanent redirect, and the RFC's avoid list applies.

Internal links to the duplicate

Internal links should agree with the canonical as well. “When linking within your site, link to the canonical URL rather than a duplicate URL.”[G-CANON] The checker reports pages whose internal links lead to URLs that redirect or declare a different canonical.

What this checker cannot see

It reads the HTML your server sends. A canonical that JavaScript adds or rewrites after load needs a rendered crawl, which the command-line scanner offers and the web tools do not. Meta descriptions often travel with canonicals in the same template, and the meta description checker reads them across the same set of pages.

The rules this checker runs

Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.

  • SEO-CANON-01The canonical link is inside the headC1 Mechanism· w8
  • SEO-CANON-02Canonical URLs are absoluteC2 Recommended· w4
  • SEO-CANON-08One canonical URL per pageC2 Recommended· w4
  • SEO-CANON-04The canonical target is a usable targetC3 Standard· w2
  • SEO-CANON-05Internal links and hreflang point to canonical URLsC2 Recommended· w4
  • SEO-CANON-07noindex and robots.txt are not used to choose canonicalsC2 Recommended· w4
  • SEO-HEAD-05The head is not closed earlyC2 Recommended· w4
  • SEO-HTTP-07HTTPS without certificate errors or downgradesC2 Recommended· w4
  • SEO-URL-01Host and protocol variants redirect to the preferred originC2 Recommended· w4
  • SEO-CANON-03Canonical inventoryC4 Informational· w0
  • SEO-CANON-06Canonical headers on non-HTML filesC4 Informational· w0

Primary sources

The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.

This page checks one thing. The full audit runs all 70 scored rules.

It covers 16 categories on up to 200 pages, orders the fixes by the strength of the evidence behind them and prints its score with the arithmetic. Every finding cites the primary source it rests on. No account is needed.

Run the full 70-rule audit