SEOscanner.app

Methodology · Evidence

Evidence classes and citations

The scanner does not have opinions about SEO. It has sources, and a fixed procedure for turning what a source says into a class, a weight and a sentence a reader can check.

Only primary sources

A source qualifies when it is owned by the engine, standards body or operator that owns the system being described: Google Search Central for Google, the Bing Webmaster Blog and Microsoft Support for Bing, the IETF, WHATWG, W3C and sitemaps.org for the standards, and each AI operator's own crawler documentation. Secondary commentary, however respected, never qualifies. The registry currently holds 98 sources, and the ruleset carries 202 citations to them, 186 with a stored quote.

From book badge to scanner class

The Ground-Truth Guide to Modern SEO tags every requirement with a badge that records how strongly the source speaks. The scanner maps those badges onto four classes and one “never” bucket. The mapping is mechanical by design; a person argues about the badge once, in the knowledge base, and the code inherits the decision.

Book badgeMeaning in the bookScanner class
REQUIREDAn engine or standard states the mechanism does not work otherwiseC1
ELIGIBILITY REQUIREMENTA documented condition for indexing or for a featureC1, often with a gate
DOCUMENTED RECOMMENDATIONThe source explicitly recommends itC2
DOCUMENTED SIGNALThe engine says the concept contributes to rankingC2
STANDARDA web standard requires or recommends itC3; C1 only when the standard that defines the mechanism itself states that non-conforming input is ignored, dropped or invalid
SUPPORTEDOfficially supported, not necessarily beneficialC4
OPTIONALMay be usedC4
EXPLICITLY DISMISSEDThe source says the claim does not holdNever scored
NOT A DOCUMENTED RANKING FACTORNo reviewed source establishes itNever scored
UNKNOWN / NOT PUBLICLY DOCUMENTEDThe public record is silentNever scored
SEARCH-ENGINE-SPECIFICModifier: only the named engine documents itSets the rule's engines
house conventionStricter than any sourceC4, -HC suffix

A documented engine behavior with an adverse effect is C1 even when the book gives it only the engine-specific modifier, because the source describes a mechanism that stops working. The same applies when a site relies on a mechanism the engine explicitly dismisses, such as noindex placed in robots.txt: the dismissed mechanism silently does nothing, so the site's intent fails. Google and Bing are the selectable engines. A rule supported only by one engine's documentation is not applicable when that engine is not selected. AI crawler operators appear only in informational rules, because blocking or allowing them is a policy choice, not a defect.

How a reason is written

Reasons are never generated by a language model. Each rule condition has a reason template with named placeholders, and the evaluator fills those placeholders from observed values only, using fixed formats. A good reason does three things in two or three sentences: states what was observed with the specific values, states what the source says about it, and states the documented consequence.

A lint keeps every template within its source. Consequence phrases such as “is ignored” or “is not eligible” may appear only in C1 templates unless they sit inside a quotation. C2 templates attribute the statement to its engine with a reporting verb: recommends, advises, asks, notes. C3 templates name the standard. C4 text is a neutral observation. Banned everywhere: “will improve rankings”, “boosts SEO”, “Google penalizes”, “best practice” without a named source, and any number that is not in a cited source or in the observed values.

Citations travel with the finding

Each finding cites at least one source by its registry ID, with the section name, the verbatim quote stored in the knowledge base, and the source's last-updated date as displayed when it was verified. Quotes are stored once and compiled into the ruleset; code never retypes them, and a test holds every emitted quote to the knowledge base character for character.

Outcome statements follow the same discipline. Fixing a noindex on an intended page makes it eligible; it does not make it indexed. Adding valid breadcrumb markup makes a page eligible for the feature, and the report says so in Google's words:

“Google does not guarantee that your structured data will show up in search results, even if your page is marked up correctly according to the Rich Results Test.”

[G-SD-POLICIES] General structured data guidelines

Keeping the record current

Search documentation changes often. The knowledge base is the authority for rule logic, classes, reason templates and quotes; the registry is the authority for source identity and dates. A scheduled job re-fetches every source page monthly and diffs the text. A change opens a review task listing the affected rules. Nothing updates automatically; a person re-verifies the wording and the class. A source that softens, from “must” to “recommend”, lowers the class. A source that disappears retires the rule or moves it to the myth guard. Ruleset versions follow semantic versioning: major when a class, weight or formula changes, minor when rules or conditions are added, patch for wording that does not change outcomes. Every report shows the ruleset version and research cutoff it was produced with.