Why it matters
An audit that can disagree with itself cannot be trusted to gate a deploy. If a score moves between two runs, the reader needs to know whether the site changed, the rules changed, or the tool simply wobbled. The scanner removes the third possibility by construction, and records enough about the first two to tell them apart: every report carries the identifier of the snapshot it read and the version of the ruleset it applied.
One network pass, then a frozen snapshot
start URL + config ──> ACQUIRE ──> SNAPSHOT ──> EVALUATE ──> REPORT
(network) (immutable, (pure (canonical JSON
hashed) function) + rendered views)Acquisition is the only phase allowed to touch the network. Its job is to collect everything evaluation will need and freeze it. It runs in fixed phases, and within a phase URLs are processed in ascending byte order of their normalized form, so the set of fetched URLs does not depend on network timing or concurrency. robots.txt for an origin is always fetched before any other request to that origin. The crawl proceeds breadth-first from the home page, finishing each depth level completely before sorting and deduplicating the next. Requests carry fixed headers, no cookies and no referer, so the server returns its default variant on every run. Status codes are never retried; the snapshot reflects what the server actually answered.
The scanner identifies itself honestly. It never impersonates Googlebot or any other crawler; whether a crawler may fetch a URL is decided by parsing robots.txt with the rules the engines document, not by spoofing. In the default mode it also obeys robots.txt for its own token, so a disallowed URL is recorded as blocked by policy rather than fetched.
The snapshot
The snapshot is the boundary between the unpredictable world and the deterministic evaluator. A manifest lists every exchange: a sequence number assigned after acquisition, the request, the outcome, the status, the response headers, the sizes, a truncation flag and the SHA-256 of the body. Bodies are stored content-addressed by that hash. TLS certificate details and DNS answers are stored per origin. Timing data may be kept for diagnostics but is never read by evaluation. One clock reading, the completion time, is stored and used wherever a rule needs “now”, such as a lastmod in the future or an expired certificate. The snapshot identifier is the SHA-256 of the manifest's canonical serialization.
Evaluation as a pure function
Inside evaluation the following are forbidden: network access, reading files outside the snapshot and the compiled ruleset, the system clock, random numbers, locale-dependent formatting or collation, floating-point values in anything that reaches the output, model or language-model calls of any kind, and iteration over unordered collections without an explicit sort. Scores are computed as exact rationals and reduced before they are printed; the scoring page shows the arithmetic. The report is emitted as canonical JSON (RFC 8785), so two equal reports are equal as bytes, and the Markdown and web views are rendered from that JSON rather than computed separately.
The advisory AI review is the one place a model is called. It runs after evaluation, reads the same snapshot, and never feeds back into the report; the review page explains how it is kept apart.
How the contract is proven
These tests run in continuous integration before any ruleset ships.
| Test | Method |
|---|---|
| Golden reports | Evaluate each fixture snapshot and compare with the committed expected report byte for byte. |
| Idempotence | Evaluate the same snapshot twice in one process and once in a fresh process; all outputs identical. |
| Order independence | Shuffle the storage order of manifest entries (sequence numbers retained); output identical. |
| Environment independence | Run under different TZ, LANG and LC_ALL values; output identical. |
| Platform independence | CI matrix of at least two operating systems; output identical. |
| Rule coverage | Every scored rule has at least one passing fixture and one fixture per fail condition. CI fails if a rule or condition lacks fixtures. |
| Citation integrity | Every source ID cited in the knowledge base exists in the registry; every quote emitted by the code matches the knowledge base exactly. |
What the contract does not cover
Acquisition cannot be fully deterministic, because the site and the network change. The contract therefore covers evaluation. Acquisition is made as stable as practical, and every scan's snapshot identifier lets two reports be compared knowing exactly whether their inputs differed. When a limit truncates a crawl, the snapshot records it, and rules that need a complete crawl, orphan detection for instance, report as not evaluated rather than produce a misleading result.
The web scanner runs without rendering, Chrome UX Report data or the parity check, so the rules that need those inputs report as not evaluated. The command-line scanner supports them, stores the full snapshot, and can re-evaluate it later for a byte-identical report or diff two reports keyed by rule, condition and URL.