SEOscanner.app

Free check · AI crawlers

AI crawler checker

See whether a site's robots.txt allows or blocks GPTBot, ClaudeBot, PerplexityBot, Google-Extended and other AI crawlers, with each operator's own account of what its token controls.

Reads /robots.txt, the sitemaps it declares (for a sample of URLs) and the home page, then applies each crawler's documented token chain. Free, no account needed.

What the checker reads

It parses your robots.txt and, for each crawler token below, applies the token chain the operator documents. Most tokens fall back to the * group when the file does not name them. Applebot is the notable exception: without a group of its own it follows Googlebot's rules, just as bingbot falls back to msnbot's. The home page and up to ten other intended URLs are tested, and the results show, token by token, whether they may be fetched, which group decided, and what the operator says the token is for.

Allowing or blocking these crawlers is a policy decision rather than a technical defect, so nothing on this page is scored, and the full audit leaves it out of the score too, for the reasons the AI crawler methodology sets out. Googlebot and bingbot are different: blocking them from pages you want indexed fails an indexing gate in the full audit.

Search, training and user-triggered crawlers

The tokens differ more than their names suggest. Some fetch pages for a search index. Some govern whether content may be used to train models. Two are control tokens with no crawler of their own, and a few fetch pages only when a person asks an assistant to. Each answers to its own token, so blocking one leaves the others untouched, with the documented exception of Applebot, which follows Googlebot's group when it has none of its own. The groups below follow each operator's documented purpose, quoted as the operator wrote it.

  • Googlebot (Google): Google Search crawler (smartphone and desktop). robots.txt: Common crawlers "always obey robots.txt rules when crawling automatically". [G-CRAWLERS]
  • Googlebot-Image (Google): Image crawling for Google Search. robots.txt: As above. [G-CRAWLERS]
  • bingbot (Microsoft): Bing search crawler. robots.txt: Honors one group only: the bingbot group if present, else msnbot, else * (2012). [B-ROBOTS-2012]
  • OAI-SearchBot (OpenAI): "used to surface websites in search results in ChatGPT's search features". robots.txt: Controlled by its robots.txt token. [OPENAI-BOTS]
  • Claude-SearchBot (Anthropic): Search; blocking it "may reduce your site's visibility". robots.txt: Anthropic states its bots honor robots.txt. [ANTHROPIC-BOTS]
  • PerplexityBot (Perplexity): "designed to surface and link websites in search results on Perplexity"; "not used to crawl content for AI foundation models". robots.txt: Perplexity recommends allowing it; no explicit compliance sentence. [PERPLEXITY-BOTS]
  • Applebot (Apple): Apple search crawler (data may also train Apple models). robots.txt: Follows Googlebot rules when Applebot is not named. [APPLEBOT]
  • meta-webindexer (Meta): Meta AI search index. robots.txt: robots.txt used to express preferences. [META-CRAWLERS]

Training crawlers

  • GPTBot (OpenAI): Training; a disallow indicates content "should not be used in training". robots.txt: Controlled by its token. [OPENAI-BOTS]
  • ClaudeBot (Anthropic): Training; a disallow signals exclusion from "AI model training datasets". robots.txt: Honors robots.txt; supports crawl-delay. [ANTHROPIC-BOTS]
  • meta-externalagent (Meta): Training or product improvement. robots.txt: robots.txt used to express preferences. [META-CRAWLERS]

Control tokens

  • Google-Extended (Google): Control token, no separate crawler: Gemini training and grounding; no effect on Search inclusion or ranking. robots.txt: Is a robots.txt token. [G-CRAWLERS]
  • Applebot-Extended (Apple): Control token for training; "does not crawl webpages". robots.txt: "Webpages that disallow Applebot-Extended can still be included in search results.". [APPLEBOT]

User-triggered fetchers

  • ChatGPT-User (OpenAI): User-triggered fetches. robots.txt: "robots.txt rules may not apply". [OPENAI-BOTS]
  • Claude-User (Anthropic): User-triggered fetches; blocking it "may reduce your site's visibility". robots.txt: Honors robots.txt (no exception stated). [ANTHROPIC-BOTS]
  • Perplexity-User (Perplexity): User-triggered fetches. robots.txt: "generally ignores robots.txt rules". [PERPLEXITY-BOTS]

Open dataset crawlers

  • CCBot (Common Crawl): Open dataset used by third parties. robots.txt: Obeys robots.txt and crawl-delay. [CC-FAQ]

Two beliefs about these tokens fail against the operators' own documentation:

From the myth guard, which lists every claim the scanner refuses to score and why.
ClaimVerdictWhat the source says
Blocking GPTBot removes a site from ChatGPT searchContradictedOAI-SearchBot governs ChatGPT search; "Each setting is independent of the others" [OPENAI-BOTS]
Blocking Google-Extended removes a site from Google SearchContradictedGoogle-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" [G-CRAWLERS]

What robots.txt cannot promise

Training access is not search inclusion, and crawler access is not citation. No operator promises that allowed content will be used or linked.

User-triggered fetchers are a separate case. OpenAI's note on ChatGPT-User reads "robots.txt rules may not apply", and Perplexity says Perplexity-User "generally ignores robots.txt rules". A disallow should not be read as blocking either one, and the results flag both rows.

This page reports what your robots.txt asks of each token. What an operator does with that answer, including content it gathered earlier or copies held in datasets built by others, is for its own documentation to say; the table quotes that documentation and adds nothing to it.

Google's AI features follow Search's rules

Google-Extended governs Gemini training and grounding, not Google Search. For AI Overviews and AI Mode, the requirements are those of Search itself: a page “must be indexed and eligible to be shown in Google Search with a snippet”[G-AI], and “There are no additional technical requirements.”[G-AI] Among the controls Google documents, the snippet rules reach these features: of nosnippet, Google writes that it “will also prevent the content from being used as a direct input for AI Overviews and AI Mode.”[G-ROBOTS-META] The checker reports snippet rules on your home page, and Bing's noarchive and nocache controls, as information. Search Console's generative AI control is invisible from outside, so the full audit lists it on its manual checklist.

Blocking or allowing a single crawler

Give the crawler a group of its own:

User-agent: GPTBot
Disallow: /

A group that names a crawler replaces the * group for it entirely. Google uses “the group with the most specific user agent that matches”[G-ROBOTS-SPEC], and the checker applies the same RFC 9309 group selection to every token it tests. Rules from the * group that should still apply, a private directory for instance, have to be repeated inside the new group. The robots.txt checker lists the rules a named group drops, and the llms.txt checker covers the other file often discussed alongside AI crawlers.

The rules this checker runs

Each is a rule from the full ruleset, run by the same evaluator, with the same reason templates and citations. A rule's page lists its fail conditions and the passages it rests on. Informational rules (C4) never produce an issue.

  • SEO-ROBOTS-04robots.txt is reachableC1 Mechanism· w8
  • SEO-ROBOTS-05ASearch crawlers may fetch every intended URLC1 Mechanism· w8
  • SEO-AI-02AI crawler access matrixC4 Informational· w0
  • SEO-AI-01Snippet eligibility for Google's AI featuresC4 Informational· w0
  • SEO-AI-04Bing answer-use controlsC4 Informational· w0

Primary sources

The documentation this page and its rules rest on, with the date each page showed and the date its wording was last checked. The source registry lists every source the scanner may cite.

This page checks one thing. The full audit runs all 70 scored rules.

It covers 16 categories on up to 200 pages, orders the fixes by the strength of the evidence behind them and prints its score with the arithmetic. Every finding cites the primary source it rests on. No account is needed.

Run the full 70-rule audit