SEOscanner.app

Methodology · AI crawlers

AI crawlers and generative search

Whether to admit a training crawler is a policy decision, not a technical defect. The scanner shows the decision a site has made, in each operator's own words, and keeps it out of the score.

The reference below lists each crawler token with its operator's own description of what the token does and how the operator says it treats robots.txt. The entries differ more than their names suggest: some are search crawlers, some control training data, some are control tokens with no crawler of their own, and the user-triggered fetchers carry their operators' statements that robots.txt rules may not apply to them. The scanner repeats those statements and adds nothing to them. The technical rules in this scanner concern search engines; the one position it quotes on generative search is Google's.

“From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”

[G-AI-OPT] Is SEO still relevant for generative AI search?

To see how a particular robots.txt treats each of these tokens, run the AI crawler checker. The llms.txt checker reads the other file often discussed alongside them, and reports it for information only.

The AI answers readiness list

What the engines and AI operators document about getting content into AI answers, checked against this scan. Nothing here is scored: Google says optimizing for generative AI search is still SEO, and allowing or blocking an AI crawler is a policy choice. Each line says whether the documented condition is in place, so an owner can see what they could still do.

Each line is derived from findings the rules already made, so it adds no judgment of its own: whether pages are eligible to be shown with a snippet (Google's stated condition for AI Overviews and AI Mode), whether pages carry noarchive (which Bing says keeps content out of its answers), whether a sitemap with timed lastmod values exists (Bing's recrawl signal), whether each search crawler may fetch the home page, and whether /llms.txt exists (optional either way, because Google Search ignores it). Training crawlers are left out: training access is not search inclusion.

  • Pages eligible to be shown with a snippet, Google's condition for AI Overviews and AI Mode

    Google · G-AI

    “must be indexed and eligible to be shown in Google Search with a snippet”

  • Pages without noarchive or nocache, which keep content available to Bing's answers

    Microsoft · B-NOARCHIVE-2023

    “will not be included in Bing Chat answers”

  • A sitemap whose lastmod values include date and time, Bing's recrawl signal

    Microsoft · B-SITEMAP-2025

    “The lastmod field in your sitemap remains a key signal, helping Bing prioritize URLs for recrawling and reindexing, or skip them entirely if the content hasn't changed since the last crawl.”

  • Googlebot may fetch the home page

    Google · G-CRAWLERS

  • bingbot may fetch the home page

    Microsoft · B-ROBOTS-2012

  • OAI-SearchBot may fetch the home page

    OpenAI · OPENAI-BOTS

    “used to surface websites in search results in ChatGPT's search features”

  • Claude-SearchBot may fetch the home page

    Anthropic · ANTHROPIC-BOTS

    “may reduce your site's visibility”

  • PerplexityBot may fetch the home page

    Perplexity · PERPLEXITY-BOTS

    “designed to surface and link websites in search results on Perplexity”

  • Applebot may fetch the home page

    Apple · APPLEBOT

  • meta-webindexer may fetch the home page

    Meta · META-CRAWLERS

  • /llms.txt for agents that read it; Google Search ignores the file

    Google (ignores it); the llms.txt proposal · G-AI-OPT

    “Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”

The access matrix

Every report carries a matrix: for each crawler token below, the robots.txt outcome for the home page and a sample of intended URLs, computed with the same group-selection and longest-match rules the engines document. The purpose and robots.txt columns are the operators' own statements, cited to the registry. Nothing in the matrix is scored.

CrawlerOperatorDocumented purposeToken chainDocumented robots.txt behaviorSource
GooglebotGoogleGoogle Search crawler (smartphone and desktop)googlebot → *Common crawlers "always obey robots.txt rules when crawling automatically"G-CRAWLERS
Googlebot-ImageGoogleImage crawling for Google Searchgooglebot-image → googlebot → *As aboveG-CRAWLERS
Google-ExtendedGoogleControl token, no separate crawler: Gemini training and grounding; no effect on Search inclusion or rankinggoogle-extended → *Is a robots.txt tokenG-CRAWLERS
bingbotMicrosoftBing search crawlerbingbot → msnbot → *Honors one group only: the bingbot group if present, else msnbot, else * (2012)B-ROBOTS-2012
OAI-SearchBotOpenAI"used to surface websites in search results in ChatGPT's search features"oai-searchbot → *Controlled by its robots.txt tokenOPENAI-BOTS
GPTBotOpenAITraining; a disallow indicates content "should not be used in training"gptbot → *Controlled by its tokenOPENAI-BOTS
ChatGPT-Useruser-triggeredOpenAIUser-triggered fetcheschatgpt-user → *"robots.txt rules may not apply"OPENAI-BOTS
Claude-SearchBotAnthropicSearch; blocking it "may reduce your site's visibility"claude-searchbot → *Anthropic states its bots honor robots.txtANTHROPIC-BOTS
ClaudeBotAnthropicTraining; a disallow signals exclusion from "AI model training datasets"claudebot → *Honors robots.txt; supports crawl-delayANTHROPIC-BOTS
Claude-UserAnthropicUser-triggered fetches; blocking it "may reduce your site's visibility"claude-user → *Honors robots.txt (no exception stated)ANTHROPIC-BOTS
PerplexityBotPerplexity"designed to surface and link websites in search results on Perplexity"; "not used to crawl content for AI foundation models"perplexitybot → *Perplexity recommends allowing it; no explicit compliance sentencePERPLEXITY-BOTS
Perplexity-Useruser-triggeredPerplexityUser-triggered fetchesperplexity-user → *"generally ignores robots.txt rules"PERPLEXITY-BOTS
ApplebotAppleApple search crawler (data may also train Apple models)applebot → googlebot → *Follows Googlebot rules when Applebot is not namedAPPLEBOT
Applebot-ExtendedAppleControl token for training; "does not crawl webpages"applebot-extended → *"Webpages that disallow Applebot-Extended can still be included in search results."APPLEBOT
CCBotCommon CrawlOpen dataset used by third partiesccbot → *Obeys robots.txt and crawl-delayCC-FAQ
meta-externalagentMetaTraining or product improvementmeta-externalagent → *robots.txt used to express preferencesMETA-CRAWLERS
meta-webindexerMetaMeta AI search indexmeta-webindexer → *robots.txt used to express preferencesMETA-CRAWLERS
  • Training access is not search inclusion, and crawler access is not citation. No operator promises that allowed content will be used or linked.
  • ChatGPT-User and Perplexity-User rows quote their operators' statements that robots.txt may not apply; the matrix never implies that a disallow blocks them. Google's user-triggered fetchers ("these fetchers generally ignore robots.txt rules", G-USER-FETCHERS) and Meta's meta-externalfetcher ("may bypass robots.txt", META-CRAWLERS) are documented the same way and are not tracked in the matrix.
  • Bing's AdIdxBot, BingPreview and MicrosoftPreview are omitted: no first-party source retrieved for the book documented their strings or robots.txt behavior.

Operators covered: Google, Microsoft, OpenAI, Anthropic, Perplexity, Apple, Common Crawl, Meta. Sources: G-CRAWLERS (Google's common crawlers), B-ROBOTS-2012 (To crawl or not to crawl, that is BingBot's question), OPENAI-BOTS (Overview of OpenAI Crawlers), ANTHROPIC-BOTS (Does Anthropic crawl data from the web, and how can site owners block the crawler?), PERPLEXITY-BOTS (Perplexity Crawlers), APPLEBOT (About Applebot), CC-FAQ (Common Crawl FAQ), META-CRAWLERS (Meta Web Crawlers).

Rules in this category

The category holds the rules that bear on generative features in Google Search and the informational checks around AI crawler access. Each links to its full page.