The reference below lists each crawler token with its operator's own description of what the token does and how the operator says it treats robots.txt. The entries differ more than their names suggest: some are search crawlers, some control training data, some are control tokens with no crawler of their own, and the user-triggered fetchers carry their operators' statements that robots.txt rules may not apply to them. The scanner repeats those statements and adds nothing to them. The technical rules in this scanner concern search engines; the one position it quotes on generative search is Google's.
“From Google Search's perspective, optimizing for generative AI search is optimizing for the search experience, and thus still SEO.”
To see how a particular robots.txt treats each of these tokens, run the AI crawler checker. The llms.txt checker reads the other file often discussed alongside them, and reports it for information only.
The AI answers readiness list
What the engines and AI operators document about getting content into AI answers, checked against this scan. Nothing here is scored: Google says optimizing for generative AI search is still SEO, and allowing or blocking an AI crawler is a policy choice. Each line says whether the documented condition is in place, so an owner can see what they could still do.
Each line is derived from findings the rules already made, so it adds no judgment of its own: whether pages are eligible to be shown with a snippet (Google's stated condition for AI Overviews and AI Mode), whether pages carry noarchive (which Bing says keeps content out of its answers), whether a sitemap with timed lastmod values exists (Bing's recrawl signal), whether each search crawler may fetch the home page, and whether /llms.txt exists (optional either way, because Google Search ignores it). Training crawlers are left out: training access is not search inclusion.
Pages eligible to be shown with a snippet, Google's condition for AI Overviews and AI Mode
Google · G-AI
“must be indexed and eligible to be shown in Google Search with a snippet”
Pages without noarchive or nocache, which keep content available to Bing's answers
Microsoft · B-NOARCHIVE-2023
“will not be included in Bing Chat answers”
A sitemap whose lastmod values include date and time, Bing's recrawl signal
Microsoft · B-SITEMAP-2025
“The lastmod field in your sitemap remains a key signal, helping Bing prioritize URLs for recrawling and reindexing, or skip them entirely if the content hasn't changed since the last crawl.”
Googlebot may fetch the home page
Google · G-CRAWLERS
bingbot may fetch the home page
Microsoft · B-ROBOTS-2012
OAI-SearchBot may fetch the home page
OpenAI · OPENAI-BOTS
“used to surface websites in search results in ChatGPT's search features”
Claude-SearchBot may fetch the home page
Anthropic · ANTHROPIC-BOTS
“may reduce your site's visibility”
PerplexityBot may fetch the home page
Perplexity · PERPLEXITY-BOTS
“designed to surface and link websites in search results on Perplexity”
Applebot may fetch the home page
Apple · APPLEBOT
meta-webindexer may fetch the home page
Meta · META-CRAWLERS
/llms.txt for agents that read it; Google Search ignores the file
Google (ignores it); the llms.txt proposal · G-AI-OPT
“Doing so will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them.”
The access matrix
Every report carries a matrix: for each crawler token below, the robots.txt outcome for the home page and a sample of intended URLs, computed with the same group-selection and longest-match rules the engines document. The purpose and robots.txt columns are the operators' own statements, cited to the registry. Nothing in the matrix is scored.
| Crawler | Operator | Documented purpose | Token chain | Documented robots.txt behavior | Source |
|---|---|---|---|---|---|
| Googlebot | Google Search crawler (smartphone and desktop) | googlebot → * | Common crawlers "always obey robots.txt rules when crawling automatically" | G-CRAWLERS | |
| Googlebot-Image | Image crawling for Google Search | googlebot-image → googlebot → * | As above | G-CRAWLERS | |
| Google-Extended | Control token, no separate crawler: Gemini training and grounding; no effect on Search inclusion or ranking | google-extended → * | Is a robots.txt token | G-CRAWLERS | |
| bingbot | Microsoft | Bing search crawler | bingbot → msnbot → * | Honors one group only: the bingbot group if present, else msnbot, else * (2012) | B-ROBOTS-2012 |
| OAI-SearchBot | OpenAI | "used to surface websites in search results in ChatGPT's search features" | oai-searchbot → * | Controlled by its robots.txt token | OPENAI-BOTS |
| GPTBot | OpenAI | Training; a disallow indicates content "should not be used in training" | gptbot → * | Controlled by its token | OPENAI-BOTS |
| ChatGPT-Useruser-triggered | OpenAI | User-triggered fetches | chatgpt-user → * | "robots.txt rules may not apply" | OPENAI-BOTS |
| Claude-SearchBot | Anthropic | Search; blocking it "may reduce your site's visibility" | claude-searchbot → * | Anthropic states its bots honor robots.txt | ANTHROPIC-BOTS |
| ClaudeBot | Anthropic | Training; a disallow signals exclusion from "AI model training datasets" | claudebot → * | Honors robots.txt; supports crawl-delay | ANTHROPIC-BOTS |
| Claude-User | Anthropic | User-triggered fetches; blocking it "may reduce your site's visibility" | claude-user → * | Honors robots.txt (no exception stated) | ANTHROPIC-BOTS |
| PerplexityBot | Perplexity | "designed to surface and link websites in search results on Perplexity"; "not used to crawl content for AI foundation models" | perplexitybot → * | Perplexity recommends allowing it; no explicit compliance sentence | PERPLEXITY-BOTS |
| Perplexity-Useruser-triggered | Perplexity | User-triggered fetches | perplexity-user → * | "generally ignores robots.txt rules" | PERPLEXITY-BOTS |
| Applebot | Apple | Apple search crawler (data may also train Apple models) | applebot → googlebot → * | Follows Googlebot rules when Applebot is not named | APPLEBOT |
| Applebot-Extended | Apple | Control token for training; "does not crawl webpages" | applebot-extended → * | "Webpages that disallow Applebot-Extended can still be included in search results." | APPLEBOT |
| CCBot | Common Crawl | Open dataset used by third parties | ccbot → * | Obeys robots.txt and crawl-delay | CC-FAQ |
| meta-externalagent | Meta | Training or product improvement | meta-externalagent → * | robots.txt used to express preferences | META-CRAWLERS |
| meta-webindexer | Meta | Meta AI search index | meta-webindexer → * | robots.txt used to express preferences | META-CRAWLERS |
- Training access is not search inclusion, and crawler access is not citation. No operator promises that allowed content will be used or linked.
- ChatGPT-User and Perplexity-User rows quote their operators' statements that robots.txt may not apply; the matrix never implies that a disallow blocks them. Google's user-triggered fetchers ("these fetchers generally ignore robots.txt rules", G-USER-FETCHERS) and Meta's meta-externalfetcher ("may bypass robots.txt", META-CRAWLERS) are documented the same way and are not tracked in the matrix.
- Bing's AdIdxBot, BingPreview and MicrosoftPreview are omitted: no first-party source retrieved for the book documented their strings or robots.txt behavior.
Operators covered: Google, Microsoft, OpenAI, Anthropic, Perplexity, Apple, Common Crawl, Meta. Sources: G-CRAWLERS (Google's common crawlers), B-ROBOTS-2012 (To crawl or not to crawl, that is BingBot's question), OPENAI-BOTS (Overview of OpenAI Crawlers), ANTHROPIC-BOTS (Does Anthropic crawl data from the web, and how can site owners block the crawler?), PERPLEXITY-BOTS (Perplexity Crawlers), APPLEBOT (About Applebot), CC-FAQ (Common Crawl FAQ), META-CRAWLERS (Meta Web Crawlers).
Rules in this category
The category holds the rules that bear on generative features in Google Search and the informational checks around AI crawler access. Each links to its full page.