AI crawler reference: every bot that matters, and what each one actually does
There are four kinds of AI user agent and only two of them fetch pages. Confusing them is the most common reason a site is invisible to AI assistants while looking perfectly healthy in Google Search Console.
The four kinds of AI user agent
Treating all AI bots as one category is what produces the wrong robots.txt. They do different jobs, and only one category determines whether an AI assistant can cite you in a live answer.
- Retrieval crawlers. These build the index an assistant searches when answering a question.
OAI-SearchBot,PerplexityBot,Claude-SearchBotandGooglebotare in this group. These are the ones that decide whether you can be cited. - Training crawlers. These collect content that may be used to train future models.
GPTBot,ClaudeBotandCCBotare here. Blocking them affects future model weights, not today's answers. - User-triggered fetchers. These fetch a page because a person asked the assistant to look at it.
ChatGPT-User,Claude-UserandPerplexity-Userare here. They do not crawl on their own schedule, and they do not reliably obey robots.txt. - Control tokens.
Google-ExtendedandApplebot-Extendedare not crawlers at all. They have no user agent and fetch nothing. They are opt-out switches governing how content that was already crawled may be used.
The tokens, by operator
| Operator | robots.txt token | Kind | What it does | Obeys robots.txt |
|---|---|---|---|---|
| OpenAI | OAI-SearchBot | Retrieval | Surfaces sites in ChatGPT's search features | Yes |
| OpenAI | GPTBot | Training | Collects content that may train foundation models | Yes |
| OpenAI | ChatGPT-User | User-triggered | Fetches a page on a user's direct request | May not apply |
| OpenAI | OAI-AdsBot | Other | Validates the safety of ad landing pages; not used for training | Yes |
| Anthropic | Claude-SearchBot | Retrieval | Navigates the web to improve search result quality | Yes |
| Anthropic | ClaudeBot | Training | Collects web content that may contribute to training | Yes |
| Anthropic | Claude-User | User-triggered | Accesses sites when a person asks Claude a question | Yes |
| Perplexity | PerplexityBot | Retrieval | Surfaces and links sites in Perplexity results; does not train models | Yes |
| Perplexity | Perplexity-User | User-triggered | Fetches a page to answer a user's question and link it | Generally ignores |
Googlebot | Retrieval | Builds the search index that AI Overviews are drawn from | Yes | |
Google-Extended | Control token | Governs Gemini training and grounding use. No crawling. No ranking effect. | n/a | |
Google-CloudVertexBot | Other | Crawls for Vertex AI agents built by site owners; no Search impact | Yes | |
| Apple | Applebot | Retrieval | Crawls for Spotlight, Siri and Safari, and for Apple Intelligence training | Yes |
| Apple | Applebot-Extended | Control token | Governs training use only. Does not affect Siri or Spotlight discoverability. | n/a |
Two Anthropic tokens still circulate in older robots.txt templates: Claude-Web and anthropic-ai. Neither appears in Anthropic's current documentation. Leaving an Allow for them costs nothing, but a robots.txt built only from them is missing every crawler Anthropic actually runs today.
The two expensive mistakes
1. Blocking Google-Extended and expecting to disappear from AI Overviews
It does not work that way, in either direction. Google states plainly that Google-Extended does not impact inclusion in Google Search and is not used as a ranking signal. It has no separate HTTP user agent; crawling is done with existing Google user agents. AI Overviews are assembled from Google's search index, which Googlebot populates.
The same logic applies to Apple. Disallowing Applebot-Extended only withholds content from Apple's foundation-model training. Apple documents that pages disallowing it remain discoverable through Spotlight, Siri and Safari.
2. Allowing the training crawler and blocking the retrieval crawler
This is the common accident, and it is backwards. A site that allows GPTBot but blocks OAI-SearchBot has permitted its content into model training while removing itself from the index ChatGPT actually searches when answering a question. The visible symptom is being absent from answers while appearing to have an open robots.txt.
How to check your own site
robots.txt tells you what you intended. These commands tell you what actually happens. Run each one and confirm you get 200.
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://example.com
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://example.com
curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" https://example.com
curl -s -o /dev/null -w "%{http_code}\n" -A "ClaudeBot/1.0" https://example.com
A 403 or 503 here, on a site whose robots.txt allows everything, means the block is above robots.txt. That is the normal case.
robots.txt is not where most blocks happen
robots.txt is a request, not a gate. It is read by the crawler after the crawler has already been allowed to connect. Most real-world blocking happens earlier, in infrastructure that was never configured with AI crawlers in mind.
- CDN bot management. Managed rule sets classify unfamiliar user agents as bots and challenge or reject them. Several CDNs ship AI-crawler blocking that is on by default or one toggle away.
- WAF rules. Rules written to stop scrapers catch retrieval crawlers, because they look identical at the request level.
- Rate limiting. A crawler fetching many pages quickly trips thresholds tuned for human traffic.
- JavaScript challenges. An interstitial that a browser solves silently returns a challenge page to a crawler, which then indexes the challenge page.
- Country blocks. Crawler egress IPs sit in a small number of regions. A geo rule can remove you entirely.
Each operator publishes the IP ranges its crawlers use, which is how you verify a request is genuine rather than a user agent someone copied. OpenAI publishes gptbot.json, searchbot.json and chatgpt-user.json; Anthropic publishes a single bots list; Perplexity publishes one file per crawler. Verifying by IP matters because the user-agent string is trivially forged, and a bot-management rule that trusts the string alone is both blocking real crawlers and admitting fake ones.
Questions
Does blocking Google-Extended remove me from AI Overviews?
No. Google-Extended is a control token, not a crawler. It has no separate user agent and does no fetching of its own. Google documents that it does not affect inclusion in Google Search and is not a ranking signal. It governs whether already-crawled content may be used to train Gemini models and to ground Gemini and Vertex AI. AI Overviews are built on Google's search index, which Googlebot populates, so blocking Googlebot is what removes you from AI Overviews.
Which AI crawlers actually decide whether I get cited?
The retrieval crawlers, not the training crawlers. OAI-SearchBot surfaces sites in ChatGPT search, PerplexityBot surfaces and links sites in Perplexity, and Claude-SearchBot supports Claude's search results. Blocking a training crawler such as GPTBot or ClaudeBot keeps your content out of future model training but does not by itself stop you being cited in a live answer. Blocking a retrieval crawler does.
Do user-triggered AI fetchers obey robots.txt?
Not reliably. OpenAI documents that ChatGPT-User may not apply robots.txt rules because it acts on a direct user request rather than autonomous crawling. Perplexity documents that Perplexity-User generally ignores robots.txt for the same reason. Treat robots.txt as a signal to autonomous crawlers, not as an access control.
Why does a site pass robots.txt checks but still get blocked?
Because most blocking happens above robots.txt, at the CDN or WAF layer. Bot-management rules, rate limits, JavaScript challenges and country blocks reject the request before robots.txt is ever consulted. A robots.txt file that allows every AI crawler tells you nothing about whether those crawlers receive a 200 response.
Sources
- OpenAI — Overview of OpenAI crawlers
- Anthropic — Does Anthropic crawl data from the web?
- Perplexity — PerplexityBot and Perplexity-User
- Google — Common crawlers, including Google-Extended
- Apple — About Applebot and Applebot-Extended
Related
Why AI traffic lands in your Direct bucket, and how to fix it in GA4. Once AI crawlers can reach you, the next question is whether you can see the traffic they send. GA4 files it under Referral and Direct by default.
Check your own crawler access
The AI Visibility Audit includes a crawler access report across every bot listed above, run against your live site rather than your robots.txt. $750, five to seven business days.