Book a call

AI crawler reference: every bot that matters, and what each one actually does

There are four kinds of AI user agent and only two of them fetch pages. Confusing them is the most common reason a site is invisible to AI assistants while looking perfectly healthy in Google Search Console.

Last verified against operator documentation: 3 October 2026. Every token below is sourced from the operator's own docs, linked at the end.

The four kinds of AI user agent

Treating all AI bots as one category is what produces the wrong robots.txt. They do different jobs, and only one category determines whether an AI assistant can cite you in a live answer.

  1. Retrieval crawlers. These build the index an assistant searches when answering a question. OAI-SearchBot, PerplexityBot, Claude-SearchBot and Googlebot are in this group. These are the ones that decide whether you can be cited.
  2. Training crawlers. These collect content that may be used to train future models. GPTBot, ClaudeBot and CCBot are here. Blocking them affects future model weights, not today's answers.
  3. User-triggered fetchers. These fetch a page because a person asked the assistant to look at it. ChatGPT-User, Claude-User and Perplexity-User are here. They do not crawl on their own schedule, and they do not reliably obey robots.txt.
  4. Control tokens. Google-Extended and Applebot-Extended are not crawlers at all. They have no user agent and fetch nothing. They are opt-out switches governing how content that was already crawled may be used.

The tokens, by operator

Verified against each operator's published documentation, 3 October 2026.
Operatorrobots.txt tokenKindWhat it doesObeys robots.txt
OpenAIOAI-SearchBotRetrievalSurfaces sites in ChatGPT's search featuresYes
OpenAIGPTBotTrainingCollects content that may train foundation modelsYes
OpenAIChatGPT-UserUser-triggeredFetches a page on a user's direct requestMay not apply
OpenAIOAI-AdsBotOtherValidates the safety of ad landing pages; not used for trainingYes
AnthropicClaude-SearchBotRetrievalNavigates the web to improve search result qualityYes
AnthropicClaudeBotTrainingCollects web content that may contribute to trainingYes
AnthropicClaude-UserUser-triggeredAccesses sites when a person asks Claude a questionYes
PerplexityPerplexityBotRetrievalSurfaces and links sites in Perplexity results; does not train modelsYes
PerplexityPerplexity-UserUser-triggeredFetches a page to answer a user's question and link itGenerally ignores
GoogleGooglebotRetrievalBuilds the search index that AI Overviews are drawn fromYes
GoogleGoogle-ExtendedControl tokenGoverns Gemini training and grounding use. No crawling. No ranking effect.n/a
GoogleGoogle-CloudVertexBotOtherCrawls for Vertex AI agents built by site owners; no Search impactYes
AppleApplebotRetrievalCrawls for Spotlight, Siri and Safari, and for Apple Intelligence trainingYes
AppleApplebot-ExtendedControl tokenGoverns training use only. Does not affect Siri or Spotlight discoverability.n/a

Two Anthropic tokens still circulate in older robots.txt templates: Claude-Web and anthropic-ai. Neither appears in Anthropic's current documentation. Leaving an Allow for them costs nothing, but a robots.txt built only from them is missing every crawler Anthropic actually runs today.

The two expensive mistakes

1. Blocking Google-Extended and expecting to disappear from AI Overviews

It does not work that way, in either direction. Google states plainly that Google-Extended does not impact inclusion in Google Search and is not used as a ranking signal. It has no separate HTTP user agent; crawling is done with existing Google user agents. AI Overviews are assembled from Google's search index, which Googlebot populates.

The same logic applies to Apple. Disallowing Applebot-Extended only withholds content from Apple's foundation-model training. Apple documents that pages disallowing it remain discoverable through Spotlight, Siri and Safari.

2. Allowing the training crawler and blocking the retrieval crawler

This is the common accident, and it is backwards. A site that allows GPTBot but blocks OAI-SearchBot has permitted its content into model training while removing itself from the index ChatGPT actually searches when answering a question. The visible symptom is being absent from answers while appearing to have an open robots.txt.

How to check your own site

robots.txt tells you what you intended. These commands tell you what actually happens. Run each one and confirm you get 200.

curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" https://example.com

curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" https://example.com

curl -s -o /dev/null -w "%{http_code}\n" -A "Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)" https://example.com

curl -s -o /dev/null -w "%{http_code}\n" -A "ClaudeBot/1.0" https://example.com

A 403 or 503 here, on a site whose robots.txt allows everything, means the block is above robots.txt. That is the normal case.

robots.txt is not where most blocks happen

robots.txt is a request, not a gate. It is read by the crawler after the crawler has already been allowed to connect. Most real-world blocking happens earlier, in infrastructure that was never configured with AI crawlers in mind.

Each operator publishes the IP ranges its crawlers use, which is how you verify a request is genuine rather than a user agent someone copied. OpenAI publishes gptbot.json, searchbot.json and chatgpt-user.json; Anthropic publishes a single bots list; Perplexity publishes one file per crawler. Verifying by IP matters because the user-agent string is trivially forged, and a bot-management rule that trusts the string alone is both blocking real crawlers and admitting fake ones.

Questions

Does blocking Google-Extended remove me from AI Overviews?

No. Google-Extended is a control token, not a crawler. It has no separate user agent and does no fetching of its own. Google documents that it does not affect inclusion in Google Search and is not a ranking signal. It governs whether already-crawled content may be used to train Gemini models and to ground Gemini and Vertex AI. AI Overviews are built on Google's search index, which Googlebot populates, so blocking Googlebot is what removes you from AI Overviews.

Which AI crawlers actually decide whether I get cited?

The retrieval crawlers, not the training crawlers. OAI-SearchBot surfaces sites in ChatGPT search, PerplexityBot surfaces and links sites in Perplexity, and Claude-SearchBot supports Claude's search results. Blocking a training crawler such as GPTBot or ClaudeBot keeps your content out of future model training but does not by itself stop you being cited in a live answer. Blocking a retrieval crawler does.

Do user-triggered AI fetchers obey robots.txt?

Not reliably. OpenAI documents that ChatGPT-User may not apply robots.txt rules because it acts on a direct user request rather than autonomous crawling. Perplexity documents that Perplexity-User generally ignores robots.txt for the same reason. Treat robots.txt as a signal to autonomous crawlers, not as an access control.

Why does a site pass robots.txt checks but still get blocked?

Because most blocking happens above robots.txt, at the CDN or WAF layer. Bot-management rules, rate limits, JavaScript challenges and country blocks reject the request before robots.txt is ever consulted. A robots.txt file that allows every AI crawler tells you nothing about whether those crawlers receive a 200 response.

Sources

Related

Why AI traffic lands in your Direct bucket, and how to fix it in GA4. Once AI crawlers can reach you, the next question is whether you can see the traffic they send. GA4 files it under Referral and Direct by default.

Check your own crawler access

The AI Visibility Audit includes a crawler access report across every bot listed above, run against your live site rather than your robots.txt. $750, five to seven business days.

Book a 20 minute call