AI crawlers · directory of 20

AI Crawlers: User Agents, Operators and Purposes

AI crawlers are the bots and robots.txt tokens that AI companies use to collect training data, build search indexes and fetch pages for users. This directory lists 20 of them from 9 operators, including GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended and CCBot, with each one's robots.txt token, purpose, and whether it obeys robots.txt, taken from the operator's own documentation.

Last checked against operator documentation: 25 September 2026.

Which AI crawlers exist, and who runs them?

This directory lists 20 documented AI crawlers and tokens from 9 operators: OpenAI, Anthropic, Perplexity, Google, Apple, Meta, Amazon, Common Crawl and Microsoft. OpenAI runs 4, Anthropic and Amazon run 3 each, Perplexity, Google, Apple and Meta run 2 each, and Common Crawl and Microsoft run 1 each.

Each row takes its purpose and robots.txt behaviour from the operator’s own page, quoted where the operator hedges. Open a row’s details for the full user-agent string and the way to verify the bot.

Showing 20 of 20 crawlers

20 AI crawlers and robots.txt tokens from 9 operators, with purpose and robots.txt behaviour from operator documentation (checked 25 September 2026)
TokenOperatorRolePurposeObeys robots.txtUser-agent string and verification
GPTBotOpenAITrainingCrawls content "that may be used in training" OpenAI's generative AI foundation modelsYes
GPTBot string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

IP list: openai.com/gptbot.json

OAI-SearchBotOpenAISearchSurfaces websites in search results in ChatGPT's search featuresYes
OAI-SearchBot string and verification

Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

IP list: openai.com/searchbot.json

ChatGPT-UserOpenAIUser-triggeredVisits a page when a user asks ChatGPT or a Custom GPT a question; also used by GPT ActionsNot reliably: "robots.txt rules may not apply"
ChatGPT-User string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

IP list: openai.com/chatgpt-user.json

OAI-AdsBotOpenAIAds reviewVisits only landing pages submitted as ads on ChatGPT; its data is not used to train foundation modelsNot stated by OpenAI
OAI-AdsBot string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot

IP list: openai.com/adsbot.json

ClaudeBotAnthropicTrainingCollects web content that "could potentially contribute" to training Anthropic's modelsYes; also honours Crawl-delay
ClaudeBot string and verification

Not published by Anthropic

IP list: claude.com/crawling/bots.json

Claude-SearchBotAnthropicSearchNavigates the web to improve search result quality for Claude usersYes
Claude-SearchBot string and verification

Not published by Anthropic

IP list: claude.com/crawling/bots.json

Claude-UserAnthropicUser-triggeredAccesses websites when people ask Claude questionsYes: disallowing it stops retrieval for user queries
Claude-User string and verification

Not published by Anthropic

IP list: claude.com/crawling/bots.json

PerplexityBotPerplexitySearchSurfaces and links websites in Perplexity search; not used to train foundation modelsYes
PerplexityBot string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

IP list: perplexity.com/perplexitybot.json

Perplexity-UserPerplexityUser-triggeredVisits a page to answer a user's question; not a crawler and not used for trainingNot reliably: "generally ignores robots.txt rules"
Perplexity-User string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

IP list: perplexity.com/perplexity-user.json

GooglebotGoogleSearchGoogle Search, including AI Overviews and AI ModeYes
Googlebot string and verification

Desktop: Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Googlebot/2.1; +http://www.google.com/bot.html) Chrome/W.X.Y.Z Safari/537.36

Google-published IP ranges; reverse DNS on googlebot.com

Google-ExtendedGoogleControl tokenDecides whether Google-crawled content trains future Gemini models and grounds Gemini Apps and Vertex AI; no effect on Google SearchToken only
Google-Extended string and verification

None: Google crawls with its existing user agents

Not applicable

ApplebotAppleSearchPowers search in Spotlight, Siri and Safari; its data "may also be used" to train Apple foundation modelsYes; follows Googlebot rules when robots.txt does not name it
Applebot string and verification

Safari string ending in (Applebot/0.1; +http://www.apple.com/go/applebot)

Reverse DNS on applebot.apple.com; Applebot IP CIDR file

Applebot-ExtendedAppleControl tokenOpts content out of training Apple's foundation models; pages stay in Apple search resultsToken only: it does not crawl webpages
Applebot-Extended string and verification

None

Not applicable

Meta-ExternalAgentMetaTraining"training foundation AI models or improving products by indexing content directly"Yes
Meta-ExternalAgent string and verification

meta-externalagent/1.1

No IP list on Meta's crawler page

Meta-ExternalFetcherMetaUser-triggeredFetches individual links at a user's request, including for agentic AI tasksNot reliably: "may bypass robots.txt rules"
Meta-ExternalFetcher string and verification

meta-externalfetcher/1.1

No IP list on Meta's crawler page

AmazonbotAmazonTrainingImproves Amazon products and services; "may be used to train Amazon AI models"Yes
Amazonbot string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36

IP list: developer.amazon.com/amazonbot/ip-addresses/

Amzn-SearchBotAmazonSearchSearch experiences in Amazon products such as Alexa; not used for generative AI trainingYes; follows other search bots' rules when not named
Amzn-SearchBot string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36

IP list: developer.amazon.com/amazonbot/searchbot-ip-addresses/

Amzn-UserAmazonUser-triggeredFetches live information for user actions such as Alexa questions; not used for trainingNot reliably: "may not follow all robots.txt directives"
Amzn-User string and verification

Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36

IP list: developer.amazon.com/amazonbot/live-ip-addresses/

CCBotCommon CrawlTraining (open corpus)Builds Common Crawl's open repository of web crawl dataYes
CCBot string and verification

CCBot/2.0 (https://commoncrawl.org/faq/)

Reverse DNS on crawl.commoncrawl.org; IP list: index.commoncrawl.org/ccbot.json

BingbotMicrosoftSearchCrawls pages for the Bing search indexYes
Bingbot string and verification

Documented in Bing Webmaster Tools help

Verify Bingbot tool: bing.com/toolbox/verify-bingbot

The table splits into 7 search crawlers, 5 user-triggered fetchers, 5 training crawlers including Common Crawl’s open corpus, 2 control tokens and 1 ad-review bot. Meta’s crawler page, updated 21 May 2026, also documents Meta-WebIndexer, which Meta says helps it cite and link content in Meta AI responses; it is not yet in this table.

What are AI crawlers?

AI crawlers are automated clients that AI companies send to websites to collect training data, build search indexes or fetch a page for one user's question. Each identifies itself with a user-agent token, which site owners address in robots.txt to allow or refuse it.

Two identifiers matter for every bot. The user-agent string is the full text the bot sends with each request, such as CCBot/2.0 (https://commoncrawl.org/faq/). The token is the short name inside it, such as CCBot, and robots.txt groups match on the token. robots.txt is the plain-text file at the root of each host that lists which paths each token may fetch. Operators publish the token and, in most cases, an IP list or reverse DNS pattern so site owners can tell real traffic from copies. A crawler that fetches a page is the precondition for any AI answer citing it.

How do training crawlers, search crawlers and user-triggered fetchers differ?

Training crawlers gather content for future models, search crawlers index pages that engines cite in live answers, and user-triggered fetchers open one page because a user asked. Blocking a training crawler leaves search citations intact; blocking a search crawler removes a site from that engine's answers.

Three crawler roles and the effect of blocking each
RoleExamplesWhat it feedsEffect of blocking
Training crawlerGPTBot, ClaudeBot, CCBot, Meta-ExternalAgentTraining data for future modelsFuture content leaves the training set; search answers are unaffected
Search crawlerOAI-SearchBot, Claude-SearchBot, PerplexityBot, GooglebotThe index behind cited answersThe site drops out of that engine's answers
User-triggered fetcherChatGPT-User, Claude-User, Perplexity-UserOne answer for one userDepends on the operator: some fetchers do not reliably follow robots.txt

The split is visible in traffic. Cloudflare reported on 28 August 2025 that training accounted for nearly 80% of AI bot crawling it observed. In its unfiltered data for the first week of August 2025, Anthropic’s crawlers made nearly 50,000 requests for every visitor they referred, OpenAI’s 887 and Perplexity’s 118. Operators that run search and training as separate tokens let a site accept one and refuse the other.

What is a control token such as Google-Extended?

A control token is a robots.txt name that no bot sends. It changes how content fetched by another crawler is used. The directory lists 2:

  • Google-Extended decides whether content that Google’s existing crawlers fetch trains Gemini models and grounds Gemini Apps and Vertex AI. Google states that it “does not impact a site’s inclusion in Google Search”. For AI Overviews and AI Mode, Google names robots.txt rules for Googlebot as the control.
  • Applebot-Extended decides whether content that Applebot fetches trains Apple’s foundation models. Apple states that “Applebot-Extended does not crawl webpages” and that pages disallowing it “can still be included in search results”.

Which AI crawlers ignore robots.txt?

Four documented agents do not reliably follow robots.txt: ChatGPT-User, Perplexity-User, Meta-ExternalFetcher and Amzn-User. Their operators state that user-initiated fetches fall outside robots.txt. Anthropic states the opposite for Claude-User: disallowing it stops Claude from retrieving the site for user queries.

  • ChatGPT-User (OpenAI): “Because these actions are initiated by a user, robots.txt rules may not apply.”
  • Perplexity-User (Perplexity): the fetcher “generally ignores robots.txt rules” because a user requested the page.
  • Meta-ExternalFetcher (Meta): “this crawler may bypass robots.txt rules.”
  • Amzn-User (Amazon): “it may not follow all robots.txt directives.”

Each of these fetchers acts for one person’s request, not an automatic crawl. OpenAI adds that ChatGPT-User “is not used to determine whether content may appear in Search”. A firewall rule on the operator’s published IP list refuses them where robots.txt does not.

How do you verify that a request comes from a real AI crawler?

Verify a crawler by matching the request's IP address against the operator's published IP list or by a reverse DNS lookup. OpenAI, Anthropic, Perplexity, Amazon, Google, Apple and Common Crawl publish IP lists; Google, Apple and Common Crawl also support reverse DNS. User-agent strings alone are spoofable.

Common Crawl states it is “aware of crawlers falsely identifying themselves as CCBot”, and Google warns that “the HTTP user agent string can be spoofed”. The method differs by operator:

How to verify each operator's crawlers
OperatorMethodWhere
OpenAIIP list per botopenai.com/gptbot.json, /searchbot.json, /chatgpt-user.json, /adsbot.json
AnthropicOne IP list for all 3 botsclaude.com/crawling/bots.json
PerplexityIP list per botperplexity.com/perplexitybot.json, /perplexity-user.json
GoogleIP ranges and reverse DNSHostnames on googlebot.com
AppleReverse DNS or IP CIDR fileHostnames on applebot.apple.com
AmazonIP list per botdeveloper.amazon.com/amazonbot/ip-addresses/, /searchbot-ip-addresses/, /live-ip-addresses/
Common CrawlReverse DNS or IP listHostnames on crawl.commoncrawl.org; index.commoncrawl.org/ccbot.json
MicrosoftVerify Bingbot toolbing.com/toolbox/verify-bingbot
MetaNo IP list on the crawler pageContact webmasters@meta.com

A reverse DNS check runs in 2 steps: look up the hostname for the IP address, then confirm that the hostname resolves back to the same IP. Anthropic adds a warning against using its IP list to block: IP blocking “may not work correctly or persistently guarantee an opt-out”, because it stops its bots from reading robots.txt.

Which AI user agents are retired or undocumented?

anthropic-ai and Claude-Web are absent from Anthropic's current crawler documentation, which names only ClaudeBot, Claude-SearchBot and Claude-User. Bytespider appears in many block lists but has no operator documentation on record as of September 2026, so this directory excludes it.

  • anthropic-ai: an older Anthropic name, not listed in the current article.
  • Claude-Web: an older Anthropic name, not listed in the current article.
  • Bytespider: a ByteDance user agent with no published operator documentation.

Rules for these names do no harm in robots.txt, but they control nothing that Anthropic documents today.

The list names each bot; a site’s robots.txt decides which of them get in.

Test your robots.txt against these crawlers

Paste a robots.txt file or enter a domain, choose a path, and the AI crawler checker tool reports Allowed or Blocked for 16 of the agents above, with the rule that decides each result.

Explore each crawler

Three crawlers and one file have their own pages, with rule sets, verification steps and the effect of blocking.

  • GPTBot

    OpenAI's training crawler: user-agent string, IP list and rules to block it without leaving ChatGPT search.

  • OAI-SearchBot

    The OpenAI crawler that decides whether pages appear as sources in ChatGPT search.

  • ClaudeBot

    Anthropic's training crawler and its 2 siblings, with what blocking each one does to Claude answers.

  • llms.txt proposal

    A file for agents, not a bot: format, v2 changes and what is documented about who reads it.

Frequently asked questions

Is llms.txt an AI crawler?

No. llms.txt is a Markdown file that a site publishes for AI agents to read, not a bot that visits sites. It lists key pages and grants or blocks nothing; crawler access stays under robots.txt.

How long do AI crawlers take to respect a robots.txt change?

About 24 hours for the operators that state a figure. OpenAI gives about 24 hours for ChatGPT search results, Amazon gives about 24 hours for its 3 agents, and Meta asks site owners to allow up to 24 hours because its crawlers cache robots.txt.

Sources

  1. OpenAI — Overview of OpenAI Crawlers
  2. Anthropic — Does Anthropic crawl data from the web, and how can site owners block the crawler?
  3. Perplexity — Perplexity Crawlers
  4. Google — Google's common crawlers
  5. Google — AI features and your website
  6. Apple — About Applebot (published 4 September 2026)
  7. Meta — Meta Web Crawlers (updated 21 May 2026)
  8. Amazon — About Amazonbot
  9. Common Crawl — CCBot
  10. Bing Webmaster Tools — Webmaster Guidelines
  11. Cloudflare Blog — AI crawler traffic by purpose and industry (David Belson, 28 August 2025)