Open daily 9:00 AM – 7:00 PM · WhatsApp +8801805-564599

Part of THB — The House of Business
THBSEO

Knowledge hub · Reference

Every AI crawler and user agent, explained

Verified against official sources · Last checked 6 October 2026

Short answer

AI companies use different crawlers for different jobs: training bots collect pages for future models, search bots build the index that assistants cite, and user-triggered fetchers load a page when someone asks about it. This list covers 29 officially documented AI user agents, what each does, and whether it follows robots.txt.

Which AI crawlers exist and what do they do?

29 AI crawlers and fetchers, grouped by purpose
User agentCompanyPurposeFollows robots.txtWhat it doesSource
Amzn-SearchBotAmazonSearch indexYesSupports search experiences such as Alexa; Amazon says it is not used for AI model training.Docs
Claude-SearchBotAnthropicSearch indexYesCrawls to improve the quality of search results shown to Claude users.Docs
ApplebotAppleSearch indexYesPowers search features in Spotlight, Siri and Safari; crawled data may also be used to train Apple's foundation models.Docs
Meta-WebIndexerMetaSearch indexYesIndexes web content to improve the relevance of search results in Meta AI.Docs
bingbotMicrosoftSearch indexYesBing's main crawler; Copilot answers draw on the Bing index, and Microsoft lets publishers limit chat/AI use with NOCACHE and NOARCHIVE meta tags (https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat).Docs
MistralAI-IndexMistral AISearch indexYesIndexes content for Mistral's search features; not used for training.Docs
OAI-SearchBotOpenAISearch indexYesSurfaces sites in ChatGPT search results; can be allowed while GPTBot is blocked.Docs
PerplexityBotPerplexitySearch indexYesIndexes sites to surface and link them in Perplexity answers; Perplexity says it is not used to train foundation models.Docs
YouBotYou.comSearch indexYesCrawls and indexes pages for You.com search; respects robots.txt including crawl-delay.Docs
Amzn-UserAmazonUser-triggered fetchUnclearFetches live information for user actions such as Alexa queries; Amazon says it may not follow all robots.txt directives.Docs
Claude-UserAnthropicUser-triggered fetchYesFetches pages when a Claude user asks a question that needs web access; Anthropic states it honours robots.txt.Docs
Google-AgentGoogleUser-triggered fetchNoUsed by agents hosted on Google infrastructure to navigate the web for users; as a user-triggered fetcher it generally ignores robots.txt.Docs
Google-GeminiNotebookGoogleUser-triggered fetchNoFetches URLs that users add as sources to a notebook; user-triggered, so robots.txt is generally ignored.Docs
Meta-ExternalFetcherMetaUser-triggered fetchNoFetches individual links on user request for AI features; Meta says it may bypass robots.txt.Docs
MistralAI-UserMistral AIUser-triggered fetchYesVisits pages when users ask Mistral's assistant questions; Mistral says webmasters can control it via robots.txt.Docs
ChatGPT-UserOpenAIUser-triggered fetchNoVisits pages when a ChatGPT or Custom GPT user asks; OpenAI says robots.txt rules may not apply to these user-initiated visits.Docs
Perplexity-UserPerplexityUser-triggered fetchNoVisits pages in response to user questions; Perplexity says it generally ignores robots.txt because a user requested the fetch.Docs
AmazonbotAmazonTrainingYesUsed to improve Amazon products and services and may be used to train Amazon AI models.Docs
ClaudeBotAnthropicTrainingYesGathers public web content that may contribute to model training.Docs
Applebot-ExtendedAppleTrainingYesDoes not crawl; a robots.txt token that lets publishers opt out of their content being used to train Apple's foundation models without leaving Apple search features.Docs
Google-ExtendedGoogleTrainingYesA robots.txt control token for whether content may be used to train Gemini models; blocking it does not affect Google Search ranking or inclusion.Docs
Meta-ExternalAgentMetaTrainingYesCrawls for uses such as training AI models or improving products by indexing content.Docs
MistralAI-TrainingMistral AITrainingYesCollects web content to build datasets for training Mistral's generative models.Docs
GPTBotOpenAITrainingYesCollects pages that may be used to train OpenAI's generative foundation models.Docs
CCBotCommon CrawlOtherYesBuilds Common Crawl's open web archive, a dataset widely reused by others, including for AI training.Docs
DuckAssistBotDuckDuckGoOtherYesFetches pages in real time for DuckDuckGo's AI-assisted answers; DuckDuckGo says the data is not used to train AI models.Docs
Google-CloudVertexBotGoogleOtherYesCrawls sites at the site owner's request to build Vertex AI Agents.Docs
GoogleOtherGoogleOtherYesGeneric crawler used by Google product teams for one-off crawls and internal research; not tied to a specific product. GoogleOther-Image and GoogleOther-Video are media variants.Docs
OAI-AdsBotOpenAIOtherYesChecks landing pages submitted as ads on ChatGPT; only visits submitted ad pages.Docs

What should my robots.txt say?

To be cited by AI assistants, allow the search-index and user-triggered bots. Blocking training bots is a separate business decision and does not stop citations. A starting point:

# Allow AI search and answer bots
User-agent: Amzn-SearchBot
User-agent: Claude-SearchBot
User-agent: Applebot
User-agent: Meta-WebIndexer
User-agent: bingbot
User-agent: MistralAI-Index
User-agent: OAI-SearchBot
User-agent: PerplexityBot
User-agent: YouBot
User-agent: Amzn-User
User-agent: Claude-User
User-agent: Google-Agent
User-agent: Google-GeminiNotebook
User-agent: Meta-ExternalFetcher
User-agent: MistralAI-User
User-agent: ChatGPT-User
User-agent: Perplexity-User
Allow: /

# Optional: block model-training bots
User-agent: Amazonbot
User-agent: ClaudeBot
User-agent: Applebot-Extended
User-agent: Google-Extended
User-agent: Meta-ExternalAgent
User-agent: MistralAI-Training
User-agent: GPTBot
Disallow: /

Build a custom file with our robots.txt generator. Remember that a firewall or CDN such as Cloudflare can block these bots before robots.txt is read.

Sources

  1. blogs.bing.com/webmaster/May-2012/To-crawl-or-not-to-crawl,-that-is-BingBot-s-questi/
  2. blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat
  3. commoncrawl.org/ccbot
  4. developer.amazon.com/amazonbot
  5. developers.facebook.com/docs/sharing/webmasters/web-crawlers/
  6. developers.google.com/search/docs/crawling-indexing/google-common-crawlers
  7. developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers
  8. developers.openai.com/api/docs/bots
  9. docs.mistral.ai/robots
  10. docs.perplexity.ai/guides/bots
  11. duckduckgo.com/duckduckgo-help-pages/results/duckassistbot
  12. support.apple.com/en-us/119829
  13. support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
  14. you.com/docs/youbot