Knowledge hub · Reference
Every AI crawler and user agent, explained
Verified against official sources · Last checked 6 October 2026
Short answer
AI companies use different crawlers for different jobs: training bots collect pages for future models, search bots build the index that assistants cite, and user-triggered fetchers load a page when someone asks about it. This list covers 29 officially documented AI user agents, what each does, and whether it follows robots.txt.
Which AI crawlers exist and what do they do?
| User agent | Company | Purpose | Follows robots.txt | What it does | Source |
|---|---|---|---|---|---|
| Amzn-SearchBot | Amazon | Search index | Yes | Supports search experiences such as Alexa; Amazon says it is not used for AI model training. | Docs |
| Claude-SearchBot | Anthropic | Search index | Yes | Crawls to improve the quality of search results shown to Claude users. | Docs |
| Applebot | Apple | Search index | Yes | Powers search features in Spotlight, Siri and Safari; crawled data may also be used to train Apple's foundation models. | Docs |
| Meta-WebIndexer | Meta | Search index | Yes | Indexes web content to improve the relevance of search results in Meta AI. | Docs |
| bingbot | Microsoft | Search index | Yes | Bing's main crawler; Copilot answers draw on the Bing index, and Microsoft lets publishers limit chat/AI use with NOCACHE and NOARCHIVE meta tags (https://blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat). | Docs |
| MistralAI-Index | Mistral AI | Search index | Yes | Indexes content for Mistral's search features; not used for training. | Docs |
| OAI-SearchBot | OpenAI | Search index | Yes | Surfaces sites in ChatGPT search results; can be allowed while GPTBot is blocked. | Docs |
| PerplexityBot | Perplexity | Search index | Yes | Indexes sites to surface and link them in Perplexity answers; Perplexity says it is not used to train foundation models. | Docs |
| YouBot | You.com | Search index | Yes | Crawls and indexes pages for You.com search; respects robots.txt including crawl-delay. | Docs |
| Amzn-User | Amazon | User-triggered fetch | Unclear | Fetches live information for user actions such as Alexa queries; Amazon says it may not follow all robots.txt directives. | Docs |
| Claude-User | Anthropic | User-triggered fetch | Yes | Fetches pages when a Claude user asks a question that needs web access; Anthropic states it honours robots.txt. | Docs |
| Google-Agent | User-triggered fetch | No | Used by agents hosted on Google infrastructure to navigate the web for users; as a user-triggered fetcher it generally ignores robots.txt. | Docs | |
| Google-GeminiNotebook | User-triggered fetch | No | Fetches URLs that users add as sources to a notebook; user-triggered, so robots.txt is generally ignored. | Docs | |
| Meta-ExternalFetcher | Meta | User-triggered fetch | No | Fetches individual links on user request for AI features; Meta says it may bypass robots.txt. | Docs |
| MistralAI-User | Mistral AI | User-triggered fetch | Yes | Visits pages when users ask Mistral's assistant questions; Mistral says webmasters can control it via robots.txt. | Docs |
| ChatGPT-User | OpenAI | User-triggered fetch | No | Visits pages when a ChatGPT or Custom GPT user asks; OpenAI says robots.txt rules may not apply to these user-initiated visits. | Docs |
| Perplexity-User | Perplexity | User-triggered fetch | No | Visits pages in response to user questions; Perplexity says it generally ignores robots.txt because a user requested the fetch. | Docs |
| Amazonbot | Amazon | Training | Yes | Used to improve Amazon products and services and may be used to train Amazon AI models. | Docs |
| ClaudeBot | Anthropic | Training | Yes | Gathers public web content that may contribute to model training. | Docs |
| Applebot-Extended | Apple | Training | Yes | Does not crawl; a robots.txt token that lets publishers opt out of their content being used to train Apple's foundation models without leaving Apple search features. | Docs |
| Google-Extended | Training | Yes | A robots.txt control token for whether content may be used to train Gemini models; blocking it does not affect Google Search ranking or inclusion. | Docs | |
| Meta-ExternalAgent | Meta | Training | Yes | Crawls for uses such as training AI models or improving products by indexing content. | Docs |
| MistralAI-Training | Mistral AI | Training | Yes | Collects web content to build datasets for training Mistral's generative models. | Docs |
| GPTBot | OpenAI | Training | Yes | Collects pages that may be used to train OpenAI's generative foundation models. | Docs |
| CCBot | Common Crawl | Other | Yes | Builds Common Crawl's open web archive, a dataset widely reused by others, including for AI training. | Docs |
| DuckAssistBot | DuckDuckGo | Other | Yes | Fetches pages in real time for DuckDuckGo's AI-assisted answers; DuckDuckGo says the data is not used to train AI models. | Docs |
| Google-CloudVertexBot | Other | Yes | Crawls sites at the site owner's request to build Vertex AI Agents. | Docs | |
| GoogleOther | Other | Yes | Generic crawler used by Google product teams for one-off crawls and internal research; not tied to a specific product. GoogleOther-Image and GoogleOther-Video are media variants. | Docs | |
| OAI-AdsBot | OpenAI | Other | Yes | Checks landing pages submitted as ads on ChatGPT; only visits submitted ad pages. | Docs |
What should my robots.txt say?
To be cited by AI assistants, allow the search-index and user-triggered bots. Blocking training bots is a separate business decision and does not stop citations. A starting point:
# Allow AI search and answer bots User-agent: Amzn-SearchBot User-agent: Claude-SearchBot User-agent: Applebot User-agent: Meta-WebIndexer User-agent: bingbot User-agent: MistralAI-Index User-agent: OAI-SearchBot User-agent: PerplexityBot User-agent: YouBot User-agent: Amzn-User User-agent: Claude-User User-agent: Google-Agent User-agent: Google-GeminiNotebook User-agent: Meta-ExternalFetcher User-agent: MistralAI-User User-agent: ChatGPT-User User-agent: Perplexity-User Allow: / # Optional: block model-training bots User-agent: Amazonbot User-agent: ClaudeBot User-agent: Applebot-Extended User-agent: Google-Extended User-agent: Meta-ExternalAgent User-agent: MistralAI-Training User-agent: GPTBot Disallow: /
Build a custom file with our robots.txt generator. Remember that a firewall or CDN such as Cloudflare can block these bots before robots.txt is read.
Sources
- blogs.bing.com/webmaster/May-2012/To-crawl-or-not-to-crawl,-that-is-BingBot-s-questi/
- blogs.bing.com/webmaster/september-2023/Announcing-new-options-for-webmasters-to-control-usage-of-their-content-in-Bing-Chat
- commoncrawl.org/ccbot
- developer.amazon.com/amazonbot
- developers.facebook.com/docs/sharing/webmasters/web-crawlers/
- developers.google.com/search/docs/crawling-indexing/google-common-crawlers
- developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers
- developers.openai.com/api/docs/bots
- docs.mistral.ai/robots
- docs.perplexity.ai/guides/bots
- duckduckgo.com/duckduckgo-help-pages/results/duckassistbot
- support.apple.com/en-us/119829
- support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler
- you.com/docs/youbot