Open data · CC0-1.0
AI bot registry
The robots.txt tokens AI companies document for their crawlers, fetchers and control tokens: who runs each one, what it feeds, and whether the operator says it follows robots.txt. Every entry links to the operator's own page and says when it was last checked. CodoSEO uses this list to tell site owners which AI bots can reach their pages.
| Token | Product | Purpose | Follows robots.txt | IP list | Source | Reviewed |
|---|---|---|---|---|---|---|
| OpenAI4 | ||||||
GPTBot |
GPTBotUsed to train OpenAI's generative AI foundation models. | Training | Yes | JSON | openai.com | 2026-10-09 |
OAI-SearchBot |
ChatGPT searchBlocking it removes the site from ChatGPT search answers; changes take about 24 hours. | Search | Yes | JSON | openai.com | 2026-10-09 |
ChatGPT-User |
ChatGPT and Custom GPTsUser-initiated, so robots.txt rules may not apply. | User-triggered | Partly | JSON | openai.com | 2026-10-09 |
OAI-AdsBot |
ChatGPT adsChecks pages submitted as ChatGPT ads; the documentation does not state its robots.txt behaviour. | Ads | Not stated | JSON | openai.com | 2026-10-09 |
| Anthropic3 | ||||||
ClaudeBot |
Claude trainingCollects web content for model training; also honours Crawl-delay. | Training | Yes | JSON | claude.com | 2026-10-09 |
Claude-SearchBot |
Claude searchImproves Claude search result quality; blocking it can reduce visibility in Claude search. | Search | Yes | JSON | claude.com | 2026-10-09 |
Claude-User |
Claude user fetchFetches pages when a user asks Claude; honours robots.txt, unlike most user-triggered fetchers. | User-triggered | Yes | JSON | claude.com | 2026-10-09 |
| Perplexity2 | ||||||
PerplexityBot |
Perplexity searchSurfaces and links sites in Perplexity search results. | Search | Yes | JSON | perplexity.ai | 2026-10-09 |
Perplexity-User |
Perplexity user fetchSince a user requested the fetch, this fetcher generally ignores robots.txt rules. | User-triggered | No | JSON | perplexity.ai | 2026-10-09 |
| Google3 | ||||||
Googlebot |
Google SearchGoogle's search crawler. | Search | Yes | JSON | google.com | 2026-10-09 |
Google-Extendedcontrol token |
Gemini training and groundingControl token only, with no user agent of its own; governs Gemini training and grounding and does not affect Search or AI Overviews. | Training | Yes | — | google.com | 2026-10-09 |
Google-Agent |
Project Mariner agentsUser-triggered agent on Google infrastructure; generally ignores robots.txt and is verified by IP list or Web Bot Auth. | Agent | No | JSON | google.com | 2026-10-09 |
| Apple2 | ||||||
Applebot |
Spotlight, Siri and Safari searchCrawls for Apple search features; its data may also be used to train Apple foundation models. | Search | Yes | JSON | apple.com | 2026-10-09 |
Applebot-Extendedcontrol token |
Apple foundation model trainingDoes not crawl; a usage-permission token for Apple model training only. | Training | Yes | — | apple.com | 2026-10-09 |
| Meta4 | ||||||
meta-externalagent |
Meta AI model trainingCollects content to help build Meta's AI models; no IP list published on the page. | Training | Yes | — | facebook.com | 2026-10-09 |
meta-webindexer |
Meta AI searchImproves Meta AI search results; allowing it helps Meta cite and link to the site. | Search | Yes | — | facebook.com | 2026-10-09 |
meta-externalfetcher |
Meta AI user fetchFetches links at a user's request and may bypass robots.txt rules. | User-triggered | Partly | — | facebook.com | 2026-10-09 |
meta-externalads |
Meta adsCrawls for advertising and other business products. | Ads | Yes | — | facebook.com | 2026-10-07 |
| Amazon3 | ||||||
Amazonbot |
AmazonbotUsed to improve Amazon products and for AI model training; IP list is an HTML page, not JSON. | Training | Yes | — | amazon.com | 2026-10-09 |
Amzn-SearchBot |
Alexa and Amazon searchImproves search in Alexa and other Amazon products; not used for generative AI training; IP list is an HTML page, not JSON. | Search | Yes | — | amazon.com | 2026-10-09 |
Amzn-User |
Amazon user fetchFetches pages for live user queries and may not follow all robots.txt directives; IP list is an HTML page, not JSON. | User-triggered | Partly | — | amazon.com | 2026-10-09 |
| Microsoft1 | ||||||
Bingbot |
Bing and Copilot searchMicrosoft controls Copilot use of content with page-level meta tags, not a separate robots token. | Search | Yes | JSON | bing.com | 2026-10-07 |
| DuckDuckGo1 | ||||||
DuckAssistBot |
DuckDuckGo AI-assisted answersCrawls in real time for AI-assisted answers and is not used for training; a robots.txt opt-out takes about 72 hours. | Search | Yes | JSON | duckduckgo.com | 2026-10-09 |
| Mistral3 | ||||||
MistralAI-Index |
Mistral search indexIndexes pages to answer questions in Mistral products; the documentation does not state robots.txt behaviour. | Search | Not stated | JSON | mistral.ai | 2026-10-09 |
MistralAI-User |
Mistral user fetchVisits pages when a user asks a question; the documentation does not state robots.txt behaviour. | User-triggered | Not stated | JSON | mistral.ai | 2026-10-09 |
MistralAI-Training |
Mistral model trainingBuilds training datasets; webmasters can disallow it in robots.txt. No IP list is published. | Training | Yes | — | mistral.ai | 2026-10-09 |
| Common Crawl1 | ||||||
CCBot |
Common CrawlBuilds the open Common Crawl corpus used to train many models; spoofing is common, so verify by IP. | Training | Yes | JSON | commoncrawl.org | 2026-10-09 |
No bots match. Try a company name, or a purpose such as “training”.