CdSEO

Open data · CC0-1.0

AI bot registry

The robots.txt tokens AI companies document for their crawlers, fetchers and control tokens: who runs each one, what it feeds, and whether the operator says it follows robots.txt. Every entry links to the operator's own page and says when it was last checked. CodoSEO uses this list to tell site owners which AI bots can reach their pages.

27 bots from 11 operators ai-bots.json
TokenProductPurposeFollows robots.txtIP listSourceReviewed
OpenAI4
GPTBot GPTBotUsed to train OpenAI's generative AI foundation models. Training Yes JSON openai.com 2026-10-09
OAI-SearchBot ChatGPT searchBlocking it removes the site from ChatGPT search answers; changes take about 24 hours. Search Yes JSON openai.com 2026-10-09
ChatGPT-User ChatGPT and Custom GPTsUser-initiated, so robots.txt rules may not apply. User-triggered Partly JSON openai.com 2026-10-09
OAI-AdsBot ChatGPT adsChecks pages submitted as ChatGPT ads; the documentation does not state its robots.txt behaviour. Ads Not stated JSON openai.com 2026-10-09
Anthropic3
ClaudeBot Claude trainingCollects web content for model training; also honours Crawl-delay. Training Yes JSON claude.com 2026-10-09
Claude-SearchBot Claude searchImproves Claude search result quality; blocking it can reduce visibility in Claude search. Search Yes JSON claude.com 2026-10-09
Claude-User Claude user fetchFetches pages when a user asks Claude; honours robots.txt, unlike most user-triggered fetchers. User-triggered Yes JSON claude.com 2026-10-09
Perplexity2
PerplexityBot Perplexity searchSurfaces and links sites in Perplexity search results. Search Yes JSON perplexity.ai 2026-10-09
Perplexity-User Perplexity user fetchSince a user requested the fetch, this fetcher generally ignores robots.txt rules. User-triggered No JSON perplexity.ai 2026-10-09
Google3
Googlebot Google SearchGoogle's search crawler. Search Yes JSON google.com 2026-10-09
Google-Extendedcontrol token Gemini training and groundingControl token only, with no user agent of its own; governs Gemini training and grounding and does not affect Search or AI Overviews. Training Yes — google.com 2026-10-09
Google-Agent Project Mariner agentsUser-triggered agent on Google infrastructure; generally ignores robots.txt and is verified by IP list or Web Bot Auth. Agent No JSON google.com 2026-10-09
Apple2
Applebot Spotlight, Siri and Safari searchCrawls for Apple search features; its data may also be used to train Apple foundation models. Search Yes JSON apple.com 2026-10-09
Applebot-Extendedcontrol token Apple foundation model trainingDoes not crawl; a usage-permission token for Apple model training only. Training Yes — apple.com 2026-10-09
Meta4
meta-externalagent Meta AI model trainingCollects content to help build Meta's AI models; no IP list published on the page. Training Yes — facebook.com 2026-10-09
meta-webindexer Meta AI searchImproves Meta AI search results; allowing it helps Meta cite and link to the site. Search Yes — facebook.com 2026-10-09
meta-externalfetcher Meta AI user fetchFetches links at a user's request and may bypass robots.txt rules. User-triggered Partly — facebook.com 2026-10-09
meta-externalads Meta adsCrawls for advertising and other business products. Ads Yes — facebook.com 2026-10-07
Amazon3
Amazonbot AmazonbotUsed to improve Amazon products and for AI model training; IP list is an HTML page, not JSON. Training Yes — amazon.com 2026-10-09
Amzn-SearchBot Alexa and Amazon searchImproves search in Alexa and other Amazon products; not used for generative AI training; IP list is an HTML page, not JSON. Search Yes — amazon.com 2026-10-09
Amzn-User Amazon user fetchFetches pages for live user queries and may not follow all robots.txt directives; IP list is an HTML page, not JSON. User-triggered Partly — amazon.com 2026-10-09
Microsoft1
Bingbot Bing and Copilot searchMicrosoft controls Copilot use of content with page-level meta tags, not a separate robots token. Search Yes JSON bing.com 2026-10-07
DuckDuckGo1
DuckAssistBot DuckDuckGo AI-assisted answersCrawls in real time for AI-assisted answers and is not used for training; a robots.txt opt-out takes about 72 hours. Search Yes JSON duckduckgo.com 2026-10-09
Mistral3
MistralAI-Index Mistral search indexIndexes pages to answer questions in Mistral products; the documentation does not state robots.txt behaviour. Search Not stated JSON mistral.ai 2026-10-09
MistralAI-User Mistral user fetchVisits pages when a user asks a question; the documentation does not state robots.txt behaviour. User-triggered Not stated JSON mistral.ai 2026-10-09
MistralAI-Training Mistral model trainingBuilds training datasets; webmasters can disallow it in robots.txt. No IP list is published. Training Yes — mistral.ai 2026-10-09
Common Crawl1
CCBot Common CrawlBuilds the open Common Crawl corpus used to train many models; spoofing is common, so verify by IP. Training Yes JSON commoncrawl.org 2026-10-09