Crawlers
The agent strings, robots tokens and verification methods, kept current and dated. Every page links back into the checker so you can test the rule against your own site.
-
Amazonbot
Crawling to improve Amazon products and services, including answers surfaced to customers, and it may be used to train Amazon AI models.
-
Applebot-Extended
Controlling whether content already crawled by Applebot may be used to train Apple's generative foundation models.
-
Bingbot
Crawling for the Bing index, which also backs Copilot answers and several licensed search products.
-
Bytespider
Collecting web content for ByteDance. The operator publishes no statement of purpose, so any description of what it feeds is inference.
-
CCBot
Building the free public Common Crawl archive, which many research projects and several model training pipelines read downstream.
-
ClaudeBot
Collecting public web content for model development. Search and user-directed retrieval run under separate tokens.
-
Google-Extended
Controlling whether crawled content may be used for training Gemini models and for grounding in Gemini Apps and Vertex AI.
-
Googlebot
Crawling for Google Search indexing and the Search-derived surfaces built on that index.
-
GPTBot
Collecting web content used to train OpenAI's foundation models. It does not drive ChatGPT search citations.
-
Meta-ExternalAgent
Crawling for use cases such as training foundation AI models and improving products by indexing content directly.
-
PerplexityBot
Indexing pages so they can appear as cited sources in Perplexity answers. Perplexity states it does not collect content for model training.