Know which AI crawlers are real, what they do, and how to handle them
A practical, source-checked reference that separates operator documentation from reported user-agents and bounded Trakkr observations.
What these crawlers do
Training crawler
Collects or controls access to content that can feed AI training data.
AI search crawler
Indexes pages so AI search products can retrieve, rank, cite, or summarize them.
Live fetcher
Fetches pages because a user or agent asked for a specific URL or task.
A safer allow or block decision
- 1. Start with purpose. Search, training, and user-triggered retrieval have different costs when blocked.
- 2. Use a documented token. Target the narrowest official robots.txt token. If none is published, do not invent one.
- 3. Verify traffic separately. A robots rule sets policy; IP, reverse DNS, or signatures help identify the sender.
Observed across connected sites
A real, bounded view of crawler signatures across connected sites. It is evidence of matching requests, not proof of operator origin and not a market-share ranking.
31
completed days
85
connected sites
7,963,265
classified requests
2026-07-18
through 2026-08-17
Find a crawler
AI2Bot
Training crawlerOfficially documentedAllen Institute for AI crawler used to find web content for open language model datasets.
AI2BotAi2Bot-Dolma
Training crawlerOfficially documentedAI2 crawler token associated with Dolma/open language model dataset collection.
Ai2Bot-Dolmaamazon-kendra
AI search crawlerReported, not operator verifiedAmazon Kendra crawler token for intelligent enterprise search over configured content sources.
amazon-kendraAmazonbot
Other crawlerOfficially documentedSample observedAmazon crawler used to improve its products and services and potentially train Amazon AI models.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36AmazonBuyForMe
Live fetcherReported, not operator verifiedAmazon agent token reported for Buy for Me shopping actions directed by customers.
AmazonBuyForMeAmzn-SearchBot
AI search crawlerOfficially documentedAmazon search crawler that indexes pages so they can be retrieved and cited in Amazon search and assistant answers.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36Amzn-User
Live fetcherOfficially documentedAmazon fetcher that retrieves a specific page because a person asked an Amazon assistant about it.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36anthropic-ai
Training crawlerReported, not operator verifiedLegacy Anthropic robots.txt token that predates the current ClaudeBot, Claude-User and Claude-SearchBot names.
anthropic-aiApifyBot
Other crawlerReported, not operator verifiedToken associated with crawlers run on the Apify scraping platform by its customers.
ApifyBotApplebot
AI search crawlerOfficially documentedSample observedApple crawler for search experiences across Spotlight, Siri, Safari, and related Apple surfaces.
Mozilla/5.0 (Device; OS_version) AppleWebKit/WebKit_version (KHTML, like Gecko) Version/Safari_version Safari/WebKit_version (Applebot/Applebot_version; +http://www.apple.com/go/applebot)Applebot-Extended
Training crawlerOfficially documentedSample observedRobots.txt control token for whether Applebot-crawled content may be used to train Apple foundation models.
Applebot-Extended control tokenatlassian-bot
AI search crawlerOfficially documentedAtlassian Rovo crawler used to index connected website content for AI search, assistants, and agents.
atlassian-botbedrockbot
Other crawlerOfficially documentedAmazon Bedrock web crawler connector token for customer-configured AI applications.
bedrockbotBingbot
AI search crawlerReported, not operator verifiedMicrosoft Bing crawler used to crawl and index pages for Bing and Microsoft search-powered experiences.
bingbotBravebot
AI search crawlerReported, not operator verifiedReported name for Brave Search crawling. Brave says its crawler does not advertise a differentiated user-agent.
No official user-agent publishedBrightbot
Other crawlerOfficially documentedBright Data's declared crawler for collecting public web data for its own datasets.
Brightbot 1.0Bytespider
Training crawlerReported, not operator verifiedSample observedByteDance crawler associated with training and powering AI products.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Bytespider; spider-feedback@bytedance.comCCBot
Training crawlerOfficially documentedCommon Crawl's crawler for building public web crawl datasets used by researchers and AI builders.
CCBot/2.0 (https://commoncrawl.org/faq/)ChatGPT Agent
Live fetcherOfficially documentedOpenAI agent used when ChatGPT navigates websites for user-directed tasks.
ChatGPT AgentChatGPT-User
Live fetcherOfficially documentedSample observedUser-triggered OpenAI fetcher for ChatGPT and Custom GPT actions.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/botClaude-Code
Live fetcherReported, not operator verifiedClaude Code related agent token seen in crawler/user-agent lists.
Claude-CodeClaude-SearchBot
AI search crawlerOfficially documentedSample observedAnthropic search crawler that indexes content to improve Claude search result relevance and accuracy.
Mozilla/5.0 (compatible; Claude-SearchBot/1.0; +claudebot@anthropic.com)Claude-User
Live fetcherOfficially documentedAnthropic user-triggered fetcher for Claude answers that need a specific web page.
Claude-UserClaude-Web
Live fetcherReported, not operator verifiedReported Anthropic-related token seen in public crawler registries and Trakkr detection, but absent from Anthropic's current bot documentation.
Claude-WebClaudeBot
Training crawlerOfficially documentedSample observedAnthropic crawler for public web content that could contribute to Claude model training.
Mozilla/5.0 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)Cloudflare-AutoRAG
AI search crawlerOfficially documentedCloudflare AutoRAG crawler used to index configured content for AI search applications.
Cloudflare-AutoRAGcohere-ai
Live fetcherReported, not operator verifiedSample observedCohere token reported for retrieving data in response to user-initiated prompts.
cohere-aicohere-training-data-crawler
Training crawlerReported, not operator verifiedCohere training-data crawler token reported for downloading web data for enterprise language models.
cohere-training-data-crawlerDeepSeekBot
Training crawlerReported, not operator verifiedDeepSeek crawler token reported for training language models and improving AI products.
DeepSeekBotDiffbot
Other crawlerReported, not operator verifiedSample observedDiffbot crawler for extracting structured web data and maintaining its knowledge graph.
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com)DuckAssistBot
Live fetcherReported, not operator verifiedDuckDuckGo AI assistant fetcher used by DuckAssist to retrieve content for real-time answers.
DuckAssistBotExaSearchBot
AI search crawlerOfficially documentedCrawler for Exa's search index, which AI products and agents query to retrieve source pages.
Mozilla/5.0 (compatible; ExaSearchBot/1.0; +https://crawler.exa.ai/)FacebookBot
Social crawlerReported, not operator verifiedMeta crawler historically documented for Facebook crawling and AI-related training uses.
FacebookBotfacebookexternalhit
Social crawlerOfficially documentedMeta link preview crawler used when content is shared on Meta family apps.
facebookexternalhitFirecrawlAgent
Other crawlerReported, not operator verifiedFirecrawl agent token for AI scraping and web-to-LLM data extraction workflows.
FirecrawlAgentGemini-Deep-Research
Live fetcherReported, not operator verifiedGemini Deep Research agent token reported for collecting and scanning resources used in research answers.
Gemini-Deep-ResearchGoogle-Agent
Live fetcherOfficially documentedGoogle user-triggered fetcher used by agents that act on a person's request, including Project Mariner.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Safari/537.36; compatible; Google-AgentGoogle-CloudVertexBot
Other crawlerOfficially documentedGoogle crawler used for site-owner-requested crawls related to Vertex AI Agents.
Google-CloudVertexBotGoogle-Extended
Training crawlerOfficially documentedRobots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.
Google-Extended control tokenGoogle-Firebase
Other crawlerReported, not operator verifiedGoogle Firebase AI product token reported for app-related fetches.
Google-FirebaseGoogle-Gemini-CLI
Live fetcherReported, not operator verifiedGemini CLI related token listed in AI crawler registries for coding-agent activity.
Google-Gemini-CLIGoogle-GeminiNotebook
Live fetcherOfficially documentedGoogle fetcher that reads a URL because someone added it as a source inside Gemini or NotebookLM.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Safari/537.36; compatible; Google-GeminiNotebookGoogle-NotebookLM
Live fetcherOfficially documentedSuperseded name for the Google fetcher that reads a URL when someone adds it as a source in Gemini or NotebookLM.
Google-NotebookLMGoogleAgent-Mariner
Live fetcherReported, not operator verifiedGoogle AI agent token associated with browser-style task execution.
GoogleAgent-MarinerGooglebot
AI search crawlerOfficially documentedGoogle Search crawler used to discover, crawl, render, and index pages for Google Search.
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)GoogleOther
Other crawlerOfficially documentedGoogle generic crawler used by product teams for publicly accessible content fetches outside core Googlebot.
GoogleOtherGoogleOther-Image
Other crawlerOfficially documentedGoogle product-specific image crawler token for publicly accessible content fetches.
GoogleOther-ImageGoogleOther-Video
Other crawlerOfficially documentedGoogle product-specific video crawler token for public content fetches.
GoogleOther-VideoGPTBot
Training crawlerOfficially documentedSample observedOpenAI crawler for content that may be used to improve generative AI foundation models.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbotGrokBot
Training crawlerReported, not operator verifiedxAI crawler token listed by public crawler directories for Grok-related crawling.
GrokBotiaskspider/2.0
AI search crawlerReported, not operator verifiediAsk crawler used to provide answers to user queries.
iaskspider/2.0IbouBot
AI search crawlerReported, not operator verifiedIbou crawler for building a graph representation of the web used in search.
IbouBotICC-Crawler
Training crawlerReported, not operator verifiedNICT crawler for data used in artificial intelligence technologies and third-party research/commercial uses.
ICC-CrawlerImagesiftBot
Other crawlerOfficially documentedImageSift crawler for public image and page data used in web intelligence products.
ImagesiftBotimg2dataset
Training crawlerOfficially documentedOpen-source image dataset downloader token used to collect images for machine learning datasets.
img2datasetKagibot
AI search crawlerOfficially documentedKagi search crawler that indexes pages for its search index and the answers built on it.
Mozilla/5.0 (compatible; Kagibot/1.0; +https://kagi.com/bot)Kimi-SearchBot
AI search crawlerOfficially documentedMoonshot AI search crawler that indexes pages so they can be retrieved and cited in Kimi answers.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-SearchBot/1.0; +https://www.kimi.com/policies/kimi-crawlersKimi-User
Live fetcherOfficially documentedMoonshot AI fetcher that retrieves a specific page because a person asked Kimi about it.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-User/1.0; +https://www.kimi.com/policies/kimi-crawlersKimiBot
Training crawlerOfficially documentedMoonshot AI crawler for content that may be used to improve its Kimi models.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; KimiBot/1.0; +https://www.kimi.com/policies/kimi-crawlersKlaviyoAIBot
AI search crawlerOfficially documentedKlaviyo AI crawler for indexing configured content to tailor AI experiences and recommendations.
KlaviyoAIBotLAIONDownloader
Training crawlerOfficially documentedLAION downloader token used in machine learning research dataset collection.
LAIONDownloadermeta-externalads
Other crawlerOfficially documentedMeta crawler documented alongside its other external agents, used for advertising related page reads.
meta-externalads/1.1Meta-ExternalAgent
Training crawlerOfficially documentedSample observedMeta crawler for indexing content directly for AI model training and product improvement use cases.
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)Meta-ExternalFetcher
Live fetcherOfficially documentedSample observedMeta user-requested fetcher for AI and link features across Meta products.
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)Meta-WebIndexer
AI search crawlerOfficially documentedMeta crawler for improving Meta AI search result quality and source linking.
meta-webindexerMistralAI-Index
AI search crawlerOfficially documentedMistral automated crawler for indexing content used by Mistral AI search in Vibe.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)MistralAI-Training
Training crawlerOfficially documentedMistral crawler for content that may be used to train its models, kept separate from its search index crawler.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)MistralAI-User
Live fetcherOfficially documentedSample observedMistral user-action fetcher for Vibe responses that need a source page.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)NovaAct
Live fetcherReported, not operator verifiedAmazon Nova Act agent token reported for browser-style task execution.
NovaActOAI-AdsBot
Other crawlerOfficially documentedOpenAI crawler that reviews the safety and relevance of pages submitted as ChatGPT ad landing pages.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbotOAI-SearchBot
AI search crawlerOfficially documentedSample observedOpenAI search crawler for indexing pages that can appear in ChatGPT search results.
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbotomgili
Other crawlerOfficially documentedWebz.io crawler for collecting web data sold through APIs and datasets.
omgiliomgilibot
Other crawlerReported, not operator verifiedLegacy Omgili/Webz.io crawler token for web data collection.
omgilibotPanguBot
Training crawlerReported, not operator verifiedHuawei crawler token reported for training data collection for the PanGu multimodal LLM.
PanguBotPanscient
Other crawlerOfficially documentedPanscient crawler for collecting and structuring business data with AI and machine learning.
PanscientPerplexity-User
Live fetcherOfficially documentedSample observedUser-triggered Perplexity fetcher for pages needed to answer a specific question.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)PerplexityBot
AI search crawlerOfficially documentedSample observedPerplexity crawler for surfacing and linking websites in Perplexity search results.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)PetalBot
AI search crawlerReported, not operator verifiedHuawei crawler used for recommendations, assistant features, and AI search services.
PetalBotPhindBot
AI search crawlerReported, not operator verifiedPhind crawler token associated with AI-enhanced developer search.
PhindBotSBIntuitionsBot
Training crawlerOfficially documentedSB Intuitions crawler for data used in AI development and information analysis.
SBIntuitionsBotScrapy
Other crawlerReported, not operator verifiedScrapy framework user-agent commonly used for web scraping, including AI and machine learning data extraction.
ScrapySemrushBot-OCOB
SEO tool crawlerOfficially documentedSemrush crawler for the ContentShake AI content tool.
SemrushBot-OCOBSemrushBot-SWA
SEO tool crawlerOfficially documentedSemrush crawler for SEO Writing Assistant URL checks.
SemrushBot-SWAShap-User
Live fetcherOfficially documentedParallel user-triggered fetcher that reads content at a user's direction rather than crawling the web automatically.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Shap-User/0.1.0ShapBot
AI search crawlerOfficially documentedParallel crawler for discovering and indexing websites for Parallel web APIs.
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ShapBot/0.1.0TerraCotta
Training crawlerOfficially documentedCeramic AI crawler token for downloading data used to train LLMs.
TerraCottaTikTokSpider
Social crawlerReported, not operator verifiedByteDance/TikTok crawler token reported alongside AI crawling lists.
TikTokSpiderTimpibot
AI search crawlerReported, not operator verifiedTimpi crawler reported for scraping data used in search and AI model training contexts.
TimpibotVelenPublicWebCrawler
Training crawlerOfficially documentedVelen crawler for business datasets and machine learning models.
VelenPublicWebCrawlerWebzio-Extended
Training crawlerOfficially documentedWebz.io token covering whether crawled content may be included in the datasets it resells for AI and machine learning use.
Webzio-ExtendedxAI-SearchBot
AI search crawlerReported, not operator verifiedCrawler token observed retrieving pages for xAI's Grok search and answer features.
Mozilla/5.0 (compatible; xAI-SearchBot/1.0; +https://x.ai)YandexAdditional
AI search crawlerOfficially documentedYandex crawler token for data used in YandexGPT quick answers and additional analysis.
YandexAdditionalYandexAdditionalBot
AI search crawlerOfficially documentedYandex additional crawler token for YandexGPT-related answer and analysis features.
YandexAdditionalBotYouBot
AI search crawlerReported, not operator verifiedSample observedReported You.com crawler string associated with web search and AI answer retrieval.
Mozilla/5.0 (compatible; YouBot (+http://www.you.com))Do not confuse the families
| Name | Owner | Purpose | HTTP crawler? | Evidence |
|---|---|---|---|---|
| GPTBot | OpenAI | Training crawler | Yes | Officially documented |
| OAI-SearchBot | OpenAI | AI search crawler | Yes | Officially documented |
| ChatGPT-User | OpenAI | Live fetcher | Yes | Officially documented |
| ClaudeBot | Anthropic | Training crawler | Yes | Officially documented |
| Claude-SearchBot | Anthropic | AI search crawler | Yes | Officially documented |
| Claude-User | Anthropic | Live fetcher | Yes | Officially documented |
| Google-Extended | Training crawler | No, control token | Officially documented | |
| Applebot-Extended | Apple | Training crawler | No, control token | Officially documented |
Recent verified changes
Methodology
Official
The operator publishes the name, purpose, user-agent or control token. Every verified field links to that source and has a checked date.
Observed
Trakkr counted matching request signatures in a stated time window and site sample. User-agents can be spoofed, so observation is not operator verification.
Reported, not inferred
A public crawler list or historical record names the string, but current operator documentation is missing. We do not infer ownership from the name; treat it as a detection clue.
See what visits your own site
Paste or upload a server log to find known AI crawler requests, status codes, and pages. The analysis stays in your browser.
Analyze a log locally14-day free trial · Cancel anytime