AI Crawler Observatory

Know which AI crawlers are real, what they do, and how to handle them

A practical, source-checked reference that separates operator documentation from reported user-agents and bounded Trakkr observations.

94 crawler records
60 officially documented
17 with observation samples
[01]

What these crawlers do

Training crawler

Collects or controls access to content that can feed AI training data.

AI search crawler

Indexes pages so AI search products can retrieve, rank, cite, or summarize them.

Live fetcher

Fetches pages because a user or agent asked for a specific URL or task.

A safer allow or block decision

  1. 1. Start with purpose. Search, training, and user-triggered retrieval have different costs when blocked.
  2. 2. Use a documented token. Target the narrowest official robots.txt token. If none is published, do not invent one.
  3. 3. Verify traffic separately. A robots rule sets policy; IP, reverse DNS, or signatures help identify the sender.
[02]

Observed across connected sites

Trakkr sample

A real, bounded view of crawler signatures across connected sites. It is evidence of matching requests, not proof of operator origin and not a market-share ranking.

31

completed days

85

connected sites

7,963,265

classified requests

2026-07-18

through 2026-08-17

[03]

Find a crawler

Showing 94 records

AI2Bot

Training crawlerOfficially documented

Allen Institute for AI crawler used to find web content for open language model datasets.

Allen Institute for AIActiveHonors robots.txt
AI2Bot

Ai2Bot-Dolma

Training crawlerOfficially documented

AI2 crawler token associated with Dolma/open language model dataset collection.

Allen Institute for AIActiveHonors robots.txt
Ai2Bot-Dolma

amazon-kendra

AI search crawlerReported, not operator verified

Amazon Kendra crawler token for intelligent enterprise search over configured content sources.

AmazonStatus uncertainHonors robots.txt
amazon-kendra

Amazonbot

Other crawlerOfficially documentedSample observed

Amazon crawler used to improve its products and services and potentially train Amazon AI models.

AmazonActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amazonbot/0.1) Chrome/W.X.Y.Z Safari/537.36

AmazonBuyForMe

Live fetcherReported, not operator verified

Amazon agent token reported for Buy for Me shopping actions directed by customers.

AmazonStatus uncertainCompliance unverified
AmazonBuyForMe

Amzn-SearchBot

AI search crawlerOfficially documented

Amazon search crawler that indexes pages so they can be retrieved and cited in Amazon search and assistant answers.

AmazonActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-SearchBot/0.1) Chrome/W.X.Y.Z Safari/537.36

Amzn-User

Live fetcherOfficially documented

Amazon fetcher that retrieves a specific page because a person asked an Amazon assistant about it.

AmazonActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Amzn-User/0.1) Chrome/W.X.Y.Z Safari/537.36

anthropic-ai

Training crawlerReported, not operator verified

Legacy Anthropic robots.txt token that predates the current ClaudeBot, Claude-User and Claude-SearchBot names.

AnthropicStatus uncertainCompliance unverified
anthropic-ai

ApifyBot

Other crawlerReported, not operator verified

Token associated with crawlers run on the Apify scraping platform by its customers.

ApifyStatus uncertainCompliance unverified
ApifyBot

Applebot

AI search crawlerOfficially documentedSample observed

Apple crawler for search experiences across Spotlight, Siri, Safari, and related Apple surfaces.

AppleActiveHonors robots.txt
Mozilla/5.0 (Device; OS_version) AppleWebKit/WebKit_version (KHTML, like Gecko) Version/Safari_version Safari/WebKit_version (Applebot/Applebot_version; +http://www.apple.com/go/applebot)

Applebot-Extended

Training crawlerOfficially documentedSample observed

Robots.txt control token for whether Applebot-crawled content may be used to train Apple foundation models.

AppleActiveHonors robots.txt
Applebot-Extended control token

atlassian-bot

AI search crawlerOfficially documented

Atlassian Rovo crawler used to index connected website content for AI search, assistants, and agents.

AtlassianActiveHonors robots.txt
atlassian-bot

bedrockbot

Other crawlerOfficially documented

Amazon Bedrock web crawler connector token for customer-configured AI applications.

AmazonActiveHonors robots.txt
bedrockbot

Bingbot

AI search crawlerReported, not operator verified

Microsoft Bing crawler used to crawl and index pages for Bing and Microsoft search-powered experiences.

MicrosoftStatus uncertainHonors robots.txt
bingbot

Bravebot

AI search crawlerReported, not operator verified

Reported name for Brave Search crawling. Brave says its crawler does not advertise a differentiated user-agent.

BraveStatus uncertainPartial or user-triggered
No official user-agent published

Brightbot

Other crawlerOfficially documented

Bright Data's declared crawler for collecting public web data for its own datasets.

Bright DataActiveCompliance unverified
Brightbot 1.0

Bytespider

Training crawlerReported, not operator verifiedSample observed

ByteDance crawler associated with training and powering AI products.

ByteDanceStatus uncertainPartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Bytespider; spider-feedback@bytedance.com

CCBot

Training crawlerOfficially documented

Common Crawl's crawler for building public web crawl datasets used by researchers and AI builders.

Common CrawlActiveHonors robots.txt
CCBot/2.0 (https://commoncrawl.org/faq/)

ChatGPT Agent

Live fetcherOfficially documented

OpenAI agent used when ChatGPT navigates websites for user-directed tasks.

OpenAIActiveCompliance unverified
ChatGPT Agent

ChatGPT-User

Live fetcherOfficially documentedSample observed

User-triggered OpenAI fetcher for ChatGPT and Custom GPT actions.

OpenAIActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot

Claude-Code

Live fetcherReported, not operator verified

Claude Code related agent token seen in crawler/user-agent lists.

AnthropicStatus uncertainCompliance unverified
Claude-Code

Claude-SearchBot

AI search crawlerOfficially documentedSample observed

Anthropic search crawler that indexes content to improve Claude search result relevance and accuracy.

AnthropicActiveHonors robots.txt
Mozilla/5.0 (compatible; Claude-SearchBot/1.0; +claudebot@anthropic.com)

Claude-User

Live fetcherOfficially documented

Anthropic user-triggered fetcher for Claude answers that need a specific web page.

AnthropicActiveHonors robots.txt
Claude-User

Claude-Web

Live fetcherReported, not operator verified

Reported Anthropic-related token seen in public crawler registries and Trakkr detection, but absent from Anthropic's current bot documentation.

AnthropicStatus uncertainCompliance unverified
Claude-Web

ClaudeBot

Training crawlerOfficially documentedSample observed

Anthropic crawler for public web content that could contribute to Claude model training.

AnthropicActiveHonors robots.txt
Mozilla/5.0 (compatible; ClaudeBot/1.0; +claudebot@anthropic.com)

Cloudflare-AutoRAG

AI search crawlerOfficially documented

Cloudflare AutoRAG crawler used to index configured content for AI search applications.

CloudflareActiveHonors robots.txt
Cloudflare-AutoRAG

cohere-ai

Live fetcherReported, not operator verifiedSample observed

Cohere token reported for retrieving data in response to user-initiated prompts.

CohereStatus uncertainHonors robots.txt
cohere-ai

cohere-training-data-crawler

Training crawlerReported, not operator verified

Cohere training-data crawler token reported for downloading web data for enterprise language models.

CohereStatus uncertainCompliance unverified
cohere-training-data-crawler

DeepSeekBot

Training crawlerReported, not operator verified

DeepSeek crawler token reported for training language models and improving AI products.

DeepSeekStatus uncertainPartial or user-triggered
DeepSeekBot

Diffbot

Other crawlerReported, not operator verifiedSample observed

Diffbot crawler for extracting structured web data and maintaining its knowledge graph.

DiffbotStatus uncertainHonors robots.txt
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com)

DuckAssistBot

Live fetcherReported, not operator verified

DuckDuckGo AI assistant fetcher used by DuckAssist to retrieve content for real-time answers.

DuckDuckGoStatus uncertainCompliance unverified
DuckAssistBot

ExaSearchBot

AI search crawlerOfficially documented

Crawler for Exa's search index, which AI products and agents query to retrieve source pages.

ExaActiveHonors robots.txt
Mozilla/5.0 (compatible; ExaSearchBot/1.0; +https://crawler.exa.ai/)

FacebookBot

Social crawlerReported, not operator verified

Meta crawler historically documented for Facebook crawling and AI-related training uses.

MetaStatus uncertainHonors robots.txt
FacebookBot

facebookexternalhit

Social crawlerOfficially documented

Meta link preview crawler used when content is shared on Meta family apps.

MetaActivePartial or user-triggered
facebookexternalhit

FirecrawlAgent

Other crawlerReported, not operator verified

Firecrawl agent token for AI scraping and web-to-LLM data extraction workflows.

FirecrawlStatus uncertainHonors robots.txt
FirecrawlAgent

Gemini-Deep-Research

Live fetcherReported, not operator verified

Gemini Deep Research agent token reported for collecting and scanning resources used in research answers.

GoogleStatus uncertainCompliance unverified
Gemini-Deep-Research

Google-Agent

Live fetcherOfficially documented

Google user-triggered fetcher used by agents that act on a person's request, including Project Mariner.

GoogleActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Safari/537.36; compatible; Google-Agent

Google-CloudVertexBot

Other crawlerOfficially documented

Google crawler used for site-owner-requested crawls related to Vertex AI Agents.

GoogleActiveHonors robots.txt
Google-CloudVertexBot

Google-Extended

Training crawlerOfficially documented

Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

GoogleActiveHonors robots.txt
Google-Extended control token

Google-Firebase

Other crawlerReported, not operator verified

Google Firebase AI product token reported for app-related fetches.

GoogleStatus uncertainCompliance unverified
Google-Firebase

Google-Gemini-CLI

Live fetcherReported, not operator verified

Gemini CLI related token listed in AI crawler registries for coding-agent activity.

GoogleStatus uncertainCompliance unverified
Google-Gemini-CLI

Google-GeminiNotebook

Live fetcherOfficially documented

Google fetcher that reads a URL because someone added it as a source inside Gemini or NotebookLM.

GoogleActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Safari/537.36; compatible; Google-GeminiNotebook

Google-NotebookLM

Live fetcherOfficially documented

Superseded name for the Google fetcher that reads a URL when someone adds it as a source in Gemini or NotebookLM.

GoogleDeprecatedPartial or user-triggered
Google-NotebookLM

GoogleAgent-Mariner

Live fetcherReported, not operator verified

Google AI agent token associated with browser-style task execution.

GoogleStatus uncertainCompliance unverified
GoogleAgent-Mariner

Googlebot

AI search crawlerOfficially documented

Google Search crawler used to discover, crawl, render, and index pages for Google Search.

GoogleActiveHonors robots.txt
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/W.X.Y.Z Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)

GoogleOther

Other crawlerOfficially documented

Google generic crawler used by product teams for publicly accessible content fetches outside core Googlebot.

GoogleActiveHonors robots.txt
GoogleOther

GoogleOther-Image

Other crawlerOfficially documented

Google product-specific image crawler token for publicly accessible content fetches.

GoogleActiveHonors robots.txt
GoogleOther-Image

GoogleOther-Video

Other crawlerOfficially documented

Google product-specific video crawler token for public content fetches.

GoogleActiveHonors robots.txt
GoogleOther-Video

GPTBot

Training crawlerOfficially documentedSample observed

OpenAI crawler for content that may be used to improve generative AI foundation models.

OpenAIActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot

GrokBot

Training crawlerReported, not operator verified

xAI crawler token listed by public crawler directories for Grok-related crawling.

xAIStatus uncertainCompliance unverified
GrokBot

iaskspider/2.0

AI search crawlerReported, not operator verified

iAsk crawler used to provide answers to user queries.

iAskStatus uncertainPartial or user-triggered
iaskspider/2.0

IbouBot

AI search crawlerReported, not operator verified

Ibou crawler for building a graph representation of the web used in search.

IbouStatus uncertainHonors robots.txt
IbouBot

ICC-Crawler

Training crawlerReported, not operator verified

NICT crawler for data used in artificial intelligence technologies and third-party research/commercial uses.

NICTStatus uncertainHonors robots.txt
ICC-Crawler

ImagesiftBot

Other crawlerOfficially documented

ImageSift crawler for public image and page data used in web intelligence products.

ImageSiftActiveHonors robots.txt
ImagesiftBot

img2dataset

Training crawlerOfficially documented

Open-source image dataset downloader token used to collect images for machine learning datasets.

img2datasetActiveCompliance unverified
img2dataset

Kagibot

AI search crawlerOfficially documented

Kagi search crawler that indexes pages for its search index and the answers built on it.

KagiActiveHonors robots.txt
Mozilla/5.0 (compatible; Kagibot/1.0; +https://kagi.com/bot)

Kimi-SearchBot

AI search crawlerOfficially documented

Moonshot AI search crawler that indexes pages so they can be retrieved and cited in Kimi answers.

Moonshot AIActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-SearchBot/1.0; +https://www.kimi.com/policies/kimi-crawlers

Kimi-User

Live fetcherOfficially documented

Moonshot AI fetcher that retrieves a specific page because a person asked Kimi about it.

Moonshot AIActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Kimi-User/1.0; +https://www.kimi.com/policies/kimi-crawlers

KimiBot

Training crawlerOfficially documented

Moonshot AI crawler for content that may be used to improve its Kimi models.

Moonshot AIActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; KimiBot/1.0; +https://www.kimi.com/policies/kimi-crawlers

KlaviyoAIBot

AI search crawlerOfficially documented

Klaviyo AI crawler for indexing configured content to tailor AI experiences and recommendations.

KlaviyoActiveHonors robots.txt
KlaviyoAIBot

LAIONDownloader

Training crawlerOfficially documented

LAION downloader token used in machine learning research dataset collection.

LAIONActivePartial or user-triggered
LAIONDownloader

meta-externalads

Other crawlerOfficially documented

Meta crawler documented alongside its other external agents, used for advertising related page reads.

MetaActiveHonors robots.txt
meta-externalads/1.1

Meta-ExternalAgent

Training crawlerOfficially documentedSample observed

Meta crawler for indexing content directly for AI model training and product improvement use cases.

MetaActiveHonors robots.txt
meta-externalagent/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)

Meta-ExternalFetcher

Live fetcherOfficially documentedSample observed

Meta user-requested fetcher for AI and link features across Meta products.

MetaActivePartial or user-triggered
meta-externalfetcher/1.1 (+https://developers.facebook.com/docs/sharing/webmasters/web-crawlers)

Meta-WebIndexer

AI search crawlerOfficially documented

Meta crawler for improving Meta AI search result quality and source linking.

MetaActiveCompliance unverified
meta-webindexer

MistralAI-Index

AI search crawlerOfficially documented

Mistral automated crawler for indexing content used by Mistral AI search in Vibe.

Mistral AIActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Index/1.0; +https://docs.mistral.ai/robots)

MistralAI-Training

Training crawlerOfficially documented

Mistral crawler for content that may be used to train its models, kept separate from its search index crawler.

Mistral AIActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-Training/1.0; +https://docs.mistral.ai/robots)

MistralAI-User

Live fetcherOfficially documentedSample observed

Mistral user-action fetcher for Vibe responses that need a source page.

Mistral AIActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)

NovaAct

Live fetcherReported, not operator verified

Amazon Nova Act agent token reported for browser-style task execution.

AmazonStatus uncertainCompliance unverified
NovaAct

OAI-AdsBot

Other crawlerOfficially documented

OpenAI crawler that reviews the safety and relevance of pages submitted as ChatGPT ad landing pages.

OpenAIActiveCompliance unverified
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-AdsBot/1.0; +https://openai.com/adsbot

OAI-SearchBot

AI search crawlerOfficially documentedSample observed

OpenAI search crawler for indexing pages that can appear in ChatGPT search results.

OpenAIActiveHonors robots.txt
Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot

omgili

Other crawlerOfficially documented

Webz.io crawler for collecting web data sold through APIs and datasets.

Webz.ioActiveHonors robots.txt
omgili

omgilibot

Other crawlerReported, not operator verified

Legacy Omgili/Webz.io crawler token for web data collection.

Webz.ioStatus uncertainHonors robots.txt
omgilibot

PanguBot

Training crawlerReported, not operator verified

Huawei crawler token reported for training data collection for the PanGu multimodal LLM.

HuaweiStatus uncertainCompliance unverified
PanguBot

Panscient

Other crawlerOfficially documented

Panscient crawler for collecting and structuring business data with AI and machine learning.

PanscientActiveHonors robots.txt
Panscient

Perplexity-User

Live fetcherOfficially documentedSample observed

User-triggered Perplexity fetcher for pages needed to answer a specific question.

PerplexityActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Perplexity-User/1.0; +https://perplexity.ai/perplexity-user)

PerplexityBot

AI search crawlerOfficially documentedSample observed

Perplexity crawler for surfacing and linking websites in Perplexity search results.

PerplexityActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)

PetalBot

AI search crawlerReported, not operator verified

Huawei crawler used for recommendations, assistant features, and AI search services.

HuaweiStatus uncertainHonors robots.txt
PetalBot

PhindBot

AI search crawlerReported, not operator verified

Phind crawler token associated with AI-enhanced developer search.

PhindStatus uncertainCompliance unverified
PhindBot

SBIntuitionsBot

Training crawlerOfficially documented

SB Intuitions crawler for data used in AI development and information analysis.

SB IntuitionsActiveHonors robots.txt
SBIntuitionsBot

Scrapy

Other crawlerReported, not operator verified

Scrapy framework user-agent commonly used for web scraping, including AI and machine learning data extraction.

ZyteStatus uncertainCompliance unverified
Scrapy

SemrushBot-OCOB

SEO tool crawlerOfficially documented

Semrush crawler for the ContentShake AI content tool.

SemrushActiveHonors robots.txt
SemrushBot-OCOB

SemrushBot-SWA

SEO tool crawlerOfficially documented

Semrush crawler for SEO Writing Assistant URL checks.

SemrushActiveHonors robots.txt
SemrushBot-SWA

Shap-User

Live fetcherOfficially documented

Parallel user-triggered fetcher that reads content at a user's direction rather than crawling the web automatically.

ParallelActivePartial or user-triggered
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Shap-User/0.1.0

ShapBot

AI search crawlerOfficially documented

Parallel crawler for discovering and indexing websites for Parallel web APIs.

ParallelActiveHonors robots.txt
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ShapBot/0.1.0

TerraCotta

Training crawlerOfficially documented

Ceramic AI crawler token for downloading data used to train LLMs.

Ceramic AIActiveHonors robots.txt
TerraCotta

TikTokSpider

Social crawlerReported, not operator verified

ByteDance/TikTok crawler token reported alongside AI crawling lists.

ByteDanceStatus uncertainCompliance unverified
TikTokSpider

Timpibot

AI search crawlerReported, not operator verified

Timpi crawler reported for scraping data used in search and AI model training contexts.

TimpiStatus uncertainCompliance unverified
Timpibot

VelenPublicWebCrawler

Training crawlerOfficially documented

Velen crawler for business datasets and machine learning models.

VelenActiveHonors robots.txt
VelenPublicWebCrawler

Webzio-Extended

Training crawlerOfficially documented

Webz.io token covering whether crawled content may be included in the datasets it resells for AI and machine learning use.

Webz.ioActiveHonors robots.txt
Webzio-Extended

xAI-SearchBot

AI search crawlerReported, not operator verified

Crawler token observed retrieving pages for xAI's Grok search and answer features.

xAIStatus uncertainCompliance unverified
Mozilla/5.0 (compatible; xAI-SearchBot/1.0; +https://x.ai)

YandexAdditional

AI search crawlerOfficially documented

Yandex crawler token for data used in YandexGPT quick answers and additional analysis.

YandexActiveHonors robots.txt
YandexAdditional

YandexAdditionalBot

AI search crawlerOfficially documented

Yandex additional crawler token for YandexGPT-related answer and analysis features.

YandexActiveHonors robots.txt
YandexAdditionalBot

YouBot

AI search crawlerReported, not operator verifiedSample observed

Reported You.com crawler string associated with web search and AI answer retrieval.

You.comStatus uncertainCompliance unverified
Mozilla/5.0 (compatible; YouBot (+http://www.you.com))
Primary sources for verified facts
Observed and reported evidence labelled
User-agent spoofing warnings included
[04]

Do not confuse the families

Comparison of crawler families by owner, purpose, HTTP behavior, and evidence status
NameOwnerPurposeHTTP crawler?Evidence
GPTBotOpenAITraining crawlerYesOfficially documented
OAI-SearchBotOpenAIAI search crawlerYesOfficially documented
ChatGPT-UserOpenAILive fetcherYesOfficially documented
ClaudeBotAnthropicTraining crawlerYesOfficially documented
Claude-SearchBotAnthropicAI search crawlerYesOfficially documented
Claude-UserAnthropicLive fetcherYesOfficially documented
Google-ExtendedGoogleTraining crawlerNo, control tokenOfficially documented
Applebot-ExtendedAppleTraining crawlerNo, control tokenOfficially documented
[05]

Recent verified changes

[06]

Methodology

Official

The operator publishes the name, purpose, user-agent or control token. Every verified field links to that source and has a checked date.

Observed

Trakkr counted matching request signatures in a stated time window and site sample. User-agents can be spoofed, so observation is not operator verification.

Reported, not inferred

A public crawler list or historical record names the string, but current operator documentation is missing. We do not infer ownership from the name; treat it as a detection clue.

This is a connected-site sample, not a representative sample of the web. A matching user-agent or signature does not prove that the named operator sent the request. Counts describe classified request signatures, not market share or unique pages crawled.

See what visits your own site

Paste or upload a server log to find known AI crawler requests, status codes, and pages. The analysis stays in your browser.

Analyze a log locally

14-day free trial · Cancel anytime