When AI comes to your website
A bounded analysis of crawler-label requests in connected-site logs. What did the sample record, and what can it not prove?
AI crawler monitoring starts with a request, not an assumption. Some documented bots may collect data for model improvement. Others build search indexes or fetch a page after a user action. Those purposes use different tokens and should not be combined into one policy or one performance claim.
This study analyzed 575,788+ crawler visits across 7 crawler labels. The results describe this panel and window. They do not represent the whole web, prove operator identity, or show that a fetch led to indexing, ranking, citation, or referral traffic.
Key findings
OpenAI labels led this sample. GPTBot and OAI-SearchBot together accounted for roughly 72% of matched requests in the study. The denominator is this connected-site panel, not global crawler traffic.
Homepage request rates differed by label. About 3% of GPTBot-labeled requests and 19% of ClaudeBot-labeled requests were for homepages. This does not by itself explain discovery method or intent.
88.5% of pages had one recorded request. The figure applies to pages with matching requests during the study window. It does not establish frequency outside the window or at request layers Trakkr could not observe.
Blog pages appeared as entry points. In the linked session sample, 21% of ChatGPT Search-labeled sessions began on blog pages. That is an observed landing-page mix, not evidence that AI systems prefer or cite long-form content.
Recorded requests skewed toward shallower URLs. More than half of matching requests landed on pages measured within three site-depth steps. URL structure, templates, popularity and customer mix may all contribute, so the pattern does not prove a ranking advantage.
Methodology
We analyzed server-side request data from websites using Trakkr crawler monitoring. The dataset spans 2025-06-11 to 2026-02-01 and covers matching labels including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Bytespider, and others. Google-Extended is not treated as a crawler because Google documents it as a robots.txt product token rather than a separate HTTP user-agent. Data was aggregated before publication. User-agent strings can be spoofed, collection layers can sample or miss requests, and the panel reflects participating sites. Request counts were not joined causally to citations or human traffic.
Purpose classes
Operators document these crawlers as collecting data that may be used to improve models. A request does not prove that a page entered a dataset or changed a later answer.
Claude TrainingOperators document these crawlers as supporting AI search discovery. A successful fetch does not guarantee indexing, ranking, retrieval or citation.
The distinction shapes access policy and measurement. GPTBot and OAI-SearchBot use separate controls. Keep request evidence separate from later citation and referral observations.
Observed request share
OpenAI labels led this sample
GPTBot and OAI-SearchBot labels represented about 72% of matched requests in this panel and window. ClaudeBot represented 3.8%.
Bytespider represented 9.2% of matched requests. Provider purpose and identity cannot be inferred from a user-agent label alone.
Amazon
Claude TrainingHomepage request rates
Different training philosophies
Claude visits homepages 7x more often than ChatGPT Training. It wants to understand who you are. ChatGPT's training crawler skips straight to your content.
Claude TrainingThe homepage rates differ in this sample. Treat that as a hypothesis about request paths, not proof of crawler strategy or a content recommendation.
Request timing
Weekend request ratios differed by label
GPTBot and OAI-SearchBot labels had higher weekend request averages in this sample. ClaudeBot had an 8% lower weekend average. The logs do not reveal why the schedules differed.
Claude TrainingWeekday publishes may get crawled faster by Anthropic. Weekend publishes may be picked up faster by OpenAI.
Session entry pages
Blog pages appeared as entry points
In the linked session sample, 21% of OAI-SearchBot-labeled entry pages were blogs. The request data does not reveal a user's question or prove that the page appeared in an answer.
Page type, internal links, URL inventory, demand and customer mix may all contribute. Compare the same page types with separate citation and referral evidence before drawing a product conclusion.
Revisit distribution
In this bounded window, 2.4% of URLs had at least three recorded requests. Activity outside the window or at an unobserved layer is not measured.
URL depth distribution
/2.7%/about10.3%/blog/post19.6%/blog/2024/post51.7%/docs/api/auth12.0%/docs/api/v1/...3.7%Requests skewed toward mid-depth URLs
In this sample, GPTBot-labeled requests appeared most often on mid-depth pages and less than 3% were for homepages. The data does not prove how the URLs were discovered or why deeper URLs had fewer requests.
Compare important deep pages with access, internal-link, citation and referral evidence before changing site structure.
Reach and request depth
Reach vs depth
ChatGPT Search prioritizes breadth - visiting 76% of sites in our dataset. ChatGPT Training prioritizes depth - fewer sites but 5,586 visits per site on average. Claude is the most selective at just 470 visits each.
OAI-SearchBot labels appeared on 76% of sites in this panel and GPTBot labels appeared on 70%. Reach describes matching requests, not search eligibility, training depth or citation likelihood.
Questions to investigate
Measure your own crawler to outcome funnel
See how Trakkr compares page-level request evidence with separately observed citations, AI referrals, access status and next actions.
Continue with this study
See how your brand performs in AI search
