Study 003

When AI comes to your website

A bounded analysis of crawler-label requests in connected-site logs. What did the sample record, and what can it not prove?

544K+
crawler visits analyzed
6
AI crawlers tracked
68%
requests with OpenAI labels
83.8%
of pages visited exactly once
Last updated · Feb 1, 2026

AI crawler monitoring starts with a request, not an assumption. Some documented bots may collect data for model improvement. Others build search indexes or fetch a page after a user action. Those purposes use different tokens and should not be combined into one policy or one performance claim.

This study analyzed 575,788+ crawler visits across 7 crawler labels. The results describe this panel and window. They do not represent the whole web, prove operator identity, or show that a fetch led to indexing, ranking, citation, or referral traffic.

Key findings

OpenAI labels led this sample. GPTBot and OAI-SearchBot together accounted for roughly 72% of matched requests in the study. The denominator is this connected-site panel, not global crawler traffic.

Homepage request rates differed by label. About 3% of GPTBot-labeled requests and 19% of ClaudeBot-labeled requests were for homepages. This does not by itself explain discovery method or intent.

88.5% of pages had one recorded request. The figure applies to pages with matching requests during the study window. It does not establish frequency outside the window or at request layers Trakkr could not observe.

Blog pages appeared as entry points. In the linked session sample, 21% of ChatGPT Search-labeled sessions began on blog pages. That is an observed landing-page mix, not evidence that AI systems prefer or cite long-form content.

Recorded requests skewed toward shallower URLs. More than half of matching requests landed on pages measured within three site-depth steps. URL structure, templates, popularity and customer mix may all contribute, so the pattern does not prove a ranking advantage.

Methodology

We analyzed server-side request data from websites using Trakkr crawler monitoring. The dataset spans 2025-06-11 to 2026-02-01 and covers matching labels including GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, PerplexityBot, Bytespider, and others. Google-Extended is not treated as a crawler because Google documents it as a robots.txt product token rather than a separate HTTP user-agent. Data was aggregated before publication. User-agent strings can be spoofed, collection layers can sample or miss requests, and the panel reflects participating sites. Request counts were not joined causally to citations or human traffic.

[01]

Purpose classes

Training Crawlers
Possible model improvement

Operators document these crawlers as collecting data that may be used to improve models. A request does not prove that a page entered a dataset or changed a later answer.

ChatGPT TrainingClaude Training
Sample request share
61%
Purpose
Training
Search Crawlers
Search retrieval index

Operators document these crawlers as supporting AI search discovery. A successful fetch does not guarantee indexing, ranking, retrieval or citation.

ChatGPT Search
Sample request share
15%
Purpose
Search

The distinction shapes access policy and measurement. GPTBot and OAI-SearchBot use separate controls. Keep request evidence separate from later citation and referral observations.

[02]

Observed request share

OpenAI labels led this sample

GPTBot and OAI-SearchBot labels represented about 72% of matched requests in this panel and window. ClaudeBot represented 3.8%.

Bytespider represented 9.2% of matched requests. Provider purpose and identity cannot be inferred from a user-agent label alone.

Training vs Search
3.8 : 1
OpenAI's crawl ratio
Top 2 Combined
72.3%
Both OpenAI bots
This is request share within participating sites, not global market share. The site mix, install layer, caching, WAF rules and sampling can all shape the result.
Matched request share by label
1
ChatGPT Training
329,57257.2%
2
ChatGPT Search
87,15515.1%
3
ByteSpider
52,7049.2%
4
Meta
45,4457.9%
5
Amazon
38,3356.7%
6
Claude Training
22,0743.8%
Ranked by traffic shareJune 2025 - Feb 2026
[03]

Homepage request rates

Different training philosophies

Claude visits homepages 7x more often than ChatGPT Training. It wants to understand who you are. ChatGPT's training crawler skips straight to your content.

Claude Training
19.2%
Identity focus - wants to understand who you are
ChatGPT Training
2.8%
Lower homepage share in this request sample

The homepage rates differ in this sample. Treat that as a hypothesis about request paths, not proof of crawler strategy or a content recommendation.

[04]

Request timing

Weekend request ratios differed by label

GPTBot and OAI-SearchBot labels had higher weekend request averages in this sample. ClaudeBot had an 8% lower weekend average. The logs do not reveal why the schedules differed.

Crawler
Weekday avg
Weekend avg
Change
Pattern
ChatGPT Training
1,430
1,841
+29%
More on weekends
ChatGPT Search
383
540
+41%
More on weekends
Claude Training
99
91
-8%
Less on weekends

Weekday publishes may get crawled faster by Anthropic. Weekend publishes may be picked up faster by OpenAI.

[05]

Session entry pages

Blog pages appeared as entry points

In the linked session sample, 21% of OAI-SearchBot-labeled entry pages were blogs. The request data does not reveal a user's question or prove that the page appeared in an answer.

21×higher observed blog-entry share than homepage-entry share in this sample. This is a descriptive ratio, not a ranking preference.

Page type, internal links, URL inventory, demand and customer mix may all contribute. Compare the same page types with separate citation and referral evidence before drawing a product conclusion.

Use the pattern to form a test: check whether important blog and non-blog pages are accessible, requested, cited and visited in the same window.
ChatGPT SearchEntry Point Distribution
Blog PagesArticles, guides, insights
21%
Product PagesFeatures, pricing, docs
3%
HomepageMain landing page
1%
Other PagesAll other entry points
75%
% of sessions starting on each page type
[06]

Revisit distribution

Revisit frequency
1
88.5%
2
8.3%
3-5
2.4%
6-10
0.4%
10+
0.3%
visits per unique URL
One-time visits
88.5%of URLs
one recorded request in window
Practical ceiling
5visits
even outliers (P99)

In this bounded window, 2.4% of URLs had at least three recorded requests. Activity outside the window or at an unobserved layer is not measured.

[07]

URL depth distribution

ChatGPT Training Depth Distribution
/2.7%
/about10.3%
/blog/post19.6%
/blog/2024/post51.7%
/docs/api/auth12.0%
/docs/api/v1/...3.7%
% of crawled URLs at each depth level

Requests skewed toward mid-depth URLs

In this sample, GPTBot-labeled requests appeared most often on mid-depth pages and less than 3% were for homepages. The data does not prove how the URLs were discovered or why deeper URLs had fewer requests.

Observed GPTBot URL-depth mix
/ Homepage (3%)
/blog (10%)
/blog/2024 (20%)
/blog/2024/post (52%)
/docs (15%)

Compare important deep pages with access, internal-link, citation and referral evidence before changing site structure.

[08]

Reach and request depth

Coverage vs Depth
Avg Visits / Site
03K6K
40%
65%
90%
ChatGPT Search
ChatGPT Training
Claude Training
Site Coverage %

Reach vs depth

ChatGPT Search prioritizes breadth - visiting 76% of sites in our dataset. ChatGPT Training prioritizes depth - fewer sites but 5,586 visits per site on average. Claude is the most selective at just 470 visits each.

ChatGPT Search
Wide reach, moderate depth
76%
1,362 visits
ChatGPT Training
Fewer sites, exhaustive crawl
70%
5,586 visits
Claude Training
Selective, targeted visits
56%
470 visits

OAI-SearchBot labels appeared on 76% of sites in this panel and GPTBot labels appeared on 70%. Reach describes matching requests, not search eligibility, training depth or citation likelihood.

[09]

Questions to investigate

Compare entry-page mixin the linked ChatGPT Search session sample. Check whether your own referrals show the same pattern.
21%
blog sessions
Use a complete time windowinside this study window. Avoid turning a bounded frequency into a permanent crawler rule.
88.5%
one-request pages
Audit important deep pagesof matched requests were at measured depth three or less. Compare access, internal links, content type and demand.
52%
at depth ≤3
Separate purpose classesin this sample combined a training crawler and a search crawler. Keep their policies and outcomes separate.
72%
OpenAI labels
Treat label differences as hypothesesbetween two labels in this sample. Test on your own request data before changing site structure.
7x
homepage-rate ratio

Measure your own crawler to outcome funnel

See how Trakkr compares page-level request evidence with separately observed citations, AI referrals, access status and next actions.