What is Bytespider? AI crawler guide
Bytespider is a crawler or fetcher associated with ByteDance. See source-graded user-agent, robots.txt, verification, and observation guidance.
What is Bytespider?
Bytespider is a reported crawler or fetcher string associated with ByteDance, but current operator evidence is incomplete. ByteDance crawler associated with training and powering AI products.
Evidence status
| Field | Value |
|---|---|
| Evidence | Reported, not operator verified |
| Lifecycle | Status uncertain |
| Purpose | training |
| robots.txt posture | partial |
| Source checked | 2026-06-11 |
Reported user-agent
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Bytespider; spider-feedback@bytedance.com
Bytespider collects pages for model training. Training inclusion is not the same as being cited, so measure where AI answers actually cite your site before drawing conclusions from crawl logs.
robots.txt allow example
User-agent: Bytespider Allow: /
robots.txt block example
User-agent: Bytespider Disallow: /
What the rule can and cannot do
Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.
How to verify a request
- Match the complete documented user-agent where one exists. A match is only a clue.
- Look for operator-published IP ranges, reverse DNS, or signed-request guidance. None is attached to this record.
- Keep the observed request, operator documentation, and any inference as separate fields.
Observed in Trakkr connected-site data
| Measure | Value |
|---|---|
| Matching signature | Bytespider |
| Classified requests | 1282320 |
| Sites in sample | 85 |
| Window | 2026-07-18 through 2026-08-17 |
| Method | Finalized daily crawler summaries from connected Trakkr sites, grouped by the crawler signature detected in each request. |
| Limits | This is a connected-site sample, not a representative sample of the web. A matching user-agent or signature does not prove that the named operator sent the request. Counts describe classified request signatures, not market share or unique pages crawled. |
JavaScript behavior
The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.
Related bots
- GPTBot: Also tracked as a training crawler.
- KimiBot: Also tracked as a training crawler.
- MistralAI-Training: Also tracked as a training crawler.
- DeepSeekBot: Also tracked as a training crawler.
- ClaudeBot: Also tracked as a training crawler.
- Meta-ExternalAgent: Also tracked as a training crawler.
- LAIONDownloader: Also tracked as a training crawler.
- cohere-training-data-crawler: Also tracked as a training crawler.
- PanguBot: Also tracked as a training crawler.
- AI Crawlers: Bytespider is a concrete crawler example for this concept.
- AI Training Opt-Out: Bytespider is a training crawler tied to this policy decision.
- TikTokSpider: Also operated by ByteDance.
Frequently Asked Questions
What is Bytespider?
Bytespider is a reported, not operator verified crawler or fetcher record associated with ByteDance.
What user-agent does Bytespider use?
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; Bytespider; spider-feedback@bytedance.com. Check the evidence status before attributing a matching request.
Can Bytespider be blocked in robots.txt?
The record identifies Bytespider as the token to review. Robots.txt is a request policy, not proof of model use or non-use.
How can I verify a Bytespider request?
Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.
Data & Sources
- Bytespider source reference - Source used to verify Bytespider.