What is Diffbot? AI crawler guide
Diffbot is a crawler or fetcher associated with Diffbot. See source-graded user-agent, robots.txt, verification, and observation guidance.
What is Diffbot?
Diffbot is a reported crawler or fetcher string associated with Diffbot, but current operator evidence is incomplete. Diffbot crawler for extracting structured web data and maintaining its knowledge graph.
Evidence status
| Field | Value |
|---|---|
| Evidence | Reported, not operator verified |
| Lifecycle | Status uncertain |
| Purpose | other |
| robots.txt posture | honors |
| Source checked | 2026-06-11 |
Reported user-agent
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com)
Allowing Diffbot only creates the possibility of retrieval. To find out whether your pages are selected, measure which answer engines actually cite your pages using a stable query set.
robots.txt allow example
User-agent: Diffbot Allow: /
robots.txt block example
User-agent: Diffbot Disallow: /
What the rule can and cannot do
Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.
How to verify a request
- Match the complete documented user-agent where one exists. A match is only a clue.
- Look for operator-published IP ranges, reverse DNS, or signed-request guidance. None is attached to this record.
- Keep the observed request, operator documentation, and any inference as separate fields.
Observed in Trakkr connected-site data
| Measure | Value |
|---|---|
| Matching signature | Diffbot |
| Classified requests | 94235 |
| Sites in sample | 85 |
| Window | 2026-07-18 through 2026-08-17 |
| Method | Finalized daily crawler summaries from connected Trakkr sites, grouped by the crawler signature detected in each request. |
| Limits | This is a connected-site sample, not a representative sample of the web. A matching user-agent or signature does not prove that the named operator sent the request. Counts describe classified request signatures, not market share or unique pages crawled. |
JavaScript behavior
The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.
Related bots
- Amazonbot: Also tracked as a general crawler.
- OAI-AdsBot: Also tracked as a general crawler.
- GoogleOther: Also tracked as a general crawler.
- GoogleOther-Image: Also tracked as a general crawler.
- GoogleOther-Video: Also tracked as a general crawler.
- ImagesiftBot: Also tracked as a general crawler.
- omgili: Also tracked as a general crawler.
- Google-Firebase: Also tracked as a general crawler.
- Brightbot: Also tracked as a general crawler.
- Robots.txt: Robots.txt is the control file used to allow or block Diffbot.
- Crawling: Diffbot is a concrete crawler example for this concept.
Frequently Asked Questions
What is Diffbot?
Diffbot is a reported, not operator verified crawler or fetcher record associated with Diffbot.
What user-agent does Diffbot use?
Mozilla/5.0 (compatible; Diffbot/0.1; +http://www.diffbot.com). Check the evidence status before attributing a matching request.
Can Diffbot be blocked in robots.txt?
The record identifies Diffbot as the token to review. Robots.txt is a request policy, not proof of model use or non-use.
How can I verify a Diffbot request?
Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.
Data & Sources
- Diffbot source reference - Source used to verify Diffbot.