What is Scrapy? AI crawler guide
Scrapy is a crawler or fetcher associated with Zyte. See source-graded user-agent, robots.txt, verification, and observation guidance.
What is Scrapy?
Scrapy is a reported crawler or fetcher string associated with Zyte, but current operator evidence is incomplete. Scrapy framework user-agent commonly used for web scraping, including AI and machine learning data extraction.
Evidence status
| Field | Value |
|---|---|
| Evidence | Reported, not operator verified |
| Lifecycle | Status uncertain |
| Purpose | other |
| robots.txt posture | unverified |
| Source checked | 2026-06-11 |
Reported user-agent
Scrapy
Allowing Scrapy only creates the possibility of retrieval. To find out whether your pages are selected, measure which answer engines actually cite your pages using a stable query set.
robots.txt allow example
User-agent: Scrapy Allow: /
robots.txt block example
User-agent: Scrapy Disallow: /
What the rule can and cannot do
Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.
How to verify a request
- Match the complete documented user-agent where one exists. A match is only a clue.
- Look for operator-published IP ranges, reverse DNS, or signed-request guidance. None is attached to this record.
- Keep the observed request, operator documentation, and any inference as separate fields.
JavaScript behavior
The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.
Related bots
- FirecrawlAgent: Also tracked as a general crawler.
- GoogleOther: Also tracked as a general crawler.
- ApifyBot: Also tracked as a general crawler.
- Google-CloudVertexBot: Also tracked as a general crawler.
- Amazonbot: Also tracked as a general crawler.
- bedrockbot: Also tracked as a general crawler.
- GoogleOther-Image: Also tracked as a general crawler.
- GoogleOther-Video: Also tracked as a general crawler.
- Panscient: Also tracked as a general crawler.
- Robots.txt: Robots.txt is the control file used to allow or block Scrapy.
- AI Crawlers: Scrapy is a concrete crawler example for this concept.
Frequently Asked Questions
What is Scrapy?
Scrapy is a reported, not operator verified crawler or fetcher record associated with Zyte.
What user-agent does Scrapy use?
Scrapy. Check the evidence status before attributing a matching request.
Can Scrapy be blocked in robots.txt?
The record identifies Scrapy as the token to review. Robots.txt is a request policy, not proof of model use or non-use.
How can I verify a Scrapy request?
Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.
Data & Sources
- Zyte documentation - Primary source for Scrapy crawler details.
- Scrapy source reference - Source used to verify Scrapy.