What is ICC-Crawler? AI crawler guide
ICC-Crawler is a crawler or fetcher associated with NICT. See source-graded user-agent, robots.txt, verification, and observation guidance.
What is ICC-Crawler?
ICC-Crawler is a reported crawler or fetcher string associated with NICT, but current operator evidence is incomplete. NICT crawler for data used in artificial intelligence technologies and third-party research/commercial uses.
Evidence status
| Field | Value |
|---|---|
| Evidence | Reported, not operator verified |
| Lifecycle | Status uncertain |
| Purpose | training |
| robots.txt posture | honors |
| Source checked | 2026-06-11 |
Reported user-agent
ICC-Crawler
ICC-Crawler collects pages for model training. Training inclusion is not the same as being cited, so measure where AI answers actually cite your site before drawing conclusions from crawl logs.
robots.txt allow example
User-agent: ICC-Crawler Allow: /
robots.txt block example
User-agent: ICC-Crawler Disallow: /
What the rule can and cannot do
Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.
How to verify a request
- Match the complete documented user-agent where one exists. A match is only a clue.
- Look for operator-published IP ranges, reverse DNS, or signed-request guidance. None is attached to this record.
- Keep the observed request, operator documentation, and any inference as separate fields.
JavaScript behavior
The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.
Related bots
- CCBot: Also tracked as a training crawler.
- Ai2Bot-Dolma: Also tracked as a training crawler.
- LAIONDownloader: Also tracked as a training crawler.
- AI2Bot: Also tracked as a training crawler.
- Google-Extended: Also tracked as a training crawler.
- Webzio-Extended: Also tracked as a training crawler.
- img2dataset: Also tracked as a training crawler.
- ClaudeBot: Also tracked as a training crawler.
- GPTBot: Also tracked as a training crawler.
- Robots.txt: Robots.txt is the control file used to allow or block ICC-Crawler.
- AI Training Opt-Out: ICC-Crawler is a training crawler tied to this policy decision.
Frequently Asked Questions
What is ICC-Crawler?
ICC-Crawler is a reported, not operator verified crawler or fetcher record associated with NICT.
What user-agent does ICC-Crawler use?
ICC-Crawler. Check the evidence status before attributing a matching request.
Can ICC-Crawler be blocked in robots.txt?
The record identifies ICC-Crawler as the token to review. Robots.txt is a request policy, not proof of model use or non-use.
How can I verify a ICC-Crawler request?
Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.
Data & Sources
- ICC-Crawler source reference - Source used to verify ICC-Crawler.