What is Ai2Bot-Dolma? AI crawler guide
Ai2Bot-Dolma is a crawler or fetcher associated with Allen Institute for AI. See source-graded user-agent, robots.txt, verification, and observation guidance.
What is Ai2Bot-Dolma?
Ai2Bot-Dolma is a documented Allen Institute for AI crawler or fetcher. AI2 crawler token associated with Dolma/open language model dataset collection.
Evidence status
| Field | Value |
|---|---|
| Evidence | Officially documented |
| Lifecycle | Active |
| Purpose | training |
| robots.txt posture | honors |
| Source checked | 2026-06-11 |
Documented user-agent
Ai2Bot-Dolma
Ai2Bot-Dolma collects pages for model training. Training inclusion is not the same as being cited, so measure where AI answers actually cite your site before drawing conclusions from crawl logs.
robots.txt allow example
User-agent: Ai2Bot-Dolma Allow: /
robots.txt block example
User-agent: Ai2Bot-Dolma Disallow: /
What the rule can and cannot do
Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.
How to verify a request
- Match the complete documented user-agent where one exists. A match is only a clue.
- Look for operator-published IP ranges, reverse DNS, or signed-request guidance. None is attached to this record.
- Keep the observed request, operator documentation, and any inference as separate fields.
JavaScript behavior
The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.
Related bots
- AI2Bot: Another Allen Institute for AI training crawler to compare.
- ICC-Crawler: Also tracked as a training crawler.
- img2dataset: Also tracked as a training crawler.
- LAIONDownloader: Also tracked as a training crawler.
- PanguBot: Also tracked as a training crawler.
- CCBot: Also tracked as a training crawler.
- Google-Extended: Also tracked as a training crawler.
- Webzio-Extended: Also tracked as a training crawler.
- cohere-training-data-crawler: Also tracked as a training crawler.
- AI Training Opt-Out: Ai2Bot-Dolma is a training crawler tied to this policy decision.
- Robots.txt: Robots.txt is the control file used to allow or block Ai2Bot-Dolma.
Frequently Asked Questions
What is Ai2Bot-Dolma?
Ai2Bot-Dolma is a officially documented crawler or fetcher record associated with Allen Institute for AI.
What user-agent does Ai2Bot-Dolma use?
Ai2Bot-Dolma. Check the evidence status before attributing a matching request.
Can Ai2Bot-Dolma be blocked in robots.txt?
The record identifies Ai2Bot-Dolma as the token to review. Robots.txt is a request policy, not proof of model use or non-use.
How can I verify a Ai2Bot-Dolma request?
Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.
Data & Sources
- Allen Institute for AI documentation - Primary source for Ai2Bot-Dolma crawler details.