What is AI2Bot? AI crawler guide

AI2Bot is a crawler or fetcher associated with Allen Institute for AI. See source-graded user-agent, robots.txt, verification, and observation guidance.

What is AI2Bot?

AI2Bot is a documented Allen Institute for AI crawler or fetcher. Allen Institute for AI crawler used to find web content for open language model datasets.

Evidence status

Field Value
Evidence Officially documented
Lifecycle Active
Purpose training
robots.txt posture honors
Source checked 2026-06-11

Documented user-agent

AI2Bot

AI2Bot collects pages for model training. Training inclusion is not the same as being cited, so measure where AI answers actually cite your site before drawing conclusions from crawl logs.

robots.txt allow example

User-agent: AI2Bot Allow: /

robots.txt block example

User-agent: AI2Bot Disallow: /

What the rule can and cannot do

Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.

How to verify a request

JavaScript behavior

The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.

Related bots

Frequently Asked Questions

What is AI2Bot?

AI2Bot is a officially documented crawler or fetcher record associated with Allen Institute for AI.

What user-agent does AI2Bot use?

AI2Bot. Check the evidence status before attributing a matching request.

Can AI2Bot be blocked in robots.txt?

The record identifies AI2Bot as the token to review. Robots.txt is a request policy, not proof of model use or non-use.

How can I verify a AI2Bot request?

Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.

Data & Sources