What is cohere-training-data-crawler? AI crawler guide

cohere-training-data-crawler is a crawler or fetcher associated with Cohere. See source-graded user-agent, robots.txt, verification, and observation guidance.

What is cohere-training-data-crawler?

cohere-training-data-crawler is a reported crawler or fetcher string associated with Cohere, but current operator evidence is incomplete. Cohere training-data crawler token reported for downloading web data for enterprise language models.

Evidence status

Field Value
Evidence Reported, not operator verified
Lifecycle Status uncertain
Purpose training
robots.txt posture unverified
Source checked 2026-06-11

Reported user-agent

cohere-training-data-crawler

cohere-training-data-crawler collects pages for model training. Training inclusion is not the same as being cited, so measure where AI answers actually cite your site before drawing conclusions from crawl logs.

robots.txt allow example

User-agent: cohere-training-data-crawler Allow: /

robots.txt block example

User-agent: cohere-training-data-crawler Disallow: /

What the rule can and cannot do

Robots.txt expresses an access policy to compliant automated crawlers. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content.

How to verify a request

JavaScript behavior

The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.

Related bots

Frequently Asked Questions

What is cohere-training-data-crawler?

cohere-training-data-crawler is a reported, not operator verified crawler or fetcher record associated with Cohere.

What user-agent does cohere-training-data-crawler use?

cohere-training-data-crawler. Check the evidence status before attributing a matching request.

Can cohere-training-data-crawler be blocked in robots.txt?

The record identifies cohere-training-data-crawler as the token to review. Robots.txt is a request policy, not proof of model use or non-use.

How can I verify a cohere-training-data-crawler request?

Treat the user-agent as a clue and look for operator-published IP, reverse-DNS, or signature guidance.

Data & Sources