Google

google.com

Training crawler

Google-Extended

Google-Extended is a robots.txt product control token from Google, not a fetching crawler user-agent. Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

Source checked 18 Aug 2026Officially documentedActive

Purpose

Training crawler

Evidence

Officially documented

Status

Active

robots.txt

Honors robots.txt
Last source check: 18 Aug 2026 · First verified in this record: 18 Aug 2026
[01]

What is Google-Extended?

Google-Extended is a robots.txt product control token from Google, not a fetching crawler user-agent. Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

[02]

Control token

No official HTTP user-agent is published for this record. Google-Extended is a robots.txt control token.

[03]

Allow or block in robots.txt

Use the narrow token below only after deciding whether this crawler's purpose fits your policy. An allow rule makes that choice explicit. A block rule asks compliant automated crawlers not to fetch matching paths.

Allow Google-Extended
User-agent: Google-Extended
Allow: /
Block Google-Extended
User-agent: Google-Extended
Disallow: /

Robots.txt is a request policy. It does not authenticate the sender, remove content already collected, or guarantee that a model will use or not use content. User-triggered fetchers can have different behavior from automated crawlers.

[04]

Verify a request before attributing it

01

Match the full user-agent

Use the published format where one exists. Treat the match as a clue, not proof.

02

Look for operator verification

No operator-published IP, reverse-DNS, or signature check is attached to this record.

03

Keep attribution separate from policy

Log what you observed, what the operator documents, and what you inferred as three separate fields.

[05]

JavaScript behavior

Rendering is not documented

The cited operator material does not verify JavaScript rendering. Serve useful HTML before client-side JavaScript where possible.

[08]

Frequently asked questions

Google-Extended is a robots.txt product control token from Google, not a fetching crawler user-agent. Robots.txt product token that controls eligible use of Google-crawled content for Gemini training and grounding.

Spoofing warning: any client can copy a user-agent string. Use operator-published IP ranges, reverse DNS, or request signatures where available. A name match alone does not verify ownership.

See which crawler signatures reach your site

Paste or upload a server log to find known AI crawlers, status codes, and requested pages. The analysis stays in your browser.