Rankealo

Google · Training crawlers

Google-Extended

A robots.txt-only switch for Gemini training and grounding. It is not an HTTP crawler. Disallowing it does not change Google Search crawling.

Robots token, not a crawler
Operator
Google
Traffic type
Training crawlers
Verification
Robots token, no HTTP crawler
robots.txt
Not an HTTP crawler

Usually means: Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.

User-agent

robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap Google-Extended in extra product or version text.

Google-Extended

What Google-Extended does

Google-Extended controls whether content Google already crawled may be used for Gemini model training and grounding. There is no separate Google-Extended user-agent hitting your server.

How to get discovered

Search and AI Overview discovery still go through Googlebot. Use Google-Extended only as the Gemini-training opt-out; do not expect log lines for this token.

Why Google-Extended might skip you

It will never “visit” as its own UA. If Gemini still uses you for grounding, that is a Google-Extended policy question, not a missing crawl.

robots.txt rule

This token is matched in robots.txt only. You will not see it as a separate user-agent in access logs.

# Robots.txt-only token — not a separate HTTP crawler
User-agent: Google-Extended
Disallow: /

No IP ranges: this token never hits your server as its own user-agent.

Sources

JavaScript: reads raw HTML only.

Other Google bots

Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.

Google-Extended FAQ

What is Google-Extended?

Google-Extended controls whether content Google already crawled may be used for Gemini model training and grounding. There is no separate Google-Extended user-agent hitting your server.

Should I allow Google-Extended in robots.txt?

Google-Extended is a robots.txt-only switch, not a crawler that hits your server. Use it to opt out of training-style use without expecting a new user-agent in the logs.

How do I know a request is really Google-Extended?

Google-Extended is not an HTTP crawler, so you will not see it in access logs. It only appears as a robots.txt User-agent group.

Letting Google-Extended in is the start. Getting cited is the job.

Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.

Reading is step one. Measuring your AI visibility is step two.

Rankealo tracks how often your brand is mentioned and cited across the major AI engines, then helps you publish the pages that close the gaps.