Rankealo

Google · Training crawlers

Google-CloudVertexBot

Google crawler for targeted Vertex AI agent crawls that a site owner requested. It is not the public Search crawler.

Optional training opt-out
Operator
Google
Traffic type
Training crawlers
Verification
User-agent + IP range
robots.txt
Honors robots.txt

Usually means: Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.

User-agent

robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap Google-CloudVertexBot in extra product or version text.

Google-CloudVertexBot

What Google-CloudVertexBot does

Google-CloudVertexBot fetches pages when someone building Vertex AI agents asks Google to crawl those URLs. Traffic is opt-in at the Vertex side, not a general web crawl.

How to get discovered

If you are the one requesting the Vertex crawl, allow this UA and keep those URLs public. It will not replace Googlebot for Search or AI Overviews.

Why Google-CloudVertexBot might skip you

You never requested a Vertex crawl, the target URLs 401, or robots Disallow on this token while the Vertex job still expects a 200.

robots.txt rule

This snippet opts out of training-style collection. Keep the search and user-fetch tokens from the same operator allowed if you still want citations.

# Optional training opt-out — does not remove you from AI search by itself
User-agent: Google-CloudVertexBot
Disallow: /

Allowlist so Google-CloudVertexBot can reach you

A user-agent is spoofable. Google publishes current CIDR ranges as JSON — fetch that file for WAF/CDN allowlists instead of copying ranges from a blog. Ranges rotate; a screenshot of ten prefixes will go stale. If Cloudflare “Block AI bots” is on, this is the list that gets the real crawler through.

https://developers.google.com/static/search/apis/ipranges/special-crawlers.json

Sources

JavaScript: render behavior is not documented.

Other Google bots

Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.

Google-CloudVertexBot FAQ

What is Google-CloudVertexBot?

Google-CloudVertexBot fetches pages when someone building Vertex AI agents asks Google to crawl those URLs. Traffic is opt-in at the Vertex side, not a general web crawl.

Should I allow Google-CloudVertexBot in robots.txt?

Only if you want to opt out of training-style collection. Blocking Google-CloudVertexBot does not, by itself, remove you from live AI search citations — those use the search and user-fetch tokens from Google.

How do I know a request is really Google-CloudVertexBot?

A user-agent is a claim anyone can send. Match the Google-CloudVertexBot token, then check the source IP against Google’s published ranges. Use those ranges as a WAF allowlist so a “Block AI bots” rule does not 403 the real crawler.

Letting Google-CloudVertexBot in is the start. Getting cited is the job.

Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.

Reading is step one. Measuring your AI visibility is step two.

Rankealo tracks how often your brand is mentioned and cited across the major AI engines, then helps you publish the pages that close the gaps.