Rankealo

Directory · verified September 2026

Get discovered by the bots that feed AI answers

63 crawlers, grouped by what they are actually for: live answers, search indexes, training, and the rest. The Rankealo question is not “what showed up in your logs.” It is whether these bots can find you — and why they are not visiting if they cannot.

AI answers
16
Search indexes
14
Training crawlers
20
Other AI bots
13

Why these bots are not visiting your website

Empty logs are usually an access problem, not a popularity problem. Fix the fetch path before you spend another month on content the crawler never received.

robots.txt Disallow on the wrong token

Blocking GPTBot (training) is optional. Blocking OAI-SearchBot, PerplexityBot, or Claude-SearchBot is how you vanish from AI answers. The tokens are not interchangeable.

A CDN “Block AI bots” toggle

Cloudflare and other WAFs can 403 GPTBot and friends while robots.txt still says Allow. Rankealo’s GEO audit probes the homepage as those UAs for that reason.

JavaScript-only pages

Googlebot may come back for a render pass. GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot generally read the HTML you serve. If the article appears after hydration, they skip it.

Nothing worth retrieving

The bots can reach you and still never cite you if the page has no extractable answer. Discovery is access plus a page built to be quoted.

63 crawlers

User-triggered · 16

AI answers

Live fetches when someone asks an assistant a question. If these bots cannot reach the page, you cannot be the citation in that answer.

OpenAI

ChatGPT-User

Allow to get discovered

The fetch ChatGPT makes when a person asks it to open, browse, or cite a specific page. This is a live citation request, not GPTBot training.

Anthropic

Claude-User

Allow to get discovered

Claude’s on-demand fetcher when a person needs a live page for an answer. Blocking it hides you from in-product browsing and citations.

Perplexity

Perplexity-User

Allow to get discovered

Perplexity’s live fetch for a user question and the citations shown next to the answer. This is how a page becomes a source in the moment.

Google

Google-Agent

Allow to get discovered

A Google user-triggered fetcher used by AI and product experiences. If Google-Agent cannot read you, those surfaces have nothing to quote.

Google

Google-NotebookLM

Allow to get discovered

The fetcher behind NotebookLM-style answer workflows. Discovery here means your page can be loaded into a notebook and cited as a source.

Google

Google-Read-Aloud

Allow to get discovered

Google’s fetcher for read-aloud and assistant playback. If it cannot parse the article text, the assistant has nothing to speak.

Google

GoogleAgent

Allow to get discovered

A second Google agent token seen in the wild for user-triggered AI fetches. Treat it like Google-Agent: allow it if you want Google products to read the page.

Microsoft

Copilot

Allow to get discovered

User-triggered Microsoft Copilot fetches for AI answers. Blocking Copilot (or Bing preview UAs) keeps your pages out of Copilot citations.

Mistral

MistralAI-User

Allow to get discovered

Mistral’s user-triggered fetcher. If Le Chat needs your page to answer, this is the UA that has to succeed.

Amazon

Amzn-User

Allow to get discovered

Amazon’s user-triggered fetch for fresh answers in Alexa and related products. Discovery here is being readable when Amazon needs the live page.

DuckDuckGo

DuckAssistBot

Allow to get discovered

DuckDuckGo’s real-time crawler for AI-assisted answers with citations. If DuckAssistBot is blocked, you will not be a Duck.ai source.

xAI

xAI-SearchBot

Allow to get discovered

xAI’s user-triggered search fetcher for Grok answers. Allow it if you want Grok to be able to retrieve and cite the page.

xAI

Grok-DeepSearch

Allow to get discovered

Grok’s deeper user-triggered research fetch. This is how a page gets pulled into a Grok DeepSearch-style answer.

Meta

meta-externalfetcher

Allow to get discovered

Meta’s fetcher when a person requests or shares a specific URL. Blocking it breaks Meta AI reads and some share previews.

Moonshot AI

Kimi-User

Allow to get discovered

Moonshot’s user-triggered Kimi fetcher. If Kimi cannot open the page, it cannot cite you in that answer.

Alibaba

Qwen-User

Allow to get discovered

Alibaba’s user-triggered Qwen fetcher. Discovery means Qwen can actually read the page when someone asks.

Discovery · 14

Search indexes

Crawlers that build the indexes ChatGPT search, Claude search, Perplexity, Google, and Bing draw from. Blocking them is how brands disappear from AI answers without noticing.

OpenAI

OAI-SearchBot

Allow to get discovered

OpenAI’s search-index crawler. This is the bot that makes a site eligible to appear in ChatGPT search results and citations.

Anthropic

Claude-SearchBot

Allow to get discovered

Anthropic’s crawler for Claude search discovery. If Claude-SearchBot cannot index you, Claude search has no page to retrieve.

Perplexity

PerplexityBot

Allow to get discovered

Perplexity’s index crawler. This is how pages enter the corpus Perplexity answers from — a search bot, not a training bot.

Google

Google-InspectionTool

Allow to get discovered

The UA Search Console’s URL Inspection uses. If it cannot fetch you, you cannot debug Google’s view of the page — including AI-related features that ride on Search.

Google

Googlebot

Allow to get discovered

Google’s Search crawler. Cloudflare notes Googlebot traffic can also support AI features such as AI Overviews — blocking it is not an AI-only decision.

Microsoft

Bingbot

Allow to get discovered

Microsoft’s Search crawler. Copilot and Bing AI features depend on Bing’s index — if Bingbot cannot crawl you, Copilot has a thinner source set.

Microsoft

msnbot

Allow to get discovered

Legacy Microsoft search crawler still seen in logs. Allow it unless you have a specific reason to drop old Microsoft UAs.

Mistral

MistralAI-Index

Allow to get discovered

Mistral’s search/index crawler. This is how pages become eligible for Mistral search-style answers, not the live user fetch.

Amazon

Amzn-SearchBot

Allow to get discovered

Amazon’s crawler for making content eligible in Amazon search experiences. Blocking it hides you from those surfaces even if Amazonbot still trains.

Meta

meta-webindexer

Allow to get discovered

Meta’s crawler for Meta AI search quality. If it cannot index you, Meta AI search has a weaker view of the site.

Moonshot AI

Kimi-SearchBot

Allow to get discovered

Moonshot’s Kimi search-index crawler. This is the discovery path into Kimi search, not the live user fetch.

ByteDance

TikTokSpider

Allow to get discovered

ByteDance’s search/index spider. If you want TikTok/ByteDance search surfaces to know the site exists, this bot has to succeed.

Baidu

Baiduspider

Allow to get discovered

Baidu’s search crawler. Baidu AI features sit on top of Baidu’s index — no crawl, no discovery in that ecosystem.

You.com

YouBot

Allow to get discovered

You.com’s search/index crawler. Allow it if you want You.com AI search to be able to retrieve your pages.

Model data · 20

Training crawlers

Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.

OpenAI

GPTBot

Optional training opt-out

OpenAI’s training crawler. Blocking GPTBot opts you out of future model training. It does not remove you from ChatGPT search if OAI-SearchBot stays allowed.

Anthropic

ClaudeBot

Optional training opt-out

Anthropic’s training crawler for Claude. Blocking ClaudeBot is a training opt-out. Keep Claude-SearchBot allowed if you still want to be cited.

Google

GoogleOther

Optional training opt-out

Google’s generic crawler for product teams, research, and development fetches. It is not Googlebot and does not replace Search indexing.

Google

Google-Extended

Robots token, not a crawler

A robots.txt-only switch for Gemini training and grounding. It is not an HTTP crawler. Disallowing it does not change Google Search crawling.

Google

Google-CloudVertexBot

Optional training opt-out

Google crawler for targeted Vertex AI agent crawls that a site owner requested. It is not the public Search crawler.

Apple

Applebot-Extended

Robots token, not a crawler

Apple’s robots.txt-only opt-out for AI training. Disallowing it does not remove you from Apple search, Spotlight, or Siri results that use Applebot.

Apple

Applebot

Optional training opt-out

Apple’s crawler for search and, depending on Applebot-Extended, AI training. Cloudflare classifies it across search and training use cases.

Amazon

Amazonbot

Optional training opt-out

Amazon’s broad crawler for improving Amazon products and services, including search and AI. Blocking it is a product-opt-out, not a Google ranking move.

Meta

meta-externalagent

Optional training opt-out

Meta’s crawler for indexing or improving Meta products and AI systems. Publishers have disputed how consistently it honors robots.txt.

Moonshot AI

KimiBot

Optional training opt-out

Moonshot’s public-content crawler. Blocking KimiBot is a training-style opt-out; keep Kimi-SearchBot allowed for discovery.

ByteDance

Bytespider

Optional training opt-out

ByteDance’s public-content crawler. It is widely reported to ignore robots.txt — a Disallow alone may not stop it.

Baidu

ERNIEBot

Optional training opt-out

Baidu’s ERNIE public-content crawler. Blocking it is a training opt-out; Baiduspider remains the Search discovery path.

Alibaba

QwenBot

Optional training opt-out

Alibaba’s Qwen training-style crawler. Keep Qwen-User allowed if you still want live Qwen answers to be able to fetch you.

DeepSeek

DeepSeekBot

Optional training opt-out

DeepSeek’s public-content crawler. DeepSeek’s chat product does not ship a first-party web search bot comparable to OAI-SearchBot.

Zhipu AI

ChatGLM-Spider

Optional training opt-out

Zhipu’s ChatGLM public-content spider. Treat it as a training-style crawl unless Zhipu documents a separate search UA.

Cohere

cohere-ai

Optional training opt-out

Cohere’s public-content crawler token. Blocking it opts you out of Cohere collection; it is not a ChatGPT citation switch.

Cohere

cohere-training-data-crawler

Optional training opt-out

Cohere’s named training-data crawler. A Disallow here is a clear training opt-out.

Allen AI

AI2Bot

Optional training opt-out

Allen Institute crawler for AI research datasets. Blocking it is a research-crawl opt-out, not a ChatGPT ranking move.

Common Crawl

CCBot

Optional training opt-out

Common Crawl’s CCBot builds open web datasets that many labs reuse for pretraining. Blocking it is the widest training opt-out you can make in one token.

Anthropic

anthropic-ai

Optional training opt-out

Anthropic’s legacy training token. Rankealo still tracks it because older robots.txt files and crawlers still emit anthropic-ai.

Previews and ads · 13

Other AI bots

Link-preview, ads, and product fetchers. Allow them if you want unfurls and landing-page fetches to work; they are not the citation path.

OpenAI

OAI-AdsBot

Allow to get discovered

OpenAI’s ads and landing-page fetcher. Allow it if you run ChatGPT ads or need OpenAI to preview a destination URL.

xAI

GrokBot

Allow to get discovered

An xAI Grok crawler token seen in the wild. Allow Grok UAs if you want Grok products to be able to fetch the site.

xAI

xAI-Bot

Allow to get discovered

A generic xAI bot token. If it cannot fetch you, some Grok product fetches fail even when SearchBot is allowed.

xAI

xAI-Grok

Allow to get discovered

Another xAI/Grok identity in logs. Rankealo lists it so you can allow the real token instead of guessing.

xAI

xAI-Web-Crawler

Allow to get discovered

xAI’s web-crawler token. Allow it if you want xAI systems to discover pages beyond a single user click.

xAI

Grok

Allow to get discovered

A short Grok user-agent token. Short tokens are easy to miss in robots.txt because they do not look like GPTBot.

Meta

meta-externalads

Allow to get discovered

Meta’s ads-related fetcher. Allow it on destinations you run as Facebook/Instagram ads or the preview will not match the live page.

Meta

facebookexternalhit

Allow to get discovered

Meta’s link-preview crawler for Facebook, Instagram, and Messenger. Blocking it breaks unfurls, not ChatGPT citations.

Meta

FacebookBot

Allow to get discovered

Meta crawler traffic associated with Facebook bot fetches. Allow it unless you are deliberately dropping Facebook product crawls.

ByteDance

Doubaobot

Allow to get discovered

ByteDance’s Doubao product crawler. Allow it if Doubao should be able to fetch your pages; block at the edge if you want a hard no.

Baidu

YiyanBot

Allow to get discovered

Baidu’s Ernie/Yiyan assistant crawler token. Distinct from Baiduspider (Search) and ERNIEBot (training-style collection).

Alibaba

TongyiBot

Allow to get discovered

Alibaba’s Tongyi assistant crawler. Allow it for Tongyi fetches; Qwen-User is the closer “live answer” twin.

Alibaba

AliyunBot

Allow to get discovered

Alibaba Cloud’s crawler token. Usually an infrastructure/product fetch, not ChatGPT search — still needs a 200 if you want those products to read you.

AI crawler FAQ

How is this different from a crawler-tracking analytics tool?

Those products tell you which bots already hit your logs. This directory is the GEO version: which bots you need to be discoverable to, what each token is for, and the usual reasons they never arrive — robots.txt, CDN bot fights, and JS-only HTML. Rankealo then checks access and publishes pages those engines can cite.

Should I allow every AI crawler?

Allow the search and user-triggered tokens if you want citations (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot, and their peers). Training crawlers like GPTBot, ClaudeBot, and CCBot are a separate, optional choice. You can opt out of training and still be cited.

Why aren’t AI crawlers visiting my website?

The usual stack is: a Disallow on the retrieval token, a WAF “block AI bots” rule that robots.txt cannot see, a page that is empty without JavaScript, or a URL that 404s / challenges the bot. Check robots.txt first, then fetch the homepage as GPTBot and PerplexityBot.

Does blocking GPTBot hide me from ChatGPT?

No. GPTBot is for training. ChatGPT search uses OAI-SearchBot, and in-session browsing uses ChatGPT-User. Block GPTBot if you want a training opt-out; keep the search tokens allowed if you want to be discovered.

Want the long-form policy explainer? See GPTBot and every AI crawler, explained.

Get found by these crawlers, then get cited.

Rankealo checks AI-bot access, generates robots.txt and llms.txt, and publishes pages built to be retrieved in ChatGPT, Claude, Perplexity, and Gemini — not a log tracker.

Reading is step one. Measuring your AI visibility is step two.

Rankealo tracks how often your brand is mentioned and cited across the major AI engines, then helps you publish the pages that close the gaps.