Directory · verified September 2026
Get discovered by the bots that feed AI answers
63 crawlers, grouped by what they are actually for: live answers, search indexes, training, and the rest. The Rankealo question is not “what showed up in your logs.” It is whether these bots can find you — and why they are not visiting if they cannot.
- AI answers
- 16
- Search indexes
- 14
- Training crawlers
- 20
- Other AI bots
- 13
Why these bots are not visiting your website
Empty logs are usually an access problem, not a popularity problem. Fix the fetch path before you spend another month on content the crawler never received.
robots.txt Disallow on the wrong token
Blocking GPTBot (training) is optional. Blocking OAI-SearchBot, PerplexityBot, or Claude-SearchBot is how you vanish from AI answers. The tokens are not interchangeable.
A CDN “Block AI bots” toggle
Cloudflare and other WAFs can 403 GPTBot and friends while robots.txt still says Allow. Rankealo’s GEO audit probes the homepage as those UAs for that reason.
JavaScript-only pages
Googlebot may come back for a render pass. GPTBot, OAI-SearchBot, ClaudeBot, and PerplexityBot generally read the HTML you serve. If the article appears after hydration, they skip it.
Nothing worth retrieving
The bots can reach you and still never cite you if the page has no extractable answer. Discovery is access plus a page built to be quoted.
63 crawlers
User-triggered · 16
AI answers
Live fetches when someone asks an assistant a question. If these bots cannot reach the page, you cannot be the citation in that answer.
OpenAI
ChatGPT-User
The fetch ChatGPT makes when a person asks it to open, browse, or cite a specific page. This is a live citation request, not GPTBot training.
Anthropic
Claude-User
Claude’s on-demand fetcher when a person needs a live page for an answer. Blocking it hides you from in-product browsing and citations.
Perplexity
Perplexity-User
Perplexity’s live fetch for a user question and the citations shown next to the answer. This is how a page becomes a source in the moment.
Google-Agent
A Google user-triggered fetcher used by AI and product experiences. If Google-Agent cannot read you, those surfaces have nothing to quote.
Google-NotebookLM
The fetcher behind NotebookLM-style answer workflows. Discovery here means your page can be loaded into a notebook and cited as a source.
Google-Read-Aloud
Google’s fetcher for read-aloud and assistant playback. If it cannot parse the article text, the assistant has nothing to speak.
GoogleAgent
A second Google agent token seen in the wild for user-triggered AI fetches. Treat it like Google-Agent: allow it if you want Google products to read the page.
Microsoft
Copilot
User-triggered Microsoft Copilot fetches for AI answers. Blocking Copilot (or Bing preview UAs) keeps your pages out of Copilot citations.
Mistral
MistralAI-User
Mistral’s user-triggered fetcher. If Le Chat needs your page to answer, this is the UA that has to succeed.
Amazon
Amzn-User
Amazon’s user-triggered fetch for fresh answers in Alexa and related products. Discovery here is being readable when Amazon needs the live page.
DuckDuckGo
DuckAssistBot
DuckDuckGo’s real-time crawler for AI-assisted answers with citations. If DuckAssistBot is blocked, you will not be a Duck.ai source.
xAI
xAI-SearchBot
xAI’s user-triggered search fetcher for Grok answers. Allow it if you want Grok to be able to retrieve and cite the page.
xAI
Grok-DeepSearch
Grok’s deeper user-triggered research fetch. This is how a page gets pulled into a Grok DeepSearch-style answer.
Meta
meta-externalfetcher
Meta’s fetcher when a person requests or shares a specific URL. Blocking it breaks Meta AI reads and some share previews.
Moonshot AI
Kimi-User
Moonshot’s user-triggered Kimi fetcher. If Kimi cannot open the page, it cannot cite you in that answer.
Alibaba
Qwen-User
Alibaba’s user-triggered Qwen fetcher. Discovery means Qwen can actually read the page when someone asks.
Discovery · 14
Search indexes
Crawlers that build the indexes ChatGPT search, Claude search, Perplexity, Google, and Bing draw from. Blocking them is how brands disappear from AI answers without noticing.
OpenAI
OAI-SearchBot
OpenAI’s search-index crawler. This is the bot that makes a site eligible to appear in ChatGPT search results and citations.
Anthropic
Claude-SearchBot
Anthropic’s crawler for Claude search discovery. If Claude-SearchBot cannot index you, Claude search has no page to retrieve.
Perplexity
PerplexityBot
Perplexity’s index crawler. This is how pages enter the corpus Perplexity answers from — a search bot, not a training bot.
Google-InspectionTool
The UA Search Console’s URL Inspection uses. If it cannot fetch you, you cannot debug Google’s view of the page — including AI-related features that ride on Search.
Googlebot
Google’s Search crawler. Cloudflare notes Googlebot traffic can also support AI features such as AI Overviews — blocking it is not an AI-only decision.
Microsoft
Bingbot
Microsoft’s Search crawler. Copilot and Bing AI features depend on Bing’s index — if Bingbot cannot crawl you, Copilot has a thinner source set.
Microsoft
msnbot
Legacy Microsoft search crawler still seen in logs. Allow it unless you have a specific reason to drop old Microsoft UAs.
Mistral
MistralAI-Index
Mistral’s search/index crawler. This is how pages become eligible for Mistral search-style answers, not the live user fetch.
Amazon
Amzn-SearchBot
Amazon’s crawler for making content eligible in Amazon search experiences. Blocking it hides you from those surfaces even if Amazonbot still trains.
Meta
meta-webindexer
Meta’s crawler for Meta AI search quality. If it cannot index you, Meta AI search has a weaker view of the site.
Moonshot AI
Kimi-SearchBot
Moonshot’s Kimi search-index crawler. This is the discovery path into Kimi search, not the live user fetch.
ByteDance
TikTokSpider
ByteDance’s search/index spider. If you want TikTok/ByteDance search surfaces to know the site exists, this bot has to succeed.
Baidu
Baiduspider
Baidu’s search crawler. Baidu AI features sit on top of Baidu’s index — no crawl, no discovery in that ecosystem.
You.com
YouBot
You.com’s search/index crawler. Allow it if you want You.com AI search to be able to retrieve your pages.
Model data · 20
Training crawlers
Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.
OpenAI
GPTBot
OpenAI’s training crawler. Blocking GPTBot opts you out of future model training. It does not remove you from ChatGPT search if OAI-SearchBot stays allowed.
Anthropic
ClaudeBot
Anthropic’s training crawler for Claude. Blocking ClaudeBot is a training opt-out. Keep Claude-SearchBot allowed if you still want to be cited.
GoogleOther
Google’s generic crawler for product teams, research, and development fetches. It is not Googlebot and does not replace Search indexing.
Google-Extended
A robots.txt-only switch for Gemini training and grounding. It is not an HTTP crawler. Disallowing it does not change Google Search crawling.
Google-CloudVertexBot
Google crawler for targeted Vertex AI agent crawls that a site owner requested. It is not the public Search crawler.
Apple
Applebot-Extended
Apple’s robots.txt-only opt-out for AI training. Disallowing it does not remove you from Apple search, Spotlight, or Siri results that use Applebot.
Apple
Applebot
Apple’s crawler for search and, depending on Applebot-Extended, AI training. Cloudflare classifies it across search and training use cases.
Amazon
Amazonbot
Amazon’s broad crawler for improving Amazon products and services, including search and AI. Blocking it is a product-opt-out, not a Google ranking move.
Meta
meta-externalagent
Meta’s crawler for indexing or improving Meta products and AI systems. Publishers have disputed how consistently it honors robots.txt.
Moonshot AI
KimiBot
Moonshot’s public-content crawler. Blocking KimiBot is a training-style opt-out; keep Kimi-SearchBot allowed for discovery.
ByteDance
Bytespider
ByteDance’s public-content crawler. It is widely reported to ignore robots.txt — a Disallow alone may not stop it.
Baidu
ERNIEBot
Baidu’s ERNIE public-content crawler. Blocking it is a training opt-out; Baiduspider remains the Search discovery path.
Alibaba
QwenBot
Alibaba’s Qwen training-style crawler. Keep Qwen-User allowed if you still want live Qwen answers to be able to fetch you.
DeepSeek
DeepSeekBot
DeepSeek’s public-content crawler. DeepSeek’s chat product does not ship a first-party web search bot comparable to OAI-SearchBot.
Zhipu AI
ChatGLM-Spider
Zhipu’s ChatGLM public-content spider. Treat it as a training-style crawl unless Zhipu documents a separate search UA.
Cohere
cohere-ai
Cohere’s public-content crawler token. Blocking it opts you out of Cohere collection; it is not a ChatGPT citation switch.
Cohere
cohere-training-data-crawler
Cohere’s named training-data crawler. A Disallow here is a clear training opt-out.
Allen AI
AI2Bot
Allen Institute crawler for AI research datasets. Blocking it is a research-crawl opt-out, not a ChatGPT ranking move.
Common Crawl
CCBot
Common Crawl’s CCBot builds open web datasets that many labs reuse for pretraining. Blocking it is the widest training opt-out you can make in one token.
Anthropic
anthropic-ai
Anthropic’s legacy training token. Rankealo still tracks it because older robots.txt files and crawlers still emit anthropic-ai.
Previews and ads · 13
Other AI bots
Link-preview, ads, and product fetchers. Allow them if you want unfurls and landing-page fetches to work; they are not the citation path.
OpenAI
OAI-AdsBot
OpenAI’s ads and landing-page fetcher. Allow it if you run ChatGPT ads or need OpenAI to preview a destination URL.
xAI
GrokBot
An xAI Grok crawler token seen in the wild. Allow Grok UAs if you want Grok products to be able to fetch the site.
xAI
xAI-Bot
A generic xAI bot token. If it cannot fetch you, some Grok product fetches fail even when SearchBot is allowed.
xAI
xAI-Grok
Another xAI/Grok identity in logs. Rankealo lists it so you can allow the real token instead of guessing.
xAI
xAI-Web-Crawler
xAI’s web-crawler token. Allow it if you want xAI systems to discover pages beyond a single user click.
xAI
Grok
A short Grok user-agent token. Short tokens are easy to miss in robots.txt because they do not look like GPTBot.
Meta
meta-externalads
Meta’s ads-related fetcher. Allow it on destinations you run as Facebook/Instagram ads or the preview will not match the live page.
Meta
facebookexternalhit
Meta’s link-preview crawler for Facebook, Instagram, and Messenger. Blocking it breaks unfurls, not ChatGPT citations.
Meta
FacebookBot
Meta crawler traffic associated with Facebook bot fetches. Allow it unless you are deliberately dropping Facebook product crawls.
ByteDance
Doubaobot
ByteDance’s Doubao product crawler. Allow it if Doubao should be able to fetch your pages; block at the edge if you want a hard no.
Baidu
YiyanBot
Baidu’s Ernie/Yiyan assistant crawler token. Distinct from Baiduspider (Search) and ERNIEBot (training-style collection).
Alibaba
TongyiBot
Alibaba’s Tongyi assistant crawler. Allow it for Tongyi fetches; Qwen-User is the closer “live answer” twin.
Alibaba
AliyunBot
Alibaba Cloud’s crawler token. Usually an infrastructure/product fetch, not ChatGPT search — still needs a 200 if you want those products to read you.
AI crawler FAQ
How is this different from a crawler-tracking analytics tool?
Those products tell you which bots already hit your logs. This directory is the GEO version: which bots you need to be discoverable to, what each token is for, and the usual reasons they never arrive — robots.txt, CDN bot fights, and JS-only HTML. Rankealo then checks access and publishes pages those engines can cite.
Should I allow every AI crawler?
Allow the search and user-triggered tokens if you want citations (OAI-SearchBot, ChatGPT-User, PerplexityBot, Claude-SearchBot, and their peers). Training crawlers like GPTBot, ClaudeBot, and CCBot are a separate, optional choice. You can opt out of training and still be cited.
Why aren’t AI crawlers visiting my website?
The usual stack is: a Disallow on the retrieval token, a WAF “block AI bots” rule that robots.txt cannot see, a page that is empty without JavaScript, or a URL that 404s / challenges the bot. Check robots.txt first, then fetch the homepage as GPTBot and PerplexityBot.
Does blocking GPTBot hide me from ChatGPT?
No. GPTBot is for training. ChatGPT search uses OAI-SearchBot, and in-session browsing uses ChatGPT-User. Block GPTBot if you want a training opt-out; keep the search tokens allowed if you want to be discovered.
Want the long-form policy explainer? See GPTBot and every AI crawler, explained.
Get found by these crawlers, then get cited.
Rankealo checks AI-bot access, generates robots.txt and llms.txt, and publishes pages built to be retrieved in ChatGPT, Claude, Perplexity, and Gemini — not a log tracker.
