Reference, verified July 2026
GPTBot and every AI crawler, explained
There are three kinds of AI bot hitting your site: crawlers that train models, bots that power AI search answers, and robots.txt switches that are not crawlers at all. Confuse them and you can accidentally block yourself out of AI answers. For the per-bot playbooks — how to get discovered, and why a bot skips you — use the AI crawler directory.
Your robots.txt policy
OAI-SearchBotallowPerplexityBotallowClaude-SearchBotallowGPTBotblockCCBotblockA common setup: allow the search bots so you stay citable, block the training crawlers to opt out of model training.
Three things to get right first
Before you touch a single line of robots.txt, hold these three facts. They prevent the most common and most damaging mistake: blocking the bots that would have cited you.
To be cited, allow the search bots
If you want to show up in ChatGPT search, Perplexity, or Claude answers, the retrieval bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot) must be allowed. Blocking them is the fastest way to disappear from AI answers.
Training is a separate, optional choice
Blocking GPTBot, ClaudeBot, or CCBot only opts you out of model training. It does not remove you from AI answers, because answers come from live search, not the training set. You can block training and still be cited.
Some bots ignore the rules
A robots.txt directive is a request, not a wall. User-triggered fetchers and a few crawlers are reported to ignore it. For hard enforcement you need firewall or edge rules, not just robots.txt.
Every AI bot, by what it does
User-agent tokens below are the official names you put in robots.txt, verified against vendor docs in July 2026. The full user-agent strings are longer, but the token is what a rule matches on.
Search and answer bots, allow to stay citable
| User-agent | Operator | What it does | Honors robots.txt |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Powers ChatGPT search results and the citations shown with them. | Yes |
| ChatGPT-User | OpenAI | Fetches a page when a user asks ChatGPT to open or browse it. | User action |
| PerplexityBot | Perplexity | Builds the index Perplexity answers from. This is a search bot, not training. | Yes |
| Perplexity-User | Perplexity | Fetches a page to answer a live user question in the moment. | Often ignored, user-initiated |
| Claude-SearchBot | Anthropic | Indexes pages to improve the quality of Claude’s in-product search. | Yes |
| Claude-User | Anthropic | Fetches a page when a Claude user’s request needs it. | Yes |
Training crawlers, block to opt out of training
| User-agent | Operator | What it does | Honors robots.txt |
|---|---|---|---|
| GPTBot | OpenAI | Collects web data to train OpenAI models. | Yes |
| ClaudeBot | Anthropic | Collects web data to train Claude models. | Yes |
| CCBot | Common Crawl | Builds the open Common Crawl dataset, reused by many labs for pretraining. | Yes |
| Amazonbot | Amazon | Broad crawler feeding search and, per Amazon, model training. | Yes |
| Meta-ExternalAgent | Meta | Crawls for model training and Meta’s own indexing. | Disputed by publishers |
| Bytespider | ByteDance | Collects training data for ByteDance and TikTok models. | Reportedly ignored, undocumented |
Control tokens, not crawlers
A robots.txt-only switch. It controls whether content Googlebot already crawled can train Gemini or power grounding. It does not change Search crawling or ranking.
Apple
A robots.txt-only switch. Disallowing it opts your Applebot-crawled content out of training Apple’s models, without removing you from Apple search or Siri results.
Copy-paste robots.txt rules
Two common policies. Paste one at the root of your site as robots.txt, then confirm it with the crawler checker. Remember these are requests: well-behaved bots comply, a few do not.
Opt out of training, keep citations
# Opt out of AI training, stay eligible for citations
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Applebot-Extended
Disallow: /
# Keep the search/answer bots allowed so you can still be cited
User-agent: OAI-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
User-agent: Claude-SearchBot
Allow: /Block AI bots entirely
# Block AI training and AI search fetchers
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: Claude-User
User-agent: PerplexityBot
User-agent: CCBot
User-agent: Amazonbot
User-agent: Meta-ExternalAgent
User-agent: Bytespider
User-agent: Google-Extended
User-agent: Applebot-Extended
Disallow: /Prefer not to hand-edit? Generate a clean file with the robots.txt generator, then verify AI-bot rules with the AI crawler checker.
Set your AI-crawler rules in a few minutes
Four free tools cover the workflow: see what you allow today, generate the rules you want, read the WordPress default file, and point assistants at your best pages.
AI Crawler Checker
Paste your domain to see how your robots.txt currently treats GPTBot, ClaudeBot, PerplexityBot, and more.
Robots.txt Generator
Build a clean robots.txt with the exact AI-bot rules you want, then copy or download it.
WordPress robots.txt
The default WordPress file, line by line, plus how to edit it and the mistakes that deindex a site.
llms.txt Generator
Point AI assistants at your source-of-truth pages with an llms.txt file, generated from your site.
AI crawler FAQ
What is GPTBot?
GPTBot is OpenAI’s web crawler for collecting data to train its models. It is separate from OAI-SearchBot, which powers ChatGPT search, and from ChatGPT-User, which fetches a page when a user browses inside ChatGPT. You can block GPTBot in robots.txt with "User-agent: GPTBot" followed by "Disallow: /" while still allowing the search bots.
Should I block AI crawlers in robots.txt?
It depends on your goal. If you want AI visibility, do not block the search and answer bots (OAI-SearchBot, PerplexityBot, Claude-SearchBot), because those are what let you be cited. Blocking the training crawlers (GPTBot, ClaudeBot, CCBot) is a reasonable, independent choice if you do not want your content used to train models, and it does not remove you from AI answers.
Will blocking GPTBot remove me from ChatGPT?
No. GPTBot is for training. ChatGPT search is powered by OAI-SearchBot, and in-session browsing uses ChatGPT-User. If you block GPTBot but keep OAI-SearchBot allowed, you opt out of training while remaining eligible to appear in ChatGPT search answers with a citation.
What is the difference between Google-Extended and Googlebot?
Googlebot is the crawler that indexes pages for Google Search. Google-Extended is not a separate crawler at all; it is a robots.txt-only token that controls whether content Google already crawled can be used to train Gemini or power grounding. Disallowing Google-Extended has no effect on your Search crawling or ranking.
Do AI crawlers actually obey robots.txt?
The major training and search bots from OpenAI, Anthropic, Perplexity, and Common Crawl document that they honor robots.txt. Some user-initiated fetchers and a few crawlers, including reports about ByteDance’s Bytespider and Meta’s crawler, are said to ignore or inconsistently follow it. If you need guaranteed blocking, enforce it at the firewall or CDN edge rather than relying on robots.txt alone.
Related guides
Keep going deeper on AI search visibility across the rest of the guide series.
LLM Knowledge Cutoff Dates
A verified table of knowledge cutoff dates and live web access for every major AI model.
Read the guideHow AI Search Engines Work
The retrieve, rank, synthesize, and cite pipeline behind ChatGPT, Perplexity, Gemini, and Claude answers.
Read the guideAI Search Optimization
The broader system behind modern AI search visibility across prompts, pages, and engines.
Read the guide