Baidu · Search indexes
Baiduspider
Baidu’s search crawler. Baidu AI features sit on top of Baidu’s index — no crawl, no discovery in that ecosystem.
- Operator
- Baidu
- Traffic type
- Search indexes
- Verification
- Not independently verifiable
- robots.txt
- Usually honors robots.txt
Usually means: Crawlers that build the indexes ChatGPT search, Claude search, Perplexity, Google, and Bing draw from. Blocking them is how brands disappear from AI answers without noticing.
User-agent
robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap Baiduspider in extra product or version text.
Baiduspider
What Baiduspider does
Baiduspider indexes the web for Baidu Search. ERNIEBot and YiyanBot are separate Baidu AI tokens.
How to get discovered
Allow Baiduspider if that market matters, host crawlable HTML, and avoid geo-blocking Baidu’s crawler IPs.
Why Baiduspider might skip you
Great Firewall assumptions (blocking Baidu from a Western WAF), missing ICP-unrelated crawl errors, or robots Disallow.
robots.txt rule
Allow this token on public pages you want retrieved or cited. A CDN “Block AI bots” toggle can still 403 it after robots.txt says Allow.
# Keep Baiduspider eligible to fetch this site
User-agent: Baiduspider
Allow: /Baidu has not published a machine-readable range file for this token. Do not treat the user-agent alone as proof of identity.
JavaScript: render behavior is not documented.
Other Baidu bots
Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.
Baiduspider FAQ
What is Baiduspider?
Baiduspider indexes the web for Baidu Search. ERNIEBot and YiyanBot are separate Baidu AI tokens.
Should I allow Baiduspider in robots.txt?
Yes if you want Baidu to be able to fetch and cite this site. Blocking Baiduspider is how pages stay invisible to that product even when Googlebot still crawls them.
How do I know a request is really Baiduspider?
Baidu has not published a range file for this token. Treat the user-agent as a claimed identity, and do not build a hard IP allowlist from blog posts or screenshots.
Letting Baiduspider in is the start. Getting cited is the job.
Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.
