Alibaba · Training crawlers
QwenBot
Alibaba’s Qwen training-style crawler. Keep Qwen-User allowed if you still want live Qwen answers to be able to fetch you.
- Operator
- Alibaba
- Traffic type
- Training crawlers
- Verification
- Not independently verifiable
- robots.txt
- Usually honors robots.txt
Usually means: Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.
User-agent
robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap QwenBot in extra product or version text.
QwenBot
What QwenBot does
QwenBot collects public content for Qwen systems. Qwen-User is the live answer fetch.
How to get discovered
Citations come from Qwen-User (and any index bot), not from QwenBot. Split the robots groups the same way you split GPTBot vs OAI-SearchBot.
Why QwenBot might skip you
A Qwen* Disallow that also matched Qwen-User, or regional blocking.
robots.txt rule
This snippet opts out of training-style collection. Keep the search and user-fetch tokens from the same operator allowed if you still want citations.
# Optional training opt-out — does not remove you from AI search by itself
User-agent: QwenBot
Disallow: /Alibaba has not published a machine-readable range file for this token. Do not treat the user-agent alone as proof of identity.
JavaScript: reads raw HTML only.
Other Alibaba bots
Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.
QwenBot FAQ
What is QwenBot?
QwenBot collects public content for Qwen systems. Qwen-User is the live answer fetch.
Should I allow QwenBot in robots.txt?
Only if you want to opt out of training-style collection. Blocking QwenBot does not, by itself, remove you from live AI search citations — those use the search and user-fetch tokens from Alibaba.
How do I know a request is really QwenBot?
Alibaba has not published a range file for this token. Treat the user-agent as a claimed identity, and do not build a hard IP allowlist from blog posts or screenshots.
Letting QwenBot in is the start. Getting cited is the job.
Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.
