Cohere · Training crawlers
cohere-training-data-crawler
Cohere’s named training-data crawler. A Disallow here is a clear training opt-out.
- Operator
- Cohere
- Traffic type
- Training crawlers
- Verification
- Not independently verifiable
- robots.txt
- Usually honors robots.txt
Usually means: Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.
User-agent
robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap cohere-training-data-crawler in extra product or version text.
cohere-training-data-crawler
What cohere-training-data-crawler does
cohere-training-data-crawler collects public pages for Cohere training data. It is not an answer-index bot.
How to get discovered
You do not get consumer-assistant citations via this crawler. Use it only as a robots policy decision.
Why cohere-training-data-crawler might skip you
An explicit Disallow, which is the usual intent when this token is named at all.
robots.txt rule
This snippet opts out of training-style collection. Keep the search and user-fetch tokens from the same operator allowed if you still want citations.
# Optional training opt-out — does not remove you from AI search by itself
User-agent: cohere-training-data-crawler
Disallow: /Cohere has not published a machine-readable range file for this token. Do not treat the user-agent alone as proof of identity.
JavaScript: reads raw HTML only.
Other Cohere bots
Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.
cohere-training-data-crawler FAQ
What is cohere-training-data-crawler?
cohere-training-data-crawler collects public pages for Cohere training data. It is not an answer-index bot.
Should I allow cohere-training-data-crawler in robots.txt?
Only if you want to opt out of training-style collection. Blocking cohere-training-data-crawler does not, by itself, remove you from live AI search citations — those use the search and user-fetch tokens from Cohere.
How do I know a request is really cohere-training-data-crawler?
Cohere has not published a range file for this token. Treat the user-agent as a claimed identity, and do not build a hard IP allowlist from blog posts or screenshots.
Letting cohere-training-data-crawler in is the start. Getting cited is the job.
Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.
