Cohere · Training crawlers
cohere-ai
Cohere’s public-content crawler token. Blocking it opts you out of Cohere collection; it is not a ChatGPT citation switch.
- Operator
- Cohere
- Traffic type
- Training crawlers
- Verification
- Not independently verifiable
- robots.txt
- Usually honors robots.txt
Usually means: Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.
User-agent
robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap cohere-ai in extra product or version text.
cohere-ai
What cohere-ai does
cohere-ai identifies Cohere fetches of public pages. cohere-training-data-crawler is the more explicit training crawler name.
How to get discovered
Cohere is not a consumer answer engine Rankealo measures. Allow or block based on training preference, not GEO citations.
Why cohere-ai might skip you
Disallow, WAF, or simply low crawl priority.
robots.txt rule
This snippet opts out of training-style collection. Keep the search and user-fetch tokens from the same operator allowed if you still want citations.
# Optional training opt-out — does not remove you from AI search by itself
User-agent: cohere-ai
Disallow: /Cohere has not published a machine-readable range file for this token. Do not treat the user-agent alone as proof of identity.
JavaScript: reads raw HTML only.
Other Cohere bots
Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.
cohere-ai FAQ
What is cohere-ai?
cohere-ai identifies Cohere fetches of public pages. cohere-training-data-crawler is the more explicit training crawler name.
Should I allow cohere-ai in robots.txt?
Only if you want to opt out of training-style collection. Blocking cohere-ai does not, by itself, remove you from live AI search citations — those use the search and user-fetch tokens from Cohere.
How do I know a request is really cohere-ai?
Cohere has not published a range file for this token. Treat the user-agent as a claimed identity, and do not build a hard IP allowlist from blog posts or screenshots.
Letting cohere-ai in is the start. Getting cited is the job.
Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.
