Allen AI · Training crawlers
AI2Bot
Allen Institute crawler for AI research datasets. Blocking it is a research-crawl opt-out, not a ChatGPT ranking move.
- Operator
- Allen AI
- Traffic type
- Training crawlers
- Verification
- Not independently verifiable
- robots.txt
- Honors robots.txt
Usually means: Bots that collect public pages for future models. Blocking them opts you out of training. It does not, by itself, remove you from live AI search citations.
User-agent
robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap AI2Bot in extra product or version text.
AI2Bot
What AI2Bot does
AI2Bot finds documents for Allen AI research systems. It is closer to Common Crawl than to PerplexityBot.
How to get discovered
Allow AI2Bot if you want research corpora to include you. It will not put you in ChatGPT search.
Why AI2Bot might skip you
robots Disallow, or research crawlers skipping low-authority domains.
robots.txt rule
This snippet opts out of training-style collection. Keep the search and user-fetch tokens from the same operator allowed if you still want citations.
# Optional training opt-out — does not remove you from AI search by itself
User-agent: AI2Bot
Disallow: /Allen AI has not published a machine-readable range file for this token. Do not treat the user-agent alone as proof of identity.
Sources
JavaScript: reads raw HTML only.
AI2Bot FAQ
What is AI2Bot?
AI2Bot finds documents for Allen AI research systems. It is closer to Common Crawl than to PerplexityBot.
Should I allow AI2Bot in robots.txt?
Only if you want to opt out of training-style collection. Blocking AI2Bot does not, by itself, remove you from live AI search citations — those use the search and user-fetch tokens from Allen AI.
How do I know a request is really AI2Bot?
Allen AI has not published a range file for this token. Treat the user-agent as a claimed identity, and do not build a hard IP allowlist from blog posts or screenshots.
Letting AI2Bot in is the start. Getting cited is the job.
Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.
