Meta · Search indexes
meta-webindexer
Meta’s crawler for Meta AI search quality. If it cannot index you, Meta AI search has a weaker view of the site.
- Operator
- Meta
- Traffic type
- Search indexes
- Verification
- Not independently verifiable
- robots.txt
- Usually honors robots.txt
Usually means: Crawlers that build the indexes ChatGPT search, Claude search, Perplexity, Google, and Bing draw from. Blocking them is how brands disappear from AI answers without noticing.
User-agent
robots.txt and most WAF rules match the token, not the full Mozilla string. Operators often wrap meta-webindexer in extra product or version text.
meta-webindexer
What meta-webindexer does
meta-webindexer crawls to improve Meta AI search results. It is an index bot, not facebookexternalhit (link previews).
How to get discovered
Allow meta-webindexer on public content. Keep entity-clear titles and schema so Meta can associate the page with the right brand.
Why meta-webindexer might skip you
Blanket meta-* Disallows aimed at training (meta-externalagent) that also hit the indexer, or bot fights against Facebook’s ASN.
robots.txt rule
Allow this token on public pages you want retrieved or cited. A CDN “Block AI bots” toggle can still 403 it after robots.txt says Allow.
# Keep meta-webindexer eligible to fetch this site
User-agent: meta-webindexer
Allow: /Meta has not published a machine-readable range file for this token. Do not treat the user-agent alone as proof of identity.
JavaScript: render behavior is not documented.
Other Meta bots
Training, search, and user-fetch tokens from the same operator are not interchangeable. Allow the discovery path even when you opt out of training.
meta-webindexer FAQ
What is meta-webindexer?
meta-webindexer crawls to improve Meta AI search results. It is an index bot, not facebookexternalhit (link previews).
Should I allow meta-webindexer in robots.txt?
Yes if you want Meta to be able to fetch and cite this site. Blocking meta-webindexer is how pages stay invisible to that product even when Googlebot still crawls them.
How do I know a request is really meta-webindexer?
Meta has not published a range file for this token. Treat the user-agent as a claimed identity, and do not build a hard IP allowlist from blog posts or screenshots.
Letting meta-webindexer in is the start. Getting cited is the job.
Rankealo checks whether AI crawlers can reach you, then publishes pages built to be retrieved and quoted in ChatGPT, Claude, Perplexity, and Gemini.
