An AI crawler is an automated bot that fetches web content on behalf of an AI product — to ground search results, power real-time answers, or train models. Each major AI company operates one or more named crawlers identified by user-agent.
In plain terms
Crawler tokens can serve different purposes: search indexing, live user-initiated retrieval, grounding, dataset collection, or model training.
A blocked crawler cannot use that crawl path, but the associated product may still know a page through a separate search index, previous training, licensed data, or another documented retrieval mechanism.
Crawler policy should therefore be based on each provider's current documentation, not on the assumption that one token controls every product from that company.
Why it matters
Knowing a bot's documented purpose helps a site owner separate search indexing, live retrieval, grounding, and training policy.
Crawler access is one part of eligibility and policy, not proof that a product has fetched, indexed, or cited the site.