Technical AI discoverability

Updated: 23.07.2026

AI crawler

An AI crawler is an automated bot that fetches web content on behalf of an AI product — to ground search results, power real-time answers, or train models. Each major AI company operates one or more named crawlers identified by user-agent.

In plain terms

Crawler tokens can serve different purposes: search indexing, live user-initiated retrieval, grounding, dataset collection, or model training.

A blocked crawler cannot use that crawl path, but the associated product may still know a page through a separate search index, previous training, licensed data, or another documented retrieval mechanism.

Crawler policy should therefore be based on each provider's current documentation, not on the assumption that one token controls every product from that company.

Why it matters

Knowing a bot's documented purpose helps a site owner separate search indexing, live retrieval, grounding, and training policy.

Crawler access is one part of eligibility and policy, not proof that a product has fetched, indexed, or cited the site.

How it relates to GetCited.me

GetCited.me's audit reports access status for a fixed list of named bots and the wildcard block. It does not infer that every allowed bot actually fetched the site or that every blocked token removes the site from all AI answers.

Where this shows up in practice

Per-engine reference pages: what the vendor documents about attribution, and what is not documented.

Free tools for this term

Stop reading, start checking. No signup, no credit card.

Related terms

robots.txt (for AI crawlers)robots.txt controls which compliant crawlers may fetch specified URL paths; different AI and search uses may rely on different crawler tokens.CrawlabilityCrawlability is how easily bots, including AI crawlers, can access and parse a site's content.

See how AI answers about your market

Free scan: market map, AI visibility snapshot, and a GEO score in about a minute. No card, no call.