Technical AI discoverability

Updated: 23.07.2026

Content-Signal

Content-Signal is a Cloudflare-originated, emerging robots.txt policy directive. Its core signals are search, ai-input, and ai-train; omitted signals express no preference. An additional use field for retention or reuse is still an evolving extension. These signals state policy preferences and require crawler cooperation or separate enforcement.

In plain terms

search covers building a search index and returning links or short excerpts; it does not include AI-generated summaries.

ai-input covers using content as real-time input for RAG, grounding, or generated search answers; ai-train covers training or fine-tuning models.

A Content-Signal line does not itself block network access. Access control still comes from Allow/Disallow rules and, when needed, edge or application enforcement.

Why it matters

It lets brands invite the uses they want, like citation and grounding, while expressing preferences about others, without shutting crawlers out entirely.

As content-usage norms mature, an explicit signal documents your intent for the agents that respect it.

How it relates to GetCited.me

GetCited.me publishes Content-Signal in its own robots.txt, but the current customer-site GEO audit does not parse or score Content-Signal directives. It audits named crawler Allow/Disallow rules and fetches robots.txt, /llms.txt, and /ai.txt.

Free tools for this term

Stop reading, start checking. No signup, no credit card.

Related terms

robots.txt (for AI crawlers)robots.txt controls which compliant crawlers may fetch specified URL paths; different AI and search uses may rely on different crawler tokens.llms.txtllms.txt is a proposed Markdown file that gives compatible AI tools a curated map of a site's important content.

See how AI answers about your market

Free scan: market map, AI visibility snapshot, and a GEO score in about a minute. No card, no call.