robots.txt and the Content-Signal AI-preference extension — researched

robots.txt is the Robots Exclusion Protocol (RFC 9309, Sept 2022, Proposed Standard): a crawl-access file of User-agent groups with Allow/Disallow rules, MUST-level wildcards (*, $), octet-based longest-match precedence with an allow-wins SHOULD tie-break, a 500 KiB parsing floor, and explicit non-goals (it is not access authorization). Riding inside it is a newer, unstandardized AI-preference layer — Cloudflare's Content-Signal (search/ai-input/ai-train, plus a testing content-use field) and its IETF successor draft (Content-Usage, AIPREF) — which states what a crawler may do with content after fetching it, has no announced enforcing consumer, and is evidenced (Perplexity dispute, Consent in Crisis study, Google's 'no effects whatsoever') to be a stated preference rather than a technical control. Drawn from two /dr-authored references (RFC 9309 spec analysis and the Content-Signal/AIPREF layer), 34 sourced footnotes.

Definitions

Structure and components

How it works

How-to and procedures

Measurements and reference values

Problems, failure modes and limitations

Comparisons and alternatives

Changes and history

Facts and statements

Related concepts