EU AI Act Article 53(1)(c) TDM Opt-Out — researched
EU AI Act Article 53(1)(c) is the GPAI-provider duty to maintain a copyright policy that identifies and honours text-and-data-mining (TDM) rights reservations made under DSM Directive Art. 4(3). Machine-readable means (Recital 18) is the operative standard for online content; robots.txt/RFC 9309 is the only protocol named in the GPAI Code of Practice's Copyright chapter (Measure 1.3), while W3C TDMRep, Cloudflare Content Signals and IETF aipref are credible but unrecognised candidates, and llms.txt/ai.txt are not opt-outs at all. Kneschke v LAION (OLG Hamburg) held a reservation must be machine-actionable, not just machine-readable; GEMA v OpenAI shows the TDM exception does not shelter memorisation/regurgitation. Sourced from 3 files (1 skill + 2 references) in the eu-ai-act-tdm-opt-out skill, covering the DSM/AI-Act legal framework, the Code of Practice, case law, and the protocol layer.
Definitions
- **The one-line mental model:** DSM Art. 4(3) creates the *right to reserve*. AI Act Art. 53(1)(c) creates a *public-law duty on the model provider to go looking for that reservation and honour it* — and extends that duty extraterritorially. The Code of Practice tells you what "looking" concretely means. Nobody has yet agreed what a reservation must *look like*, and that unresolved gap is the whole story. [source] — core one-line mental model
- The AI Preferences WG is the standards-track attempt to unify the space: `draft-ietf-aipref-vocab` (a vocabulary with `train-ai`/`search` categories and y/n values) and `draft-ietf-aipref-attach` (a `Content-Usage` HTTP header and a `Content-Usage` `robots.txt` rule).[^21][^22] `verified-as-of: 2026-09-02` — vocab at rev -06 (Proposed Standard track, updated April 2026); attach adopted with an August 2026 IESG milestone but its published revision has lapsed under the six-month rule. [source]
Structure and components
- | Measure | Commitment | |---|---| | **1.1** Copyright policy | Draw up, keep up to date and implement a copyright policy, in **one single document**; assign internal responsibility. Publishing a summary is encouraged, not required.[^10][^11] | | **1.2** Lawful crawling only | Do not circumvent effective technological measures under Art. 6(3) of Dir. 2001/29/EC — expressly incl. *"technological denial or restriction of access imposed by subscription models or paywalls"*; exclude sites *"recognised as persistently and repeatedly infringing copyright ... on a commercial scale by courts or public [source] — the 5 measures
How it works
- - **"expressly reserved"** — silence is consent. TDM is lawful *by default*; the burden of action sits on the rightsholder. This is an **opt-out, not an opt-in**, regime. - **"by their rightholders"** — only the rightsholder (or their agent) can reserve. A platform host reserving on content it does not own is doing something else. - **"in an appropriate manner"** — the general standard. Offline/non-public content can be reserved by contract or unilateral declaration. - **"such as machine-readable means"** — grammatically an *example*, but Recital 18 turns it into the operative requirement for [source]
- The Commission has run the process Code Measure 1.3(b) anticipates: a Dec 2025–Jan 2026 consultation (153 responses) followed by stakeholder workshops from **2 June 2026**, co-chaired by the AI Office and the Media Policy directorate. Three front-runners were identified — **TDMRep, C2PA-based assertions, and JPEG Trust Core Foundation V2** — and critics argue **none meets the threshold of technical maturity, scalability, and open governance** that would justify even an interim endorsement.[^26] `verified-as-of: 2026-09-02; no protocol had been endorsed as of this date.` [source]
- German scholarship argues Art. 53(1)(c) is a *Schutzgesetz* (protective statute) under §823(2) BGB: a public-law product-safety provision with **third-party protective effect**, because it protects individual rightsholders' interests (not just market order) and Recital 106 shows protective intent. That opens a **tort claim by rightsholders directly against the provider**, independent of the AI Office.[^18] `[QUALIFIED — a well-argued academic position, not yet a decided holding.]` Practical upshot: compliance exposure is not capped at the regulator's appetite for enforcement. [source]
- The lesson for both sides: **a reservation is a legal act, not a technical control.** It creates liability when ignored; it does not prevent the crawl. Enforcement needs WAF/bot management (technical) plus the Art. 53(1)(c) / Art. 4 claim (legal). [source]
- Art. 2(2) defines TDM broadly: *"any automated analytical technique aimed at analysing text and data in digital form in order to generate information which includes but is not limited to patterns, trends and correlations"*.[^1] Web-scraping to train a model is TDM.[^15] [source]
Parameters and configuration
- - **GPAI model**: presumed where training compute exceeds **10²³ FLOP** and the model can generate language/image/video and competently perform a wide range of distinct tasks (Commission GPAI Guidelines, 18 July 2025).[^33] - **Downstream modifiers** become providers in their own right where the modification's training compute exceeds roughly **one third** of the GPAI presumption threshold.[^33] Fine-tuning at scale is not a safe harbour. - **Open-source models are NOT exempt from 53(1)(c).** Art. 53(2) exempts only 53(1)(a) and (b) (technical documentation) for free-and-open-source models — a [source] — who is bound thresholds
- | Date | What happens | |---|---| | 1 Aug 2024 | AI Act enters into force | | **2 Aug 2025** | **GPAI obligations incl. Art. 53(1)(c) and (d) apply** to models placed on the market from this date[^2][^4][^13] | | **2 Aug 2026** | Commission/AI Office **enforcement powers** over GPAI begin[^13][^34] | | **2 Aug 2027** | Models placed on the market **before** 2 Aug 2025 must be brought into compliance[^13][^34] | [source] — compliance date deadlines
How-to and procedures
- 1. **Write the policy** — one document (Measure 1.1), owner named, covering national implementations, not just the Directive. 2. **Gate on lawful access** — no TPM/paywall circumvention; exclude the EU piracy list (Measure 1.2). Note this is a *prior* gate: no lawful access, no Art. 4 exception at all. 3. **Honour `robots.txt` per RFC 9309** at crawl time, with a declared, documented user-agent. 4. **Read supplementary signals** — TDMRep (all three methods), Content Signals, `X-Robots-Tag`. "State of the art" is a ratchet; reading only `robots.txt` gets weaker every year. 5. **Snapshot the evi [source] — GPAI compliant ingest pipeline
- 1. **`robots.txt` (mandatory).** The only named protocol. Enumerate AI training agents (§4.2) and re-review quarterly as new agents appear. 2. **Cloudflare Content Signals** (or the equivalent directive by hand) — adds explicit use-scoping (`ai-train=no`) that `robots.txt` cannot express, and is self-declared as an Art. 4 reservation. 3. **W3C TDMRep**, all three methods where feasible — the only purpose-built Art. 4(3) instrument, and the strongest evidence of an *express* reservation. Header form is the spec's preference; add `/.well-known/tdmrep.json` for path scoping. 4. **`X-Robots-Tag` / [source] — rightsholder reservation stack
- **Practical rule:** `llms.txt` and a TDM reservation are orthogonal. If you want both discoverability for inference *and* a training reservation, publish `llms.txt` **and** a `robots.txt` training block **and** a TDMRep/Content-Signals reservation. Never let an `llms.txt` stand in for an opt-out. See `document-formats` → `references/llms-txt.md` for the file itself. [source] — layer llms.txt + robots.txt + TDMRep
Measurements and reference values
- Art. 101: the **Commission** (via the AI Office) may fine GPAI providers up to **3 % of total worldwide annual turnover or €15,000,000, whichever is higher**, for intentional or negligent infringement, for failing to supply requested information (Art. 91), for refusing Art. 93 measures, or for denying model access for evaluation (Art. 92).[^6] For GPAI models, enforcement is **exclusively** the Commission's — not national market-surveillance authorities. [source]
Problems, failure modes and limitations
- 1. **For online content, machine-readable is effectively mandatory** — "only ... appropriate". 2. **Recital 18 explicitly lists "terms and conditions of a website"** among machine-readable means. This is the textual hook for the argument that natural-language T&Cs qualify — the argument that won at first instance and lost on appeal in *Kneschke v LAION* (§5). 3. **A TDM reservation is use-scoped, not access-scoped**: *"Other uses should not be affected"*. Reserving TDM rights does not reserve search indexing, quotation, or any other use. This is precisely the distinction `robots.txt` cannot ex [source]
- | Anti-pattern | Why it is wrong | |---|---| | "We published `llms.txt`, so we've opted out." | `llms.txt` is an inference-time index, explicitly not an opt-out; it arguably signals the opposite.[^25] §4.8 | | "The Code of Practice names `llms.txt`." | It does not. Only `robots.txt`/RFC 9309 is named — verified against the Code's full text and three independent specialist readings.[^9] §3.3 | | "Our T&Cs say no AI training, that's enough." | Rejected on appeal in *Kneschke v LAION*; needs to be machine-**actionable**.[^17] §5.1 | | "We're open-source, so Art. 53 doesn't apply." | Art. 53(2) ex [source]
- - It expresses **crawl access**, not **use rights**. Art. 4(3) reserves a *use*; Recital 18 says other uses are unaffected. `robots.txt` cannot say "search yes, training no" natively — one bit per crawler. - It is **enumerative**: you must name every agent, and new ones appear constantly. Agents that do not declare a distinct UA cannot be addressed at all. - It requires **control of the domain root**. A photographer whose work is syndicated to third-party sites cannot reserve rights there. This is the "domain-based opt-out" problem.[^31] - It is **prospective only**: it blocks a future crawl, [source]
- Two lessons: (1) the TDM exception covers the *mining*, and does not automatically launder what ends up **inside the model** or **in its output** — which is exactly the risk Code Measure 1.4 targets; (2) AI Act compliance and national copyright liability are separate tracks. `[QUALIFIED — first-instance national ruling; appeal and CJEU divergence possible.]` [source]
- Art. 4 was implemented member state by member state — Germany's `UrhG §44b` (with the research exception in `§60d`) predates and modelled the EU text. A GPAI provider's copyright policy must comply with **national implementations**, not just the Directive; the Code of Practice makes signatories responsible for verifying this themselves.[^11] Divergence in national wording is a live compliance risk. [source]
- Strengths: says *TDM specifically* (not crawl access), carries a licensing pointer (opt-out plus a route to a deal), works per-asset. Weaknesses: **adoption is negligible** — about 45 hosts serving `tdmrep.json` as of Jan 2024, ~60 domains using the meta tag; under 3 % of EU news publishers.[^26][^31] It requires checking several layers at once, metadata is stripped in transit, and it cannot distinguish crawler types.[^26] [source]
- **Contested, and you should say so.** A recital is interpretive, not operative; it cannot by itself extend the territorial scope of substantive copyright law. The defensible framing: 53(1)(c) is a *product-compliance* obligation attaching to market placement, not a retroactive application of EU copyright to foreign acts of reproduction.[^18][^13] Commentators also note limited practical effect on non-EU developers to date.[^32] [source]
- - **Rightsholders**: a broad EU cultural/creative coalition rejected the claim that the Code strikes a fair balance, and called the Art. 53(1)(d) template *"alarmingly superficial"* — insufficient to let a rightsholder verify whether their works were used.[^32] - **Industry**: Meta declined to sign; CCIA-aligned criticism argues the Code goes beyond the AI Act's text.[^35] - **Neutral**: the Code creates process obligations without resolving the substantive question of what a valid reservation *is*.[^15][^26] [source] — criticism from both directions
- **Why it belongs in this reference:** (d) is the *evidentiary* counterpart to (c). Without it, a rightsholder cannot tell whether their reservation was honoured; with it, the crawler disclosure required by Code Measure 1.3(4) plus the (d) summary make the opt-out claim checkable. Rightsholder groups argue the template is far too shallow to serve that function.[^32] Like (c), **(d) survives the open-source carve-out**.[^2][^3] [source]
- Two honest caveats: Cloudflare's own framing is a **vendor legal assertion**, not a judicial or regulatory holding; and Cloudflare itself says signals *"express preferences; they are not technical countermeasures against scraping"* and *"some companies might simply ignore them"*. Pair with WAF/Bot Management for actual enforcement.[^23] [source] — Content Signals legal-weight caveats
- - Cloudflare (4 Aug 2025) documented **Perplexity using undeclared stealth crawlers** — impersonating Chrome-on-macOS user agents, rotating ASNs, and ignoring or not even fetching `robots.txt` — generating **3–6 million requests per day across tens of thousands of domains** against sites that had explicitly blocked its declared bots. Cloudflare de-listed Perplexity as a verified bot and shipped blocking heuristics.[^29] - Cloudflare also reports **over 2.5 million websites** have opted out of AI training via its tools.[^29] [source] — Perplexity stealth-crawler evidence
- **The Digital Omnibus did NOT delay any of this.** The omnibus (proposed 19 Nov 2025) deferred **high-risk** obligations (Annex III to 2 Dec 2027; Annex I to 2 Aug 2028). GPAI obligations under Arts. 51–56, including 53(1)(c), were expressly **unchanged**.[^27] Assuming "the AI Act got delayed" is a live and expensive error. [source]
Comparisons and alternatives
- | Signal | Standards status | Named in the Code? | Practical weight | |---|---|---|---| | `robots.txt` (RFC 9309) | **IETF RFC** (2022) | **Yes, by name + RFC no.**[^9] | **Highest.** The de-facto compliance baseline | | `X-Robots-Tag` header | De-facto, non-standard for TDM | No | Useful supplement; honoured for indexing, not TDM-specific | | `<meta name="robots">` / `noai` | De-facto (`noai` from DeviantArt) | No | Weak; per-page, narrow honouring | | **W3C TDMRep** | CG **Final Report** (10 May 2024) — *not* a W3C Recommendation | No | Purpose-built for Art. 4(3); ~45–60 deploying hosts[^26 [source] — protocol recognition scorecard
- - **`robots.txt` / RFC 9309 is the ONLY protocol named in the entire Copyright chapter.** Verified by full-text search of the final Code: zero occurrences of `llms.txt`, `ai.txt`, `TDMRep`, or `C2PA`.[^9] Corroborated independently by three specialist readings of the chapter.[^10][^11][^36] Several law-firm and press summaries nonetheless assert the Code names `llms.txt`; **it does not.** Treat any such claim as a summarisation error. `[FACT — 4 independent sources; one secondary source dissents and is assessed as wrong.]` - Limb (b) is a **forward-looking, conditional** commitment: a protocol [source]
- - **Not named in the Code of Practice.** Full-text verification of the final Code returns zero occurrences of `llms.txt`.[^9] Multiple law-firm and press summaries claim otherwise; they are wrong. - **Its own specification disclaims the role.** llmstxt.org distinguishes it from `robots.txt` — *"robots.txt lets automated tools know what access to a site is considered acceptable ... llms.txt information is instead used on demand, when an agent needs information"* — and states the expectation that it is *"mainly ... useful for inference rather than training"*. It contains no access-control, opt-o [source]
- | | **Art. 3** | **Art. 4** | |---|---|---| | Beneficiary | Research organisations & cultural heritage institutions | **Anyone**, incl. commercial actors | | Purpose | Scientific research only | Any purpose | | Precondition | *lawful access* | *lawfully accessible* | | Opt-out? | **No** — cannot be reserved away | **Yes** — Art. 4(3) | | Contract override? | Barred (Art. 7(1)) | Not barred | [source] — Art 3 vs Art 4 table
- This is the first widely-deployed signal that **separates use from access** — precisely the Art. 4(3) shape `robots.txt` alone lacks. Cloudflare states content signals are *"express reservations of rights under Article 4 of the European Union Directive 2019/790"* and pushed `search=yes, ai-train=no` into managed `robots.txt` across ~3.8 million domains.[^23] [source] — Content Signals vs robots.txt
- What is genuinely new versus DSM Art. 4(3): [source]
Open questions
- 1. **What is "machine-readable" in 2026+?** *Kneschke* judged 2021 technology. If LLMs can parse T&Cs, does the *OLG*'s machine-actionability test bend? No authority yet.[^16] 2. **Which protocol will the Commission bless?** TDMRep, C2PA and JPEG Trust are the named front-runners; critics say none is mature enough. IETF `aipref` may overtake all three.[^21][^26] 3. **Is Recital 106's extraterritorial reach legally sound?** A recital cannot itself extend substantive scope; the product-compliance framing is a workaround, not a ruling.[^18] 4. **Is Art. 53(1)(c) privately enforceable?** The Germa [source]
Facts and statements
- | If the question is about | Read | |---|---| | DSM Art. 3/4, Art. 4(3) wording, Recital 18, lawful access, national implementation | `references/eu-tdm-legal-framework.md` §1 | | AI Act Art. 53(1)(c) text, who is bound, open-source, Recital 106, dates, fines, private enforcement | `references/eu-tdm-legal-framework.md` §2 | | The GPAI Code of Practice Copyright chapter, Measures 1.1–1.5, signatories | `references/eu-tdm-legal-framework.md` §3 | | *Kneschke v LAION*, *GEMA v OpenAI*, what "machine-readable" means in case law | `references/eu-tdm-legal-framework.md` §5 | | The Art. 53(1)(d) tra [source]
- | Question | Answer | |---|---| | Does `robots.txt` count as a TDM opt-out? | **Yes — it is the only protocol named in the GPAI Code of Practice**, by name and RFC number.[^9] It is also the weakest fit conceptually (§4.2). | | Does `llms.txt` count? | **No.** It is not named in the Code, and its own spec disclaims access-control/opt-out purpose.[^9][^25] See §4.8. | | Does `ai.txt` count? | **Not recognised in any binding or quasi-binding EU instrument.** Spawning frames it as an Art. 4(3) reservation; no regulator, court, or the Code endorses it. See §4.7. | | Does W3C TDMRep count? | Purpos [source]
- Art. 4 is the one commercial AI training relies on — and the one Art. 53(1)(c) points at. Art. 4(4): *"This Article shall not affect the application of Article 3"*.[^1] That carve-out is why the LAION defence succeeded on Art. 3/§60d grounds (§5). [source]
- Art. 4(1) covers *"lawfully accessible"* works.[^1] Content behind a paywall, subscription, or acquired from a piracy source never reaches the Art. 4(3) question — the exception simply does not apply, opt-out or not. Code of Practice Measure 1.2 operationalises exactly this (§3.2). Circumventing a technological measure under Art. 6(3) InfoSoc is an independent wrong. [source]
- Per-resource, so they work for content you do not control the domain root of, and they can be applied to non-HTML assets via the header. But `noai`/`noimageai` are a de-facto convention (originating with DeviantArt), not standardised for TDM, and not named in the Code. Deploy as a supplement, never as the primary reservation. [source]
- - **A positive duty to go looking.** Art. 4(3) told rightsholders how to say no; 53(1)(c) tells providers they must actively *identify* those signals. - **A documented policy** — an artefact a regulator can demand and inspect, not just a litigable state of affairs. - **"State-of-the-art technologies"** — a moving standard. As detection technology improves, the compliance bar rises automatically. This is the clause that imports the whole protocol question in §4. - **Regulatory (not just civil) enforcement**, with fines (§2.5). - **Extraterritorial reach** (§2.3). [source]
- It asks for data modality, scale (size bands), languages, collection period, identifiers for the main public datasets, **web-crawler specifications and purpose**, whether commercially licensed data was used, and **how TDM opt-outs were honoured**. Narrative descriptions are used to balance transparency against trade secrets; summaries must be updated every six months or on material change.[^34] [source]
- This matters legally: an IETF standards-track output is exactly the kind of thing that satisfies Code Measure 1.3(b)'s "adopted by international standardisation organisations" limb — and RFC 9309's successor clause in 1.3(a) would pull a robots.txt-integrated `Content-Usage` rule straight into the named commitment. **Watch this space; it is the most likely resolution of the fragmentation problem.** [source]
- Hamann's JIPITEC analysis assessed seven proposed reservation protocols against Art. 4(3) and concluded that **only some qualify as "machine-readable" in a legal sense at all, and that "the proliferation of standards currently precludes any effective reservation of TDM rights"**.[^15] He frames Art. 4(3) as demanding two things simultaneously: *explicitness* (specific to given content and use) and *automatability* (a well-defined technical protocol). [source]
- | | Holding | |---|---| | **LG Hamburg**, 27 Sep 2024, 310 O 227/23 | LAION's use fell within the **scientific-research** TDM exception (`UrhG §60d` / DSM Art. 3), so no opt-out applied. *Obiter*, the court suggested a **natural-language reservation in website T&Cs could be machine-readable**, reasoning that LLMs can process unstructured text.[^16][^17] | | **OLG Hamburg**, 10 Dec 2025, 5 U 104/24 | Appeal dismissed — LAION could rely on the TDM/research exceptions. Critically, the court **rejected the lower court's machine-readability reasoning**: a reservation must be **machine-*actionable** [source] — Kneschke case-law holdings table
- Munich Regional Court I, **11 Nov 2025**, case 42 O 14139/24. GEMA sued over nine German song lyrics used to train GPT-4/4o. The court held that **memorisation** (retention of protected content in model weights) and **regurgitation** (reproduction in output) are both copyright-relevant reproductions **not sheltered by the TDM exceptions**, and that hallucinated additions did not render the lyrics unrecognisable. OpenAI was ordered to cease, pay damages, and provide information on the scope of use and revenue.[^28] [source]
- Spawning AI's 2023 root-level convention, explicitly framed by its author as satisfying the DSM Art. 4 machine-readable reservation requirement.[^14] But: it is not named in the Code, not adopted by any standards body, has near-zero measured adoption, and suffers a five-way name collision with unrelated `ai.txt` proposals. No regulator or court has endorsed it. **Deploy it if you like — it costs nothing — but never as your primary reservation.** File mechanics and the name-collision detail live in `document-formats` → `references/ai-txt.md`. [source]
- **Net effect:** the "state-of-the-art technologies" standard in Art. 53(1)(c) currently resolves, in practice, to `robots.txt` plus best-efforts on whatever else a provider chooses to read. [source]
- Note the deliberate split in vendor tokens — OpenAI documents **GPTBot** (training), **OAI-SearchBot** (ChatGPT search surfacing), **OAI-AdsBot** (ad landing-page validation, explicitly not used for training), and **ChatGPT-User** (user-triggered fetch, where OpenAI states robots.txt rules *may not apply*).[^24] Blocking `GPTBot` reserves training without removing you from ChatGPT search; that granularity is the point. [source]
- Announced **24 September 2025**, CC0. Extends `robots.txt` with a `Content-Signal` directive carrying three independent `yes|no` signals — `search`, `ai-input` (RAG/grounding), and `ai-train`:[^23] [source]
- **Do not over-read it.** The court judged 2021 capability. It gave little guidance on what machine-readability means in 2025+, when LLMs plainly *can* parse T&Cs. Commentators note the absence of a standardised vocabulary for "TDM" is what still blocks large-scale implementation, not the reading of text.[^16] `[QUALIFIED — first-instance obiter and one appellate ruling; not CJEU authority.]` [source]
- Published **10 July 2025**, drafted by Commission-appointed independent experts across a 1,400+ stakeholder consultation; three chapters — Transparency, Copyright, Safety & Security.[^7][^12] Under Art. 56 it is **voluntary**; signing is the presumption-of-conformity route and reduces administrative burden. Non-signatories must demonstrate compliance by other adequate means. Copyright + Transparency bind all GPAI providers; Safety & Security applies only to systemic-risk models.[^10] [source] — Code of Practice status
- Signing buys regulatory goodwill under the AI Act. It buys **no immunity** from a national copyright infringement claim — as *GEMA v OpenAI* demonstrated (§5.2). [source]
- 26+ organisations signed, incl. OpenAI, Google, Anthropic, Microsoft, Amazon, IBM, Mistral, Cohere and Aleph Alpha; endorsed by the Commission and the AI Board in August 2025. **Meta publicly refused** (July 2025, citing legal uncertainty). **xAI signed only the Safety & Security chapter** — i.e. *not* Copyright.[^35] `verified-as-of: 2026-09-02` — the signatory list is live; re-check the Commission's published list before citing. [source]
- What it *does* do: bite on every subsequent crawl, and — via Art. 53(1)(c) — attach a continuing regulatory duty to any provider placing a model on the EU market. Pre-2 Aug 2025 models get until **2 Aug 2027**. Note the separate route *GEMA* opens: even where the *mining* was covered, **retention in weights and reproduction in output** can be independently infringing (§5.2). [source]
- `[TENTATIVE — sources conflict on the omnibus's own adoption status: one law-firm tracker described it as still pending formal adoption and OJ publication, while other secondary reporting describes it as adopted and in force. Re-verify the omnibus's status before citing it. What is NOT in dispute across sources: whatever its status, it does not move the Art. 53 GPAI dates.]` [source]
- Art. 53(1)(d) requires providers to *"draw up and make publicly available a sufficiently detailed summary about the content used for training ... according to a template provided by the AI Office"*.[^2][^3] The **mandatory template was published 24 July 2025**.[^34] [source]
Quotes
- > *"(1) ... Signatories commit:* > *(a) to employ web-crawlers that read and follow instructions expressed in accordance with > the **Robot Exclusion Protocol (robots.txt), as specified in the Internet Engineering Task > Force (IETF) Request for Comments No. 9309**, and any subsequent version of this Protocol > for which the IETF demonstrates that it is technically feasible and implementable ... and* > *(b) to identify and comply with **other appropriate machine-readable protocols** to express > rights reservations pursuant to Article 4(3) of Directive (EU) 2019/790, **for example > through as [source] — Measure 1.3 verbatim
- > *"(c) put in place a policy to comply with Union law on copyright and related rights, and > in particular to identify and comply with, **including through state-of-the-art > technologies**, a reservation of rights expressed pursuant to Article 4(3) of Directive > (EU) 2019/790"*.[^2][^3][^4] [source] — Art 53(1)(c) verbatim
- > *"In the case of content that has been made publicly available online, it **should only be > considered appropriate to reserve those rights by the use of machine-readable means, > including metadata and terms and conditions of a website or a service**. Other uses should > not be affected by the reservation of rights for the purposes of text and data mining. In > other cases, it can be appropriate to reserve the rights by other means, such as > contractual agreements or a unilateral declaration."*[^1] [source] — Recital 18 verbatim
- > *"The exception or limitation provided for in paragraph 1 shall apply on condition that > the use of works and other subject matter referred to in that paragraph **has not been > expressly reserved by their rightholders in an appropriate manner, such as > machine-readable means in the case of content made publicly available online**."*[^1] [source] — Art 4(3) verbatim
- > *"Any provider placing a general-purpose AI model on the Union market should comply with > this obligation, **regardless of the jurisdiction in which the copyright-relevant acts > underpinning the training of those general-purpose AI models take place**. This is > necessary to ensure a level playing field ... where no provider should be able to gain a > competitive advantage in the Union market by applying lower copyright standards than those > provided in the Union."*[^5] [source] — Recital 106 verbatim
Related concepts
- machine-readable — is a related of EU AI Act Article 53(1)(c) TDM Opt-Out
- text and data mining — is a synonym of EU AI Act Article 53(1)(c) TDM Opt-Out
- Article 4 — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- robots.txt — is a hyponym of EU AI Act Article 53(1)(c) TDM Opt-Out
- GPAI Code of Practice — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- Article 4(3) — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- DSM Directive — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- W3C TDMRep — is a hyponym of EU AI Act Article 53(1)(c) TDM Opt-Out
- Article 53(1)(d) — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- Cloudflare Content Signals — is a hyponym of EU AI Act Article 53(1)(c) TDM Opt-Out
- llms.txt — is a hyponym of EU AI Act Article 53(1)(c) TDM Opt-Out
- Recital 106 — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- Measure 1.3 — is a part of EU AI Act Article 53(1)(c) TDM Opt-Out
- AI Office — is a related of EU AI Act Article 53(1)(c) TDM Opt-Out
- X-Robots-Tag — is a hyponym of EU AI Act Article 53(1)(c) TDM Opt-Out
- ai.txt — is a hyponym of EU AI Act Article 53(1)(c) TDM Opt-Out
- Kneschke v LAION — is a related of EU AI Act Article 53(1)(c) TDM Opt-Out
- GEMA v OpenAI — is a related of EU AI Act Article 53(1)(c) TDM Opt-Out
- GPAI provider — is a related of EU AI Act Article 53(1)(c) TDM Opt-Out
- rights reservation — is a synonym of EU AI Act Article 53(1)(c) TDM Opt-Out