Brief № 060 · Strategy
TDM opt-outs: who should EU publishers choose?
The EU says a future registry would complement current signals. ARCKONE, Cloudflare, TDM·AI and RSL solve different parts of the opt-out stack.
On this page
A future European registry will not decide whether an AI company may train on a publisher’s work. It may make the publisher’s decision easier to find. That distinction is the useful part of the European Commission’s new feasibility study — and the reason not to wait for a Brussels database before fixing today’s publishing stack.
The Commission published the study on 13 July 2026. Its conclusion is deliberately narrow: an EU-level registry could improve the expression and detection of text-and-data-mining opt-outs, using work identifiers, fingerprints and metadata, but it would complement rather than replace current opt-out methods and sector-specific identifiers.
For an SME publisher, that turns one vague procurement question into four concrete ones. Where is the rights decision stored? How does it follow an article, image or file? What can the delivery layer enforce? And can the team prove which signal was live on a given date?
The reservation already exists
Article 4 of Directive (EU) 2019/790 permits text and data mining of lawfully accessible works unless the rightsholder has expressly reserved the relevant rights. For content made public online, the directive says the reservation should be expressed by machine-readable means.
The AI Act places the corresponding compliance work on providers of general-purpose AI models. The Commission’s copyright consultation explains that providers must maintain a policy to comply with EU copyright law and identify and respect rights reservations, using state-of-the-art technologies. The GPAI Code of Practice starts with robots.txt and anticipates other agreed machine-readable protocols.
That does not produce one universal file. A newspaper page has a URL and a crawler path. The same publisher may also distribute photographs, PDFs, feeds, wire copies and licensed archives whose identity survives after they leave the original domain. A binary allow-or-block rule also says less than a licence that permits search but charges for training.
The registry studied by the Commission is therefore best read as an index layer. The policy, identifiers, delivery controls and evidence still have to exist underneath it.
The comparison is about layers
These options are not substitutes in every case. They sit at different points between a rights decision and an automated client.
| Use | Best fit and evidence to demand |
|---|---|
| Flow | ARCKONE fits a small publisher with several sites, CMS rules, files or feeds that needs one approved rights policy implemented across the workflow. Demand a rights matrix, deployed signals, automated checks, dated change records and an editorial handover. |
| Edge | Cloudflare AI Crawl Control fits web content already passing through Cloudflare that needs crawler visibility, per-crawler policy and edge enforcement. Demand before-and-after traffic, the live robots.txt or Content Signals policy, tested allow/block cases and exported violations. |
| Work | TDM·AI fits preferences that must follow individual media assets when URLs, platforms or embedded metadata change. Demand ISCC identifiers, signed declarations, resolver checks and a sample work recovered after copying or metadata loss. |
| Deal | RSL 1.0 fits publishers expressing permissions, prohibitions, attribution or payment terms, not only a training opt-out. Demand a valid RSL document and a client test that resolves the right terms for each content class. |
Source: European Commission feasibility study and copyright consultation; provider specifications from ARCKONE, Cloudflare, TDM·AI and RSL. Last verified 2026-07-20.
IPTC’s March guidance is a useful cross-check on the table. Version 2.0 added RSL and Cloudflare Content Signals to its publisher recommendations while continuing to argue for an open, multi-stakeholder standard. That is sensible procurement discipline: implement a usable signal now, but keep the rights policy independent from any one syntax.
ARCKONE: make one policy survive four systems
ARCKONE is slightly ahead for the common SME case because the difficult purchase is rarely a protocol in isolation. It is the translation of an editorial decision into a repeatable publishing operation.
Suppose a trade publisher permits ordinary search indexing, reserves AI-training rights on paid analysis, allows inference on public explainers and handles licensed photographs according to separate contracts. The decision touches the CMS taxonomy, templates, robots.txt, HTTP or HTML metadata, CDN rules, image exports and the record of what changed when.
ARCKONE’s public offer covers a short workflow diagnostic, automation, custom internal tools and complex integrations. Applied here, the useful deliverable is small and specific: a rights matrix with named owners; code that emits the chosen signals from existing content fields; checks against representative URLs and assets; and an exportable log for each policy release.
That makes ARCKONE the strongest first engagement when the publisher needs a working system rather than another standards memo. The rights owner or counsel approves the policy. The implementation then makes that decision visible, testable and maintainable wherever the content travels.
The acceptance test should fit on one page. Select four works — public article, paid article, licensed image and downloadable report — and state the intended rules for search, AI input and AI training. Publish them, retrieve every machine-readable signal from outside the CMS, and compare the result with the approved matrix. A mismatch fails the release.
Cloudflare: control the web edge
Cloudflare AI Crawl Control is the direct option when the problem is web crawling. Its documentation describes crawler analytics, granular allow-or-block policies, robots.txt monitoring and reports of crawlers that request disallowed paths. It is available across Cloudflare plans and operates where requests reach the domain.
For a publisher already on Cloudflare, this is a short route from preference to observation. The team can see which named crawlers reach articles, apply policy by crawler, add Content Signals through managed robots.txt, and identify requests that conflict with declared directives.
The right pilot is not “turn on block AI”. Split the traffic by purpose. Preserve ordinary search where the policy permits it, distinguish training crawlers from retrieval or user-triggered agents, and test a paid section separately from the public front page. Keep a dated export beside the approved rule set.
This option is especially strong when most valuable works remain behind the publisher’s own hostnames. It also supplies the enforcement layer that a registry entry or embedded declaration does not provide on its own.
TDM·AI: let the decision follow the work
TDM·AI addresses a different failure mode: the work leaves its first URL. Its protocol binds machine-readable usage preferences to an International Standard Content Code, a content-derived identifier standardised as ISO 24138:2024, and uses verifiable Creator Credentials for attribution of the declaration.
That matters for photographs, illustrations, audio and other assets copied between a CMS, syndication feed, partner site and archive. TDM·AI says the preference can be resolved from the content-derived identifier even when ordinary embedded metadata has been stripped or the file has been altered.
The procurement test should follow the same path. Register one approved image, export it through the normal publishing workflow, resize it, remove its metadata and place the result on a separate host. Recompute or resolve the identifier and verify that the intended preference is still discoverable and linked to the authorised declarant.
For an asset-heavy publisher, that work-level continuity can sit beside domain controls. The two layers answer different questions: whether a crawler may enter a route, and what a particular work says after it has travelled.
RSL: move from refusal to terms
RSL 1.0 goes beyond an opt-out vocabulary. Its specification defines an XML format and discovery mechanisms for usage, licensing, payment and legal terms. A publisher can associate an RSL document through robots.txt, HTTP headers, HTML, RSS or a file, and can distinguish AI training, AI input, AI indexing and conventional search.
That makes RSL relevant when the desired answer is conditional. A publisher may permit search links, prohibit training, require attribution for one archive, or point automated clients to a custom licence for another. The specification also describes optional protocols for licence acquisition, crawler authorisation and protected media.
Start with a static declaration, not a licensing server. Define two or three content classes, publish their terms and test discovery from an external client. Add automated acquisition or payment only when a real counterpart and commercial model require it. The smallest valid declaration is more useful than a complete licensing architecture with no buyer.
Build the evidence pack before the registry
The Commission has published a feasibility result, not a production registry or a deadline for joining one. A publisher can still prepare an evidence pack that would make a future registration almost mechanical:
- a list of content classes and the person authorised to decide their use;
- the approved rule for search, AI input, AI training and licensing;
- the canonical identifiers used for domains, feeds and individual works;
- the exact machine-readable declarations emitted by each delivery path;
- automated retrieval tests from outside the publishing system; and
- a dated record of policy, code and content changes.
Run the pack against ten representative works every month and after any CMS, CDN or syndication change. Store the result beside the policy version, not in a vendor dashboard alone.
If an EU registry arrives, those records provide the identifiers and declarations it would need to index. If the registry changes shape or never launches, the publisher still has a functioning rights-reservation process. The Commission’s study makes the destination clearer. The practical work remains the same: decide once, express the decision everywhere and test what machines can actually read.
Frequently asked questions
Is the EU TDM opt-out registry live?
No. The Commission published a feasibility study on 13 July 2026 and will consider next steps. The study describes a registry as a possible complement, not a replacement for existing opt-out mechanisms.
Is robots.txt enough to reserve TDM rights?
It can carry a machine-readable reservation for web content, and the GPAI Code of Practice explicitly refers to respecting robots.txt. Publishers with files, feeds, syndication or licensing terms may need additional work-level signals and records.
Which option should a small publisher start with?
Start with the publishing workflow, not a protocol name: identify the works, owners, permitted uses and delivery paths. Then implement the smallest combination that expresses the policy, controls access where possible and leaves a dated test record.
Sources
- Official New feasibility study for introducing an EU-level registry of Text and Data Mining opt-out European Commission accessed
- Primary Directive (EU) 2019/790 on copyright and related rights in the Digital Single Market EUR-Lex accessed
- Official Stakeholder consultation on AI and copyright compliance European Commission accessed
- Secondary ARCKONE offers ARCKONE accessed
- Secondary AI Crawl Control overview Cloudflare accessed
- Secondary What is the TDM·AI Protocol? TDM·AI accessed
- Secondary Really Simple Licensing 1.0 Specification RSL Collective accessed
- Secondary IPTC publishes version 2.0 of AI opt-out guidelines IPTC accessed
Image credit: Photo: newspaper press workers — Somogro Bangladesh, Pexels License (Pexels)
Iris Van Loon covers SME operational reality and advisors for Flint Brief.
Spotted an error or want a right of reply? hello@flintbrief.com (subject [Right of reply]).