Skip to main content
Home / Legal / Rights & Compliance Center / TDM & Model Training
Rights & Compliance Center · Domain 04 · EN

TDM & Model Training

ENTIA's documented positions on text and data mining ("TDM") and AI model training. Each node states the position, its legal basis, the clause that implements it, the objections it resolves, and what it means for a party evaluating a licence. This is not legal advice; it is ENTIA's reasoned, citable position.

Version 2.0.0 · Last updated 24 September 2026 · Español
Author and ENTIA lead

Fernando Vilches · Founder & CEO · Founder and architect of ENTIA.

View bio →
Current status · Policy V3 · evidentiary representation V3.1

The effective TDM policy is V3 from 2026-09-19T03:08:41Z. From that instant ENTIA prospectively withdrew its earlier broad positive permissions for automated search, indexing, linking, citation, fresh retrieval and grounding. From 2026-09-21T10:29:29Z, /robots.txt also has the V3.1 evidentiary first-contact representation entia:robots:v3.1:fc-001. V3.1 is not a fourth policy, does not change V3's legal rights scope and does not create a new T0.

This KB is explanatory. The canonical version registry and the current machine-readable policy determine the policy window; event-level telemetry and observed publication evidence determine which bytes were actually served in a particular request. The reservation operates only where applicable rights exist; it does not turn an access event into proof of TDM, training or infringement, does not form a contract or automatically create a fee, debt or damages, and does not affect the Article 3 exception.

See V1 → V2 → V3 → V3.1 chronology · Machine-readable history JSON.

Related document · TDM Reservation

This page documents the chronology and scope of each version. ENTIA's general TDM rights-reservation instrument is available at TDM Rights Reservation. For the machine-readable current state and applicable version, see the canonical TDM policy and the version registry.

Open general TDM Reservation → · History JSON → · Fernando Vilches bio →

Relationship with the current TDM policy. Under V3 ENTIA no longer grants the broad positive permission for automated search/indexing/linking/citation/fresh retrieval/grounding that existed under V1/V2. This is a prospective withdrawal of ENTIA permission, not an automatic conclusion of unlawfulness. TDM, copyright, database rights, exceptions, infringement, contract, debt and damages remain separate questions requiring their own evidence. Current V3 policy · History JSON.

The reserved uses discussed here are anchored to ENTIA's public instruments: the TDM Rights Reservation, the Database Rights Notice and the Data Licensing Framework. This domain distils those instruments into six load-bearing positions.

TDM policy history: V1, V2, V3 and V3.1

ENTIA distinguishes policy versions from evidentiary representation versions. V1, V2 and V3 are legal policy versions with their own temporal windows. V3.1 is not a new policy: it preserves exactly the V3 rights scope and adds a canonical, cryptographically identified first-contact representation at /robots.txt.

V1 · 2026-07-19T06:56:00Z → 2026-09-09T09:04:06Z

What it did. Express Article 4(3) reservation. Public signal search=yes, ai-input=yes, ai-train=no. ENTIA kept search crawling/indexing, linking and canonical citation, and fresh retrieval/grounding of the public layer open; where applicable rights existed it reserved TDM, pre-training, fine-tuning, incremental training, inclusion in training/evaluation datasets and database/corpus reconstruction.

What it did not do. It did not turn all crawling into TDM, prohibit search or grounding, create copyright or a sui generis right, affect Article 3 scientific-research TDM, prove training, create automatic debt, or operate before its effective date.

V2 · 2026-09-09T09:04:06Z → 2026-09-19T03:08:41Z

What it did. It retained search=yes, ai-input=yes, ai-train=no while formalising objective classification: P1 search/indexing, P2 fresh retrieval/grounding, C1 fact-dependent technical structures, and RU1–RU5 for model training/parameter update, development/evaluation datasets, independent corpus or reconstruction, cross-lane promotion and redistribution. It expressly separated statutory TDM reservation from technical-use classification.

What it did not do. UA, operator name, ASN, volume, persistence or embeddings did not by themselves decide TDM or infringement. It did not form a contract by access, create rights, price, debt or damages, or rewrite V1.

V3 · from 2026-09-19T03:08:41Z

What it does. It retains the Article 4(3) reservation and RU1–RU5, but prospectively withdraws the earlier broad positive permissions for automated uses: search=no, ai-input=no, ai-train=no. Technical reachability of a URL is not equivalent to ENTIA permission for automated use.

What it does not do. It does not automatically declare every later access unlawful; it does not by itself decide copyright, database rights, exceptions, TDM, infringement, contract, debt or damages; and it has no retroactive effect on V1/V2.

V3.1 · evidentiary boundary from 2026-09-21T10:29:29Z

What it does. It does not change V3. It fixes /robots.txt as canonical first-contact representation entia:robots:v3.1:fc-001, with sealed identity and hashes, so an observed request can be bound to the exact representation served.

What it does not do. It is not V4, creates no new legal T0 and neither expands nor narrows V3. A GET /robots.txt with HTTP 200 bound to this representation proves which representation ENTIA generated and served; it does not prove remote parsing, understanding, acknowledgement, compliance, training or downstream use.

representation_sha256=cc91f6c30438da4edc707bffe520adbfbe60f594cdea3ebc7513ba63307f1e1b
reservation_sha256=5225bee6732770760532e306709c513cd5bf259692f4d289b7bad682645a5ec1

Two clocks: legal effectiveness and observed serving

The legal effective instant and the instant at which a particular representation became observable at the edge are distinct evidentiary concepts. The registry fixes the legal policy window; request-level telemetry, sealed public captures and deployment evidence determine which representation was actually served in a particular event. During transitions, ENTIA does not replace observed evidence with a retrospective date-only inference.

Training and persistent corpus uses: conditioned reservation and authorisation route

Position

The policy depends on date. Under V1 and V2 ENTIA granted broad positive permission for discovery, public search, indexing and fresh retrieval, without extending that permission to training. From V3, effective 2026-09-19T03:08:41Z, those positive automated permissions were prospectively withdrawn. Across all versions, the Article 4(3) reservation operates only where applicable rights are engaged and does not by itself turn an access event into proof of TDM, training or infringement.

Basis
  • Article 4(3) of Directive (EU) 2019/790 — the general (commercial) TDM exception applies only where use has not been expressly reserved by the rightholder, including by machine-readable means for content made publicly available online.
  • Article 7 of Directive 96/9/EC — the maker's sui generis right against extraction and/or re-utilisation of the whole or a substantial part of a database, where there is qualifying substantial investment.
  • Article 53(1)(c) of Regulation (EU) 2024/1689 (AI Act) — providers of general-purpose AI models must put in place a policy to identify and comply with reservations of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790.
Implementing clause

The TDM Rights Reservation (expressed concurrently through tdmrep.json, the tdm-reservation HTTP header on every apex response, the Content-Signal lines in robots.txt and /ai.txt), together with the Train section of the Data Licensing Framework and Annex C of the contract stack, under which Foundational training is gated to a dedicated Annex D.

Objections resolved
  • "It was open to our crawler." — Access was granted for indexing and retrieval, not for training; tolerating a crawler is not a waiver (see Node 3).
  • "It is public data." — Public availability does not enable a third party to train on the content; the reservation and the database right operate on that use specifically.
  • "The reservation came later." — It is not retroactive (see Node 4); the basis covering prior access is the sui generis right and the contractual prohibitions, not the reservation date.
What it means for you

Training on ENTIA content is licensable through Annex D following a Source Audit. It is not offered as an included right of the discovery or retrieval layers, and it is not something the free access tier grants by implication.

See also

Node 3 — no-waiver · Node 4 — non-retroactivity · Node 5 — how to license (Annex D) · Database sui generis right · What you are licensing · AI Act.

The reservation is expressed by machine-readable means; its effect operates on applicable rights

Position

A reservation of TDM rights expressed by machine-readable means is the valid form of opt-out for content made publicly available online. Article 4(3) does not prescribe a single technology; it requires that the reservation be appropriate and, for online content, machine-readable. ENTIA expresses one reservation through several concurrent signals, and the failure of any one signal does not defeat the others.

Basis
  • Article 4(3) of Directive (EU) 2019/790 — for content made publicly available online, the reservation of rights is effective where expressed in an appropriate manner, such as machine-readable means.
  • Recital 106 of Regulation (EU) 2024/1689 (AI Act) — where rights have been expressly reserved in an appropriate manner, providers of general-purpose AI models need to obtain an authorisation from rightholders to carry out text and data mining over the works concerned.
Implementing clause

The reservation is expressed concurrently through /.well-known/tdmrep.json, HTTP headers tdm-reservation: 1 and tdm-policy, robots.txt, /ai.txt and the licence manifest. The robots.txt signal is versioned: V1/V2 served search=yes, ai-input=yes, ai-train=no; V3 serves search=no, ai-input=no, ai-train=no. From the V3.1 boundary, the first-contact representation is additionally bound to entia:robots:v3.1:fc-001.

Objections resolved
  • "There is no standard format." — Article 4(3) is technology-neutral; ENTIA deploys the recognised conventions (TDMRep, robots Content-Signal, ai.txt) in parallel precisely to be identifiable by state-of-the-art means as required of GPAI providers.
  • "One signal was missing or failed." — The reservation is single and expressed redundantly; the temporary absence, withdrawal or technical failure of one signal does not enerve the remaining signals or the declaration itself.
What it means for you

A provider crawling ENTIA is on constructive notice of the reservation through multiple detectable signals. A compliance policy that reads any one of them will surface the opt-out; ignoring all of them is not consistent with the state-of-the-art diligence Article 53(1)(c) expects.

See also

Node 1 — licence required · Node 3 — no-waiver · TDM Rights Reservation · AI Act.

Crawler allow-listing is not a training licence (no-waiver)

Position

Tolerating or not blocking a bot is not a tacit licence. Under V1/V2, where ENTIA positively allowed automated discovery, search or retrieval, that permission did not include training. Under V3, technical reachability may remain even though the earlier positive automated permissions have been withdrawn. In no version does absence of a technical block convert a reserved use into a permitted one.

Basis
  • Article 4(3) of Directive (EU) 2019/790 — the reservation persists irrespective of whether access was technically possible; lawful access does not carry a training right.
  • General principles of licence interpretation — a licence is construed by its granted scope; conduct that merely permits access does not imply a grant of the reserved uses. ENTIA's instruments state expressly that lawful access is not a licence.
Implementing clause

The "Lawful access is not a licence" statements in Sections 3 and 4 of the TDM Rights Reservation, read with the Content-Signal triad (search=yes, ai-input=yes, ai-train=no) that expressly separates discovery/indexing (allowed) from training (reserved). The Acceptable Use and Data Licensing instruments delimit permitted automated consumption.

Objections resolved
  • "You let our user agent in, so you consented." — Allow-listing scopes to indexing and retrieval; the training signal is separately and expressly set to no. Permission to access is not permission to train.
  • "We relied in good faith on the open endpoint." — The reserved-use signals are published and machine-readable; a provider on constructive notice cannot found a legitimate expectation to train on the mere reachability of the content.
What it means for you

Do not treat reachability as a licence. If your pipeline ingests ENTIA content for training, obtain Annex D coverage; an allow-listed crawler status will not support a scope defence.

See also

Node 1 — licence required · Node 2 — reservation effective · Node 6 — Art. 3 vs Art. 4.

The reservation is not retroactive — and what covers prior access

Position

The reservation takes effect from the date of its publication and applies to access after that date. It does not purport to reach back in time. For access that predates the reservation, the basis is not the reservation but the sui generis database right (in force from first publication under Directive 96/9/EC) together with the contractual prohibitions on unlicensed training already published in the licence manifest.

Basis
  • Article 7 of Directive 96/9/EC — the sui generis right subsists from the completion of the database / first making available, independent of and prior to any Article 4(3) reservation.
  • Article 4(3) of Directive (EU) 2019/790 — the reservation governs the availability of the TDM exception going forward from the moment it is appropriately expressed.
  • Contract — the prohibitions on unlicensed training published in the ENTIA licence manifest since 19 May 2026 bind parties on their own terms, separately from the reservation.
Implementing clause

Section 8 of the TDM Rights Reservation (effective 19 July 2026 for access after that date, expressly without prejudice to (a) the contractual prohibitions in force since 19 May 2026 and (b) the database rights in force since first publication). ENTIA preserves edge-log and governance evidence of publication and continuity for use in negotiation or proceedings.

Objections resolved
  • "We ingested it before your reservation, so we are clear." — The reservation is not the only basis; prior extraction of a substantial part engages the sui generis right and any applicable contractual prohibitions, which predate the reservation.
  • "You cannot apply a new rule to old crawling." — Correct, and ENTIA does not: it applies the reservation prospectively and relies on the pre-existing database right and contract for earlier access.
What it means for you

A pre-reservation ingestion date is not a safe harbour. Assess exposure under the database right and the licence manifest for any substantial extraction, and regularise training use through Annex D regardless of when the content was first crawled.

See also

Node 1 — licence required · Node 5 — how to license · Database Rights Notice.

How to request Training & Corpus authorisation

Position

Training rights are available under a defined path. They are not sold self-serve alongside the discovery or API tiers; they are granted bilaterally after ENTIA and the counterparty have established which layers, by source, may lawfully be used for training. The path is: Source Audit, then a qualified-signature addendum, then Annex D.

Basis
  • Article 4(3) of Directive (EU) 2019/790 — because the use is reserved, an authorisation (licence) is the lawful route to carry out the mining for training.
  • Recital 106 of Regulation (EU) 2024/1689 — GPAI providers need to obtain authorisation to mine reserved works; Annex D is that authorisation for ENTIA content.
  • Provenance and flow-down — per-source terms determine which layers of the corpus may be included in a training grant.
Implementing clause
  • Source Audit — a per-source determination of which layers may be used for training, so that upstream register terms are respected and flowed down.
  • Qualified-signature (QES) addendum — execution with a qualified electronic signature, giving the grant heightened evidentiary weight (see the eIDAS / evidence domain).
  • Annex D — the training grant itself: scope, permitted layers, model and use restrictions, term and consideration.

Contact: partnerships@entia.systems.

Objections resolved
  • "Why can't we just buy an API tier and train?" — The API and discovery tiers do not grant training rights; training is gated to Annex D because permitted layers vary by source and must be audited first.
  • "The audit slows us down." — The Source Audit is what makes the grant defensible: it fixes, per source, what may lawfully be trained on, protecting both parties from downstream provenance challenges.
What it means for you

To train on ENTIA content, request a Source Audit through partnerships. The output is a scoped Annex D you can rely on, executed under qualified signature, rather than an implied or best-efforts permission.

See also

Node 1 — licence required · Node 4 — prior access · What you are licensing · Deal shapes · eIDAS / evidence / sealing.

Art. 3 research exception vs Art. 4 commercial TDM

Position

Two distinct TDM regimes coexist. The Article 3 exception, for research organisations and cultural heritage institutions carrying out TDM for scientific research with lawful access, cannot be reserved. The Article 4 exception, for TDM more generally including commercial use, can be reserved — and it is the regime that applies to a commercial AI laboratory. ENTIA's reservation operates on Article 4 and does not, and cannot, restrict genuine Article 3 research.

Basis
  • Article 3 of Directive (EU) 2019/790 — mandatory TDM exception for research organisations and cultural heritage institutions, for scientific research, on works to which they have lawful access; this exception is not subject to the rightholder's opt-out.
  • Article 4 of Directive (EU) 2019/790 — general TDM exception, available for any purpose including commercial, but subject to express reservation of use under Article 4(3).
Implementing clause

Section 4 of the TDM Rights Reservation records that TDM by research organisations and cultural heritage institutions for scientific research under Article 3 is not affected by the reservation, since that exception cannot be reserved. All other, commercial TDM falls under the Article 4 reservation and requires a licence for training (Nodes 1 and 5).

Objections resolved
  • "We are doing research, so Article 3 covers us." — Article 3 is confined to research organisations and cultural heritage institutions acting for scientific research with lawful access. A commercial AI laboratory training a product model is generally outside that scope and falls under Article 4.
  • "Our research arm crawled it." — The exemption follows the qualifying entity and the scientific-research purpose, not a label; onward commercial training use is not sheltered by Article 3.
What it means for you

If you are a commercial provider, Article 4 governs and the reservation applies: license training through Annex D. Genuine Article 3 scientific research by a qualifying institution is unaffected and needs no licence for that reserved-use question.

See also

Node 1 — licence required · Node 3 — no-waiver · Node 5 — how to license · AI Act.


Rightholder: PrecisionAI Marketing OÜ · Sepapaja tn 4, Tallinn, 11415, Estonia.
Legal: legal@entia.systems · Licensing: partnerships@entia.systems · Back to the Rights & Compliance Center.

Contents
↑ Back to top