Studies / S16
Text-and-data-mining reservations in the EU-27 press, channel by channel
Do EU news publishers state text-and-data-mining reservations that an AI crawler can read, whatever its name?
Abstract
Since 2 August 2025 the AI Act (Art. 53(1)(c)) has required providers of general-purpose AI models to identify and comply with reservations of text and data mining expressed under Art. 4(3) of Directive (EU) 2019/790, "such as machine-readable means" for content online. Models placed on the market before that date have until 2 August 2027 (Art. 111(3)). The Commission can fine providers from 2 August 2026.
We read what a crawler can read on 1,511 news websites in 27 Member States: newspapers and news sites from Wikidata whose host is in the top 100,000 of their country's Chrome UX Report list, with official gazettes excluded. The channels were robots.txt (RFC 9309), TDMRep (file, HTTP header, <meta>), Content-Signal, the IETF Content-Usage rule, noai and llms.txt. We requested a URL only where robots.txt allowed it.
On the 1,356 sites we could read in full, 99 of 1,356 (7.3%), 95% CI 6.0–8.8%, state a reservation that binds a crawler whatever its name.
- 558 of 1,356 (41.2%) block at least one named AI crawler at the root. Of those, 480 of 558 (86.0%), CI 82.9–88.7%, state nothing that a crawler with a new name would read.
- 777 of 1,356 (57.3%) state no reservation in any channel we read.
- Across all 1,510 sites that answered for
robots.txt, the agnostic count is at least 110 of 1,510 (7.3%). Named blockers without an agnostic reservation are at most 621 of 699 (88.8%). - By country band, agnostic among fully read sites: large countries 74 of 800 (9.2%), medium 15 of 431 (3.5%), small 10 of 125 (8.0%).
- Sensitivity to our own rule change (D5). Under the access rule frozen before collection, without the D5 re-scan, the result is 93 of 1,282 (7.3%), CI 6.0–8.8%, fully read sites with an agnostic reservation, and 0 contradictions.
Channels.
- TDMRep is on 71 of 1,356 (5.2%), CI 4.2–6.6%. For 58 of the 71 it is only in a header or a
<meta>element (meta 41, header 22, file 13). 53 of 192 fully read French sites (27.6%) use it. - Content-Signal
ai-train=noappears on 18 sites. The IETFtrain-ai=nappears on 0. - 17 sites send
Content-Usage: ai=n. That is the AI-training label of draft-ietf-aipref-vocab-01; the current draft dropped it, so parsers must ignore it. Read as a training reservation, it would raise the agnostic share to 116 of 1,356 (8.6%). - Contradictions: 2 contradictions among the 1,356 under the current rule, 0 under the frozen rule. Both are Corriere della Sera sites, which were read only in the D5 re-scan. Each sends Content-Signal
ai-train=yestogether with a TDMRep<meta>reservation. - 218 of 1,510 (14.4%) write a reservation or a general prohibition in
robots.txtcomments, which an RFC 9309 parser discards.
Access, and what we got wrong. For 154 sites we read robots.txt only, because a deliberately conservative detector flags their comments as a general prohibition of automated access. The study broke its own access rule five times; every breach is declared and its data deleted:
- D0: the scouting pilot.
- D1: 7 requests to a host before its
robots.txt. - D2: 245 requests to 92 sites.
- D4: 3 requests to one site.
- D7: 72 requests to 24 sites.
The D7 requests were sent in a re-scan (D5) made after collection, after the study's coordinating AI agent changed the rule for what counts as a prohibition. The study's instructions misclassified the 23 Mediahuis sites' notice as purpose-specific. EasyxLab's lab-wide rule (3 October 2026) treats robots.txt as binding at the legal minimum. S16 deliberately keeps a stricter rule, because its subject is those reservations.
Detector audits (all by AI agents).
- Author's own sample, detectors v1/v2: the reservation detector was right on 54 of 54 flags; the v2 prohibition detector was right on 35 of 62.
- First independent reviewer, v1/v2, 44 other sites: prohibition right on 15 of 21 flags and found 15 of 15; reservation right on 20 of 20 flags and found 20 of 21.
- Re-reviewer, v3, 43 further sites: prohibition 15 of 15 in the sample, but about 84% recall (130 of 154) once the 24 D7 misses are counted; reservation 15 of 16.
- v4 changes exactly those two failure modes. No audit of v4 is claimed.
Providers. We could read the crawler documentation of 8 of 14 providers on 2026-10-03; all 8 name robots.txt. None names TDMRep, Content-Signal, Content-Usage or noai.
Per-site results state what each site's files say. They are not an assessment of the site: the obligation is on AI providers. The checker is scripts/optout_check.py <domain>, which may become the free tool optout-lint.
Cite this study
EasyxLab (2026). Text-and-data-mining reservations in the EU-27 press, channel by channel. Study S16. EasyByte Hub S. Coop. Mad. https://github.com/easybytehub/easyxlab/tree/main/studies/s16-eu-press-tdm-reservations@techreport{easyxlab_s16,
title = {Text-and-data-mining reservations in the EU-27 press, channel by channel},
author = {{EasyxLab}},
institution = {EasyByte Hub S. Coop. Mad.},
number = {S16},
year = {2026},
url = {https://github.com/easybytehub/easyxlab/tree/main/studies/s16-eu-press-tdm-reservations}
}