The legal regime for using protected works to train a model, based in Europe on an exception coupled with a right to opt out.
Training a model requires technically reproducing protected works, which falls under copyright, and the answers differ profoundly between legal systems. Union law created a text and data mining exception, with no restriction of purpose, but coupled for uses other than scientific research with an opt-out the rightholder may exercise in an appropriate manner, notably by machine-readable means. The American system has no equivalent exception and reasons through fair use, a case-by-case assessment whose criteria, in particular transformative character and effect on the market for the work, are precisely what current litigation is testing. The economic stake goes beyond training alone, since an adverse ruling would render unlawful models already deployed in thousands of applications, without any technically feasible way of selectively removing the data. For insurance the exposure sits in intellectual property cover and in the indemnities providers have begun to grant their customers.
Directive (EU) 2019/790 of 17 April 2019 introduced in Articles 3 and 4 the text and data mining exceptions, the second coupled with a rightholder opt-out. The New York Times separately sued OpenAI and Microsoft on 27 December 2023 before a United States federal court, alleging reproduction of its articles in the training data.
text and data mining, TDM, exception de fouille, opt-out éditeur