The EU AI Act training data summary, and which lines of your list feed it
A provider of a general-purpose AI model draws up and publishes a sufficiently detailed summary of the content used for training, following a template the EU AI Office provides. The duty applies to providers of general-purpose models, not to a company fine-tuning a model for its own use; the checker shows Article 53 only when that box is ticked. The template itself is named, not quoted.
The lines of the list that feed it
| Column | What it gives the summary |
|---|---|
| Source | the origin of each dataset, grouped by the eight source groups the checker places it in: internal, customer-generated, vendor-licensed, public open data, web-scraped, synthetic, model-generated and partner-shared |
| Terms | the terms family of each dataset, and every line where the terms are not recorded or may not permit the use |
| Collected | the collection period, so the summary can say when the data was gathered |
| Web-scraped lines | the scraped sources and whether an opt-out or reservation of rights was recorded against them (the EU copyright directive, Article 4, named, not quoted) |
| Personal data and declared categories | which datasets carry personal data, declared or assumed from the source class, and which declare Article 9 special category data, criminal offence data or children's personal data |
| Version | which snapshot trained the model, so the summary describes the data that was actually used |
The clause
EU AI Act Art. 53Obligations for providers of general-purpose AI models providers of general-purpose modelsProviders of general-purpose AI models must draw up and keep up to date the technical documentation of the model, including its training and testing process and the results of its evaluation, containing at least the Annex XI information, for provision on request to the AI Office and the national competent authorities; draw up, keep up to date and make available to providers who intend to integrate the model information and documentation containing at least the Annex XII elements and sufficient to let them understand the model's capabilities and limitations and meet their own obligations; put in place a policy to comply with Union law on copyright and related rights, including identifying and complying, through state of the art technologies, with a reservation of rights expressed under Art.4(3) of Directive (EU) 2019/790; and draw up and make publicly available a sufficiently detailed summary of the content used for training, following the template provided by the AI Office. The two documentation duties do not apply to models released under a free and open source licence meeting the stated conditions, unless the model has systemic risk. Providers must cooperate with the Commission and the national competent authorities, and where they neither adhere to an approved code of practice nor comply with a European harmonised standard they must demonstrate alternative adequate means of compliance for assessment by the Commission.
Named, not quoted: the EU copyright directive, Article 4 (text and data mining, and the rightholder's opt-out); US copyright fair use; the EU AI Office template for the public summary of training content. The checker does not write the summary and does not say a summary meets the template; it gives the provider the list, grouped and with the gaps marked, that the summary is written from.