AI Training Data Provenance Checker
For providers of general-purpose models

The EU AI Act training data summary, and which lines of your list feed it

A provider of a general-purpose AI model draws up and publishes a sufficiently detailed summary of the content used for training, following a template the EU AI Office provides. The duty applies to providers of general-purpose models, not to a company fine-tuning a model for its own use; the checker shows Article 53 only when that box is ticked. The template itself is named, not quoted.

The lines of the list that feed it

ColumnWhat it gives the summary
Sourcethe origin of each dataset, grouped by the eight source groups the checker places it in: internal, customer-generated, vendor-licensed, public open data, web-scraped, synthetic, model-generated and partner-shared
Termsthe terms family of each dataset, and every line where the terms are not recorded or may not permit the use
Collectedthe collection period, so the summary can say when the data was gathered
Web-scraped linesthe scraped sources and whether an opt-out or reservation of rights was recorded against them (the EU copyright directive, Article 4, named, not quoted)
Personal data and declared categorieswhich datasets carry personal data, declared or assumed from the source class, and which declare Article 9 special category data, criminal offence data or children's personal data
Versionwhich snapshot trained the model, so the summary describes the data that was actually used

The clause

EU AI Act Art. 53Obligations for providers of general-purpose AI models providers of general-purpose models

Providers of general-purpose AI models must draw up and keep up to date the technical documentation of the model, including its training and testing process and the results of its evaluation, containing at least the Annex XI information, for provision on request to the AI Office and the national competent authorities; draw up, keep up to date and make available to providers who intend to integrate the model information and documentation containing at least the Annex XII elements and sufficient to let them understand the model's capabilities and limitations and meet their own obligations; put in place a policy to comply with Union law on copyright and related rights, including identifying and complying, through state of the art technologies, with a reservation of rights expressed under Art.4(3) of Directive (EU) 2019/790; and draw up and make publicly available a sufficiently detailed summary of the content used for training, following the template provided by the AI Office. The two documentation duties do not apply to models released under a free and open source licence meeting the stated conditions, unless the model has systemic risk. Providers must cooperate with the Commission and the national competent authorities, and where they neither adhere to an approved code of practice nor comply with a European harmonised standard they must demonstrate alternative adequate means of compliance for assessment by the Commission.

What an auditor asks to see: Model technical documentation checked against every Annex XI element, with training, testing and evaluation results included; The downstream integrator pack checked against Annex XII, with evidence it is actually supplied to integrators; The copyright compliance policy, naming the technologies used to identify and honour text and data mining reservations of rights; The published training content summary following the AI Office template, with its publication location and date; Where the open source exemption is claimed, evidence the licence and the published parameters, architecture and usage information meet the Art.53(2) conditions; Where neither an approved code of practice nor a harmonised standard is followed, the documented alternative adequate means of compliance
What an auditor will probe: A training content summary published at a level of generality that is not sufficiently detailed against the template; A copyright policy that states intent but names no technology for identifying reservations of rights; The open source exemption claimed for a model with systemic risk, where it does not apply; Downstream documentation limited to an interface reference, omitting the capability and limitation information integrators need to meet their own duties
Source: EU AI Act

Named, not quoted: the EU copyright directive, Article 4 (text and data mining, and the rightholder's opt-out); US copyright fair use; the EU AI Office template for the public summary of training content. The checker does not write the summary and does not say a summary meets the template; it gives the provider the list, grouped and with the gaps marked, that the summary is written from.

See the specimen run The template CSV