Paste the list of datasets behind your AI models. See which ones you cannot account for.
AI Training Data Provenance Checker reads your dataset list line by line for provenance, labelling and lineage gaps against EU AI Act Article 10 and ISO/IEC 42001. Paste the list of datasets behind your AI models, the list and never the data. Every dataset is checked for where it came from, the terms it came under, personal and special category data, who labelled it, which version it is, and whether the same data sits in train and test, each finding with the EU AI Act Article 10 or ISO/IEC 42001 clause behind it.
Paste the list of your datasets, never the data. One line per dataset: name, source, terms, personal data, who labelled it, version. The list is read in your browser and nothing is stored until you save. A cell that looks like record content, such as an email address or an identifier number, is removed before anything is shown; review the list before you save it.
It reads what your list says about each dataset. It does not open or scan the data, it does not match your files against other collections, and it gives no score.
Every finding names the column that raised it and the clause behind it, and says what it could not determine: terms not recorded, source not recorded.
It never says a dataset is lawful, licensed, cleared or compliant.
- D-01trainClaims notes 2019-2024internal claims system; internal; labelled in house; snapshot 2024-12-31
- D-04trainMotor forum postsweb scrape; [no terms recorded]; labelled by a crowd; [no version]
- D-08trainImage set for vehicle damagevendor; vendor licence; labelled by a vendor; [no version]; [provider not named]
- D-30trainVehicle images from assessors[origin not recorded]; [no terms recorded]; labelled by a crowd; v3
- D-10testEval set Q3copy of claims notes, Q3 cut; internal; labelled in house; eval-q3
and 6 more datasets in this chain
32 datasets, 7 with incomplete provenance records.
12 finding types, 2 declared potential overlaps, 31 of 32 placed.

Paste the list as it sits in the spreadsheet
One dataset per row: dataset | source | terms at least, or with any of feeds, split, personal data, sensitive categories, lawful basis, labelled by, label check, version, collected, provider and bias examined. Each line is placed on one of 42 published source classes and 26 terms families; a line that fits none is marked unplaced and never guessed.
Read the chain for each model
Every model with the datasets that feed it in split order, each written origin; terms; labelled by; version. A link your list does not record is printed in square brackets and marked, and the chain breaks there. Declared potential overlaps between train and test are drawn as pairs.
Take the datasheet pack to the dataset owners
Twelve findings in fixed order, each with the datasets it names, the clause from ISO/IEC 42001, the NIST AI RMF, the EU AI Act, GDPR or UK GDPR, and the question to put to the dataset owner or to legal.
Why a list, and not a scan
Data lineage software connects to your systems and traces tables. The head of data governance at a 900-person company is asked a simpler question first: which datasets trained this model, where did each come from, and under what terms. The answer sits in a spreadsheet kept by whoever ran the last project. Before anyone buys a lineage tool, somebody has to show which lines of that spreadsheet have no origin, no terms, no version or no check on the labels, and which clause asks for each. That is what this reads, in your browser, from the list as it stands.