AI Training Data Provenance Checker
For whoever is asked where the training data came from

Paste the list of datasets behind your AI models. See which ones you cannot account for.

AI Training Data Provenance Checker reads your dataset list line by line for provenance, labelling and lineage gaps against EU AI Act Article 10 and ISO/IEC 42001. Paste the list of datasets behind your AI models, the list and never the data. Every dataset is checked for where it came from, the terms it came under, personal and special category data, who labelled it, which version it is, and whether the same data sits in train and test, each finding with the EU AI Act Article 10 or ISO/IEC 42001 clause behind it.

Paste the list of your datasets, never the data. One line per dataset: name, source, terms, personal data, who labelled it, version. The list is read in your browser and nothing is stored until you save. A cell that looks like record content, such as an email address or an identifier number, is removed before anything is shown; review the list before you save it.

It reads what your list says about each dataset. It does not open or scan the data, it does not match your files against other collections, and it gives no score.

Every finding names the column that raised it and the clause behind it, and says what it could not determine: terms not recorded, source not recorded.

It never says a dataset is lawful, licensed, cleared or compliant.

Check your own listEight lines free, no account. A published dictionary of 42 source classes and 26 terms families and the clause text load with the page; every check runs in your browser.
Specimen: an invented insurer1 of its 3 chains
Claims triage model11 datasets, 4 with a gapproduction · insurance use: classification requires review

and 6 more datasets in this chain

32 datasets, 7 with incomplete provenance records.

12 finding types, 2 declared potential overlaps, 31 of 32 placed.

An analyst at a desk holding a printed page of tables beside a laptop showing more tables
Walk into the model review knowing which datasets have no origin, no terms or no version on record, and which clause asks for each. It works from the spreadsheet you already keep, without handing over a single record or connecting anything to your data systems.
01

Paste the list as it sits in the spreadsheet

One dataset per row: dataset | source | terms at least, or with any of feeds, split, personal data, sensitive categories, lawful basis, labelled by, label check, version, collected, provider and bias examined. Each line is placed on one of 42 published source classes and 26 terms families; a line that fits none is marked unplaced and never guessed.

02

Read the chain for each model

Every model with the datasets that feed it in split order, each written origin; terms; labelled by; version. A link your list does not record is printed in square brackets and marked, and the chain breaks there. Declared potential overlaps between train and test are drawn as pairs.

03

Take the datasheet pack to the dataset owners

Twelve findings in fixed order, each with the datasets it names, the clause from ISO/IEC 42001, the NIST AI RMF, the EU AI Act, GDPR or UK GDPR, and the question to put to the dataset owner or to legal.

One dataset per row: Dataset | Source | Terms, or a header row with any of Feeds, Split, Personal data, Sensitive categories, Lawful basis, Labelled by, Label check, Version, Collected, Provider, Bias examined. Tabs, pipes, commas or double spaces. A first line such as EU AI Act: yes | role: provider | GDPR: yes | UK GDPR: no | high-risk use: life or health insurance pricing | older than: 3 years | as at: 2026-09-26 sets the scope and the date. Write the stage after the model: Pricing model (production).
Nothing is sent anywhere until you choose to save.
Scope (each regime follows its own trigger)

Why a list, and not a scan

Data lineage software connects to your systems and traces tables. The head of data governance at a 900-person company is asked a simpler question first: which datasets trained this model, where did each come from, and under what terms. The answer sits in a spreadsheet kept by whoever ran the last project. Before anyone buys a lineage tool, somebody has to show which lines of that spreadsheet have no origin, no terms, no version or no check on the labels, and which clause asks for each. That is what this reads, in your browser, from the list as it stands.

The dictionary is ours and published in full: every source class in eight groups, every terms family and its triage status, the clauses each regime attaches, the columns a dataset list needs and which lines feed the EU training data summary. It reads labels only: a cell that looks like record content is removed before anything is shown, and a finding is a question for the dataset owner or for legal, not a ruling.