A datasheet for datasets template, one line per dataset
Download the columns as a CSV, fill it in, one line per dataset, with labels and never the data, then paste it back and the checker reads it. It is the list form of a datasheet for datasets: where each dataset came from, its terms, collection, labelling and version. The list you already keep works too: paste it in whatever shape it is in; a column the checker cannot find is simply reported as not recorded.
Columns the checker reads
header names are matched loosely| Column | What goes in it | Needed |
|---|---|---|
| Dataset | the name your team uses for it | required |
| Source | the system, vendor, site or partner it came from; "copy of X" when it was cut from another line | required |
| Terms | the licence or terms as recorded: internal, CC BY 4.0, vendor licence, data sharing agreement, none recorded | required |
| Feeds | the model it trains, evaluates or grounds, with its stage: Pricing model (production) | strongly advised |
| Split | train, validation, test or retrieval | strongly advised |
| Personal data | yes, no or unknown; blank is read from the source class and labelled as assumed | strongly advised |
| Sensitive categories | as declared, in three separate kinds: Article 9 special category data (health, genetic, biometric for identification, racial or ethnic origin, political opinions, religious or philosophical beliefs, trade union membership, sex life or sexual orientation); criminal offence data (GDPR Art. 10); children's personal data (heightened protection, not Art. 9) | advised |
| Lawful basis | the basis for training on it; add "compatibility assessed" where a new purpose was tested | advised |
| Labelled by | in house, vendor, crowd, model, rules or none | advised |
| Label check | yes where a quality check on the labels is recorded | advised |
| Version | the version or snapshot that trained the model | strongly advised |
| Collected | the collection period: 2023-01 to 2026-06, or a date | advised |
| Provider | the company that supplied bought or licensed data, and the contract | for vendor data |
| Bias examined | yes where a bias examination is recorded; read for models in an Annex III area | for Annex III uses |
| Notes | anything else, as a label; record content is removed | optional |
A first line for the company
EU AI Act: yes | role: provider | GDPR: yes | UK GDPR: no | high-risk use: life or health insurance pricing | general-purpose model provider: no | older than: 3 years | as at: 2026-09-26
Each regime follows its own trigger. EU AI Act: the model is placed on the EU market, put into service in the EU, or its output is used in the EU; the role says whether the provider or the deployer articles apply. GDPR: established in the EU, or processing personal data of people in the EU. UK GDPR: the same test for the UK. The high-risk use decides where Article 10's bias line is read, if the system is high-risk under Annex III; a general-purpose model provider sees Article 53; the threshold is your policy threshold, not a legal deadline.
The four rows in the CSV
invented| Dataset | Source | Terms | Feeds | Split | Version |
|---|---|---|---|---|---|
| Claims notes | internal claims system | internal | Claims triage model (production) | train | snapshot 2026-06-30 |
| Claims notes holdout | copy of claims notes | internal | Claims triage model (production) | test | holdout-2026-06 |
| Address reference file | data vendor | vendor licence, commercial, AI training permitted | Life and health pricing model (production) | train | 2026-03 |
| Flood risk grid | public open data | CC BY 4.0 | Life and health pricing model (production) | train | v4.1 |