Synthetic data, generated in house
Synthetic data inherits the provenance of what it was generated from; the seed data is the missing line.
What to record for it
The generator, the seed data it was fitted on, the version, and the checks that it does not reproduce real records.
Personal data is unlikely: where your personal data column is blank the checker reads no, and says it assumed so.
A line that places here
exampleSynthetic claims set | generated in house | internal
What the checker reads on these lines
7 of the 12 findings- Origin not recorded: Where did this dataset come from, and who in the company can show it?
- Terms not recorded: What terms did this data come under, and where is the copy?
- Labelled by a vendor, a crowd or a model with no quality check recorded: Who checked these labels, on what sample, and where is the result?
- No version or snapshot date: Which version of this dataset trained the model, and is that copy kept?
- Declared potential overlap between train and test: Were these two sets separated before training, and on what key?
- High-risk use with no bias examination recorded: Has this dataset been examined for bias against the people the model decides about, and where is the record?
- Vendor data with no provider named: Which company supplied this data, under which contract?
Clauses
3 regimes| Regime | Clause |
|---|---|
| ISO/IEC 42001 | ISO/IEC 42001 A.7.4 Quality of data for AI systems ISO/IEC 42001 A.7.6 Data preparation ISO/IEC 42001 A.7.5 Data provenance |
| NIST AI RMF | NIST AI RMF MP-2.3 Data collection, selection and TEVV |
| EU AI Act | EU AI Act Art. 10 Data and data governance |
The clauses, set out
ISO/IEC 42001 A.7.4Quality of data for AI systemsThe organization shall define and document quality requirements for data and ensure they are met.
ISO/IEC 42001 A.7.6Data preparationThe organization shall define and document criteria for selecting data preparations and the data preparation methods used.
ISO/IEC 42001 A.7.5Data provenanceThe organization shall document the provenance of data used in AI systems to enable evaluation and traceability.
NIST AI RMF MP-2.3Data collection, selection and TEVVScientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation. The evaluation design is documented as a scientific claim: what was measured, on what data, and whether the measure validly stands for the property claimed.
EU AI Act Art. 10Data and data governance applies if this system is high-risk under Annex IIIHigh-risk AI systems that make use of techniques involving the training of AI models shall use training, validation and testing data that meet the quality criteria in Art.10(2)-(5): appropriate data governance, examination for possible biases, identification of data gaps/shortcomings, statistically relevant datasets to the intended purpose, and considerations specific to the geographical, contextual, behavioural or functional setting of intended use.