Pre-trained model or embeddings
A pre-trained model brings the provenance of its own training data, which is rarely known in full.
What to record for it
The model, its version and licence, and what its provider publishes about its training data.
Personal data is unlikely: where your personal data column is blank the checker reads no, and says it assumed so.
A line that places here
exampleBase model | pre-trained model | internal
What the checker reads on these lines
7 of the 12 findings- Origin not recorded: Where did this dataset come from, and who in the company can show it?
- Terms not recorded: What terms did this data come under, and where is the copy?
- Terms that need review for this use: Do the terms, read in full, cover training this model, and who has read them?
- Labelled by a vendor, a crowd or a model with no quality check recorded: Who checked these labels, on what sample, and where is the result?
- No version or snapshot date: Which version of this dataset trained the model, and is that copy kept?
- Declared potential overlap between train and test: Were these two sets separated before training, and on what key?
- High-risk use with no bias examination recorded: Has this dataset been examined for bias against the people the model decides about, and where is the record?
Clauses
3 regimes| Regime | Clause |
|---|---|
| ISO/IEC 42001 | ISO/IEC 42001 A.7.6 Data preparation ISO/IEC 42001 A.7.4 Quality of data for AI systems |
| NIST AI RMF | NIST AI RMF MN-3.2 Pre-trained models monitored NIST AI RMF MP-4.2 Internal controls for third-party components |
| EU AI Act | EU AI Act Art. 10 Data and data governance EU AI Act Art. 15 Accuracy, robustness and cybersecurity |
The clauses, set out
ISO/IEC 42001 A.7.6Data preparationThe organization shall define and document criteria for selecting data preparations and the data preparation methods used.
ISO/IEC 42001 A.7.4Quality of data for AI systemsThe organization shall define and document quality requirements for data and ensure they are met.
NIST AI RMF MN-3.2Pre-trained models monitoredPre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance. Pre-trained and transfer-learned models are treated as a monitored component in their own right, since their provenance and behaviour are not controlled by the deploying organisation.
NIST AI RMF MP-4.2Internal controls for third-party componentsInternal risk controls for components of the AI system including third-party AI technologies are identified and documented. For each component carrying risk, the internal control applied to it is identified and written down, so the risk mapping produces controls rather than a list.
EU AI Act Art. 10Data and data governance applies if this system is high-risk under Annex IIIHigh-risk AI systems that make use of techniques involving the training of AI models shall use training, validation and testing data that meet the quality criteria in Art.10(2)-(5): appropriate data governance, examination for possible biases, identification of data gaps/shortcomings, statistically relevant datasets to the intended purpose, and considerations specific to the geographical, contextual, behavioural or functional setting of intended use.
EU AI Act Art. 15Accuracy, robustness and cybersecurity a provider dutyHigh-risk AI systems shall be designed and developed in such a way that they achieve an appropriate level of accuracy, robustness, and cybersecurity, and shall perform consistently in those respects throughout their lifecycle. Resilience to errors, faults and inconsistencies; protection against attempts by unauthorised third parties to alter use, output or performance (incl data poisoning, model poisoning, adversarial examples and confidentiality attacks).