AI Training Data Provenance Checker
Model-generated · source class

Labels produced by a model

A model that labels training data passes its own errors and bias into the next model.

What to record for it

The labelling model and its version, the thresholds, and the human check on a sample.

Personal data is unlikely: where your personal data column is blank the checker reads no, and says it assumed so.

A line that places here

example

Fraud flags | model labels | internal

Check this line

What the checker reads on these lines

7 of the 12 findings

Clauses

3 regimes
RegimeClause
ISO/IEC 42001ISO/IEC 42001 A.7.6 Data preparation
ISO/IEC 42001 A.7.4 Quality of data for AI systems
NIST AI RMFNIST AI RMF MN-3.2 Pre-trained models monitored
NIST AI RMF MP-4.2 Internal controls for third-party components
EU AI ActEU AI Act Art. 10 Data and data governance
EU AI Act Art. 15 Accuracy, robustness and cybersecurity

The clauses, set out

ISO/IEC 42001 A.7.6Data preparation

The organization shall define and document criteria for selecting data preparations and the data preparation methods used.

What an auditor asks to see: Data preparation procedures; Transformation scripts; Approval records; Preparation methods (cleaning, labeling, augmentation, balancing); Decisions and rationale; Reproducibility evidence
What an auditor will probe: Are labeling decisions reviewed for bias?
Source: ISO/IEC 42001:2023
ISO/IEC 42001 A.7.4Quality of data for AI systems

The organization shall define and document quality requirements for data and ensure they are met.

What an auditor asks to see: Data quality standards; Quality assessment reports; Quality dimensions defined (accuracy, completeness, representativeness, bias); Measurement evidence; Remediation records
What an auditor will probe: Is representativeness and bias assessed for training data?
Source: ISO/IEC 42001:2023
NIST AI RMF MN-3.2Pre-trained models monitored

Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance. Pre-trained and transfer-learned models are treated as a monitored component in their own right, since their provenance and behaviour are not controlled by the deploying organisation.

What an auditor asks to see: Identification of pre-trained models used, with version and provenance; Monitoring records covering those models within regular maintenance; Assessment of risks carried over from the pre-training data and objective; The procedure followed when the upstream model is updated or withdrawn
What an auditor will probe: Pre-trained model treated as a fixed dependency and excluded from monitoring; Provenance of pre-training data unknown and unrecorded; No procedure for an upstream model being deprecated
Source: NIST AI Risk Management Framework
NIST AI RMF MP-4.2Internal controls for third-party components

Internal risk controls for components of the AI system including third-party AI technologies are identified and documented. For each component carrying risk, the internal control applied to it is identified and written down, so the risk mapping produces controls rather than a list.

What an auditor asks to see: The internal controls identified for each AI system component; Controls specific to third-party and open-source AI technologies; The pre-adoption evaluation practice for third-party material; Evidence the controls named are actually in place
What an auditor will probe: Risks identified for components with no control named against them; Freely available third-party material adopted outside the evaluation practice; Controls documented centrally but absent in the deployed pipeline
Source: NIST AI Risk Management Framework
EU AI Act Art. 10Data and data governance applies if this system is high-risk under Annex III

High-risk AI systems that make use of techniques involving the training of AI models shall use training, validation and testing data that meet the quality criteria in Art.10(2)-(5): appropriate data governance, examination for possible biases, identification of data gaps/shortcomings, statistically relevant datasets to the intended purpose, and considerations specific to the geographical, contextual, behavioural or functional setting of intended use.

What an auditor asks to see: Data governance procedures; Bias examination records and remediation; Data-quality assessment per dataset
What an auditor will probe: Training data used without bias examination; Datasets not representative of the deployment context
Source: EU AI Act
EU AI Act Art. 15Accuracy, robustness and cybersecurity a provider duty

High-risk AI systems shall be designed and developed in such a way that they achieve an appropriate level of accuracy, robustness, and cybersecurity, and shall perform consistently in those respects throughout their lifecycle. Resilience to errors, faults and inconsistencies; protection against attempts by unauthorised third parties to alter use, output or performance (incl data poisoning, model poisoning, adversarial examples and confidentiality attacks).

What an auditor asks to see: Accuracy/robustness measurements relevant to the intended purpose; Adversarial/data-poisoning threat modelling and mitigation; Cybersecurity controls aligned with state-of-the-art
What an auditor will probe: No adversarial-attack threat modelling; Accuracy claims not supported by test evidence
Source: EU AI Act

Other sources in model-generated