AI Training Data Provenance Checker
Regime

NIST AI Risk Management Framework: what it asks of a training dataset list

The NIST AI Risk Management Framework. Its MAP, MEASURE, MANAGE and GOVERN subcategories name the legal risk of third-party data (MP-4.1), privacy risk (MS-2.10), bias (MS-2.11), test sets (MS-2.1) and third-party resources (MN-3.1).

Cited for every list.

Findings that cite it

FindingClause
Origin not recordedNIST AI RMF GV-1.6
Terms not recordedNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Terms that need review for this useNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Personal data with no lawful basis recorded, or reused from another purposeNIST AI RMF MS-2.10
Labelled by a vendor, a crowd or a model with no quality check recordedNIST AI RMF MP-4.2
No version or snapshot dateNIST AI RMF MS-2.1
Declared potential overlap between train and testNIST AI RMF MS-2.1
NIST AI RMF MP-2.3
High-risk use with no bias examination recordedNIST AI RMF MS-2.11
Older than your N-year policy threshold (a threshold you set, not a legal deadline)NIST AI RMF MP-2.3
Content where a label belongsNIST AI RMF MS-2.10
Vendor data with no provider namedNIST AI RMF MN-3.1
NIST AI RMF GV-6.1

Source classes anchored here

42
Source classClause
Claims system extractNIST AI RMF MS-2.10
Policy and underwriting systemNIST AI RMF MS-2.10
CRM exportNIST AI RMF MS-2.10
ERP and finance systemNIST AI RMF MS-2.10
HR records and staff surveysNIST AI RMF MS-2.10
Service desk and ticketingNIST AI RMF MS-2.10
Application logs and telemetryNIST AI RMF MS-2.10
Internal document storeNIST AI RMF MS-2.10
Email archiveNIST AI RMF MS-2.10
Data warehouse or lake extractNIST AI RMF MS-2.10
Internal, system not namedNIST AI RMF MS-2.10
Support and call transcriptsNIST AI RMF MS-2.10
Chat logsNIST AI RMF MS-2.10
Reviews on your own siteNIST AI RMF MS-2.10
Survey and complaint free textNIST AI RMF MS-2.10
Call recordingsNIST AI RMF MS-2.10
Licensed vendor datasetNIST AI RMF MP-4.1
NIST AI RMF MN-3.1
NIST AI RMF GV-6.1
Data broker listNIST AI RMF MP-4.1
NIST AI RMF MN-3.1
NIST AI RMF GV-6.1
Annotation vendor outputNIST AI RMF MP-4.1
NIST AI RMF MN-3.1
NIST AI RMF GV-6.1
Licensed image or media libraryNIST AI RMF MP-4.1
NIST AI RMF MN-3.1
NIST AI RMF GV-6.1
Credit reference extractNIST AI RMF MP-4.1
NIST AI RMF MN-3.1
NIST AI RMF GV-6.1
Government open dataNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Academic or research datasetNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Open benchmarkNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Public register or recordsNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Open source codeNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Forum postsNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
News articlesNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Social media postsNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
General web crawlNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Specific website scrapeNIST AI RMF MP-4.1
NIST AI RMF GV-6.1
Synthetic data, generated in houseNIST AI RMF MP-2.3
Synthetic data from a vendorNIST AI RMF MP-2.3
NIST AI RMF MN-3.1
Simulation outputNIST AI RMF MP-2.3
Labels produced by a modelNIST AI RMF MN-3.2
NIST AI RMF MP-4.2
Text generated by a modelNIST AI RMF MN-3.2
NIST AI RMF MP-4.2
Distillation outputsNIST AI RMF MN-3.2
NIST AI RMF MP-4.2
Pre-trained model or embeddingsNIST AI RMF MN-3.2
NIST AI RMF MP-4.2
Data shared under an agreementNIST AI RMF MN-3.1
Joint venture dataNIST AI RMF MN-3.1
Reinsurer, broker or bank feedNIST AI RMF MN-3.1
Industry consortium dataNIST AI RMF MN-3.1

NIST AI Risk Management Framework: every clause cited

10 of the 72 held

The requirement text is our statement of each clause, read against the copy we hold and cited to it; it is not the instrument verbatim.

NIST AI RMF GV-1.6Inventory of AI systems

Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities. A maintained inventory identifies the AI systems and models in use, and the resource given to maintaining it is proportionate to the risk the inventoried systems carry.

What an auditor asks to see: The AI system and model inventory with the fields it captures per entry; The intake process that causes a new system or model to be added; Evidence of periodic reconciliation between the inventory and systems actually in production; Named owner and resourcing for inventory maintenance
What an auditor will probe: Inventory covers models built in house but omits embedded and third-party AI features; Compiled once for an audit and not maintained since; No link from an inventory entry to its documentation, so the entry carries no usable detail
Source: NIST AI Risk Management Framework
NIST AI RMF GV-6.1Third-party and intellectual property risk policy

Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third party’s intellectual property or other rights. Third-party AI risk is addressed by policy covering data, models, software and services obtained externally, including the rights position on training data and model outputs.

What an auditor asks to see: Third-party AI policy covering data, pre-trained models, software and services; Due diligence records for third-party AI components in use; Contract terms addressing intellectual property, data rights and liability for AI components; The intellectual property position recorded for training data and model outputs
What an auditor will probe: Standard vendor due diligence applied with no AI-specific questions; Open source models adopted with no review of the licence or the training data provenance; Policy addresses suppliers but not freely obtained models and datasets
Source: NIST AI Risk Management Framework
NIST AI RMF MN-3.1Third-party resources monitored

AI risks and benefits from third-party resources are regularly monitored, and risk controls are applied and documented. Third-party data, model, software and hardware dependencies are monitored on an ongoing basis, not assessed once at procurement, and the controls applied are recorded.

What an auditor asks to see: The inventory of third-party resources the AI system depends on; Monitoring records showing regular review of those dependencies; The risk controls applied to each and evidence they are in place; The route by which a third-party change reaches the risk owner
What an auditor will probe: Third parties assessed at onboarding and not monitored afterwards; Model or API version changes by the provider go unnoticed; Controls documented in the contract with no operational evidence
Source: NIST AI Risk Management Framework
NIST AI RMF MN-3.2Pre-trained models monitored

Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance. Pre-trained and transfer-learned models are treated as a monitored component in their own right, since their provenance and behaviour are not controlled by the deploying organisation.

What an auditor asks to see: Identification of pre-trained models used, with version and provenance; Monitoring records covering those models within regular maintenance; Assessment of risks carried over from the pre-training data and objective; The procedure followed when the upstream model is updated or withdrawn
What an auditor will probe: Pre-trained model treated as a fixed dependency and excluded from monitoring; Provenance of pre-training data unknown and unrecorded; No procedure for an upstream model being deprecated
Source: NIST AI Risk Management Framework
NIST AI RMF MP-2.3Data collection, selection and TEVV

Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation. The evaluation design is documented as a scientific claim: what was measured, on what data, and whether the measure validly stands for the property claimed.

What an auditor asks to see: Documented experimental design for evaluation of the system; Data collection and selection decisions, with availability, representativeness and suitability addressed; Construct validation showing the metric stands for the property claimed; The TEVV considerations identified and how each is handled
What an auditor will probe: Benchmark accuracy reported with no argument that it measures the property claimed; Data selection undocumented, so representativeness cannot be assessed; Evaluation designed by the same people optimising against it
Source: NIST AI Risk Management Framework
NIST AI RMF MP-4.1Legal risks of components and third-party data

Approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented, as are risks of infringement of a third-party’s intellectual property or other rights. There is a followed approach for mapping the technology and legal risk carried by each component, including data and software obtained from third parties and the rights position attached to them.

What an auditor asks to see: The documented approach for mapping component technology and legal risk; Component inventory identifying third-party data, models and software; Intellectual property and rights analysis for each third-party component; Evidence the approach was followed for the components actually in use
What an auditor will probe: Approach documented but not applied to components adopted since; Pre-trained models used with no analysis of the provenance of their training data; Rights reviewed for commercial components only, not for freely obtained ones
Source: NIST AI Risk Management Framework
NIST AI RMF MP-4.2Internal controls for third-party components

Internal risk controls for components of the AI system including third-party AI technologies are identified and documented. For each component carrying risk, the internal control applied to it is identified and written down, so the risk mapping produces controls rather than a list.

What an auditor asks to see: The internal controls identified for each AI system component; Controls specific to third-party and open-source AI technologies; The pre-adoption evaluation practice for third-party material; Evidence the controls named are actually in place
What an auditor will probe: Risks identified for components with no control named against them; Freely available third-party material adopted outside the evaluation practice; Controls documented centrally but absent in the deployed pipeline
Source: NIST AI Risk Management Framework
NIST AI RMF MS-2.1Test sets documented

Test sets, metrics, and details about the tools used during test, evaluation, validation, and verification (TEVV) are documented. The TEVV record is complete enough for the evaluation to be repeated: which test sets, which metrics, which tools and which versions.

What an auditor asks to see: Documentation of the test sets used, including provenance and composition; The metrics computed and their definitions; Tools and versions used during evaluation; Sufficient detail for the evaluation to be repeated by someone else
What an auditor will probe: Results reported with the test set identified only by filename; Metric named but not defined, so results are not comparable across runs; Tooling versions unrecorded, so a result cannot be reproduced
Source: NIST AI Risk Management Framework
NIST AI RMF MS-2.10Privacy risk examined

Privacy risk of the AI system – as identified in the MAP function – is examined and documented. Privacy examination covers what the AI system makes possible, inference and re-identification from training data and outputs, not only the lawfulness of the input data.

What an auditor asks to see: Privacy risk examination covering training data, inference and outputs; Assessment of re-identification and memorisation risk; Privacy-enhancing measures applied and their assessed effect; Documentation traced to the privacy risks mapping identified
What an auditor will probe: Privacy assessed as lawful basis for input data only; Memorisation and training data extraction not considered; Assessment completed for the original data set and not repeated after retraining
Source: NIST AI Risk Management Framework
NIST AI RMF MS-2.11Fairness and bias evaluated

Fairness and bias – as identified in the MAP function – is evaluated and results are documented. Fairness evaluation states which fairness definition was applied and why, evaluates against it, and records the results including where the definition itself is contested.

What an auditor asks to see: The fairness definition or definitions applied and the reason for choosing them; Evaluation results disaggregated across the relevant groups; Identification of the groups assessed and the basis for that selection; Documentation of trade-offs where fairness definitions conflict
What an auditor will probe: A single fairness metric applied with no argument that it fits the context; Groups assessed limited to those for which attribute data happened to exist; Disparity measured and documented with no decision recorded about it
Source: NIST AI Risk Management Framework