GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-013 Training Data Provenance

Tier 2+AIProviderDeployerGPAI Model ProviderManaged Service Provider

Description

The provenance of every dataset used to train, fine-tune or evaluate a production model is recorded. The record names the origin of each source, whether internal, third-party, licensed, web-sourced or synthetic, its licence and copyright status, the date it was collected or acquired, the transformations applied to it and the lawful basis relied on for any personal data it holds. Each provenance record is linked to the model registry entry for the models trained on it and is retained for as long as those models remain in use or under a retention obligation.

Rationale

Provenance is what makes a dispute survivable. A dataset with no recorded origin cannot be removed from a training set when a licence is challenged or a subject exercises a right, because nobody can say which models it reached. It also underpins bias auditing, since a fairness result is only interpretable against the population the data came from. The copyright compliance policy that a general-purpose AI model provider owes, including the reservation of rights under a text and data mining opt-out, is a duty of that role and belongs to the GPAI provider bundle rather than here. AIG-010 holds the model registry this record links to and AIG-012 the quality and bias properties of the dataset itself. The public summary of training content a general-purpose model provider publishes is AIG-048 and is derived from the provenance records here; the crawl-side conduct that decides what enters the datasets is AIG-049.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredsatisfied by provider

Obtain the provider's provenance statement for the training data (AIG-015). Provenance of data the deployer fine-tunes or evaluates on is its own.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredrisk class duty

Art.10(2) makes the collection processes and the origin of each dataset content of the documented data governance for a high-risk system, so the provenance record is read as part of the Annex IV technical documentation and is retained with it for the Art.18(1) ten years after the system is placed on the market, which is longer than the control's own floor of as long as the models are in use.

Public Body Deployer (EU)stablerequiredsatisfied by provider

Obtain the provider's provenance statement for the training data (AIG-015). Provenance of data the deployer fine-tunes or evaluates on is its own.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (29)

DSP-20Data Provenance and Transparencypartial
EU-AI-Art.10.2Data Governance — Data Preparation and Bias Managementfull
GDPR-Art.14.3Timing of Indirect Collection Noticeinformative
GDPR-Art.6.1Lawfulness of Processing: the Six Legal Basesinformative
COP-C-1.2Reproduce and extract only lawfully accessible copyright-protected content when crawling the World Wide Webinformative
COP-C-1.3Identify and comply with rights reservations when crawling the World Wide Webinformative
A.7.5Data provenancefull
AML.M0023AI Bill of Materialsinformative
AML.M0025Maintain AI Dataset Provenancefull
SR-4Provenanceinformative
GV-1.2-001Trustworthy AI Characteristics Integration | GV-1.2-001partial
GV-1.6-003AI System Inventory | GV-1.6-003informative
GV-6.1-008Third-Party AI Risk Policies | GV-6.1-008partial
MG-2.2-002Deployed AI System Value Maintenance | MG-2.2-002full
MG-3.1-004Third-Party AI Risk Monitoring and Controls | MG-3.1-004partial
MG-3.2-003Pre-Trained Model Monitoring | MG-3.2-003informative
MG-4.1-006Post-Deployment AI System Monitoring | MG-4.1-006partial
MP-2.1-001AI System Task and Method Definition | MP-2.1-001full
MP-2.1-002AI System Task and Method Definition | MP-2.1-002informative
MP-4.1-006AI Technology and Legal Risk Mapping | MP-4.1-006partial
MP-4.1-010AI Technology and Legal Risk Mapping | MP-4.1-010full
MS-1.1-001AI Risk Measurement Approach Selection | MS-1.1-001informative
MS-2.11-005AI Fairness and Bias Evaluation | MS-2.11-005partial
MS-2.5-005AI System Validity and Reliability | MS-2.5-005full
MS-2.6-002AI System Safety Risk Evaluation | MS-2.6-002informative
MS-2.9-002AI Model Explainability and Validation | MS-2.9-002informative
GOVERN 6.1Third-Party AI Risk Policiespartial
LLM04Supply Chaininformative
LLM05Data and Model Poisoninginformative

Evidence (3)

recorddocumentmanual

Training data provenance record for each production model, documenting origin (internal, third-party, web-scraped, synthetic), licence and copyright status, collection date, transformations applied, and legal basis for use.

Example: Data Provenance Record · LLM Fine-Tune Dataset v2 (Confluence), listing 4 source datasets: internal CRM exports (contract basis), licensed Common Crawl subset (licence agreement #CC-2024-07), synthetic augmentation (internal generation), with copyright review completed by legal 2025-05-10

Test: Request the provenance records for a sample of datasets used by production models. Verify: (1) the origin of each source is recorded, (2) licence and copyright status is recorded for each source, (3) the lawful basis relied on for any personal data is stated, (4) the transformations applied are described, (5) the record is linked to the model registry entry for the models trained on it, (6) records for retired model versions are retained while those versions remain in use or under a retention obligation.

contractdocumentmanual

Licence agreements or data processing agreements for third-party or licensed training datasets, confirming the organisation has the legal right to use the data for AI training purposes.

Example: Data Licence Agreement with DataProvider Ltd (executed 2024-07-15), explicitly granting rights to use dataset for model training, specifying permitted use scope and restrictions on redistribution of derivative models

Test: Request licence or data processing agreements for all third-party training datasets identified in provenance records. Verify: (1) the agreement explicitly permits use for AI/ML model training, (2) any restrictions on derivative models are identified and assessed against current use, (3) agreements are current (not expired), (4) agreements are stored in a retrievable contract repository.

system_exporttechnicalautomated

Dataset lineage export from the data catalogue linking each production model to the datasets it was trained on and the recorded source of each.

Example: Data catalogue lineage export, 2026-08-31: 9 production models, 14 datasets, each carrying source, acquisition route, licence reference and version

Test: Export the dataset lineage for every model marked production in the model registry. Verify: (1) each production model resolves to at least one dataset, (2) each dataset carries a recorded source and acquisition route, (3) each dataset carries the version the model was trained against rather than the current version alone, (4) a dataset with no recorded source appears as a gap in the export rather than being absent from it, (5) the datasets in the export reconcile with the provenance records held for the same models.

Questions (2)

boolean

Is the provenance of every dataset used to train, fine-tune or evaluate a production model recorded?

Answer for the datasets behind the models currently in production. A record that covers the most recent dataset but not the ones earlier versions were trained on does not meet the control, because the obligation attaches to the model for as long as it is in use.

multi

What does your training data provenance record include for each data source?

Origin of the source, whether internal, third-party, licensed, web-sourced or syntheticLicence and copyright statusCollection or acquisition dateTransformations applied to the dataLawful basis relied on for any personal data the source holdsA link to the model registry entry for the models trained on itNone of the above

All six elements are expected for any source used by a production model. The link to the model registry is the element most often absent and the one that decides whether a disputed source can be traced to the models that consumed it.