GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-012 Training Data Management and Quality

Tier 2+AIProviderDeployerGPAI Model ProviderManaged Service Provider

Description

Data used to train, fine-tune, or evaluate AI models is subject to documented data management practices. These include: documented acquisition and selection criteria, quality requirements (completeness, representativeness, accuracy, freshness), labelling and annotation procedures, bias identification and mitigation steps, and handling of data from underrepresented subgroups. Data quality is validated before use. Training datasets are versioned and referenced from the model registry.

Rationale

Data quality is the single largest determinant of AI system quality; undocumented or unvalidated training data is unauditable.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredsatisfied by provider

Obtain the provider's training data summary (AIG-015, released under AIG-034). Where the deployer fine-tunes or evaluates a model on its own data, the practices apply to that data and the row is its own.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredrisk class duty

Art.10(1) makes the quality criteria a condition of development rather than a target and Art.10(3) and (4) state them as a testable standard: the training, validation and testing sets are relevant, sufficiently representative and, to the best extent possible, free of errors and complete in view of the intended purpose, with appropriate statistical properties for the affected persons or groups and an account of the geographical, contextual, behavioural or functional characteristics of the setting the system is used in. Art.10(2) adds the documented content, the collection processes and origins, the preparation operations, the assumptions about what the data represents and the bias examination. That last item makes the setting of the system, not only the data, something the validation record has to name.

Public Body Deployer (EU)stablerequiredsatisfied by provider

Obtain the provider's training data summary (AIG-015, released under AIG-034). Where the deployer fine-tunes or evaluates a model on its own data, the practices apply to that data and the row is its own.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (27)

DSP-21Data Poisoning Prevention & Detectioninformative
DSP-23Data Integrity Checkpartial
DSP-24Data Differentiation and Relevancepartial
GRC-11Bias and Fairness Assessmentinformative
EU-AI-Art.10.1Data Governance — Training, Validation and Testing Dataset Qualityfull
EU-AI-Art.10.2Data Governance — Data Preparation and Bias Managementfull
EU-AI-Art.10.3Data Governance — Dataset Representativeness and Completenessfull
A.7.2Data for development and enhancement of AI systemfull
A.7.3Acquisition of datafull
A.7.4Quality of data for AI systemsfull
A.7.6Data preparationfull
AML.M0007Sanitize Training Datainformative
MG-2.2-004Deployed AI System Value Maintenance | MG-2.2-004informative
MP-1.2-002Interdisciplinary AI Team Composition | MP-1.2-002partial
MP-2.3-002Scientific Integrity and Testing Considerations | MP-2.3-002full
MP-4.1-004AI Technology and Legal Risk Mapping | MP-4.1-004full
MP-4.1-005AI Technology and Legal Risk Mapping | MP-4.1-005partial
MS-1.1-002AI Risk Measurement Approach Selection | MS-1.1-002informative
MS-1.1-007AI Risk Measurement Approach Selection | MS-1.1-007partial
MS-2.10-003AI Privacy Risk Examination | MS-2.10-003partial
MS-2.11-004AI Fairness and Bias Evaluation | MS-2.11-004partial
MS-2.11-005AI Fairness and Bias Evaluation | MS-2.11-005informative
MS-2.2-001Human Subject Evaluation Requirements | MS-2.2-001informative
MS-2.6-002AI System Safety Risk Evaluation | MS-2.6-002partial
MS-2.8-002AI Transparency and Accountability Risks | MS-2.8-002partial
MAP 2.3Scientific Integrity and Testing Considerationspartial
LLM05Data and Model Poisoninginformative

Evidence (2)

recorddocumentmanual

Training dataset documentation record for each AI model, covering acquisition criteria, quality requirements, labelling procedures, bias identification steps, and reference to the versioned dataset in the model registry.

Example: Training Data Card · Customer Intent Dataset v3 (MLflow artefact tag: dataset-card), documenting source, selection criteria, quality validation results, annotator agreement scores, bias review finding, and link to versioned S3 dataset

Test: Request training dataset documentation for a sample of production models. Verify: (1) acquisition and selection criteria are documented, (2) quality validation results are present (completeness, accuracy, representativeness checks), (3) bias identification step and outcome are recorded, (4) dataset is versioned and the version is referenced in the model registry entry, (5) handling of underrepresented subgroups is addressed.

tool_outputtechnicalautomated

Data quality validation report from automated data pipeline tooling (e.g. Great Expectations, dbt tests, Soda) confirming that training datasets passed defined quality checks before model training commenced.

Example: Great Expectations validation result (HTML report, run 2026-01-10) for customer-intent-dataset-v3, showing 97.4% completeness, no null rate violations, and schema conformance pass across all 14 expectations

Test: Request the data quality validation report for a recent training dataset. Verify: (1) expectations cover completeness, accuracy, and representativeness dimensions, (2) all critical expectations passed, (3) report timestamp predates the model training run timestamp, (4) any failed expectations have a documented remediation or waiver.

Questions (2)

boolean

Are the data used to train, fine-tune or evaluate AI models subject to documented data management practices?

Training data quality is the single largest determinant of AI system quality. Practices should include documented quality requirements, bias identification steps, and validation before use.

multi

Which of the following training data management practices are applied before model training begins?

Documented acquisition and selection criteriaQuality validation (completeness, representativeness, accuracy)Labelling or annotation procedures with quality controlsBias identification and mitigation reviewHandling documented for underrepresented subgroupsDataset versioned and referenced in the model registryNone of the above

All six practices are expected for a mature data governance programme. Missing bias identification or subgroup handling documentation creates exposure to fairness failures that surface after deployment.