GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

DAT-016 Data Masking and Pseudonymisation

Tier 2+ProviderDeployerGPAI Model ProviderManaged Service Provider

Description

Sensitive and personal data is masked, pseudonymised or tokenised in non-production environments, analytics workloads, support tooling and logs where exposure of real data is not required. Masking rules are defined per data classification. Re-identification controls prevent linking pseudonymised data back to individuals without authorisation.

Rationale

Using production personal data in non-production environments significantly expands the attack surface with no business necessity. Pseudonymisation in production reduces breach severity and can satisfy GDPR risk reduction requirements.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredcore
GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredcore
Public Body Deployer (EU)stablerequiredcore
DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (17)

DSP-15Limitation of Production Data Usefull
DSP-17Sensitive Data Protectionpartial
DSP-22Privacy Enhancing Technologiesinformative
LOG-08Audit Logs Sanitizationpartial
DSP-15Limitation of Production Data Usefull
DSP-17Sensitive Data Protectionpartial
LOG-08Audit Logs Sanitizationpartial
GDPR-Art.32.1Technical and Organisational Security Measurespartial
GDPR-Art.5.1fIntegrity and Confidentiality (Security Principle)informative
8.11Data maskingfull
8.33Test informationfull
NIS2-CIR-6.2Secure Development Life Cycleinformative
PM-25Minimization of Personally Identifiable Information Used in Testing, Training, and Researchfull
SI-19De-identificationfull
MG-2.2-009Deployed AI System Value Maintenance | MG-2.2-009partial
MS-2.2-002Human Subject Evaluation Requirements | MS-2.2-002partial
MS-2.2-004Human Subject Evaluation Requirements | MS-2.2-004full

Evidence (2)

configurationtechnicalautomated

Non-production environment configuration or data masking tool settings confirming that production personal data is masked, pseudonymised or replaced with synthetic data.

Example: Data masking configuration export from Delphix / AWS Macie / custom masking scripts, showing transformation rules applied to personally identifiable fields (name, email, phone, national ID) when copying data to staging or dev environments, with sample before/after records redacted

Test: Review the non-production data configuration. Verify: (1) a masking or synthetic data process runs before any production data is copied to non-production, (2) masking rules cover all PII fields identified in the data classification (at minimum: name, email, DOB, national IDs), (3) re-identification from pseudonymised fields is not possible without a separately held key, (4) process is documented and tested.

policydocumentmanual

Non-production data handling policy prohibiting use of real personal data in development, staging or test environments without documented exemption and masking controls.

Example: Non-Production Data Policy (Confluence), approved by DPO and CISO, defining: prohibition on real PII in non-prod, approved masking methods, exemption process (risk assessment + DPO approval), and audit obligations for any exempted use

Test: Request the non-production data handling policy. Verify: (1) explicitly prohibits use of unmasked production personal data in non-production environments, (2) defines an exception process requiring DPO approval, (3) specifies approved masking techniques, (4) approved within 24 months.

Questions (2)

boolean

Is sensitive and personal data masked, pseudonymised or replaced with synthetic data in non-production environments (development, staging, test) and in analytics, support tooling and logs?

Production personal data must not appear unmasked in non-production environments or in internal tooling unless there is a documented and DPO-approved exception with appropriate compensating controls.

select

What approach is used to protect personal data in non-production environments?

Automated masking pipeline runs before production data is copied to any non-production environmentSynthetic or generated test data is used exclusively: no production data in non-productionManual masking applied by engineers on an as-needed basis with no automated enforcementProduction data is used in non-production environments with access restrictions as the primary controlNo controls in place: production data is freely available in non-production environments

Automated masking or exclusive use of synthetic data are the strongest controls. Manual masking without automated enforcement is unreliable. Using production data in non-production environments with access-only controls is not an acceptable substitute for masking.