GASP AICF

Search controls and profiles

Search by control ID, name, domain or profile

AIG-019 AI Model Performance and Drift Detection

Tier 2+AIPredictiveProviderDeployerGPAI Model ProviderManaged Service Provider

Description

Deployed AI models are evaluated for performance degradation and distribution shift (data drift, concept drift) on a scheduled basis. The schedule is defined relative to the velocity of the underlying domain (at minimum quarterly for stable domains, monthly for high-velocity domains). Evaluation uses held-out test data or shadow deployments. When performance falls below defined thresholds or drift is detected, a documented escalation path is triggered, ranging from investigation to retraining or decommissioning.

Rationale

Model drift is an AI-specific failure mode that has no equivalent in conventional software; without scheduled evaluation, degraded models operate undetected.

Applicability (9 profiles)

SaaS AI Providerstablerequiredcore
Enterprise AI Deployerstablerequiredcore

Performance on the deployer's own population and input distribution is measured by the deployer. Retraining is the provider's escalation, reached through the AIG-032 change notification terms.

GPAI Model Providerstablerequiredcore
High-Risk Provider (EU)stablerequiredcore
Public Body Deployer (EU)stablerequiredcore

Performance on the deployer's own population and input distribution is measured by the deployer. Retraining is the provider's escalation, reached through the AIG-032 change notification terms.

DORA ICT Provider (EU)stablerequiredcore
NIS2 Cloud Provider (EU)stablerequiredcore

Framework Mappings (16)

DSP-21Data Poisoning Prevention & Detectioninformative
MDS-10Model Continuous Monitoringinformative
EU-AI-Art.15.2Accuracy, Robustness and Cybersecurity — Resilience and Fail-Safe Designpartial
AML.M0008Validate AI Modelinformative
MG-2.2-008Deployed AI System Value Maintenance | MG-2.2-008informative
MG-2.4-004AI System Deactivation and Override Mechanisms | MG-2.4-004informative
MG-3.2-009Pre-Trained Model Monitoring | MG-3.2-009full
MG-4.1-004Post-Deployment AI System Monitoring | MG-4.1-004partial
MP-4.1-008AI Technology and Legal Risk Mapping | MP-4.1-008full
MS-2.6-003AI System Safety Risk Evaluation | MS-2.6-003informative
MS-4.2-002Trustworthiness Measurement with Expert Input | MS-4.2-002full
MANAGE 2.2Deployed AI System Value Maintenancefull
MANAGE 3.2Pre-Trained Model Monitoringfull
MEASURE 1.2AI Metrics and Control Effectiveness Assessmentpartial
MEASURE 2.13Measurement Effectiveness Evaluationpartial
MEASURE 4.3Performance Improvement and Decline Trackingfull

Evidence (2)

reportdocumentmanual

Periodic model performance and drift evaluation report demonstrating that deployed models were assessed for performance degradation and distribution shift on the defined schedule, with comparison against baseline metrics.

Example: Quarterly Drift Report · Customer Churn Predictor (Weights & Biases artefact, Q1 2026), showing PSI score for input features, model accuracy vs baseline, concept drift F1 delta, and escalation decision: 'no action required, within threshold'

Test: Request drift evaluation reports for a sample of production models covering the last two evaluation periods. Verify: (1) reports are dated at the scheduled frequency, (2) both data drift and concept/performance drift are evaluated, (3) results are compared to documented thresholds, (4) escalation decision is recorded (no action / investigation / retrain / decommission), (5) where thresholds were breached, a documented escalation action was taken.

tool_outputtechnicalautomated

Automated drift detection tool output from model monitoring platform (e.g. Evidently AI, Arize, WhyLabs) showing scheduled drift metric computation for production models.

Example: Evidently AI drift report JSON export for fraud-model-prod (weekly run 2026-04-14): feature drift detected on 2/18 features (PSI > 0.2 threshold), dataset drift test: PASS, target drift: PASS, alert fired to ml-monitoring Slack channel

Test: Request automated drift detection tool output for a production model. Verify: (1) drift metrics are computed automatically on the defined schedule (check run timestamps), (2) alert thresholds are configured and alert firing is evidenced, (3) the tool output is linked to the escalation process (Slack/PagerDuty alert or ticket creation), (4) tool is monitoring the live production model (not a shadow environment).

Questions (2)

boolean

Are deployed AI models evaluated on a scheduled basis for performance degradation and distribution shift (data drift, concept drift)?

Model drift is an AI-specific failure mode with no equivalent in conventional software. Without scheduled evaluation, degraded models operate undetected. Evaluation should use held-out test data or shadow deployments and trigger a documented escalation path when thresholds are breached.

select

How frequently are your production AI models evaluated for drift or performance degradation?

Continuously or weekly through automated toolingMonthlyQuarterlyAnnuallyAd hoc, only when an issue is reported

Options run from strongest to weakest. Frequency should match the velocity of the underlying domain: a fast-moving domain such as fraud or content moderation needs monthly evaluation or better. Annual evaluation is insufficient wherever the data environment changes.