AIG-027 AI Output Validation and Confidence Controls
Description
AI systems that produce outputs acted upon by users or automated processes have defined acceptable output ranges or confidence thresholds. Outputs below the minimum confidence threshold trigger a defined fallback: human review queue, abstention, or escalation, not silent degradation. Output validation logic is documented and version-controlled. For classification tasks, threshold calibration is tested and its impact on precision/recall documented. Output ranges and thresholds are reviewed after any model update.
Rationale
AI systems that act on low-confidence outputs without disclosure or fallback create uncontrolled risk; confidence-gating is a structural quality control unique to probabilistic systems.
Applicability (9 profiles)
Threshold calibration and validation logic are the provider's. The deployer sets the fallback in its own workflow, in the AIG-022 oversight design, using the thresholds the instructions for use give.
Art.15(1) requires an appropriate level of accuracy achieved by design and held consistently across the lifecycle and Art.13(3) requires the accuracy level and the metrics it was measured against to be stated in the instructions for use. Between them they turn the confidence threshold and the calibration result from an internal tuning parameter into a declared figure a deployer relies on and an authority can test the system against, which is also why a threshold changed after a model update has to move in the instructions.
Threshold calibration and validation logic are the provider's. The deployer sets the fallback in its own workflow, in the AIG-022 oversight design, using the thresholds the instructions for use give.
Framework Mappings (18)
| AIS-10 | Output Validation | partial |
| TVM-13 | Guardrails | partial |
| EU-AI-Art.13.3 | Transparency — Mandatory Content of Instructions for Use | informative |
| EU-AI-Art.14.2 | Human Oversight — Capabilities Assigned to Oversight Persons | partial |
| EU-AI-Art.15.1 | Accuracy, Robustness and Cybersecurity — Performance Standards | partial |
| EU-AI-Art.15.4 | Accuracy, Robustness and Cybersecurity — Benchmarks and Measurement Methodologies | informative |
| EU-AI-Art.15.5 | Accuracy, Robustness and Cybersecurity — Declaration of Accuracy Levels and Metrics | informative |
| A.6.2.4 | AI system verification and validation | partial |
| AML.M0020 | Generative AI Guardrails | informative |
| AML.M0033 | Input and Output Validation for AI Agent Components | informative |
| SI-15 | Information Output Filtering | informative |
| MG-2.2-001 | Deployed AI System Value Maintenance | MG-2.2-001 | partial |
| MG-3.2-008 | Pre-Trained Model Monitoring | MG-3.2-008 | informative |
| MS-2.6-004 | AI System Safety Risk Evaluation | MS-2.6-004 | full |
| MANAGE 2.4 | AI System Deactivation and Override Mechanisms | informative |
| MEASURE 2.3 | AI System Performance Measurement | informative |
| LLM07 | Misinformation | full |
| LLM10 | Improper Output Handling | partial |
Evidence (2)
Output validation configuration for AI systems, documenting defined confidence thresholds, fallback behaviour triggered below threshold, and version-controlled validation logic.
Example: Model serving configuration · fraud-classifier-prod (exported from BentoML or Seldon, YAML): confidence_threshold: 0.82, low_confidence_action: route_to_human_review_queue, abstain_below: 0.60, threshold_version: v3 (git commit abc123), last_reviewed: 2026-01-20
Test: Request the output validation configuration for a sample of AI systems acting on outputs. Verify: (1) confidence thresholds are defined per use case (not a single global default), (2) fallback behaviour is configured (human review queue, abstention, or escalation, not silent pass-through), (3) configuration is version-controlled with a dated review record, (4) for classification tasks, threshold calibration results are documented showing precision/recall impact, (5) thresholds were reviewed after the last model update.
Low-confidence output routing logs demonstrating that outputs below the defined confidence threshold are actually being routed to the defined fallback, rather than passed through silently.
Example: Datadog log query result for fraud-classifier-prod (last 30 days): 2,341 events with confidence < 0.82, action=human_review_queue; 0 events with confidence < 0.82 and action=auto_approve, confirms fallback routing is functioning
Test: Query AI event logs for low-confidence output routing events over a 30-day period. Verify: (1) events with confidence below the configured threshold are present in logs, (2) all such events show the correct fallback action (human review / abstention), (3) no events show auto-approval or silent pass-through below threshold, (4) the volume of low-confidence events is reviewed periodically to inform threshold calibration.
Questions (2)
Do AI systems whose outputs are acted upon have a defined acceptable output range or confidence threshold?
Net-new control: confidence-gating is a structural quality control unique to probabilistic AI systems, not addressed by existing frameworks at an operational level. Outputs acted upon without confidence validation create uncontrolled downstream risk.
What action is taken when an AI output falls below the defined confidence threshold?
Options run from strongest to weakest. Mandatory escalation, human review and abstention all meet the control. Passing a low-confidence output through unchanged does not, because nothing downstream can tell it apart from a confident one.