AIG-022 Human Oversight of AI Outputs
Description
AI systems whose outputs are used in decisions affecting individuals have a documented oversight design naming, for each system, the oversight measure in force, whether it applies before the output is acted on or after, the point in the process at which it applies and the role that performs it. The design states the competencies the oversight role requires, the training it receives including automation bias and the time allowed for each review. Oversight activity is recorded with the reviewer, the timestamp and the decision. The override rate is monitored against the range the design states as expected.
Rationale
Oversight is the last barrier between a wrong output and a consequence for a person. It also fails quietly: a reviewer with sixty cases an hour and a default of agree produces a complete audit trail and no oversight at all. That is why the time allowance and the override rate are in the control. An override rate at or near zero over a long period is the signal to investigate. Whether review must precede the action is an assurance-depth decision carried by the tier model and the profile, not by this text. AIG-023 holds the override and deactivation mechanisms this oversight relies on and AIG-017 the explanation the reviewer reads.
Applicability (9 profiles)
Assignment of oversight to competent, trained persons and the record of oversight activity are the deployer's (Art.26.2). The oversight measures available are designed by the provider (Art.14) and described in the instructions for use.
Art.14(1) and (3) move oversight upstream from an operating procedure to a design obligation: the system is designed and developed, human-machine interface tools included, so that natural persons can effectively oversee it while it is in use, with the measures commensurate with the risks, the level of autonomy and the context. The oversight design is therefore drawn before the system is placed on the market and travels with it and Art.14(4) fixes the capabilities it has to deliver to the person holding oversight. For remote biometric identification under Annex III, point 1(a), Art.14(5) requires that no action or decision is taken on an identification result unless at least two natural persons with the necessary competence, training and authority have separately verified and confirmed it (extract row EU-AI-Art.14.3); the deployer performs that verification and the provider designs the system so it can be performed.
Art.26(2) binds both seats: oversight is assigned to natural persons who have the competence, the training and the authority to perform it, with the support resources to do so. Art.14(5) adds a step the provider designs and the deployer performs: no action or decision is taken on an identification result produced by a remote biometric identification system under Annex III point 1(a) unless at least two natural persons with the necessary competence, training and authority have separately verified and confirmed it. The extract row, EU-AI-Art.14.3, maps partial to this control since migration 063 with the two-person step named as the gap, so the oversight design states that step for any such system. HRS-013 carries the competence the persons need.
Framework Mappings (18)
| GRC-15 | Human supervision | partial |
| IAM-17 | Output Modification and Special Authorization | partial |
| EU-AI-Art.14.1 | Human Oversight — System Design for Oversight | full |
| EU-AI-Art.14.2 | Human Oversight — Capabilities Assigned to Oversight Persons | full |
| EU-AI-Art.14.3 | Human Oversight — Dual Verification for Biometric Identification | partial |
| EU-AI-Art.26.2 | Deployer Obligations — Human Oversight Assignment | full |
| EU-AI-Art.5.3 | Prohibited Practices — Prior Authorisation of Each Real-Time Remote Biometric Identification Use | informative |
| GDPR-Art.22 | Automated Decision-Making and Profiling | informative |
| AML.M0029 | Human In-the-Loop for AI Agent Actions | informative |
| MG-3.2-008 | Pre-Trained Model Monitoring | MG-3.2-008 | full |
| MP-3.4-005 | Operator Proficiency Processes | MP-3.4-005 | full |
| MS-3.3-002 | User and Community Feedback Processes | MS-3.3-002 | informative |
| MS-4.2-004 | Trustworthiness Measurement with Expert Input | MS-4.2-004 | full |
| GOVERN 3.2 | Human-AI Configuration Roles | full |
| MAP 3.4 | Operator Proficiency Processes | informative |
| MAP 3.5 | Human Oversight Process Definition | full |
| ASI09 | Human-Agent Trust Exploitation | partial |
| LLM07 | Misinformation | informative |
Evidence (2)
Oversight design or operational procedure for each production AI system whose outputs are used in decisions affecting individuals, naming the oversight measure in force, the role that performs it, the competencies required, the training given and the time allowed per review.
Example: Human Oversight Procedure · AI Credit Decisioning System (Confluence), specifying that all AI-flagged decline decisions require human review within 4 hours, reviewer qualification requirements (credit underwriting certification), automation bias awareness training requirement, and escalation path for reviewer disagreement
Test: Request the oversight design for each production AI system whose outputs are used in decisions affecting individuals. Verify: (1) the oversight measure in force is described for that system, (2) the design states whether review precedes the action or follows it and names the point in the process at which it applies, (3) competency requirements for the oversight role are defined, (4) automation bias is addressed in the training or the operator guidance, (5) the time allowed per review is stated and is consistent with the volume of cases the role receives.
Audit trail records showing human review and override events for AI-generated outputs, demonstrating that oversight is operationally active and not merely nominal.
Example: AI-Credit-System override log (Splunk, last 90 days): 1,247 AI decisions reviewed, 89 overrides recorded with reviewer ID, timestamp, and override reason category; override rate 7.1%, consistent with expected range 5–10%
Test: Request the oversight design and the override and review event logs for a 90-day sample. Verify: (1) every review event records the reviewer, the timestamp and the decision, (2) every override event records a reason category, (3) the oversight design states the override range it expects for the system, (4) the measured override rate over the sample is compared against that range and a rate outside it carries a recorded investigation, (5) the review volume in the log matches the population of outputs the design says are reviewed.
Questions (2)
Do AI systems that produce outputs used in decisions affecting individuals have documented human oversight mechanisms proportionate to their risk level?
Human oversight is the last line of defence against harmful AI outputs. It must be substantively designed, not nominal. Oversight persons must have defined competencies, training, and sufficient time to conduct meaningful review.
Which of the following are true of the human oversight of your AI systems used in decisions affecting individuals?
Every item applies to each system in scope. An override rate at or near zero sustained over a long period is a signal that review is nominal rather than substantive, which is why monitoring the rate against an expected range matters more than the rate itself.