INF-013 Infrastructure Redundancy
Description
Production infrastructure is deployed with redundancy to eliminate single points of failure for critical components. Availability architecture (multi-zone, multi-region, or equivalent) is documented and aligned to RTO/RPO targets. Redundancy configurations are tested at defined intervals.
Rationale
Single points of failure in cloud infrastructure result in outages that breach SLA commitments. Verified redundancy is the technical foundation of availability guarantees.
Applicability (9 profiles)
The deployer's own infrastructure. The provider's availability architecture is evidenced under VND-006 and BCM-010.
The deployer's own infrastructure. The provider's availability architecture is evidenced under VND-006 and BCM-010.
Framework Mappings (14)
| BCR-11 | Equipment Redundancy | full |
| DCS-18 | Datacenter Operations Resilience | partial |
| MDS-11 | Model Failure | partial |
| BCR-11 | Equipment Redundancy | full |
| DCS-18 | Datacenter Operations Resilience | partial |
| EU-AI-Art.15.2 | Accuracy, Robustness and Cybersecurity — Resilience and Fail-Safe Design | informative |
| 8.14 | Redundancy of information processing facilities | full |
| NIS2-CIR-13.1 | Supporting Utilities | informative |
| NIS2-CIR-4.2 | Backup and Redundancy Management | informative |
| CP-7(1) | Alternate Processing Site | Separation from Primary Site | informative |
| CP-8(2) | Telecommunications Services | Single Points of Failure | full |
| SC-36 | Distributed Processing and Storage | full |
| SI-13 | Predictable Failure Prevention | partial |
| A1.2 | Environmental Protections, Software, Data Back-Up Processes, and Recovery Infrastructure | partial |
Evidence (2)
Infrastructure deployment configuration showing multi-zone or multi-region redundancy for critical production components, with no single points of failure for services subject to availability SLAs.
Example: AWS CloudFormation template, Terraform configuration, or cloud provider console screenshot showing auto-scaling groups spanning multiple availability zones, load balancer configuration, and multi-AZ database configuration for production services
Test: Review infrastructure-as-code or cloud console configuration for production services. Verify: (1) critical compute services are deployed across at least two availability zones; (2) database services use multi-AZ or equivalent replication; (3) load balancers are configured to route around failed zones; (4) verify the redundancy configuration matches the documented RTO/RPO commitments.
Redundancy test record demonstrating that failover between zones or regions was tested and recovery met RTO targets.
Example: Chaos engineering test report or availability failover drill results (e.g., AWS Fault Injection Simulator run log, or equivalent) documenting the test scenario, results, and measured recovery time
Test: Request the most recent redundancy or failover test record. Verify: (1) the test covered the failure of a primary availability zone or equivalent component; (2) recovery time was measured and documented; (3) measured recovery time is at or below the defined RTO; (4) the test was conducted within the defined interval.
Questions (3)
Is production infrastructure deployed with redundancy for the components whose failure would stop the service?
Where the service runs in a cloud region, multi-zone deployment of compute, database and load balancing is the minimum expected redundancy posture.
What level of infrastructure redundancy is implemented for production services?
Multi-AZ within a single region is the baseline expectation. Multi-region is expected where SLAs commit to recovery times that a single-region failure would breach.
Which of the following apply to your availability architecture?
Options run from the most commonly in place to the least. Redundancy that has never been exercised is a design rather than a capability. The failover most likely to fail is the one that has only ever been drawn.