Design Highly Available and Fault-Tolerant Architectures (Task 2.2)
Design Highly Available and Fault-Tolerant Architectures
Source: https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/design-for-recovery.htmlHigh availability is an ENGINEERING BUDGET problem: the lower your RTO/RPO, the more the design costs to run. The exam gives you requirements and a budget; the senior skill is placing them on the ladder honestly.
Two Vocabulary Words That Anchor Everything
- RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
- RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
- Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
- Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
- Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
- Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
- Aurora: 6 storage copies across 3 AZs; storage auto-heals
- S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
- AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
- Snapshots copy AMIs/EBS cross-Region for reconstruction
- CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
- AWS X-Ray traces requests through tiers to find the slow/failing hop
- Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
- Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
- Region pair: us-east-1 (primary), eu-west-1 (secondary)
- Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
- Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
- Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
- Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
The Four DR Strategies (Ladder of Cost)
| Strategy | Standby Side Runs | RTO | RPO | Relative Cost |
|---|---|---|---|---|
| Backup & restore | Nothing — restore from copies on demand | Hours | Hours (backup cadence) | $ |
| Pilot light | Core (replicated DB) only; everything else provisioned at failover | ~10s of minutes | Minutes (replication) | $$ |
| Warm standby | Scaled-down FULL stack always running | Minutes | Seconds | $$$ |
| Active-active | FULL stack live in 2+ Regions serving traffic | Seconds | ~Zero | $$$$ |
How to answer scenario questions: "Data is critical, budget tight, can tolerate 4-hour downtime" → backup & restore. "Must keep serving within minutes; cost-sensitive" → warm standby. "Cannot lose a transaction and cannot be down" → multi-Region active-active.
Pilot light anatomy
How it works: the core data layer (cross-Region RDS replication) runs continuously; everything else exists as READY-TO-DEPLOY artifacts — AMIs/container images copied to the target Region, CloudFormation templates in version control, a Route 53 failover record pre-created. On disaster: deploy the full stack around the already-warm database and flip DNS. RTO is tens of minutes because infrastructure provisioning is the slow part.
Real use-case: A mid-market SaaS runs its full prod in us-east-1 and a single replicated Aurora + stored templates in us-west-2 — standby cost is ~10% of prod. In the annual game day, the team deploys the stack and fails over in 41 minutes, inside their 60-minute RTO. The drill, not the design, proved the number.
Warm standby vs active-active
Warm standby: a complete mini version (small ASG, small Aurora) runs idle or low-traffic; on failover you scale it up and shift Route 53. Faster than pilot light because everything already runs. Active-active: Route 53 latency/geoproximity routes users to the NEAREST Region; Aurora Global Database (<1 s lag) or DynamoDB Global Tables replicates data; a Regional failure degrades capacity, not availability.
Gotchas & interview notes: in active-active with async replication, a Region failure can lose the last sub-second of writes — true zero-RPO needs synchronous replication (same-Region Multi-AZ) or accepted trade-offs. Route 53 failover speed is bounded by DNS TTL — set low TTLs (30–60 s) on failover records BEFORE you need them.
Availability Building Blocks Within a Region
- RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
- RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
- Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
- Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
- Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
- Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
- Aurora: 6 storage copies across 3 AZs; storage auto-heals
- S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
- AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
- Snapshots copy AMIs/EBS cross-Region for reconstruction
- CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
- AWS X-Ray traces requests through tiers to find the slow/failing hop
- Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
- Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
- Region pair: us-east-1 (primary), eu-west-1 (secondary)
- Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
- Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
- Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
- Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
Gotchas & interview notes: the standby in Multi-AZ is NOT readable — read scaling is replicas, failover is Multi-AZ (the most-tested pairing in SAA). One NAT GW per AZ costs more but removes a cross-AZ failure dependency; the exam tests whether you can ARGUE both sides. Instance-level vs ELB health checks: only ELB checks detect a broken APPLICATION on a healthy OS.
Data Durability Mechanics
- RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
- RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
- Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
- Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
- Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
- Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
- Aurora: 6 storage copies across 3 AZs; storage auto-heals
- S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
- AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
- Snapshots copy AMIs/EBS cross-Region for reconstruction
- CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
- AWS X-Ray traces requests through tiers to find the slow/failing hop
- Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
- Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
- Region pair: us-east-1 (primary), eu-west-1 (secondary)
- Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
- Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
- Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
- Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
Observability and Automation for Resilience
- RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
- RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
- Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
- Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
- Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
- Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
- Aurora: 6 storage copies across 3 AZs; storage auto-heals
- S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
- AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
- Snapshots copy AMIs/EBS cross-Region for reconstruction
- CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
- AWS X-Ray traces requests through tiers to find the slow/failing hop
- Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
- Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
- Region pair: us-east-1 (primary), eu-west-1 (secondary)
- Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
- Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
- Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
- Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
Legacy Applications (Not Cloud-Native)
When the app cannot be rewritten: AWS Elastic Disaster Recovery (DRS) continuously replicates block-level disk changes to AWS, enabling failover of UNMODIFIED servers; warm standby on AWS Outposts for on-prem latency needs. The exam rewards "improve reliability WITHOUT changing the application."
Worked Example: Choosing DR for a Payments Platform
Requirements: zero data loss (RPO ≈ 0), max 60-second outage (RTO ≈ 1 min), global users, cost-conscious CFO.
- RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
- RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
- Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
- Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
- Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
- Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
- Aurora: 6 storage copies across 3 AZs; storage auto-heals
- S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
- AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
- Snapshots copy AMIs/EBS cross-Region for reconstruction
- CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
- AWS X-Ray traces requests through tiers to find the slow/failing hop
- Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
- Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
- Region pair: us-east-1 (primary), eu-west-1 (secondary)
- Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
- Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
- Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
- Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored