Domain 2: Design Resilient Architectures

Design Highly Available and Fault-Tolerant Architectures (Task 2.2)

Route 53 Multi-AZ RDS Aurora Global Database S3 Cross-Region Replication AWS Backup CloudWatch AWS X-Ray EC2 Auto Scaling RDS Proxy
Exam Tip
Memorize the DR strategy ladder with RTO/RPO/cost: backup & restore (hours/days, $) → pilot light (minutes-hours, $$) → warm standby (minutes, $$$) → active-active (seconds, $$$$). RPO = acceptable DATA loss; RTO = acceptable DOWNTIME. Multi-AZ RDS standby is NOT readable (that’s read replicas). Aurora Global = <1s cross-Region lag. Route 53 failover + health checks = DNS-level DR. Watch service quotas in the standby Region!

Design Highly Available and Fault-Tolerant Architectures

Source: https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/design-for-recovery.html

High availability is an ENGINEERING BUDGET problem: the lower your RTO/RPO, the more the design costs to run. The exam gives you requirements and a budget; the senior skill is placing them on the ladder honestly.

Two Vocabulary Words That Anchor Everything

  • RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
  • RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
  • Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
  • Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
  • Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
  • Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
  • Aurora: 6 storage copies across 3 AZs; storage auto-heals
  • S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
  • AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
  • Snapshots copy AMIs/EBS cross-Region for reconstruction
  • CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
  • AWS X-Ray traces requests through tiers to find the slow/failing hop
  • Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
  • Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
  • Region pair: us-east-1 (primary), eu-west-1 (secondary)
  • Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
  • Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
  • Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
  • Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
Every DR strategy is chosen by placing requirements on the RTO/RPO grid. A nightly backup gives RPO = 24 hours; continuous cross-Region replication gives RPO ≈ 0. The gap between those two is money.

The Four DR Strategies (Ladder of Cost)

Strategy Standby Side Runs RTO RPO Relative Cost
Backup & restore Nothing — restore from copies on demand Hours Hours (backup cadence) $
Pilot light Core (replicated DB) only; everything else provisioned at failover ~10s of minutes Minutes (replication) $$
Warm standby Scaled-down FULL stack always running Minutes Seconds $$$
Active-active FULL stack live in 2+ Regions serving traffic Seconds ~Zero $$$$


How to answer scenario questions: "Data is critical, budget tight, can tolerate 4-hour downtime" → backup & restore. "Must keep serving within minutes; cost-sensitive" → warm standby. "Cannot lose a transaction and cannot be down" → multi-Region active-active.

Pilot light anatomy

How it works: the core data layer (cross-Region RDS replication) runs continuously; everything else exists as READY-TO-DEPLOY artifacts — AMIs/container images copied to the target Region, CloudFormation templates in version control, a Route 53 failover record pre-created. On disaster: deploy the full stack around the already-warm database and flip DNS. RTO is tens of minutes because infrastructure provisioning is the slow part.

Real use-case: A mid-market SaaS runs its full prod in us-east-1 and a single replicated Aurora + stored templates in us-west-2 — standby cost is ~10% of prod. In the annual game day, the team deploys the stack and fails over in 41 minutes, inside their 60-minute RTO. The drill, not the design, proved the number.

Warm standby vs active-active

Warm standby: a complete mini version (small ASG, small Aurora) runs idle or low-traffic; on failover you scale it up and shift Route 53. Faster than pilot light because everything already runs. Active-active: Route 53 latency/geoproximity routes users to the NEAREST Region; Aurora Global Database (<1 s lag) or DynamoDB Global Tables replicates data; a Regional failure degrades capacity, not availability.

Gotchas & interview notes: in active-active with async replication, a Region failure can lose the last sub-second of writes — true zero-RPO needs synchronous replication (same-Region Multi-AZ) or accepted trade-offs. Route 53 failover speed is bounded by DNS TTL — set low TTLs (30–60 s) on failover records BEFORE you need them.

Availability Building Blocks Within a Region

  • RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
  • RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
  • Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
  • Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
  • Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
  • Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
  • Aurora: 6 storage copies across 3 AZs; storage auto-heals
  • S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
  • AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
  • Snapshots copy AMIs/EBS cross-Region for reconstruction
  • CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
  • AWS X-Ray traces requests through tiers to find the slow/failing hop
  • Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
  • Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
  • Region pair: us-east-1 (primary), eu-west-1 (secondary)
  • Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
  • Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
  • Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
  • Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
How a Multi-AZ failover actually plays out (real use-case): an AZ impairment hits a fleet — the ALB deregisters that AZ's targets (health checks fail), the ASG launches replacements in healthy AZs, RDS promotes the standby in 60–120 s (RDS Proxy keeps application connections alive through it), and ElastiCache fails over its primary. Users see elevated errors for a minute; nobody sees data loss. Designing is predicting this sequence, not hoping for it.

Gotchas & interview notes: the standby in Multi-AZ is NOT readable — read scaling is replicas, failover is Multi-AZ (the most-tested pairing in SAA). One NAT GW per AZ costs more but removes a cross-AZ failure dependency; the exam tests whether you can ARGUE both sides. Instance-level vs ELB health checks: only ELB checks detect a broken APPLICATION on a healthy OS.

Data Durability Mechanics

  • RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
  • RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
  • Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
  • Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
  • Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
  • Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
  • Aurora: 6 storage copies across 3 AZs; storage auto-heals
  • S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
  • AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
  • Snapshots copy AMIs/EBS cross-Region for reconstruction
  • CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
  • AWS X-Ray traces requests through tiers to find the slow/failing hop
  • Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
  • Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
  • Region pair: us-east-1 (primary), eu-west-1 (secondary)
  • Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
  • Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
  • Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
  • Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored

Observability and Automation for Resilience

  • RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
  • RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
  • Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
  • Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
  • Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
  • Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
  • Aurora: 6 storage copies across 3 AZs; storage auto-heals
  • S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
  • AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
  • Snapshots copy AMIs/EBS cross-Region for reconstruction
  • CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
  • AWS X-Ray traces requests through tiers to find the slow/failing hop
  • Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
  • Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
  • Region pair: us-east-1 (primary), eu-west-1 (secondary)
  • Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
  • Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
  • Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
  • Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
Real use-case: A failover drill stalls: the standby Region's EC2 quota (20 vCPUs) blocks the fleet launch. The quota raise that takes one day to request calmly is impossible to get at 3 a.m. during an outage — quotas are now part of the deployment checklist, not the incident retro.

Legacy Applications (Not Cloud-Native)

When the app cannot be rewritten: AWS Elastic Disaster Recovery (DRS) continuously replicates block-level disk changes to AWS, enabling failover of UNMODIFIED servers; warm standby on AWS Outposts for on-prem latency needs. The exam rewards "improve reliability WITHOUT changing the application."

Worked Example: Choosing DR for a Payments Platform

Requirements: zero data loss (RPO ≈ 0), max 60-second outage (RTO ≈ 1 min), global users, cost-conscious CFO.

  • RPO (Recovery Point Objective) — how much DATA LOSS is tolerable (measured backwards from the incident)
  • RTO (Recovery Time Objective) — how much DOWNTIME is tolerable
  • Multi-AZ everything: ASG across 3 AZs behind an ALB; RDS Multi-AZ (synchronous standby in another AZ, automatic failover 60–120 s, standby NOT readable); ElastiCache Redis with automatic failover; NAT gateway per AZ
  • Eliminate single points of failure: one NAT gateway shared across AZs is a cost choice with an availability price — know both sides
  • Health checks + self-healing: ALB target checks + ASG replacement; Route 53 health checks for DNS failover; RDS Proxy shields failovers by pooling and preserving connections
  • Immutable infrastructure: bake updates into new AMIs/containers rather than patching live servers — rollback = redeploy previous image
  • Aurora: 6 storage copies across 3 AZs; storage auto-heals
  • S3: 11 nines durability; cross-Region replication for Regional disaster; versioning against human error
  • AWS Backup: orchestrated plans with cross-Region AND cross-account copies (isolation from compromise), retention lock (compliance)
  • Snapshots copy AMIs/EBS cross-Region for reconstruction
  • CloudWatch alarms + dashboards: track availability metrics (error rate, latency p99, saturation)
  • AWS X-Ray traces requests through tiers to find the slow/failing hop
  • Auto-remediation patterns: EventBridge on GuardDuty/health events → Lambda
  • Service quotas: request limit increases in the STANDBY Region BEFORE disaster — you cannot launch 500 instances during failover if the quota is 20
  • Region pair: us-east-1 (primary), eu-west-1 (secondary)
  • Aurora Global Database: writes in us-east-1, <1 s replication to eu-west-1 → RPO under a second
  • Full app stack in both Regions behind Route 53 latency routing with health checks; both serve traffic (active-active) → RTO seconds (failover only shifts DNS weight)
  • Backups: AWS Backup hourly cross-account copy (ransomware isolation), 35-day retention
  • Quotas pre-raised in eu-west-1 for 3x burst capacity; quarterly game-day drills exercise the failover with production traffic mirrored
The senior summary: RTO/RPO are business decisions the architecture then satisfies at minimum cost — and every number is only real if a drill has proven it.