Domain 4: Design Cost-Optimized Architectures

Cost-Optimized Compute Solutions (Task 4.2)

Spot Instances Reserved Instances Savings Plans EC2 Auto Scaling Lambda Fargate Compute Optimizer EC2 Hibernation
Exam Tip
Commitment ladder: On-Demand (0%) → Compute Savings Plans (up to 66%, EC2+Lambda+Fargate flexible) → EC2 Instance SP/RI (up to 72%, family+Region locked). Spot = up to 90% for interruptible. Steady 24/7 → commit; spiky → on-demand/elastic; batch → Spot. Rightsize with Compute Optimizer/Cost Explorer. Lambda = pay only per request+duration (idle free). Hibernation saves boot state for intermittent workloads. Production vs non-prod: turn off dev/test nights+weekends.

Cost-Optimized Compute Solutions

Source: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/

Compute spend follows one formula: (price per unit) × (units running) × (hours running). Every optimization attacks one of the three terms — purchasing options attack price, right-sizing attacks units, elasticity attacks hours.

Purchasing-Option Economics

Option Discount Locks Best Workload Profile
On-Demand — nothing Unpredictable, short-lived, canary
Compute Savings Plans up to ~66% $/hour spend, 1–3 yr Steady ANY compute: EC2 any family/Region/OS + Lambda + Fargate
EC2 Instance Savings Plans / Standard RI up to ~72% instance family + Region + OS Very stable footprint
Convertible RI ~54% family/Region, can exchange Need flexibility to change
Spot up to 90% none (2-min interruption) Batch, CI, rendering, stateless fleets
Capacity Reservation $0 discount capacity in AZ Certainty to launch; not a savings tool


How the commitment trade works: discounts are AWS paying you for PREDICTABILITY. A 3-year Compute Savings Plan says "I will spend $X/hour on compute" — any family, any Region, EC2+Lambda+Fargate apply it automatically. The deeper 72% discounts lock deeper (instance family + Region + OS). Spot is the mirror image: you sell your flexibility — the 2-minute interruption warning — for up to 90% off.

Matching exercise (exam style):

  • 24/7 production API on 30 instances, family could change → Compute Savings Plan
  • A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
  • Nightly Spark jobs, restartable → Spot
  • Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
  • AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
  • Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
  • Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
  • BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
  • Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
  • Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
  • EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
  • Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
  • Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
  • Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
  • Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
  • Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
  • Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
  • 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
  • Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
  • Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
  • Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
Real use-case: A fleet audit: 150 steady web/API nodes → 3-yr Compute Savings Plans (−$22k/mo); 80 restartable batch nodes → Spot with checkpointing + ASG relaunch (−$32k/mo); 50 dev/test nodes → scheduled 12×5 instead of 24×7 (−$3k/mo). Same architecture, same SLAs — the commitment decisions alone cut ~45%, because each workload was priced by its ACTUAL flexibility.

Gotchas & interview notes: "could migrate to Graviton or another family" → Compute SP (flexible), not instance-locked. Spot for stateless/async/batch ONLY — the exam rejects Spot for long in-memory stateful jobs UNLESS checkpointed. Capacity Reservations guarantee LAUNCH capacity but discount nothing — they are an availability tool mislisted as a savings tool in wrong answers.

Right-sizing and Utilization

  • 24/7 production API on 30 instances, family could change → Compute Savings Plan
  • A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
  • Nightly Spark jobs, restartable → Spot
  • Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
  • AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
  • Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
  • Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
  • BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
  • Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
  • Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
  • EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
  • Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
  • Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
  • Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
  • Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
  • Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
  • Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
  • 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
  • Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
  • Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
  • Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
How right-sizing pays twice: an instance at 8% CPU costs 100% of its price — downsizing it 4x saves 75% with zero user-visible impact. The org-wide pattern: rightsizing is a QUARTERLY discipline (drift is constant), not a one-time project — new code changes utilization, so yesterday's right size becomes today's waste.

Gotchas & interview notes: memory-bound workloads (JVMs, caches) can show LOW CPU while memory-swapping — rightsizing decisions need both metrics. T-family instances are cheapest for SPiky small workloads; a T-instance running hot EXHAUSTS CPU credits and throttles — sustained load wants M/C/R families.

Serverless and Containers Economics

  • 24/7 production API on 30 instances, family could change → Compute Savings Plan
  • A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
  • Nightly Spark jobs, restartable → Spot
  • Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
  • AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
  • Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
  • Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
  • BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
  • Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
  • Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
  • EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
  • Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
  • Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
  • Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
  • Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
  • Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
  • Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
  • 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
  • Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
  • Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
  • Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
How the Lambda/Fargate/EC2 curve works: at LOW/INTERMITTENT load, Lambda wins — you pay per invocation and idle is free. At MODERATE steady load, Fargate's per-task pricing beats paying for idle cluster nodes. At HIGH steady utilization, EC2 bin-packing wins — a packed m5.4xlarge running 30 tasks is cheaper per task-hour than 30 Fargate tasks at the same utilization. The crossover points are workload-specific; the DIRECTION is universal.

Real use-case: An event-driven thumbnail service ran 2 always-on Fargate services for 40k images/day in bursts. Migrating to Lambda: idle time (nights) now costs $0, and the 2 Fargate services were deleted (−$1.2k/mo). The memory knob was also tuned UP (512 → 1,024 MB): 2x per-invocation cost but half the duration — net cheaper AND faster.

Gotchas & interview notes: "scale to zero" is the serverless superpower — the answer whenever utilization is low or intermittent. "Hibernation" appears for long-running VMs used intermittently (RAM preserved, resume without re-warming). High steady utilization + cost focus → EC2 with bin-packing, NOT Fargate.

Scaling and Availability Tiers (Cost Angle)

  • 24/7 production API on 30 instances, family could change → Compute Savings Plan
  • A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
  • Nightly Spark jobs, restartable → Spot
  • Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
  • AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
  • Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
  • Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
  • BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
  • Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
  • Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
  • EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
  • Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
  • Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
  • Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
  • Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
  • Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
  • Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
  • 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
  • Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
  • Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
  • Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
How elasticity converts peak-to-average: a fleet provisioned for peak runs at ~15% average utilization; an ASG tracking demand runs at ~60%. Same peak performance, one-quarter of the hours — elasticity is the highest-leverage cost feature in AWS, and it is also FREE (Auto Scaling itself costs nothing).

Gotchas & interview notes: non-prod schedules (nights/weekends off) recover ~70% of non-prod compute cost — 24×7 dev environments are pure waste unless CI needs them. One ALB with path rules hosting 12 services beats 12 ALBs — "consolidate load balancers" is a legit exam answer.

Worked Example: Reining in a Runaway Compute Bill

Findings (monthly): 400 EC2 On-Demand $86k — 55% CPU idle fleet-wide.

  • 24/7 production API on 30 instances, family could change → Compute Savings Plan
  • A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
  • Nightly Spark jobs, restartable → Spot
  • Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
  • AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
  • Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
  • Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
  • BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
  • Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
  • Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
  • EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
  • Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
  • Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
  • Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
  • Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
  • Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
  • Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
  • 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
  • Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
  • Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
  • Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
The senior summary: price each workload by its flexibility (commitments for the steady, Spot for the interruptible), size each unit by its utilization, and let elasticity zero out the idle hours — then keep it there with budgets and anomaly detection.