Cost-Optimized Compute Solutions (Task 4.2)
Cost-Optimized Compute Solutions
Source: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/Compute spend follows one formula: (price per unit) × (units running) × (hours running). Every optimization attacks one of the three terms — purchasing options attack price, right-sizing attacks units, elasticity attacks hours.
Purchasing-Option Economics
| Option | Discount | Locks | Best Workload Profile |
|---|---|---|---|
| On-Demand | — | nothing | Unpredictable, short-lived, canary |
| Compute Savings Plans | up to ~66% | $/hour spend, 1–3 yr | Steady ANY compute: EC2 any family/Region/OS + Lambda + Fargate |
| EC2 Instance Savings Plans / Standard RI | up to ~72% | instance family + Region + OS | Very stable footprint |
| Convertible RI | ~54% | family/Region, can exchange | Need flexibility to change |
| Spot | up to 90% | none (2-min interruption) | Batch, CI, rendering, stateless fleets |
| Capacity Reservation | $0 discount | capacity in AZ | Certainty to launch; not a savings tool |
How the commitment trade works: discounts are AWS paying you for PREDICTABILITY. A 3-year Compute Savings Plan says "I will spend $X/hour on compute" — any family, any Region, EC2+Lambda+Fargate apply it automatically. The deeper 72% discounts lock deeper (instance family + Region + OS). Spot is the mirror image: you sell your flexibility — the 2-minute interruption warning — for up to 90% off.
Matching exercise (exam style):
- 24/7 production API on 30 instances, family could change → Compute Savings Plan
- A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
- Nightly Spark jobs, restartable → Spot
- Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
- AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
- Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
- Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
- BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
- Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
- Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
- EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
- Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
- Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
- Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
- Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
- Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
- Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
- 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
- Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
- Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
- Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
Gotchas & interview notes: "could migrate to Graviton or another family" → Compute SP (flexible), not instance-locked. Spot for stateless/async/batch ONLY — the exam rejects Spot for long in-memory stateful jobs UNLESS checkpointed. Capacity Reservations guarantee LAUNCH capacity but discount nothing — they are an availability tool mislisted as a savings tool in wrong answers.
Right-sizing and Utilization
- 24/7 production API on 30 instances, family could change → Compute Savings Plan
- A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
- Nightly Spark jobs, restartable → Spot
- Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
- AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
- Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
- Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
- BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
- Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
- Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
- EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
- Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
- Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
- Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
- Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
- Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
- Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
- 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
- Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
- Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
- Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
Gotchas & interview notes: memory-bound workloads (JVMs, caches) can show LOW CPU while memory-swapping — rightsizing decisions need both metrics. T-family instances are cheapest for SPiky small workloads; a T-instance running hot EXHAUSTS CPU credits and throttles — sustained load wants M/C/R families.
Serverless and Containers Economics
- 24/7 production API on 30 instances, family could change → Compute Savings Plan
- A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
- Nightly Spark jobs, restartable → Spot
- Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
- AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
- Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
- Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
- BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
- Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
- Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
- EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
- Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
- Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
- Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
- Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
- Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
- Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
- 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
- Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
- Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
- Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
Real use-case: An event-driven thumbnail service ran 2 always-on Fargate services for 40k images/day in bursts. Migrating to Lambda: idle time (nights) now costs $0, and the 2 Fargate services were deleted (−$1.2k/mo). The memory knob was also tuned UP (512 → 1,024 MB): 2x per-invocation cost but half the duration — net cheaper AND faster.
Gotchas & interview notes: "scale to zero" is the serverless superpower — the answer whenever utilization is low or intermittent. "Hibernation" appears for long-running VMs used intermittently (RAM preserved, resume without re-warming). High steady utilization + cost focus → EC2 with bin-packing, NOT Fargate.
Scaling and Availability Tiers (Cost Angle)
- 24/7 production API on 30 instances, family could change → Compute Savings Plan
- A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
- Nightly Spark jobs, restartable → Spot
- Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
- AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
- Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
- Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
- BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
- Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
- Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
- EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
- Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
- Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
- Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
- Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
- Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
- Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
- 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
- Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
- Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
- Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there
Gotchas & interview notes: non-prod schedules (nights/weekends off) recover ~70% of non-prod compute cost — 24×7 dev environments are pure waste unless CI needs them. One ALB with path rules hosting 12 services beats 12 ALBs — "consolidate load balancers" is a legit exam answer.
Worked Example: Reining in a Runaway Compute Bill
Findings (monthly): 400 EC2 On-Demand $86k — 55% CPU idle fleet-wide.
- 24/7 production API on 30 instances, family could change → Compute Savings Plan
- A fixed Oracle database host for 3 years → EC2 Instance SP / Standard RI
- Nightly Spark jobs, restartable → Spot
- Monday-9am internal tool → scheduled Auto Scaling (scale to zero off-hours)
- AWS Compute Optimizer recommends instance sizes from CloudWatch utilization ML
- Cost Explorer rightsizing flags under-utilized EC2 (idle <10% CPU 14+ days)
- Downsize over-provisioned m5.2xlarge → m5.large (≈75% saving on that node)
- BURSTABLE T4g for spiky small workloads (CPU credits) instead of paying for constant headroom
- Lambda: pay per request + GB-second; idle costs zero — perfect for spiky/low-traffic; watch duration×memory: optimize memory (shorter runs can net-cheaper)
- Fargate: pay per vCPU+GB-second of TASK — no idle cluster nodes; vs EC2-backed ECS/EKS cheaper at high steady utilization (packed bin-packing)
- EC2 hibernation: pause with RAM saved to EBS — resume in minutes without re-warming caches; for long-running VMs used intermittently
- Elasticity is a cost feature: target-tracking ASG tracks demand instead of provisioning for peak
- Multi-AZ production (required SLA) vs single-AZ dev/test — availability class by environment; schedule non-prod down nights/weekends (Instance Scheduler on Systems Manager)
- Cross-AZ data transfer: AZ-affinity in latency-tolerant tiers
- Load balancer choice: NLB vs ALB pricing differs; one shared ALB with path rules hosts many services (host/path-based routing) instead of an ALB each
- Compute Optimizer: 120 instances <15% CPU → downsize (−$11k)
- Steady 150 web/API nodes → 3-yr Compute Savings Plans (−~$22k on those)
- 80 batch nodes → Spot with checkpointing + ASG relaunch (−~$32k)
- Dev/test 50 nodes: scheduled 12×5 instead of 24×7 (−~$3k)
- Event-driven thumbnail service: migrate to Lambda — 2 idle Fargate services deleted (−$1.2k)
- Total ≈ 45% reduction with no SLA impact; Budgets + anomaly detection keep it there