Domain 3: Design High-Performing Architectures

High-Performing and Elastic Compute Solutions (Task 3.2)

Amazon EC2 EC2 Auto Scaling AWS Lambda AWS Fargate Amazon ECS Amazon EKS AWS Batch Amazon EMR Spot Instances Placement Groups
Exam Tip
Scale: horizontal (more instances — the AWS default, requires statelessness) vs vertical (bigger box — simple, has ceiling, brief downtime). Autoscaling signals: CPU, ALB request count per target, SQS backlog (queue depth), custom CloudWatch metrics. Placement groups: Cluster = lowest latency/tight packing; Spread = isolate failures (max 7 per AZ); Partition = big distributed systems (Hadoop). Lambda: more memory = proportionally more CPU — tuning memory tunes speed. Batch = queues for compute jobs; EMR = big data frameworks.

High-Performing and Elastic Compute Solutions

Source: https://docs.aws.amazon.com/wellarchitected/latest/performance-efficiency-pillar/

Elastic compute design has three axes: the SHAPE of scaling (horizontal vs vertical), the SIGNAL that triggers it (what metric actually represents load), and the PLACEMENT of instances (how physical topology affects latency and failure). Get all three right and the fleet bends instead of breaking.

Horizontal vs Vertical Scaling

Approach How Pros Cons
Vertical ("scale up") Bigger instance type No app changes Hard ceiling, cost grows steeply, restart during resize
Horizontal ("scale out") More instances behind a load balancer Near-unlimited, fault-isolated, cheap units Requires stateless app, more moving parts


How they trade off: vertical scaling is a one-line change with a ceiling — the biggest instance is finite, costs grow super-linearly, and every resize is a restart. Horizontal scaling needs statelessness (session data in ElastiCache/DynamoDB, files in EFS/S3) but then has no practical ceiling: 10 instances or 10,000 are the same architecture, and losing one unit is a non-event.

Real use-case: A legacy billing monolith (sticky sessions, local file scratch) hits its ceiling every quarter-end. The bridge: vertical scale to r6i.8xlarge NOW (survives unchanged), while sessions move to ElastiCache and scratch to EFS — after which the ASG scales horizontally and quarter-end costs drop 60%. Vertical is a bridge, not a destination.

Gotchas & interview notes: the exam's "requires application changes" phrasing marks vertical as the no-change answer; "fault isolation" or "no downtime while scaling" marks horizontal. A stateful app that cannot be changed → vertical scaling is legitimate — know when, not just that.

Auto Scaling Design (Metrics and Conditions)

Brief: the scaling SIGNAL must be a proxy for user pain — the metric that actually saturates — or the fleet scales on the wrong thing.

How it works: target tracking ("keep CPU at 60%," "keep ALB requests-per-target at 1,000") is the simplest: you declare the SLO, the policy does the math. Step scaling adds bigger responses to bigger breaches. Scheduled scaling pre-positions capacity for KNOWN peaks (batch windows, business hours). Predictive scaling ML-forecasts cyclical traffic and warms capacity BEFORE the curve arrives — capacity leads load instead of chasing it.

The decoupled signal (senior pattern): for async workers, CPU is a poor proxy (a worker can be blocked on I/O at 5% CPU while the queue grows). Scale on SQS backlog per instance (ApproximateNumberOfMessages / fleet size) — the metric that directly measures how far behind the system is.

Real use-case: A video-transcoding fleet scaled on CPU and thrashed: encoders use GPU+network, so CPU stayed under 40% while queue age hit an hour. Switching to backlog-per-instance scaling made the fleet track WORK, not CPU — queue age returned to single-digit minutes and the p99 user wait with it.

Gotchas & interview notes: scale-in flapping is prevented by cooldowns/warm-up and step adjustments — know why, not just the knob. Predictive scaling pairs with target tracking (forecast the floor, react to spikes). Choose the metric by what the workload is short of: CPU (compute-bound), request-count (web), queue depth (async), memory (caching/JVM — often a custom CloudWatch metric, which is a valid answer).

Compute Portfolio — Match the Workload

Service Sweet Spot
EC2 (ASG) Steady/variable web apps, full OS control, GPUs
Lambda Event-driven bursts, <15 min, pay-per-use; more memory = more CPU (single knob performance tuning)
Fargate Containers, no cluster ops
ECS/EKS Cluster cost-tuning at scale, GPUs, K8s ecosystem
AWS Batch Queued batch jobs; auto-provisions right compute; Spot-friendly
Amazon EMR Spark/Hadoop/Hive big data processing
EC2 Spot 90% discount for interruptible (batch, rendering, CI fleets)


How the Lambda memory knob works: CPU and network share scale WITH memory — at 512 MB a function gets ~fractional vCPU; at 1,769 MB it gets a full vCPU; at 10,240 MB, ~6 vCPU. Doubling memory can HALVE duration, and since you pay for MEMORY × TIME, a faster run can cost LESS in total. There is no separate CPU dial — benchmark memory upward until cost-per-invocation bottoms out.

Gotchas & interview notes: "Lambda is slow and costs too much" → tune memory up (counterintuitive but constantly correct). "checkpointing" and "interruption tolerant" language → Spot. Batch on Spot with automatic retry of interrupted jobs is the canonical cheap-batch answer. EMR vs Glue: EMR when you manage the cluster/tuning/framework versions; Glue when serverless is worth it.

Placement Groups (Frequently Tested)

Type Behavior Use
Cluster Same rack/high-bandwidth switch — lowest latency, shared failure domain HPC tightly coupled, ultra-low latency needs
Spread Max 7 instances per AZ, distinct hardware Small HA clusters (quorum databases)
Partition Divided across partitions (dedicated racks), up to 7 per AZ HDFS/Cassandra/Hadoop — rack awareness


How to reason it: the question is "do instances need to be CLOSE together or FAR apart?" Close (low inter-node latency, high bisection bandwidth) → cluster — accepting a shared failure domain. Far (isolate correlated hardware failure) → spread for small critical sets (7 per AZ limit), partition for large distributed data systems that understand rack awareness (HDFS replica placement, Cassandra ring).

Real use-case: A 3-node etcd quorum in spread placement survives the rack failure that would have taken out all three cluster-placed nodes — spread exists precisely when "2 of 3 down" means outage. Meanwhile the 200-node Spark EMR cluster sits in partition placement so one rack failure degrades one partition, not the cluster.

Distributed and Edge Compute

  • AWS Wavelength/Local Zones: millisecond latency for metro/5G users
  • Outposts: AWS hardware on-prem for hybrid low-latency
  • Global Accelerator + edge: route to nearest healthy Region over the AWS backbone
  • API Gateway → Lambda enqueues jobs to SQS
  • AWS Batch on Spot Fargate/EC2 (or an ECS service) drains the queue, scaling on backlog depth — Spot keeps cost 60–90% lower, checkpointing tolerates interruptions
  • Failures retry ×3 → DLQ
  • Steady front-end API on Lambda with memory tuned 512 MB → 1,024 MB (2x speed for 2x per-invocation cost — cheaper overall due to shorter runs)

Worked Example: Image-Processing Pipeline

Requirements: 2M images/day, bursty (10x spikes at holidays), each takes ~60 s.

  • AWS Wavelength/Local Zones: millisecond latency for metro/5G users
  • Outposts: AWS hardware on-prem for hybrid low-latency
  • Global Accelerator + edge: route to nearest healthy Region over the AWS backbone
  • API Gateway → Lambda enqueues jobs to SQS
  • AWS Batch on Spot Fargate/EC2 (or an ECS service) drains the queue, scaling on backlog depth — Spot keeps cost 60–90% lower, checkpointing tolerates interruptions
  • Failures retry ×3 → DLQ
  • Steady front-end API on Lambda with memory tuned 512 MB → 1,024 MB (2x speed for 2x per-invocation cost — cheaper overall due to shorter runs)
The senior summary: pick the scaling shape (stateless → horizontal), pick the honest signal (queue depth for async, not CPU), pick the placement (close vs far), and let Spot absorb everything interruptible — each choice is a lever with a specific failure it prevents.