Domain 3: Cloud Technology and Services

AWS Compute Services (Task 3.3)

Amazon EC2 Amazon ECS Amazon EKS AWS Fargate AWS Lambda EC2 Auto Scaling Elastic Load Balancing
Exam Tip
EC2 = virtual servers you manage (pick instance type by workload). Lambda = serverless functions, 15-min max, event-driven. ECS = AWS-flavored containers; EKS = Kubernetes. Fargate = serverless containers (no EC2 to manage). Auto Scaling = elasticity. Load balancer = distribute traffic across targets. Scenario "consistent heavy compute, cheapest at scale" → EC2 Reserved/Savings; "unpredictable short bursts" → Lambda; "containerized microservices, no infra mgmt" → Fargate.

AWS Compute Services

Source: https://docs.aws.amazon.com/whitepapers/latest/aws-overview/compute-services.html

Compute choice is a spectrum of control: EC2 (you own the OS) → containers (you own the image) → serverless (you own only code). The exam maps scenarios onto that spectrum; senior engineers also weigh the operations model and the unit of scale (instance vs task vs request).

Amazon EC2 — Virtual Servers

Brief: Resizable virtual machines. You pick the AMI (OS + software template), instance type (CPU/RAM/network recipe), security group, key pair, and storage (EBS or instance store).

How it works: EC2 instances run on the Nitro hypervisor with hardware-isolated CPU, memory, and (optionally) enclaves. An instance lives in ONE AZ; surviving AZ failure requires an ASG spanning AZs. Billing is per-second (Linux, 60s minimum) or per-hour (Windows), and pricing models trade commitment for discount: On-Demand (no commitment), Savings Plans/Reserved (1–3 year commitment, up to ~72% off), Spot (spare capacity, up to 90% off, 2-minute interruption warning).

Instance families — match the workload:

Family Name Meaning Use Case
General purpose (M, T) Balanced Web servers, small DBs; T3/T4g burstable
Compute optimized (C) CPU heavy Batch processing, gaming servers, HPC
Memory optimized (R, X, z) RAM heavy In-memory caches (Redis), SAP HANA, analytics
Storage optimized (I, D, H) Fast local disk NoSQL (Cassandra), data warehousing, HDFS
Accelerated (P, G, Inf, Trn) GPUs/FPGAs/ML chips ML training, video encoding, graphics


Example: a Cassandra cluster needs massive local IOPS → I-series (storage optimized), not a general purpose type. A GPU-based ML training job → P-series (or Trn for Trainium).

Real use-case: A web tier runs m5.xlarge at 18% average CPU; a burstable t3.xlarge (cheaper, with CPU credits banked during idle) or a smaller m5.large covers the load — classic rightsizing from the Cloud Economics topic.

Gotchas & interview notes: "cheapest for a steady 3-year workload" → Savings Plans/Reserved, never On-Demand. "Batch, can be interrupted" → Spot. Stopped instances still bill their EBS volumes — terminate to fully stop billing. An instance store is EPHEMERAL (lost on stop/terminate) — see the Storage topic.

Containers: ECS, EKS, ECR, Fargate

Brief: Containers package apps with their dependencies; AWS orchestrates them four ways.

How it works: ECS is AWS-native orchestration (simpler, deeply integrated — ALB service discovery, IAM task roles). EKS is managed Kubernetes — the choice when the team already runs K8s or needs its ecosystem. ECR is the private Docker registry. Fargate removes the servers entirely: you specify CPU/memory per task and AWS runs it — "serverless containers."

Real use-case: A microservices team with Kubernetes experience chooses EKS to keep their existing Helm charts and operators; a smaller team ships the same microservices on ECS Fargate with zero nodes to patch. Both are correct — the deciding factor was the operations model, not the technology.

Gotchas & interview notes: with ECS/EKS on EC2 you patch the HOSTS and the images; with Fargate, AWS manages hosts and you own only the image — a shared-responsibility question in disguise. "Containers, no infrastructure management" → Fargate, always.

AWS Lambda — Serverless Functions

Brief: Upload code; AWS runs it event-driven (S3 upload, API Gateway call, schedule, queue message) with zero server management.

How it works: Lambda provisions an execution environment on demand, runs your handler, and bills per request + GB-second of duration — idle costs nothing. Limits to memorize: 15-minute max execution, 10 GB memory, 1,000 concurrent executions per Region (adjustable via reserved concurrency). The first request after idle pays an init/cold-start penalty (environment boot); provisioned concurrency pre-warms environments for latency-sensitive paths.

Example: thumbnail generator — every image uploaded to S3 triggers a Lambda that resizes it and writes back. Zero cost when no images arrive.

Real use-case: An API with 4 million requests/month, each running 50 ms: Lambda costs a few dollars; a 24/7 EC2 fleet for the same traffic costs hundreds — the spikier and shorter the work, the more Lambda wins. But the nightly 40-minute ETL does NOT fit Lambda (15-minute cap) — it moves to Fargate/EC2.

Gotchas & interview notes: anything longer than 15 minutes, stateful, or latency-critical at cold start is a Lambda anti-pattern. Concurrency is the failure mode: a downstream DB can be overwhelmed by 1,000 parallel functions — reserve/throttle to protect it. SQS-driven Lambdas: set the queue visibility timeout LONGER than the function timeout.

EC2 Auto Scaling and Load Balancing

Brief: Auto Scaling adds/removes instances to track demand (elasticity); Elastic Load Balancing distributes traffic across targets and health-checks them.

How it works: An ASG launches from a launch template and enforces min/desired/max counts. Scaling policies: scheduled (known peaks), simple/step (threshold alarms), target-tracking (keep CPU at 50% — the thermostat model). The ELB family: ALB (Layer 7, routes HTTP by path/host, integrates WAF), NLB (Layer 4, ultra-low latency, static IPs, handles millions of requests/sec), GWLB (Layer 3, for third-party virtual appliances). The classic HA recipe: ALB in front of an ASG spanning 3 AZs, with ELB health checks so the ASG replaces unhealthy instances.

Real use-case: Flash-sale traffic hits a retail site: target-tracking scales web instances 4 → 40 in ~8 minutes; the ALB spreads load and drains terminating instances gracefully. After the sale, scale-in returns the fleet to 4 — the whole event cost hours of compute, not months.

Gotchas & interview notes: scale-out takes minutes (boot + warmup) — pre-scale with scheduled scaling for known events. Health-check type matters: EC2 checks see the OS, ELB checks see the APPLICATION (an app that returns 500s is replaced only under ELB health checks). Connection draining ensures in-flight requests finish during termination.

Choosing a Compute Option (Decision Table)

Scenario Best Fit
Legacy enterprise app needing full OS control EC2
Containerized microservices, team knows Kubernetes EKS
Simple container workloads, minimal ops ECS (or ECS Fargate)
Containers, zero infra management Fargate
Event-driven, short-lived, spiky, pay-per-use Lambda
24/7 steady workload, lowest per-hour cost EC2 Reserved/Savings Plans
Interruptible batch/rendering, cheapest per hour EC2 Spot

Worked Example: Three-Tier Architecture

  • Web tier: ALB → Auto Scaling group of general-purpose EC2 across 3 AZs (elasticity + HA)
  • Image-processing tier: SQS queue → Lambda functions (spiky workload, serverless)
  • Batch analytics tier: nightly job on Spot Instances in an ECS cluster with Fargate launch type (interruptible, cheap, no servers)