AWS Compute Services (Task 3.3)
AWS Compute Services
Source: https://docs.aws.amazon.com/whitepapers/latest/aws-overview/compute-services.htmlCompute choice is a spectrum of control: EC2 (you own the OS) → containers (you own the image) → serverless (you own only code). The exam maps scenarios onto that spectrum; senior engineers also weigh the operations model and the unit of scale (instance vs task vs request).
Amazon EC2 — Virtual Servers
Brief: Resizable virtual machines. You pick the AMI (OS + software template), instance type (CPU/RAM/network recipe), security group, key pair, and storage (EBS or instance store).
How it works: EC2 instances run on the Nitro hypervisor with hardware-isolated CPU, memory, and (optionally) enclaves. An instance lives in ONE AZ; surviving AZ failure requires an ASG spanning AZs. Billing is per-second (Linux, 60s minimum) or per-hour (Windows), and pricing models trade commitment for discount: On-Demand (no commitment), Savings Plans/Reserved (1–3 year commitment, up to ~72% off), Spot (spare capacity, up to 90% off, 2-minute interruption warning).
Instance families — match the workload:
| Family | Name Meaning | Use Case |
|---|---|---|
| General purpose (M, T) | Balanced | Web servers, small DBs; T3/T4g burstable |
| Compute optimized (C) | CPU heavy | Batch processing, gaming servers, HPC |
| Memory optimized (R, X, z) | RAM heavy | In-memory caches (Redis), SAP HANA, analytics |
| Storage optimized (I, D, H) | Fast local disk | NoSQL (Cassandra), data warehousing, HDFS |
| Accelerated (P, G, Inf, Trn) | GPUs/FPGAs/ML chips | ML training, video encoding, graphics |
Example: a Cassandra cluster needs massive local IOPS → I-series (storage optimized), not a general purpose type. A GPU-based ML training job → P-series (or Trn for Trainium).
Real use-case: A web tier runs m5.xlarge at 18% average CPU; a burstable t3.xlarge (cheaper, with CPU credits banked during idle) or a smaller m5.large covers the load — classic rightsizing from the Cloud Economics topic.
Gotchas & interview notes: "cheapest for a steady 3-year workload" → Savings Plans/Reserved, never On-Demand. "Batch, can be interrupted" → Spot. Stopped instances still bill their EBS volumes — terminate to fully stop billing. An instance store is EPHEMERAL (lost on stop/terminate) — see the Storage topic.
Containers: ECS, EKS, ECR, Fargate
Brief: Containers package apps with their dependencies; AWS orchestrates them four ways.
How it works: ECS is AWS-native orchestration (simpler, deeply integrated — ALB service discovery, IAM task roles). EKS is managed Kubernetes — the choice when the team already runs K8s or needs its ecosystem. ECR is the private Docker registry. Fargate removes the servers entirely: you specify CPU/memory per task and AWS runs it — "serverless containers."
Real use-case: A microservices team with Kubernetes experience chooses EKS to keep their existing Helm charts and operators; a smaller team ships the same microservices on ECS Fargate with zero nodes to patch. Both are correct — the deciding factor was the operations model, not the technology.
Gotchas & interview notes: with ECS/EKS on EC2 you patch the HOSTS and the images; with Fargate, AWS manages hosts and you own only the image — a shared-responsibility question in disguise. "Containers, no infrastructure management" → Fargate, always.
AWS Lambda — Serverless Functions
Brief: Upload code; AWS runs it event-driven (S3 upload, API Gateway call, schedule, queue message) with zero server management.
How it works: Lambda provisions an execution environment on demand, runs your handler, and bills per request + GB-second of duration — idle costs nothing. Limits to memorize: 15-minute max execution, 10 GB memory, 1,000 concurrent executions per Region (adjustable via reserved concurrency). The first request after idle pays an init/cold-start penalty (environment boot); provisioned concurrency pre-warms environments for latency-sensitive paths.
Example: thumbnail generator — every image uploaded to S3 triggers a Lambda that resizes it and writes back. Zero cost when no images arrive.
Real use-case: An API with 4 million requests/month, each running 50 ms: Lambda costs a few dollars; a 24/7 EC2 fleet for the same traffic costs hundreds — the spikier and shorter the work, the more Lambda wins. But the nightly 40-minute ETL does NOT fit Lambda (15-minute cap) — it moves to Fargate/EC2.
Gotchas & interview notes: anything longer than 15 minutes, stateful, or latency-critical at cold start is a Lambda anti-pattern. Concurrency is the failure mode: a downstream DB can be overwhelmed by 1,000 parallel functions — reserve/throttle to protect it. SQS-driven Lambdas: set the queue visibility timeout LONGER than the function timeout.
EC2 Auto Scaling and Load Balancing
Brief: Auto Scaling adds/removes instances to track demand (elasticity); Elastic Load Balancing distributes traffic across targets and health-checks them.
How it works: An ASG launches from a launch template and enforces min/desired/max counts. Scaling policies: scheduled (known peaks), simple/step (threshold alarms), target-tracking (keep CPU at 50% — the thermostat model). The ELB family: ALB (Layer 7, routes HTTP by path/host, integrates WAF), NLB (Layer 4, ultra-low latency, static IPs, handles millions of requests/sec), GWLB (Layer 3, for third-party virtual appliances). The classic HA recipe: ALB in front of an ASG spanning 3 AZs, with ELB health checks so the ASG replaces unhealthy instances.
Real use-case: Flash-sale traffic hits a retail site: target-tracking scales web instances 4 → 40 in ~8 minutes; the ALB spreads load and drains terminating instances gracefully. After the sale, scale-in returns the fleet to 4 — the whole event cost hours of compute, not months.
Gotchas & interview notes: scale-out takes minutes (boot + warmup) — pre-scale with scheduled scaling for known events. Health-check type matters: EC2 checks see the OS, ELB checks see the APPLICATION (an app that returns 500s is replaced only under ELB health checks). Connection draining ensures in-flight requests finish during termination.
Choosing a Compute Option (Decision Table)
| Scenario | Best Fit |
|---|---|
| Legacy enterprise app needing full OS control | EC2 |
| Containerized microservices, team knows Kubernetes | EKS |
| Simple container workloads, minimal ops | ECS (or ECS Fargate) |
| Containers, zero infra management | Fargate |
| Event-driven, short-lived, spiky, pay-per-use | Lambda |
| 24/7 steady workload, lowest per-hour cost | EC2 Reserved/Savings Plans |
| Interruptible batch/rendering, cheapest per hour | EC2 Spot |
Worked Example: Three-Tier Architecture
- Web tier: ALB → Auto Scaling group of general-purpose EC2 across 3 AZs (elasticity + HA)
- Image-processing tier: SQS queue → Lambda functions (spiky workload, serverless)
- Batch analytics tier: nightly job on Spot Instances in an ECS cluster with Fargate launch type (interruptible, cheap, no servers)