Design Scalable and Loosely Coupled Architectures (Task 2.1)
Design Scalable and Loosely Coupled Architectures
Source: https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/Coupling is the hidden variable in every scaling question: tightly coupled systems fail and scale as one unit; loosely coupled ones communicate through messages and fail/scale independently. A checkout outage in a loosely coupled store delays orders; in a monolith it takes the storefront with it.
The Messaging Toolkit
| Service | Pattern | Signature Fact |
|---|---|---|
| Amazon SQS | Queue — consumers pull | Buffers spikes; visibility timeout; DLQ for poison messages; FIFO = exactly-once ordering |
| Amazon SNS | Pub/sub — topic pushes | Fan-out to SQS/Lambda/email/SMS/mobile; filter policies per subscriber |
| Amazon EventBridge | Event bus + rules | Route by JSON pattern; schedule; 90+ SaaS/AWS event sources; archive/replay |
| AWS Step Functions | Workflow state machine | Retries, branching, human approval steps, up to 1-year runs |
SQS mechanics (where scenarios are won)
Brief: A durable buffer between producers and workers — producers PUT, workers PULL, spikes get absorbed instead of propagated.
How it works: Once a worker receives a message, it becomes invisible for the visibility timeout (default 30 s). If the worker crashes or fails to delete the message in time, it REAPPEARS — delivery is at-least-once, so workers must be idempotent (processing twice = same result). Set the visibility timeout LONGER than your worker's maximum processing time (Lambda: longer than the function timeout) or healthy-but-slow workers trigger duplicate processing. A dead-letter queue receives messages that failed maxReceiveCount times — debugging and replay without blocking the main queue. Long polling (20 s) waits for messages instead of returning empty: cheaper, and near-real-time. FIFO queues add strict ordering + exactly-once within message groups (300 msg/s, 3,000 with batching).
Real use-case: Concert on-sale: 100x traffic for 20 minutes. Orders enter SQS — the purchase API returns instantly while ticket-assignment workers drain the queue at their own pace. A failed assignment retries; after 3 tries it lands in the DLQ for staff review. No lost orders, no cascade failure, and the queue depth metric (approximate age of oldest message) becomes the scaling signal for the worker fleet.
Gotchas & interview notes: "messages are being processed twice" → visibility timeout too short (or non-idempotent worker). "poison message blocks the queue" → DLQ with maxReceiveCount. "order matters" → FIFO with message groups. SQS scales workers on queue depth (CloudWatch metric ApproximateNumberOfMessagesVisible) — a better signal than CPU for async work.
SNS + SQS fan-out (exam classic)
How it works: publisher → SNS topic → N SQS queues (each subscribed). Every downstream team consumes at its own pace, independently; filter policies let subscribers receive only relevant subsets; adding a consumer = subscribing a new queue — zero producer changes. SNS also pushes directly to Lambda, email, SMS, and mobile push.
Real use-case: an "OrderPlaced" event fans out to fulfillment, email, and analytics queues. Analytics is down for an hour — its queue quietly accumulates 40k messages and catches up overnight. The other consumers never noticed, and no data was lost: the fan-out pattern turns consumer outages into delays.
Gotchas & interview notes: SNS alone retries failed deliveries then can DLQ them, but SQS adds durable buffering — "fan out AND buffer" is always SNS → SQS. EventBridge replaces SNS when you need content-based rules, SaaS sources, or archive/replay.
Event-driven architecture (EDA)
How it works: producers emit FACTS ("OrderPlaced"); EventBridge rules match JSON patterns and route to targets (Lambda, SQS, Step Functions, API destinations). Because events are facts (not commands), producers do not know consumers — the ultimate loose coupling. EventBridge archives events for replay (reprocess after a bug fix) and Schema Registry documents contracts.
Real use-case: A pricing bug corrupts a day of invoice generations. The team fixes the Lambda, then REPLAYS the archived events from EventBridge — invoices regenerate correctly without touching producers or re-collecting data.
Serverless, Containers, and Microservices
Stateless vs stateful — the scaling enabler: if no session/user data lives on the compute node, ANY instance can serve ANY request and the fleet scales freely. Keep state in ElastiCache/DynamoDB/RDS, not on local disk.
| Compute Choice | Pick When |
|---|---|
| Lambda | Event-driven, <15 min, spiky, zero admin |
| Fargate | Containers without cluster management |
| ECS/EKS on EC2 | Need cluster control, cost tuning at scale, GPU nodes |
| EC2 ASG | Legacy apps, full OS access, steady state |
API Gateway is the front door of loosely coupled designs: managed REST/HTTP APIs with throttling (rate + burst limits — protecting backends from traffic floods), auth (Cognito/IAM/custom authorizers), response caching, canary deploys, and usage plans per API key (monetization/partner tiers).
Microservices vs monolith (senior trade-off): microservices buy independent scaling/deployment/failure isolation and pay in distributed-systems complexity (eventual consistency, network failure modes, observability). Start modular-monolith; extract services when a TEAM boundary or SCALING SIGNAL demands it — never "because Netflix did."
Caching and Edge Acceleration
- ElastiCache (Redis/Memcached): cache DB reads and sessions — Redis adds persistence, replication, and rich data structures (sorted sets for leaderboards)
- DynamoDB DAX: microsecond reads for DynamoDB
- CloudFront: cache static AND (with care) dynamic content at edge; origins can be S3/ALB/API Gateway
- Read-heavy (90% SELECTs) → RDS/Aurora read replicas (up to 15; async; promotable; cross-Region capable)
- Connection storms from Lambda/serverless → RDS Proxy (connection pooling — thousands of short-lived functions share hundreds of real connections)
- Write/both-heavy at scale → DynamoDB with on-demand or provisioned capacity
- Tier 1 — Edge: Route 53 + CloudFront (static assets), WAF
- Tier 2 — Web/API: ALB → stateless ASG/ECS across 3 AZs (session state in ElastiCache)
- Tier 3 — Async work: SQS between API and heavy processors; Step Functions for multi-step flows
- Tier 4 — Data: Aurora with read replica for reporting; DynamoDB for session/cart; S3 for objects
Database Scaling Signals
- ElastiCache (Redis/Memcached): cache DB reads and sessions — Redis adds persistence, replication, and rich data structures (sorted sets for leaderboards)
- DynamoDB DAX: microsecond reads for DynamoDB
- CloudFront: cache static AND (with care) dynamic content at edge; origins can be S3/ALB/API Gateway
- Read-heavy (90% SELECTs) → RDS/Aurora read replicas (up to 15; async; promotable; cross-Region capable)
- Connection storms from Lambda/serverless → RDS Proxy (connection pooling — thousands of short-lived functions share hundreds of real connections)
- Write/both-heavy at scale → DynamoDB with on-demand or provisioned capacity
- Tier 1 — Edge: Route 53 + CloudFront (static assets), WAF
- Tier 2 — Web/API: ALB → stateless ASG/ECS across 3 AZs (session state in ElastiCache)
- Tier 3 — Async work: SQS between API and heavy processors; Step Functions for multi-step flows
- Tier 4 — Data: Aurora with read replica for reporting; DynamoDB for session/cart; S3 for objects
Multi-Tier Architecture Blueprint
The canonical loosely coupled web app:
- ElastiCache (Redis/Memcached): cache DB reads and sessions — Redis adds persistence, replication, and rich data structures (sorted sets for leaderboards)
- DynamoDB DAX: microsecond reads for DynamoDB
- CloudFront: cache static AND (with care) dynamic content at edge; origins can be S3/ALB/API Gateway
- Read-heavy (90% SELECTs) → RDS/Aurora read replicas (up to 15; async; promotable; cross-Region capable)
- Connection storms from Lambda/serverless → RDS Proxy (connection pooling — thousands of short-lived functions share hundreds of real connections)
- Write/both-heavy at scale → DynamoDB with on-demand or provisioned capacity
- Tier 1 — Edge: Route 53 + CloudFront (static assets), WAF
- Tier 2 — Web/API: ALB → stateless ASG/ECS across 3 AZs (session state in ElastiCache)
- Tier 3 — Async work: SQS between API and heavy processors; Step Functions for multi-step flows
- Tier 4 — Data: Aurora with read replica for reporting; DynamoDB for session/cart; S3 for objects