Domain 4: Design Cost-Optimized Architectures

Cost-Optimized Network Architectures (Task 4.4)

NAT Gateway VPC Endpoints AWS Direct Connect AWS VPN VPC Peering AWS Transit Gateway Amazon CloudFront AWS Global Accelerator Route 53
Exam Tip
NAT: gateway-per-AZ (HA, pricier) vs single shared (cheap, SPOF) vs NAT instance (cheapest, you manage). Cross-AZ & cross-Region transfer bills BOTH directions — same-AZ private IP is free. S3/DynamoDB from VPC → gateway endpoint (FREE). Other services → interface endpoint (cheap vs NAT data charges). Many VPC full-mesh → Transit Gateway. Serving users → CloudFront egress pricing beats EC2 egress. Minimize Region hops; keep data flows AZ-local where latency allows.

Cost-Optimized Network Architectures

Source: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/

Network cost is invisible in the console and enormous in the bill — data transfer charges hide behind ordinary architecture. The discipline: know the price of every PATH, then design traffic to take cheap paths.

Data Transfer Economics (Where the Money Hides)

Path Charge
Same AZ, private IP FREE
Cross-AZ (either direction) per-GB both ways
Cross-Region per-GB (most expensive common path)
To internet from EC2 per-GB egress (highest)
To internet via CloudFront per-GB at (cheaper) edge rates
Into AWS FREE


How to reason it: pricing tracks DISTANCE. Same-AZ private IP is free; cross-AZ bills BOTH directions (a 10 GB cross-AZ exchange costs 20 GB of charges); cross-Region is priciest of the private paths; internet egress tops everything. Since INTO-AWS is always free, designs that PULL data toward the consumer (replica reads local, edge caches) beat designs that PUSH across hops.

Design consequences:

  • Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
  • Serve end users through CloudFront — edge egress pricing + offload
  • Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
  • VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
  • Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
  • DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
  • Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
  • CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
  • Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
  • S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
  • API Gateway usage plans/throttling protect backend from runaway (cost) loops
  • Rate-limit expensive endpoints; back-pressure via SQS depth alarms
  • Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
  • Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
  • User downloads moved behind CloudFront: egress at edge rates + origin offload
  • 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
  • Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
  • Result ≈ 70% reduction; Cost anomaly detection guards regressions
Real use-case: A bill analysis: 60% of a $9k/month surprise was cross-AZ chatter (app tier in AZ-a calling Redis in AZ-b, all day). Making service pairs AZ-affine (each AZ's app uses its own local replica) cut cross-AZ GB by 80% — an availability trade made deliberately: per-AZ replication replaced cross-AZ calls.

Gotchas & interview notes: cross-AZ billing BOTH ways means a request-response pair pays twice — chatty protocols amplify it. Same-AZ free transfer is private-IP only; through an internet-facing ALB or public IP it bills as internet.

NAT Gateway Cost Patterns

Option Cost Availability Exam Cue
NAT gateway per AZ highest (device-hours × N + data) No cross-AZ dependency (best practice) Production HA
One shared NAT gateway 1 device-hour charge Cross-AZ traffic + SPOF "Most cost-effective" scenarios
NAT instance EC2 price only You patch/manage; script failover Cheapest DIY, legacy pattern


How the NAT trap forms: a NAT gateway charges per device-hour (per AZ where deployed) PLUS per-GB processing — and private subnets reaching S3/DynamoDB through it pay that processing on EVERY GB. The escape hatches: gateway endpoints (S3, DynamoDB) are FREE — route-table entries that bypass the NAT entirely. Interface endpoints (PrivateLink) cost hourly+GB but beat NAT processing for heavy traffic to other AWS services.

Real use-case: Private-subnet ECS tasks pulling containers and layers from S3/ECR through NAT: 3 TB/month × NAT processing ≈ $135+ on top of the gateway-hours. Gateway endpoints for S3 in every route table (plus an ECR interface endpoint) cut NAT processing 70% — free route-table entries against a per-GB meter.

Gotchas & interview notes: "private subnet → S3 without internet and without NAT charges" → gateway endpoint (free) — the single most repeatable network-cost answer on the exam. One NAT gateway is the "cost-effective but SPOF" answer; NAT-per-AZ is the production answer — the exam tests that you know WHICH cue selects which.

Inter-VPC and Hybrid Connectivity Choices

  • Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
  • Serve end users through CloudFront — edge egress pricing + offload
  • Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
  • VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
  • Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
  • DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
  • Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
  • CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
  • Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
  • S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
  • API Gateway usage plans/throttling protect backend from runaway (cost) loops
  • Rate-limit expensive endpoints; back-pressure via SQS depth alarms
  • Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
  • Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
  • User downloads moved behind CloudFront: egress at edge rates + origin offload
  • 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
  • Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
  • Result ≈ 70% reduction; Cost anomaly detection guards regressions
How to reason the DX-vs-VPN math: VPN pays hourly ($0.05/tunnel-hour) — nearly free for light, occasional traffic. DX pays a fixed port fee (hundreds/month) but LOW per-GB rates — the crossover is volume: past tens of TB/month, the port fee amortizes and DX wins per byte. Production hybrid runs DX with VPN as failover — availability and economics together.

Real use-case: Nightly 300 GB replication to HQ over VPN: 9 TB/month, fine. The company added a nightly 2 TB data-lake sync (60+ TB/month) — the DX crossover was crossed: a 1 Gbps DX port cost less per month than the VPN's internet egress for that volume, with consistent latency as a bonus.

Gotchas & interview notes: peering is NON-TRANSITIVE — 40 VPCs need either 780 peerings (absurd) or a Transit Gateway (1 attachment each). TGW bills per-attachment + per-GB — at LOW inter-VPC volume, a small mesh of peerings can be cheaper; the exam usually signals "many VPCs + simplification" → TGW.

Edge and DNS Economics

  • Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
  • Serve end users through CloudFront — edge egress pricing + offload
  • Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
  • VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
  • Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
  • DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
  • Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
  • CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
  • Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
  • S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
  • API Gateway usage plans/throttling protect backend from runaway (cost) loops
  • Rate-limit expensive endpoints; back-pressure via SQS depth alarms
  • Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
  • Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
  • User downloads moved behind CloudFront: egress at edge rates + origin offload
  • 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
  • Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
  • Result ≈ 70% reduction; Cost anomaly detection guards regressions

Throttling and Bandwidth Shaping

  • Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
  • Serve end users through CloudFront — edge egress pricing + offload
  • Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
  • VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
  • Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
  • DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
  • Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
  • CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
  • Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
  • S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
  • API Gateway usage plans/throttling protect backend from runaway (cost) loops
  • Rate-limit expensive endpoints; back-pressure via SQS depth alarms
  • Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
  • Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
  • User downloads moved behind CloudFront: egress at edge rates + origin offload
  • 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
  • Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
  • Result ≈ 70% reduction; Cost anomaly detection guards regressions
Gotchas & interview notes: a retry storm (client bug × 100x retries) is a COST incident, not just a reliability one — throttling is a budget control. CloudFront for user-facing egress is the rare "better AND cheaper" answer.

Worked Example: Fixing a $9k/month Surprise Network Bill

Bill analysis: 60% = cross-AZ chatter, 25% = NAT processing, 15% = EC2 internet egress.

  • Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
  • Serve end users through CloudFront — edge egress pricing + offload
  • Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
  • VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
  • Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
  • DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
  • Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
  • CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
  • Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
  • S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
  • API Gateway usage plans/throttling protect backend from runaway (cost) loops
  • Rate-limit expensive endpoints; back-pressure via SQS depth alarms
  • Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
  • Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
  • User downloads moved behind CloudFront: egress at edge rates + origin offload
  • 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
  • Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
  • Result ≈ 70% reduction; Cost anomaly detection guards regressions
The senior summary: price every path before the traffic takes it — same-AZ free, gateway endpoints for S3, edge pricing for users, commitments for volume — and let anomaly detection catch the paths that find new ways to bill you.