Cost-Optimized Network Architectures (Task 4.4)
Cost-Optimized Network Architectures
Source: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/Network cost is invisible in the console and enormous in the bill — data transfer charges hide behind ordinary architecture. The discipline: know the price of every PATH, then design traffic to take cheap paths.
Data Transfer Economics (Where the Money Hides)
| Path | Charge |
|---|---|
| Same AZ, private IP | FREE |
| Cross-AZ (either direction) | per-GB both ways |
| Cross-Region | per-GB (most expensive common path) |
| To internet from EC2 | per-GB egress (highest) |
| To internet via CloudFront | per-GB at (cheaper) edge rates |
| Into AWS | FREE |
How to reason it: pricing tracks DISTANCE. Same-AZ private IP is free; cross-AZ bills BOTH directions (a 10 GB cross-AZ exchange costs 20 GB of charges); cross-Region is priciest of the private paths; internet egress tops everything. Since INTO-AWS is always free, designs that PULL data toward the consumer (replica reads local, edge caches) beat designs that PUSH across hops.
Design consequences:
- Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
- Serve end users through CloudFront — edge egress pricing + offload
- Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
- VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
- Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
- DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
- Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
- CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
- Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
- S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
- API Gateway usage plans/throttling protect backend from runaway (cost) loops
- Rate-limit expensive endpoints; back-pressure via SQS depth alarms
- Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
- Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
- User downloads moved behind CloudFront: egress at edge rates + origin offload
- 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
- Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
- Result ≈ 70% reduction; Cost anomaly detection guards regressions
Gotchas & interview notes: cross-AZ billing BOTH ways means a request-response pair pays twice — chatty protocols amplify it. Same-AZ free transfer is private-IP only; through an internet-facing ALB or public IP it bills as internet.
NAT Gateway Cost Patterns
| Option | Cost | Availability | Exam Cue |
|---|---|---|---|
| NAT gateway per AZ | highest (device-hours × N + data) | No cross-AZ dependency (best practice) | Production HA |
| One shared NAT gateway | 1 device-hour charge | Cross-AZ traffic + SPOF | "Most cost-effective" scenarios |
| NAT instance | EC2 price only | You patch/manage; script failover | Cheapest DIY, legacy pattern |
How the NAT trap forms: a NAT gateway charges per device-hour (per AZ where deployed) PLUS per-GB processing — and private subnets reaching S3/DynamoDB through it pay that processing on EVERY GB. The escape hatches: gateway endpoints (S3, DynamoDB) are FREE — route-table entries that bypass the NAT entirely. Interface endpoints (PrivateLink) cost hourly+GB but beat NAT processing for heavy traffic to other AWS services.
Real use-case: Private-subnet ECS tasks pulling containers and layers from S3/ECR through NAT: 3 TB/month × NAT processing ≈ $135+ on top of the gateway-hours. Gateway endpoints for S3 in every route table (plus an ECR interface endpoint) cut NAT processing 70% — free route-table entries against a per-GB meter.
Gotchas & interview notes: "private subnet → S3 without internet and without NAT charges" → gateway endpoint (free) — the single most repeatable network-cost answer on the exam. One NAT gateway is the "cost-effective but SPOF" answer; NAT-per-AZ is the production answer — the exam tests that you know WHICH cue selects which.
Inter-VPC and Hybrid Connectivity Choices
- Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
- Serve end users through CloudFront — edge egress pricing + offload
- Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
- VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
- Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
- DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
- Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
- CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
- Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
- S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
- API Gateway usage plans/throttling protect backend from runaway (cost) loops
- Rate-limit expensive endpoints; back-pressure via SQS depth alarms
- Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
- Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
- User downloads moved behind CloudFront: egress at edge rates + origin offload
- 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
- Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
- Result ≈ 70% reduction; Cost anomaly detection guards regressions
Real use-case: Nightly 300 GB replication to HQ over VPN: 9 TB/month, fine. The company added a nightly 2 TB data-lake sync (60+ TB/month) — the DX crossover was crossed: a 1 Gbps DX port cost less per month than the VPN's internet egress for that volume, with consistent latency as a bonus.
Gotchas & interview notes: peering is NON-TRANSITIVE — 40 VPCs need either 780 peerings (absurd) or a Transit Gateway (1 attachment each). TGW bills per-attachment + per-GB — at LOW inter-VPC volume, a small mesh of peerings can be cheaper; the exam usually signals "many VPCs + simplification" → TGW.
Edge and DNS Economics
- Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
- Serve end users through CloudFront — edge egress pricing + offload
- Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
- VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
- Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
- DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
- Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
- CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
- Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
- S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
- API Gateway usage plans/throttling protect backend from runaway (cost) loops
- Rate-limit expensive endpoints; back-pressure via SQS depth alarms
- Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
- Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
- User downloads moved behind CloudFront: egress at edge rates + origin offload
- 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
- Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
- Result ≈ 70% reduction; Cost anomaly detection guards regressions
Throttling and Bandwidth Shaping
- Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
- Serve end users through CloudFront — edge egress pricing + offload
- Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
- VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
- Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
- DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
- Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
- CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
- Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
- S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
- API Gateway usage plans/throttling protect backend from runaway (cost) loops
- Rate-limit expensive endpoints; back-pressure via SQS depth alarms
- Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
- Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
- User downloads moved behind CloudFront: egress at edge rates + origin offload
- 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
- Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
- Result ≈ 70% reduction; Cost anomaly detection guards regressions
Worked Example: Fixing a $9k/month Surprise Network Bill
Bill analysis: 60% = cross-AZ chatter, 25% = NAT processing, 15% = EC2 internet egress.
- Put latency-tolerant chatty tiers in the same AZ on private IPs (trade resilience consciously)
- Serve end users through CloudFront — edge egress pricing + offload
- Minimize cross-Region chatter: replicate once, consume locally (Aurora Global reads local secondary)
- VPC peering: no transit, non-transitive; full mesh of N VPCs = N(N−1)/2 connections + per-connection route tables — the link is free but operational sprawl grows quadratically; data transfer over peering across Regions/AZs still bills
- Transit Gateway: hub — per-attachment hourly + per-GB data processing; simplifies routing massively at scale; inter-Region TGW peering available
- DX vs VPN: DX port + data-out is cheaper per GB at high volume and consistent; VPN cheap hourly — a math question ("500 TB/month office transfer") → DX; occasional admin access → VPN
- Global Accelerator: fixed IPs + backbone routing — cheaper than building multi-Region front doors; but for cacheable content CloudFront wins
- CloudFront: reduces origin egress AND serves users from edge pricing — nearly always a cost WIN for user-facing content
- Route 53: pennies per zone/query — health checks billed; routing policy choice is availability, not cost, driven
- S3 Transfer Acceleration: only pay when distance makes it faster than plain upload
- API Gateway usage plans/throttling protect backend from runaway (cost) loops
- Rate-limit expensive endpoints; back-pressure via SQS depth alarms
- Microservice pairs rewritten to be AZ-affine (client + its cache co-located): cross-AZ GB ↓ 80%
- Private subnets reaching S3 through NAT → add gateway endpoints in every route table: NAT processing ↓ 70%
- User downloads moved behind CloudFront: egress at edge rates + origin offload
- 40-VPC full mesh → Transit Gateway (route tables collapse, per-GB processing acceptable after traffic fixes)
- Nightly 300 GB batch to HQ over DX renegotiated to off-peak; VPN kept as failover only
- Result ≈ 70% reduction; Cost anomaly detection guards regressions