Domain 4: Design Cost-Optimized Architectures

Cost-Optimized Storage Solutions (Task 4.1)

S3 Storage Classes S3 Lifecycle Policies S3 Intelligent-Tiering S3 Requester Pays EBS gp3 AWS DataSync AWS Storage Gateway AWS Transfer Family AWS Backup
Exam Tip
Storage cost ladder: S3 Standard > IA > Glacier IR > Glacier Flexible > Deep Archive. Lifecycle policies = automatic tiering by age (THE answer for aging data). Unknown access patterns = Intelligent-Tiering. EBS: gp3 decouples IOPS from size (cheaper than gp2 for same performance); rightsize or delete unattached volumes. Batch vs individual uploads — multipart + batching reduces request costs. Requester Pays = bucket owner pays storage, downloader pays transfer.

Cost-Optimized Storage Solutions

Source: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/

Storage cost is a DATA LIFECYCLE problem: bytes are written once, read at a decaying rate, and sit for years. Paying Standard rates for a decade of cold data is the single most common waste pattern in AWS — and the most fixable.

The Storage Cost Ladder (Per GB-Month, Roughly)

Tier Relative Cost Access Retrieval Fee
S3 Standard highest of the S3 tiers instant none
S3 Standard-IA / One Zone-IA ~45% less instant per-GB
S3 Glacier Instant Retrieval ~68% less instant per-GB
S3 Glacier Flexible Retrieval ~82% less minutes–hours per-GB
S3 Deep Archive ~95% less ~12 hours per-GB


How to reason the ladder: the trade is STORAGE price vs ACCESS price. Cold tiers charge less per GB but per-GB retrieval fees and (for Glacier) minimum storage durations. If data is touched rarely, retrieval fees are irrelevant and the GB price dominates — the colder the better. If data is touched monthly, per-GB retrieval can erase the savings: Glacier IR is priced for data accessed once a QUARTER, not once a week.

Decision heuristics:

  • Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
  • Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
  • Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
  • Reproducible temporary data → delete aggressively (lifecycle expire)
  • Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
  • Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
  • S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
  • S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
  • gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
  • Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
  • io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
  • Snapshot → AMI sprawl: audit and prune with tagging policies
  • AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
  • RDS automated backups vs manual snapshots retention tuning
  • Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
  • Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
  • The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
  • Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
  • Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Real use-case: A 2 PB media archive uploaded once, ~1% touched monthly, all in Standard. Lifecycle policy (IA at 30 d, Glacier Flexible at 90 d, Deep Archive at 1 yr) cut aged-segment cost ~90% — the same bytes, correctly priced for their temperature. The most recent 90 days (unpredictable access) went to Intelligent-Tiering, paying a small monitoring fee to avoid GUESSING.

Gotchas & interview notes: early-deletion minimums: 30 d (Standard/IA), 90 d (Glacier Flexible), 180 d (Deep Archive) — deleting sooner still bills the minimum, so lifecycle transitions must respect them. "Access patterns unknown" → Intelligent-Tiering, always (no retrieval fees = no way to lose). One Zone-IA saves more but dies with the AZ — only for re-creatable data.

S3 Cost Mechanics That Show Up in Questions

  • Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
  • Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
  • Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
  • Reproducible temporary data → delete aggressively (lifecycle expire)
  • Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
  • Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
  • S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
  • S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
  • gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
  • Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
  • io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
  • Snapshot → AMI sprawl: audit and prune with tagging policies
  • AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
  • RDS automated backups vs manual snapshots retention tuning
  • Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
  • Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
  • The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
  • Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
  • Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Real use-case: A genomics partner pulls 200 TB/month from a public dataset bucket. Enabling Requester Pays shifted the transfer line to the downloader's bill (they accepted — the data is worth it), while the owner kept only the (tiny, Glacier-tiered) storage cost. For user-facing distribution instead of partner pulls, CloudFront is the better shape (edge egress pricing).

Gotchas & interview notes: request charges are why "many small objects" is an anti-pattern — compact files into larger objects where access allows. Requester Pays downloads require the requester's AWS credentials — it is for identified partners/teams, not anonymous public data (that is CloudFront's job).

EBS Cost Control

  • Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
  • Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
  • Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
  • Reproducible temporary data → delete aggressively (lifecycle expire)
  • Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
  • Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
  • S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
  • S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
  • gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
  • Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
  • io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
  • Snapshot → AMI sprawl: audit and prune with tagging policies
  • AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
  • RDS automated backups vs manual snapshots retention tuning
  • Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
  • Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
  • The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
  • Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
  • Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Real use-case: A decommissioned batch fleet left 40 unattached volumes and 900 snapshots billing silently for 8 months (~$3k) — found via Cost Explorer's resource-level filtering. DLM policies plus a tag-on-create rule now expire unattached volumes after 7 days: the waste pattern was engineered out, not just cleaned up.

Gotchas & interview notes: unattached EBS volumes are the classic "zombie cost" — they bill hourly while attached to NOTHING. gp3's decoupled IOPS means a small volume can keep performance — pay for size OR speed independently.

Backup and Archival Economics

  • Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
  • Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
  • Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
  • Reproducible temporary data → delete aggressively (lifecycle expire)
  • Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
  • Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
  • S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
  • S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
  • gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
  • Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
  • io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
  • Snapshot → AMI sprawl: audit and prune with tagging policies
  • AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
  • RDS automated backups vs manual snapshots retention tuning
  • Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
  • Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
  • The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
  • Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
  • Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
How to reason it: retention math is multiplicative — daily backups kept 365 days = 365 copies; tiered retention (daily-7, weekly-30, monthly-365) = ~26 copies for the same protection window. Every cross-Region copy doubles storage spend — pay it only where the DR plan actually reads from that Region.

Hybrid Transfer Cost Choices

Need Cheapest Good Answer
Scheduled file sync on-prem ↔ AWS AWS DataSync (vs DIY scripts)
Partner SFTP drop AWS Transfer Family (pay per GB/hour, no servers)
On-prem apps needing local cache Storage Gateway (file/volume/tape)
One-shot 100 TB, slow link Snowball (vs paying weeks of DX)


Data transfer into AWS is free; egress to internet is the expensive direction — serve users via CloudFront (edge egress pricing) rather than EC2 direct.

Gotchas & interview notes: the Snowball math: 100 TB over a 100 Mbps dedicated link ≈ 3 months; a Snowball round trip ≈ a week. "How do we move 80 TB once" → Snow; "ongoing nightly sync" → DataSync; "partners drop files" → Transfer Family.

Worked Example: Trimming a 2-PB Media Archive

Current state: all 2 PB in S3 Standard, uploaded once, ~1% touched monthly, monthly bill extreme.

  • Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
  • Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
  • Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
  • Reproducible temporary data → delete aggressively (lifecycle expire)
  • Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
  • Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
  • S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
  • S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
  • gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
  • Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
  • io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
  • Snapshot → AMI sprawl: audit and prune with tagging policies
  • AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
  • RDS automated backups vs manual snapshots retention tuning
  • Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
  • Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
  • The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
  • Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
  • Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
The senior summary: know your data's temperature curve, let lifecycle policies price each era correctly, and audit for zombie resources — the savings are structural, not one-time.