Cost-Optimized Storage Solutions (Task 4.1)
Cost-Optimized Storage Solutions
Source: https://docs.aws.amazon.com/wellarchitected/latest/cost-optimization-pillar/Storage cost is a DATA LIFECYCLE problem: bytes are written once, read at a decaying rate, and sit for years. Paying Standard rates for a decade of cold data is the single most common waste pattern in AWS — and the most fixable.
The Storage Cost Ladder (Per GB-Month, Roughly)
| Tier | Relative Cost | Access | Retrieval Fee |
|---|---|---|---|
| S3 Standard | highest of the S3 tiers | instant | none |
| S3 Standard-IA / One Zone-IA | ~45% less | instant | per-GB |
| S3 Glacier Instant Retrieval | ~68% less | instant | per-GB |
| S3 Glacier Flexible Retrieval | ~82% less | minutes–hours | per-GB |
| S3 Deep Archive | ~95% less | ~12 hours | per-GB |
How to reason the ladder: the trade is STORAGE price vs ACCESS price. Cold tiers charge less per GB but per-GB retrieval fees and (for Glacier) minimum storage durations. If data is touched rarely, retrieval fees are irrelevant and the GB price dominates — the colder the better. If data is touched monthly, per-GB retrieval can erase the savings: Glacier IR is priced for data accessed once a QUARTER, not once a week.
Decision heuristics:
- Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
- Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
- Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
- Reproducible temporary data → delete aggressively (lifecycle expire)
- Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
- Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
- S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
- S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
- gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
- Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
- io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
- Snapshot → AMI sprawl: audit and prune with tagging policies
- AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
- RDS automated backups vs manual snapshots retention tuning
- Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
- Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
- The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
- Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
- Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Gotchas & interview notes: early-deletion minimums: 30 d (Standard/IA), 90 d (Glacier Flexible), 180 d (Deep Archive) — deleting sooner still bills the minimum, so lifecycle transitions must respect them. "Access patterns unknown" → Intelligent-Tiering, always (no retrieval fees = no way to lose). One Zone-IA saves more but dies with the AZ — only for re-creatable data.
S3 Cost Mechanics That Show Up in Questions
- Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
- Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
- Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
- Reproducible temporary data → delete aggressively (lifecycle expire)
- Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
- Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
- S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
- S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
- gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
- Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
- io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
- Snapshot → AMI sprawl: audit and prune with tagging policies
- AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
- RDS automated backups vs manual snapshots retention tuning
- Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
- Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
- The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
- Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
- Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Gotchas & interview notes: request charges are why "many small objects" is an anti-pattern — compact files into larger objects where access allows. Requester Pays downloads require the requester's AWS credentials — it is for identified partners/teams, not anonymous public data (that is CloudFront's job).
EBS Cost Control
- Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
- Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
- Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
- Reproducible temporary data → delete aggressively (lifecycle expire)
- Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
- Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
- S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
- S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
- gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
- Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
- io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
- Snapshot → AMI sprawl: audit and prune with tagging policies
- AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
- RDS automated backups vs manual snapshots retention tuning
- Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
- Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
- The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
- Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
- Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Gotchas & interview notes: unattached EBS volumes are the classic "zombie cost" — they bill hourly while attached to NOTHING. gp3's decoupled IOPS means a small volume can keep performance — pay for size OR speed independently.
Backup and Archival Economics
- Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
- Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
- Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
- Reproducible temporary data → delete aggressively (lifecycle expire)
- Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
- Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
- S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
- S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
- gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
- Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
- io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
- Snapshot → AMI sprawl: audit and prune with tagging policies
- AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
- RDS automated backups vs manual snapshots retention tuning
- Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
- Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
- The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
- Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
- Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest
Hybrid Transfer Cost Choices
| Need | Cheapest Good Answer |
|---|---|
| Scheduled file sync on-prem ↔ AWS | AWS DataSync (vs DIY scripts) |
| Partner SFTP drop | AWS Transfer Family (pay per GB/hour, no servers) |
| On-prem apps needing local cache | Storage Gateway (file/volume/tape) |
| One-shot 100 TB, slow link | Snowball (vs paying weeks of DX) |
Data transfer into AWS is free; egress to internet is the expensive direction — serve users via CloudFront (edge egress pricing) rather than EC2 direct.
Gotchas & interview notes: the Snowball math: 100 TB over a 100 Mbps dedicated link ≈ 3 months; a Snowball round trip ≈ a week. "How do we move 80 TB once" → Snow; "ongoing nightly sync" → DataSync; "partners drop files" → Transfer Family.
Worked Example: Trimming a 2-PB Media Archive
Current state: all 2 PB in S3 Standard, uploaded once, ~1% touched monthly, monthly bill extreme.
- Access frequency known and declining → lifecycle policy (e.g., Standard → IA at 30 d → Glacier at 90 d → Deep Archive at 365 d)
- Access patterns unknown/changing → S3 Intelligent-Tiering (auto-moves objects between access tiers on usage; no retrieval fees; small monitoring fee; skips objects <128 KB)
- Compliance archives (7-year retention) → Glacier + Object Lock/Vault Lock
- Reproducible temporary data → delete aggressively (lifecycle expire)
- Costs = storage GB + requests (PUT/GET/LIST) + retrieval + data transfer + early-deletion
- Multipart upload + batching: fewer, larger operations = lower request charges and faster; the "batch uploads instead of individual" scenario answer — a million 1-PUT objects cost more in REQUESTS than a thousand 1,000-object batches
- S3 Requester Pays: the bucket owner pays storage; the DOWNLOADER pays transfer — share large public datasets without transfer bills
- S3 Transfer Acceleration is faster but adds cost — only for genuinely distant uploads
- gp3 over gp2: 3,000 IOPS baseline independent of size — rightsize without losing performance; provision extra IOPS/throughput only when needed
- Delete unattached volumes and obsolete snapshots (billing continues while they exist) — automate tagging + lifecycle via Data Lifecycle Manager (DLM)
- io1/io2 only where latency SLAs demand; HDD types (st1/sc1) for big sequential cold data
- Snapshot → AMI sprawl: audit and prune with tagging policies
- AWS Backup centralizes plans: frequency vs retention trade (daily-7d, weekly-30d, monthly-1y tiers); cold-tier older backups; cross-Region copies only where DR requires
- RDS automated backups vs manual snapshots retention tuning
- Compare archive (Glacier) vs backup (AWS Backup) vs DR (replication) — each has a distinct cost profile matched to RPO/RTO
- Apply lifecycle policy: transition to Standard-IA at 30 days, Glacier Flexible at 90 days, Deep Archive at 1 year → ~90% reduction on aged segments
- The 1% "hot unknown" recent month → Intelligent-Tiering for the first 90 days
- Partner downloads (200 TB/month egress) → enable Requester Pays or move distribution behind CloudFront
- Old EBS: 40 unattached volumes found via Cost Explorer rightsizing → snapshot the 3 useful ones, delete the rest