AWS Storage Services (Task 3.6)
AWS Storage Services
Source: https://docs.aws.amazon.com/whitepapers/latest/aws-overview/storage-services.htmlThree storage types, one question: HOW is it accessed? Object = HTTP API; block = a mounted disk on one machine; file = a shared folder many machines mount. Every storage scenario on the exam resolves to this triple.
Three Storage Types — the Core Exam Distinction
| Type | Service | Access Pattern | Key Trait |
|---|---|---|---|
| Object | Amazon S3 / S3 Glacier | HTTP API (GET/PUT) | Unlimited capacity, objects up to 5 TB, metadata-rich |
| Block | Amazon EBS, instance store | Mounted disk on ONE EC2 | Low-latency disk I/O, like a physical drive |
| File | Amazon EFS, Amazon FSx | NFS/SMB shared folder | Many machines mount the SAME filesystem simultaneously |
Amazon S3 — Object Storage
Brief: Buckets (globally named) hold objects with 99.999999999% (11 nines) durability — replication across multiple facilities, versioning, replication (CRR/SRR), static hosting, presigned URLs.
How it works: Objects are stored by key with metadata and served over HTTP(S). Durability comes from redundant writes across multiple facilities; availability (99.95% for Standard) is a separate promise. Lifecycle policies move objects between classes automatically by age/prefix/tag; Intelligent-Tiering moves them by observed access pattern instead.
| Class | Use | Retrieval |
|---|---|---|
| S3 Standard | Frequent access | Instant |
| S3 Standard-IA / One Zone-IA | Infrequent access, ms access | Instant, per-GB retrieval fee |
| S3 Intelligent-Tiering | Unknown/changing patterns; auto-moves tiers | Instant, no retrieval fees, small monitoring fee |
| S3 Glacier Instant Retrieval | Archive read 1x/quarter | Instant, retrieval fee |
| S3 Glacier Flexible Retrieval | Archive, minutes–12 hours | Expedited/Standard/Bulk |
| S3 Glacier Deep Archive | Archive, 1–2x/year | ~12 hours standard |
Real use-case: 400 TB of video: new uploads land in Standard, a lifecycle rule moves objects to Standard-IA at 30 days and Deep Archive at 365 — storage cost drops ~85% with zero application changes. The one-clip-per-quarter retrieval still returns instantly (it sits in Deep Archive — 12-hour wait, planned for).
Gotchas & interview notes: retrieval FEES and minimum-billing days are the trap: Standard-IA bills 30 days minimum and per-GB retrieval — objects deleted after a week cost MORE in IA than Standard. Glacier Flexible/Deep Archive charge for restores. "Cheapest for archive accessed yearly" → Deep Archive; "archive but instant access" → Glacier Instant Retrieval; "unknown access pattern" → Intelligent-Tiering. Moving between AZs/Regions or restoring versions generates request + transfer costs — model before bulk operations.
Block Storage — EBS and Instance Store
Brief: EBS is network-attached block storage (gp3 general SSD, io1/io2 high IOPS, st1/sc1 HDD) attached to ONE instance at a time, living in ONE AZ. The instance store is physical disk on the host — fastest, but ephemeral.
How it works: EBS snapshots are incremental backups to S3 (only changed blocks), usable to rebuild or migrate a volume across AZs/Regions (snapshot → copy → restore). gp3 delivers 3,000 IOPS baseline independent of size; io2 Block Express reaches 256,000 IOPS for transactional databases.
Real use-case: A database's redo-log volume needs consistent sub-millisecond writes: io2 with pre-provisioned IOPS. Its nightly snapshot chain restores point-in-time copies in another AZ during DR drills — proving RTO with data, not hope.
Gotchas & interview notes: instance store data is LOST on stop/termination/host migration — scratch space and buffers only, never databases. EBS is AZ-bound: "move a volume to another AZ" = snapshot and restore (an exam favorite). Stopped instances still bill EBS volumes.
File Storage — EFS and FSx
Brief: EFS is managed NFS for Linux — mountable on hundreds of EC2s across AZs simultaneously, petabyte-scale, grows/shrinks automatically. The FSx family is purpose-built: Windows File Server (SMB + Active Directory), Lustre (HPC; can link an S3 bucket as its repository), NetApp ONTAP, OpenZFS.
Real use-case: An Auto Scaling fleet of WordPress servers needs shared uploads: EFS mounted on every node — any instance can serve any image, scale events are invisible to users. A genomics team links FSx for Lustre to the S3 data lake: jobs read at HPC speeds, results sync back.
Gotchas & interview notes: EBS = one instance, one AZ; EFS = many instances, many AZs — that contrast IS the question. Windows/SMB or AD integration → FSx for Windows, never EFS. HPC parallel file access → FSx for Lustre.
Cached/Hybrid Storage — AWS Storage Gateway
Brief: A virtual appliance that connects on-premises apps to cloud storage: local cache in front, durable storage in S3 (File, Volume, and Tape gateway types).
Real use-case: An on-prem video editor reads proxies from the local cache while masters live in S3 Glacier — local speed on the hot files, cloud durability and price on the cold ones.
Gotchas & interview notes: the exam phrase "on-premises applications needing low-latency access to cloud storage" → Storage Gateway. Snow Family MOVES data once; Storage Gateway keeps it connected.
AWS Backup
Brief: A centralized, policy-driven backup service: define plans (frequency, retention, cold-storage tiering, cross-Region/cross-account copies) covering supported resources — EC2, EBS, RDS, Aurora, DynamoDB, EFS, FSx, Storage Gateway volumes — in one auditable place.
Real use-case: A compliance rule demands "all databases backed up daily, 90-day retention, monthly copy to the DR Region": one AWS Backup plan with a tag-based selection enforces it across 60 resources — instead of 60 per-service configuration pages — and the audit report is generated from the same tool.
Gotchas & interview notes: AWS Backup replaces per-service backup configuration — "centralized backup management" is always the answer. It does NOT replace DR architecture (backups are the last line, not the first).