High-Performing and Scalable Storage Solutions (Task 3.1)
High-Performing and Scalable Storage Solutions
Source: https://docs.aws.amazon.com/wellarchitected/latest/performance-efficiency-pillar/Performance storage design is a two-question discipline: (1) what ACCESS pattern does the workload have (block I/O, shared files, HTTP objects, parallel scratch)? (2) what SPEED does each layer need (IOPS, throughput MB/s, latency)? Every service choice below falls out of those answers.
Choosing the Storage Type (The Root Decision)
| Requirement | Storage Type | Service |
|---|---|---|
| Shared file system across many servers | File | EFS (Linux/NFS), FSx (Windows/SMB) |
| Boot disk / transactional DB disk on one server | Block | EBS, instance store |
| Massive scale, HTTP access, static assets, archives | Object | S3 |
| Extreme parallel throughput for compute jobs | File (parallel) | FSx for Lustre |
| HPC scratch with local NVMe | Block (ephemeral) | Instance store (i-series) |
How to reason it: ONE server that owns its disk → block (EBS) — the OS sees a raw volume and manages filesystem, caching, and I/O scheduling. MANY servers needing one shared namespace → file (EFS/FSx) — parallel access with standard protocols. HTTP-scale distribution, writes-once-read-many, unbounded growth → object (S3) — flat keys, not a filesystem. Trying to use S3 as a database filesystem or EBS as a shared drive fails on both performance and semantics — matching access pattern to storage TYPE is the root decision everything else builds on.
Real use-case: A render farm needs 50 nodes reading the same 2 TB of scene files. EBS is 1:1 (instance-to-volume) — wrong shape. EFS mounts everywhere, but metadata-heavy small-file reads over NFS add latency. FSx for Lustre links to the S3 archive, caches scene files at hundreds of GB/s, and releases scratch space after the render — the access pattern (parallel reads of shared large files) pointed to parallel file storage.
Gotchas & interview notes: "database on EFS" — RDS does not support EFS for data volumes (block only). "shared home directories across an ASG" → EFS (the classic). "instance store" is EPHEMERAL — fastest I/O in EC2, but data dies with the instance; only for scratch, replicas, or rebuild-able data.
Block Performance — EBS Deep Dive
| Volume | IOPS | Throughput | Notes |
|---|---|---|---|
| gp3 | 3,000 base (up to 16,000) | 125 MB/s base (up to 1,000) | Baseline INDEPENDENT of volume size; cheapest general SSD |
| gp2 | 3 IOPS/GB (scales with size) | 250 MB/s max | Legacy; size-tied performance |
| io2/io2 Block Express | up to 256,000 IOPS | up to 4,000 MB/s | 99.999% durability, sub-ms latency; mission-critical DBs |
| st1 (throughput HDD) | — | 500 MB/s | Big sequential reads (logs, data lakes) |
| sc1 (cold HDD) | — | 250 MB/s | Cheapest infrequent access |
How performance is bought (levers in order): gp3 lets you provision IOPS and throughput SEPARATELY from size (a 100 GB volume can have 16,000 IOPS — impossible on gp2, where IOPS = 3 × GB). If 16,000 IOPS is still not enough: RAID 0 across multiple EBS volumes multiplies IOPS and throughput linearly (2 × 32,000-IOPS volumes ≈ 64,000). For the extreme tier, io2 Block Express reaches 256,000 IOPS, 4,000 MB/s with sub-millisecond latency. Instance-side prerequisites matter: EBS-optimized instances, and the instance's own IOPS ceiling (an m5.large cannot consume what an io2 delivers).
Real use-case: A PostgreSQL primary on a 1 TB gp2 volume sustained 3,000 IOPS (3 × 1,000 GB) and stalled under a reporting join. Migration to gp3 at 16,000 IOPS / 1,000 MB/s — same size — cut the query from 40 s to 8 s. Cost rose ~15%; a RAID-0 pair was the next lever if needed.
Gotchas & interview notes: the gp3 decoupling is a favorite exam nuance — "increase IOPS WITHOUT growing the volume" → gp3. st1/sc1 cannot be boot volumes. Throughput (MB/s) matters for big sequential scans (logs, media); IOPS matters for transactional random I/O — pick the volume type by which one the workload needs.
File Performance
- EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
- FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
- FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
- Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
- S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
- S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
- CloudFront in front of S3 for global read latency and cheaper egress
- Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
- AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
- AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
- AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
- AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
- 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
- Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
- Edited masters archived: lifecycle to Glacier Deep Archive
- CDN delivery: CloudFront + signed URLs
- On-prem broadcast playout cache: Storage Gateway file gateway
- Database redo on EC2: io2 Block Express RAID-0 pair
Gotchas & interview notes: "shared, grows and shrinks automatically, Linux" → EFS Elastic throughput. "Windows/SMB/AD-integrated" → FSx for Windows. "HPC reading an S3 data lake fast" → FSx for Lustre with the S3 repository link — data is CACHED into the scratch filesystem at job start and results can sync back.
Object Performance — S3 at Speed
How it works:
- EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
- FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
- FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
- Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
- S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
- S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
- CloudFront in front of S3 for global read latency and cheaper egress
- Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
- AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
- AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
- AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
- AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
- 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
- Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
- Edited masters archived: lifecycle to Glacier Deep Archive
- CDN delivery: CloudFront + signed URLs
- On-prem broadcast playout cache: Storage Gateway file gateway
- Database redo on EC2: io2 Block Express RAID-0 pair
Gotchas & interview notes: "large object uploads are slow" → multipart; "slow from far away" → Transfer Acceleration (test it first — S3 provides a speed comparison tool); "downloads are slow" → byte-range fetches or CloudFront. Encryption with SSE-KMS adds a KMS API call per object — high-throughput workloads use S3 Bucket Keys to reduce that.
Hybrid and Transfer Performance
- EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
- FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
- FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
- Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
- S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
- S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
- CloudFront in front of S3 for global read latency and cheaper egress
- Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
- AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
- AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
- AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
- AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
- 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
- Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
- Edited masters archived: lifecycle to Glacier Deep Archive
- CDN delivery: CloudFront + signed URLs
- On-prem broadcast playout cache: Storage Gateway file gateway
- Database redo on EC2: io2 Block Express RAID-0 pair
Worked Example: Media Company Pipeline
- EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
- FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
- FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
- Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
- S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
- S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
- CloudFront in front of S3 for global read latency and cheaper egress
- Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
- AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
- AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
- AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
- AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
- 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
- Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
- Edited masters archived: lifecycle to Glacier Deep Archive
- CDN delivery: CloudFront + signed URLs
- On-prem broadcast playout cache: Storage Gateway file gateway
- Database redo on EC2: io2 Block Express RAID-0 pair