Domain 3: Design High-Performing Architectures

High-Performing and Scalable Storage Solutions (Task 3.1)

Amazon S3 Amazon EBS Amazon EFS Amazon FSx AWS Storage Gateway AWS DataSync AWS Transfer Family
Exam Tip
Object vs file vs block decision drives everything. Performance answers: EBS io2 Block Express (256k IOPS), gp3 (16k IOPS decoupled from size), RAID-0 for more, EFS for shared access at scale, FSx Lustre for HPC + S3 link, S3 + CloudFront for static. Transfer performance: multipart upload for S3 (parallel), S3 Transfer Acceleration (edge upload), DataSync for scheduled/hybrid moves.

High-Performing and Scalable Storage Solutions

Source: https://docs.aws.amazon.com/wellarchitected/latest/performance-efficiency-pillar/

Performance storage design is a two-question discipline: (1) what ACCESS pattern does the workload have (block I/O, shared files, HTTP objects, parallel scratch)? (2) what SPEED does each layer need (IOPS, throughput MB/s, latency)? Every service choice below falls out of those answers.

Choosing the Storage Type (The Root Decision)

Requirement Storage Type Service
Shared file system across many servers File EFS (Linux/NFS), FSx (Windows/SMB)
Boot disk / transactional DB disk on one server Block EBS, instance store
Massive scale, HTTP access, static assets, archives Object S3
Extreme parallel throughput for compute jobs File (parallel) FSx for Lustre
HPC scratch with local NVMe Block (ephemeral) Instance store (i-series)


How to reason it: ONE server that owns its disk → block (EBS) — the OS sees a raw volume and manages filesystem, caching, and I/O scheduling. MANY servers needing one shared namespace → file (EFS/FSx) — parallel access with standard protocols. HTTP-scale distribution, writes-once-read-many, unbounded growth → object (S3) — flat keys, not a filesystem. Trying to use S3 as a database filesystem or EBS as a shared drive fails on both performance and semantics — matching access pattern to storage TYPE is the root decision everything else builds on.

Real use-case: A render farm needs 50 nodes reading the same 2 TB of scene files. EBS is 1:1 (instance-to-volume) — wrong shape. EFS mounts everywhere, but metadata-heavy small-file reads over NFS add latency. FSx for Lustre links to the S3 archive, caches scene files at hundreds of GB/s, and releases scratch space after the render — the access pattern (parallel reads of shared large files) pointed to parallel file storage.

Gotchas & interview notes: "database on EFS" — RDS does not support EFS for data volumes (block only). "shared home directories across an ASG" → EFS (the classic). "instance store" is EPHEMERAL — fastest I/O in EC2, but data dies with the instance; only for scratch, replicas, or rebuild-able data.

Block Performance — EBS Deep Dive

Volume IOPS Throughput Notes
gp3 3,000 base (up to 16,000) 125 MB/s base (up to 1,000) Baseline INDEPENDENT of volume size; cheapest general SSD
gp2 3 IOPS/GB (scales with size) 250 MB/s max Legacy; size-tied performance
io2/io2 Block Express up to 256,000 IOPS up to 4,000 MB/s 99.999% durability, sub-ms latency; mission-critical DBs
st1 (throughput HDD) — 500 MB/s Big sequential reads (logs, data lakes)
sc1 (cold HDD) — 250 MB/s Cheapest infrequent access


How performance is bought (levers in order): gp3 lets you provision IOPS and throughput SEPARATELY from size (a 100 GB volume can have 16,000 IOPS — impossible on gp2, where IOPS = 3 × GB). If 16,000 IOPS is still not enough: RAID 0 across multiple EBS volumes multiplies IOPS and throughput linearly (2 × 32,000-IOPS volumes ≈ 64,000). For the extreme tier, io2 Block Express reaches 256,000 IOPS, 4,000 MB/s with sub-millisecond latency. Instance-side prerequisites matter: EBS-optimized instances, and the instance's own IOPS ceiling (an m5.large cannot consume what an io2 delivers).

Real use-case: A PostgreSQL primary on a 1 TB gp2 volume sustained 3,000 IOPS (3 × 1,000 GB) and stalled under a reporting join. Migration to gp3 at 16,000 IOPS / 1,000 MB/s — same size — cut the query from 40 s to 8 s. Cost rose ~15%; a RAID-0 pair was the next lever if needed.

Gotchas & interview notes: the gp3 decoupling is a favorite exam nuance — "increase IOPS WITHOUT growing the volume" → gp3. st1/sc1 cannot be boot volumes. Throughput (MB/s) matters for big sequential scans (logs, media); IOPS matters for transactional random I/O — pick the volume type by which one the workload needs.

File Performance

  • EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
  • FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
  • FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
  • Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
  • S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
  • S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
  • CloudFront in front of S3 for global read latency and cheaper egress
  • Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
  • AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
  • AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
  • AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
  • AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
  • 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
  • Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
  • Edited masters archived: lifecycle to Glacier Deep Archive
  • CDN delivery: CloudFront + signed URLs
  • On-prem broadcast playout cache: Storage Gateway file gateway
  • Database redo on EC2: io2 Block Express RAID-0 pair
How the EFS throughput modes work: Bursting mode gives a credit pool — small file systems accrue burst credits for short high-throughput windows (great for spiky small workloads). Sustained heavy read needs Provisioned (pay for guaranteed MB/s) or Elastic (auto-scales with usage — the modern default). Max I/O mode trades per-file latency for aggregate parallelism — it exists for thousands of concurrent clients, not for faster single-file access.

Gotchas & interview notes: "shared, grows and shrinks automatically, Linux" → EFS Elastic throughput. "Windows/SMB/AD-integrated" → FSx for Windows. "HPC reading an S3 data lake fast" → FSx for Lustre with the S3 repository link — data is CACHED into the scratch filesystem at job start and results can sync back.

Object Performance — S3 at Speed

How it works:

  • EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
  • FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
  • FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
  • Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
  • S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
  • S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
  • CloudFront in front of S3 for global read latency and cheaper egress
  • Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
  • AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
  • AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
  • AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
  • AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
  • 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
  • Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
  • Edited masters archived: lifecycle to Glacier Deep Archive
  • CDN delivery: CloudFront + signed URLs
  • On-prem broadcast playout cache: Storage Gateway file gateway
  • Database redo on EC2: io2 Block Express RAID-0 pair
Real use-case: A genome pipeline in Singapore uploads 40 GB samples to us-east-1 nightly over a 200 Mbps link with 250 ms RTT — a single TCP stream tops out around 5 MB/s. Multipart (16 parts) + Transfer Acceleration pushes effective throughput past 60 MB/s: parallel streams beat RTT, and the edge hop avoids congested public paths.

Gotchas & interview notes: "large object uploads are slow" → multipart; "slow from far away" → Transfer Acceleration (test it first — S3 provides a speed comparison tool); "downloads are slow" → byte-range fetches or CloudFront. Encryption with SSE-KMS adds a KMS API call per object — high-throughput workloads use S3 Bucket Keys to reduce that.

Hybrid and Transfer Performance

  • EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
  • FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
  • FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
  • Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
  • S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
  • S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
  • CloudFront in front of S3 for global read latency and cheaper egress
  • Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
  • AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
  • AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
  • AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
  • AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
  • 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
  • Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
  • Edited masters archived: lifecycle to Glacier Deep Archive
  • CDN delivery: CloudFront + signed URLs
  • On-prem broadcast playout cache: Storage Gateway file gateway
  • Database redo on EC2: io2 Block Express RAID-0 pair
Gotchas & interview notes: the Snow-family math is an exam staple: 70 TB over a dedicated 100 Mbps link ≈ 65 days; Snowball Edge ≈ a week door-to-door. DataSync over Direct Connect is the steady hybrid workhorse; Storage Gateway when on-prem apps need LOCAL latency against cloud storage.

Worked Example: Media Company Pipeline

  • EFS: elastic capacity; performance modes — General Purpose (default, low latency) vs Max I/O (higher aggregate throughput, higher latency, parallel workloads); throughput modes — Bursting vs Provisioned vs Elastic
  • FSx for Lustre: hundreds of GB/s, millions of IOPS; links to S3 as durable repository (classic HPC pattern: S3 data lake + Lustre scratch)
  • FSx for Windows File Server: SMB, DFS, SSD/HDD options, multi-AZ deployments
  • Multipart upload: parallelize large uploads (required >5 GB, works for any size >5 MB) — the #1 "upload is slow" answer; parts upload concurrently and retry independently
  • S3 Transfer Acceleration: upload via the nearest edge location over AWS backbone (fast cross-continent transfers)
  • S3 byte-range fetches: parallel GETs of ranges for downloads — the mirror-image of multipart
  • CloudFront in front of S3 for global read latency and cheaper egress
  • Prefix design: S3 now scales far beyond past per-prefix limits, but spreading prefixes still helps at extreme request rates
  • AWS DataSync: agent-based parallel, validated, scheduled transfers (NFS/SMB/HDFS → S3/EFS/FSx); up to 10 Gbps per agent
  • AWS Storage Gateway: on-prem cache with cloud durability (file/volume/tape) — low-latency reads of cloud-resident data
  • AWS Transfer Family: managed SFTP/FTPS/FTP fronting S3/EFS — partner file exchange at scale
  • AWS Snow Family: when network is the bottleneck (tens of TB+, slow link) — physical beats digital
  • 8 TB daily ingest of raw footage from studios (slow links): DataSync agents at each studio → S3 Standard
  • Editors need shared workspace: FSx for Lustre linked to the S3 media bucket; render farm (HPC) reads scratch at 100+ GB/s
  • Edited masters archived: lifecycle to Glacier Deep Archive
  • CDN delivery: CloudFront + signed URLs
  • On-prem broadcast playout cache: Storage Gateway file gateway
  • Database redo on EC2: io2 Block Express RAID-0 pair
The senior summary: classify the access pattern first (block/file/object/parallel), then buy exactly the speed needed at each layer — and remember transfer performance is a first-class design axis, not an afterthought.