Cloud Economics: Fixed vs Variable Costs (Task 1.4)
Cloud Economics
Source: https://docs.aws.amazon.com/whitepapers/latest/aws-overview/cost-savings.htmlThe exam tests the three levers of cloud economics — pay-as-you-go, economies of scale, and rightsizing — plus how managed services and automation change the cost STRUCTURE, not just the amount. Senior engineers can build a TCO model, pick a licensing strategy, and defend "cloud vs on-prem" with numbers. Each concept: brief, how it works, example, real use-case, gotchas.
Fixed Costs vs Variable Costs
Brief: Moving to AWS converts large fixed investments (CapEx) into variable costs (OpEx) that track actual usage. If the business shrinks, the bill shrinks.
How it works: On-premises, capacity is bought in step-functions (a $500k refresh every 4 years) and depreciated; utilization averages 10–20% because you size for peak-plus-headroom. On AWS, the same workload bills per second/GB/request, so cost becomes a linear function of demand. TCO models capture: hardware refresh avoidance, facilities (power/cooling/space), admin labor, over-provisioning waste, and the opportunity cost of slow procurement.
| Cost Type | On-Premises Example | AWS Example |
|---|---|---|
| Fixed (CapEx) | $500,000 server purchase, 5-year depreciation | None — no upfront hardware purchase |
| Variable (OpEx) | Electricity, admin salaries (semi-fixed) | Per-second EC2 billing, per-request Lambda, per-GB S3 |
Example: A company that would have bought $500,000 of servers instead runs the workload for $8,000/month — and pays $0 the month it shuts the project down. On-premises, that abandoned hardware would still be depreciating.
Real use-case: A video startup scaled from 10 to 400 instances during pandemic demand, then back to 20 — its cost curve followed demand exactly, with no stranded hardware when usage fell.
Gotchas & interview notes: "Cloud is cheaper" is only true with elasticity and managed services in the model. A 24/7 steady-state workload at high utilization can cost MORE at on-demand rates — which is why Savings Plans/Reserved exist. Interviewers probe TCO: if a candidate's model omits egress fees, staff retraining, or migration cost, that's a flag.
Costs That Disappear on AWS
Brief: Entire cost categories are absorbed into the service price.
- Data center facilities: rent, cooling, power, physical security
- Hardware maintenance contracts and spare parts inventory
- Over-provisioned "just in case" capacity — typically 30–60% of on-premises servers sit idle
- Long procurement cycles (opportunity cost of waiting)
- Included licenses: the license is metered into the service price — e.g., RDS for SQL Server bundles the SQL Server license; you pay only while the database runs.
- Bring Your Own License (BYOL): reuse licenses you already own. When license terms tie the license to specific physical hardware (common with Microsoft/Oracle), you need Dedicated Hosts/Instances so you can track which physical machine your licenses run on.
- License mobility: some Microsoft agreements permit running on shared multi-tenant hardware under Software Assurance; when terms are unclear, Dedicated Hosts satisfy the strictest interpretation.
- On-premises: hardware refresh $400k every 4 years + power/cooling $60k/year + admin salaries $250k/year = ~$410k/year fully loaded.
- AWS: 100 right-sized instances ≈ $95k/year + RDS managed databases ≈ total ~$150k/year, with elasticity for peak seasons included.
- Decision levers in order: rightsizing first (utilization data shows 60 of 100 VMs are oversized), managed services second (repurpose 1 admin to application work), automation third (temporary environments exist for hours, not forever).
Rightsizing
Brief: Matching resource size to actual workload need — the highest-ROI cost skill in cloud.
How it works: Collect utilization history (CPU, memory, network, disk IOPS via CloudWatch), then match the observed percentile profile to the smallest instance family/size that meets it with headroom. AWS Compute Optimizer does this with ML across four recommendation types (over-provisioned, under-provisioned, optimized, none). Memory is the frequent limiter: an instance at 8% CPU but 75% memory is NOT a downsize candidate by CPU alone.
Example: An EC2 m5.2xlarge (8 vCPU / 32 GB) averaging 10% CPU and 30% memory → an m5.large (2 vCPU / 8 GB) at ~25% of the price. A spiky dev server at 5% average CPU → a burstable t3.large that accrues CPU credits during idle.
Real use-case: A company runs Compute Optimizer for 30 days across 400 instances, then applies recommendations in waves (dev first, then prod with canary validation): compute spend drops 28% with zero performance regressions measured.
Gotchas & interview notes: Rightsize from REAL metrics over a full business cycle (include month-end peaks), never from a single day. Stateful services (databases) downsize differently — changing instance class triggers a failover/maintenance window. The exam phrase "CPU has remained below 20%..." → smaller instance or burstable T-family. "Batch job finishes too slowly" → it is under-provisioned; rightsizing works in both directions.
Licensing Strategies
Brief: Software licenses are often a bigger line item than infrastructure — three ways to handle them on AWS.
- Data center facilities: rent, cooling, power, physical security
- Hardware maintenance contracts and spare parts inventory
- Over-provisioned "just in case" capacity — typically 30–60% of on-premises servers sit idle
- Long procurement cycles (opportunity cost of waiting)
- Included licenses: the license is metered into the service price — e.g., RDS for SQL Server bundles the SQL Server license; you pay only while the database runs.
- Bring Your Own License (BYOL): reuse licenses you already own. When license terms tie the license to specific physical hardware (common with Microsoft/Oracle), you need Dedicated Hosts/Instances so you can track which physical machine your licenses run on.
- License mobility: some Microsoft agreements permit running on shared multi-tenant hardware under Software Assurance; when terms are unclear, Dedicated Hosts satisfy the strictest interpretation.
- On-premises: hardware refresh $400k every 4 years + power/cooling $60k/year + admin salaries $250k/year = ~$410k/year fully loaded.
- AWS: 100 right-sized instances ≈ $95k/year + RDS managed databases ≈ total ~$150k/year, with elasticity for peak seasons included.
- Decision levers in order: rightsizing first (utilization data shows 60 of 100 VMs are oversized), managed services second (repurpose 1 admin to application work), automation third (temporary environments exist for hours, not forever).
Real use-case: An Oracle estate review finds half the databases are candidates for open-source migration (to Aurora PostgreSQL) and half must stay on Oracle; the team splits the estate — eliminating licenses where possible and using Dedicated Hosts with BYOL where not.
Gotchas & interview notes: License compliance is YOUR responsibility, not AWS's (shared responsibility extends to software licensing). The exam keyword pairing to memorize: BYOL ↔ Dedicated Hosts. And always model the switch: metered included-licenses sometimes BEAT BYOL when utilization is low, because you only pay while running.
Benefits of Automation
Brief: Manual provisioning is slow, inconsistent, and error-prone; infrastructure-as-code makes environments cheap, identical, and disposable.
How it works: CloudFormation (and Terraform) are declarative — you describe the desired end state, the engine computes and applies the diff. Templates are version-controlled, peer-reviewed like code, and reusable: the same template deploys identical stacks to dev, staging, and 5 Regions.
Example: A full 3-tier environment (VPC, ALB, ASG, RDS) provisions from one template in ~15 minutes. The experiment ends, the stack is deleted, and NOTHING keeps billing — no forgotten load balancers or EBS volumes.
Real use-case: A team's "environment spin-up" lead time drops from 3 weeks (tickets + manual builds, always slightly different) to 20 minutes (pipeline + template, always identical). Debugging time drops too, because "works on my machine" disappears when every environment is the same.
Gotchas & interview notes: IaC has its own failure modes: drift (manual console changes that templates don't know about — detect with drift detection or config rules) and deletion ordering (resources with dependencies can fail stack deletion). Senior practice: console changes in prod are disabled/audited; all changes flow through pipelines.
Managed Services Reduce Operational Cost
Brief: Managed services shift operational burden (patching, backups, failover, scaling) from your team to AWS.
| If You Run Yourself | Managed Equivalent | What AWS Takes Over |
|---|---|---|
| MySQL on EC2 | Amazon RDS | OS/DB patching, backups, Multi-AZ failover |
| Kubernetes control plane on EC2 | Amazon EKS | Control plane HA, upgrades |
| Hadoop cluster ops | Amazon EMR | Cluster provisioning, tuning |
| NoSQL at scale | Amazon DynamoDB | Sharding, replication, throughput management |
How it works: the operational labor is amortized across all customers of the service — AWS runs one patching pipeline for a million RDS databases; you'd run one for your twenty. That is economies of scale applied to LABOR, not just hardware.
Example: A self-managed MySQL on EC2 needs: a DBA doing patch windows, a backup script + restore drills, a hand-built failover mechanism, and sizing homework. RDS delivers all four as configuration options; the DBA time redirects to schema design and query optimization.
Real use-case: A 40-person company runs its entire data layer (RDS, ElastiCache, DynamoDB) with ZERO dedicated database administrators; the two engineers who used to do DBA work now build data pipelines that generate revenue.
Gotchas & interview notes: The trade-off is control and portability (lock-in): managed services limit OS-level access and tuning knobs, and switching engines later is a migration project. Senior answer: use managed services for undifferentiated workloads, self-manage where the database/engine IS your competitive advantage. The exam always favors the managed-service answer for "reduce operational overhead."
Worked Example: TCO Comparison
A company runs 100 on-premises VMs at 20% average utilization, with 2 FTEs doing patching. Cloud comparison:
- Data center facilities: rent, cooling, power, physical security
- Hardware maintenance contracts and spare parts inventory
- Over-provisioned "just in case" capacity — typically 30–60% of on-premises servers sit idle
- Long procurement cycles (opportunity cost of waiting)
- Included licenses: the license is metered into the service price — e.g., RDS for SQL Server bundles the SQL Server license; you pay only while the database runs.
- Bring Your Own License (BYOL): reuse licenses you already own. When license terms tie the license to specific physical hardware (common with Microsoft/Oracle), you need Dedicated Hosts/Instances so you can track which physical machine your licenses run on.
- License mobility: some Microsoft agreements permit running on shared multi-tenant hardware under Software Assurance; when terms are unclear, Dedicated Hosts satisfy the strictest interpretation.
- On-premises: hardware refresh $400k every 4 years + power/cooling $60k/year + admin salaries $250k/year = ~$410k/year fully loaded.
- AWS: 100 right-sized instances ≈ $95k/year + RDS managed databases ≈ total ~$150k/year, with elasticity for peak seasons included.
- Decision levers in order: rightsizing first (utilization data shows 60 of 100 VMs are oversized), managed services second (repurpose 1 admin to application work), automation third (temporary environments exist for hours, not forever).