Graviton, Spot and right-sizing
Compute is usually the largest line on an AWS bill, and it is also where the easiest savings hide. Three levers — Graviton processors, Spot capacity, and right-sizing — can cut compute cost substantially, often by half or more, with little downside when applied sensibly. Knowing them turns "the bill is high" into "here is where to cut it." This lesson is the cost of compute, and how to reduce it. It closes the compute module, and cost is treated as first-class here as throughout.
Graviton: ARM processors, cheaper and often faster
AWS builds its own Graviton processors — ARM-based CPUs — and instances using them cost meaningfully less than the equivalent Intel/AMD (x86) instances, typically around 20% cheaper for similar or better performance, and often with better energy efficiency. For most modern workloads, switching to Graviton is close to free money:
- Most software runs on ARM now. Modern language runtimes (Node, Python, Java, Go, .NET), databases, and official container base images all support ARM64. If your app is in one of these, it very likely runs on Graviton unchanged.
- The switch is often just picking a Graviton instance type (e.g. a
t4g/m7ginstead of at3/m6i), and building your container images for ARM64 (a multi-arch or ARM build — the containers module's Docker builds). - The caveat: anything with x86-specific native dependencies, or an old binary you cannot rebuild, may not run on ARM — test before assuming. But for the large majority of workloads, Graviton is a straightforward ~20% compute saving.
Preferring Graviton by default is one of the highest-return, lowest-effort cost decisions on AWS, which is why it is worth making a standing habit: choose ARM instances unless something specifically requires x86.
Spot: spare capacity at a huge discount
Spot instances are AWS's spare capacity, sold at up to ~90% off the on-demand price — with one catch: AWS can reclaim them with about two minutes' notice when it needs the capacity back. That interruption risk makes Spot unsuitable for some workloads and superb for others:
- Great for interruptible, fault-tolerant work: batch jobs, data processing, CI/CD runners, and — crucially — stateless application workloads that run in multiple replicas (an ECS service or EKS deployment with several tasks/pods). If one Spot instance is reclaimed, the orchestrator reschedules the work elsewhere and the service stays up, because it was designed to lose an instance anyway (the self-healing you built).
- Not for: stateful single-instance workloads that cannot tolerate sudden termination (a lone database), or anything where a two-minute-notice interruption is unacceptable.
- The common pattern: run a baseline of on-demand (or reserved) capacity for stability, and a portion on Spot for the savings — so you get most of the discount while staying resilient to reclamation. EKS and ECS both support mixing on-demand and Spot in a node group / capacity provider.
Spot is where the biggest compute savings live for the right workloads. Because so much modern compute is stateless and replicated (exactly what this course builds), a large share of a fleet can often run on Spot, turning a huge discount into real money saved — with the orchestrator's self-healing absorbing the interruptions.
Right-sizing: stop paying for idle
The third lever needs no new technology, just attention: most instances are bigger than they need to be. Teams over-provision "to be safe", then never revisit it, and pay for CPU and memory that sit idle. Right-sizing is matching the instance size to actual usage:
- Look at real utilisation (CloudWatch metrics, or AWS Compute Optimizer, which recommends right-sizes from observed usage). An instance averaging 10% CPU is three or four sizes too big.
- Scale down to fit, leaving sensible headroom — not the huge margins people default to. Halving an over-provisioned instance halves its cost directly.
- Prefer horizontal scaling with autoscaling over one huge instance: run several right-sized instances/tasks and let autoscaling add capacity under load, rather than paying for a large fixed instance sized for a peak that rarely comes. This also improves resilience.
- Scale non-production to zero when idle — dev and staging do not need to run overnight or at weekends; scheduling them off (the cost module returns to this) can halve their cost for no downside.
Right-sizing is the least glamorous lever and often the largest single saving, because over-provisioning is so common. It costs nothing but the discipline to look at utilisation and adjust.
Putting the levers together
Applied together, these compound: run right-sized instances, prefer Graviton, and put the interruptible, replicated portion on Spot — and you can often cut a compute bill by half or more versus the naive default of over-provisioned, on-demand, x86 instances. The habit to build, consistent with this whole course, is to treat compute cost as a design input, not an afterthought: pick ARM by default, size to real usage, use Spot for what can tolerate it, and keep a baseline of stable capacity for what cannot. That is how production AWS is run cost-effectively — deliberately, with the levers understood, rather than discovering the bill and wondering where it went.
Check your work
Graviton (ARM): AWS's ARM CPUs, ~20% cheaper for similar/better performance. Most modern runtimes/images support ARM64, so switching is often just choosing a Graviton instance type + building ARM images; caveat: x86-specific native deps may not run (test). Prefer ARM by default — high return, low effort.
Spot: spare capacity at up to ~90% off, but AWS can reclaim with ~2 min notice. Great for interruptible/ fault-tolerant work and stateless replicated services (the orchestrator reschedules on reclamation); not for stateful single instances. Pattern: on-demand/reserved baseline + a Spot portion for savings with resilience. Biggest savings for the right workloads.
Right-sizing: most instances are over-provisioned; match size to real usage (CloudWatch / Compute Optimizer — 10% avg CPU = too big), scale down with headroom, prefer horizontal autoscaling over one huge instance, scale non-prod to zero when idle. No new tech — often the largest single saving.
Together: right-sized + Graviton + Spot for the interruptible/replicated portion, with a stable baseline — often halves a compute bill vs over-provisioned/on-demand/x86. Treat compute cost as a design input.
Practice
- Explain what Graviton is, the typical saving, and how you would switch a containerised app to it (and the caveat).
- Explain Spot instances, their discount and their catch, and which workloads suit them.
- Explain the on-demand-baseline-plus-Spot pattern and why it is both cheap and resilient.
- Explain right-sizing and how you would use utilisation data to find over-provisioned instances.
- Explain why horizontal scaling with autoscaling is often cheaper than one large fixed instance.
- Combine all three levers for a stateless, replicated web service and estimate the direction of the saving.
Official documentation
- AWS — Graviton processors — ARM-based instances and their price/performance.
- AWS — Spot instances — Spare capacity at a discount, and interruption handling.
- AWS — Compute Optimizer — Right-sizing recommendations from real usage.
Next: the Kubernetes on EKS module.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship