RizTech Academy logo
RizTech Academy
Cost OptimisationLesson 3 of 535 min

Right-sizing and scale-to-zero for non-production

Over-provisioning is the largest ongoing source of AWS waste (the common-waste lesson), and right-sizing is the fix: matching resources to what workloads actually use. Combined with scale-to-zero for non-production — turning off what nobody is using — these two techniques often cut a bill substantially without changing what you run. This lesson is right-sizing and scale-to-zero, the highest-return cost work most teams can do.

Right-sizing: match capacity to actual usage

Right-sizing means adjusting a resource's size to fit its real workload, rather than the (usually over-generous) size someone picked once. Because teams over-provision "to be safe" and never revisit it, most resources are bigger than needed — so right-sizing is often the single largest saving available.

The method:

  • Measure real usage. Look at actual CPU, memory, network and disk utilisation over a representative period (CloudWatch metrics — the observability module; or AWS Compute Optimizer, which analyses usage and recommends right-sizes automatically). An instance averaging 10% CPU with occasional small peaks is several sizes too large.
  • Resize to fit, with sensible headroom. Drop to an instance/task size that comfortably covers the real peak plus a modest margin — not the huge margins people default to. Halving an over-provisioned instance halves its cost directly.
  • Right-size everything, not just instances. EKS pod requests (the resources-and-limits lesson) — accurate requests pack more pods per node, so you run fewer nodes; RDS instances and storage; Lambda memory (which also affects its speed and cost). Right-sizing applies across the stack.
  • Re-check periodically. Workloads change; a right-size is not once-and-done. The cost review (next lesson) revisits it.

Right-sizing is unglamorous and enormously effective — because over-provisioning is so widespread, a systematic right-sizing pass, guided by Compute Optimizer and real metrics, reliably finds large savings. It costs nothing but the discipline to look at utilisation and adjust down.

Prefer autoscaling over fixed over-provisioning

A structural improvement on right-sizing: instead of running a large fixed capacity sized for a rare peak, run a right-sized baseline and autoscale for the peaks (the scaling lessons). This way you pay for the peak capacity only when the peak actually happens, not all the time:

  • Horizontal autoscaling (HPA + node autoscaling / Karpenter on EKS, or ECS service autoscaling) adds capacity under load and removes it after — so average cost tracks average load, not peak load.
  • This is both cheaper and more resilient than one big fixed instance, and it means you can right-size the baseline aggressively because autoscaling handles the surges.

So right-sizing and autoscaling work together: size the baseline to normal load, and let autoscaling cover the rest — rather than paying for peak capacity 24/7. This is the cost-optimal shape for variable workloads.

Scale-to-zero for non-production

The highest-return, lowest-risk cost action of all: non-production environments do not need to run when nobody is using them. Dev, staging, test and demo environments typically sit idle overnight and at weekends — yet many run 24/7, billing for time when no one is working. Turning them off when idle can roughly halve their cost with zero downside:

  • Schedule non-production off outside working hours. A dev environment used ~50 hours a week (of 168) can be off for the other ~118 — a ~70% saving on that environment. Schedule instances/clusters to stop in the evening and start in the morning (AWS Instance Scheduler, an EventBridge rule, or a simple Lambda).
  • Scale to zero. Stop EC2 instances, scale EKS node groups / Karpenter to zero, pause non-critical RDS — whatever your environment uses — outside the hours it is needed.
  • This is safe for non-production — nobody is using it, and it starts again in the morning (or on demand). It is not for production, which must stay up.

Scale-to-zero for non-production is often the biggest single easy saving in an account, because non-production environments frequently rival production in cost yet are needed only a fraction of the time. It is the clearest example of the course's cost principle — don't pay for what you're not using — and it costs only a schedule to set up.

The combined effect

Put together: right-size everything to real usage, autoscale the variable parts so you pay for peaks only when they occur, and scale non-production to zero outside working hours. For most accounts this trio — no new technology, just measuring usage and turning things off — cuts the bill more than any single discount, because it removes the pervasive over-provisioning and idle-when-unused waste that dominates real bills. It is the highest-return cost work available, and it is exactly the "cost-optimised by default" discipline this course has argued for throughout: match spend to actual need, deliberately and continuously, rather than paying for generous margins and idle time out of habit.

Check your work

Right-sizing: match a resource's size to real usage (most are over-provisioned "to be safe"). Method: measure real CPU/mem/etc. (CloudWatch / Compute Optimizer recommendations) — 10% avg CPU = too big; resize to real peak + modest headroom (halving an over-sized instance halves its cost); right-size everything (instances, EKS pod requests → fewer nodes, RDS, Lambda memory); re-check periodically. Unglamorous, largest single saving.

Autoscale over fixed over-provisioning: right-sized baseline + autoscaling (HPA/Karpenter/ECS) → pay for peak capacity only when the peak happens (cheaper + more resilient than one big fixed instance). Size baseline to normal load; autoscaling covers surges.

Scale-to-zero (highest return): non-prod doesn't need to run when idle — schedule dev/staging/test off outside working hours (~50/168 hrs used → big saving), stop instances / scale node groups to zero / pause RDS. Safe for non-prod (nobody using it), never prod. Often the biggest single easy saving.

Combined: right-size + autoscale variable parts + scale non-prod to zero — no new tech, just measure and turn off — cuts the bill more than any single discount. "Don't pay for what you're not using."

Practice

  1. Explain right-sizing and how you'd use metrics/Compute Optimizer to find over-provisioned resources.
  2. Explain how right-sizing EKS pod requests reduces node count and cost.
  3. Explain why a right-sized baseline plus autoscaling beats a large fixed instance on both cost and resilience.
  4. Estimate the saving from scheduling a dev environment off outside a 50-hour working week.
  5. Explain why scale-to-zero is safe for non-production but not production.
  6. Combine right-sizing, autoscaling and scale-to-zero for a system with prod and non-prod, and describe the effect.

Official documentation

Next: Savings Plans and Reserved Instances.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship