RizTech Academy logo
RizTech Academy
Kubernetes on EKSLesson 4 of 635 min

Horizontal pod autoscaling and cluster autoscaling

Scaling on EKS happens at two levels, and confusing them is a common source of "why won't my app scale?" frustration. Pod autoscaling adds more copies of your app; node autoscaling adds more machines to run them on. Both must work together for real elasticity, and on EKS the node side has AWS-specific tools worth knowing. This lesson is horizontal pod autoscaling and cluster autoscaling, building on the Foundation's scaling lesson.

Two levels of scaling

When load rises, you may need two different things, and they are separate mechanisms:

  • More pods — run additional replicas of your application to handle more requests. This is pod autoscaling (the Horizontal Pod Autoscaler), the Foundation's HPA.
  • More nodes — if there is not enough room on the existing nodes to place those extra pods, you need more machines. This is node (cluster) autoscaling.

The interaction is the key insight: the HPA decides it wants 10 pods, but if the nodes are full, the new pods sit Pending (unschedulable — the Foundation's debugging lesson) until a node autoscaler adds a machine for them. So real elasticity needs both: pod autoscaling to want more pods, and node autoscaling to provide the capacity to run them. A cluster with HPA but no node autoscaling will scale pods only until the nodes fill, then stall.

Pod autoscaling: the HPA (and beyond)

The Horizontal Pod Autoscaler (HPA) works on EKS exactly as in the Foundation: it watches a metric (CPU, memory, or a custom metric) and adjusts the replica count to keep it near a target — "keep CPU around 60%, between 2 and 20 replicas." It relies on resource requests being set (the next lesson), because it measures usage against them.

Two EKS-relevant additions:

  • Custom and external metrics. Beyond CPU/memory, the HPA can scale on application metrics (requests per second, queue length) via a metrics adapter — useful when CPU is not the real signal of load.
  • KEDA (Kubernetes Event-Driven Autoscaling) — a popular add-on to scale on external signals like an SQS queue depth or other event sources, including scaling to zero when idle. Handy for workers processing a queue.

Pod autoscaling is unchanged conceptually; what EKS adds is the ability to drive it from AWS-native signals (a queue, a custom metric) rather than only CPU.

Node autoscaling: Cluster Autoscaler vs Karpenter

The AWS-specific part is scaling the nodes. Two tools:

  • Cluster Autoscaler — the traditional tool. It watches for Pending pods that cannot be scheduled and adds nodes (by growing an EKS managed node group / EC2 Auto Scaling group); when nodes are underused, it removes them. It works within pre-defined node groups of fixed instance types.
  • Karpenter — AWS's newer, more flexible node autoscaler, increasingly the recommended choice. Instead of fixed node groups, Karpenter looks at the pending pods' actual requirements and launches right-sized, well-chosen instances to fit them — picking instance types and sizes (and Spot vs on-demand) to satisfy the pending pods efficiently. It provisions faster and packs workloads more cost-effectively, and it makes using Spot (the compute module) for interruptible workloads much easier.

Karpenter's advantage ties directly to the cost module: because it right-sizes the instances it launches to the actual pending workload and can prefer Spot and Graviton, it tends to run your cluster on less, cheaper capacity than fixed node groups sized for a peak. For cost-conscious EKS (which this course insists on), Karpenter is often the better node autoscaler.

Making scaling actually work

A few practical points that determine whether scaling works in practice:

  • Set resource requests (the next lesson) — the HPA needs them to measure load, and the node autoscaler needs them to know how much room a pending pod needs. Without requests, both scale poorly.
  • Spread across AZs — run nodes and replicas across multiple Availability Zones (the networking module) so scaling adds resilient capacity, not more eggs in one basket.
  • Scale interruptible workloads on Spot — with a stable baseline on on-demand (the compute module), autoscale the burst on Spot for cost; Karpenter makes this straightforward.
  • Test that pending pods actually trigger nodes — a common misconfiguration is HPA scaling pods that then sit Pending because node autoscaling is not set up. Verify the whole chain: load → more pods → (if needed) more nodes → pods scheduled.

The combined picture: the HPA scales your pods on load; Karpenter (or Cluster Autoscaler) scales the nodes to fit them; resource requests make both work; and Spot plus right-sized instances keep it cheap. That is elastic, cost-aware scaling on EKS — the thing managed Kubernetes is supposed to give you, assembled correctly.

Check your work

Two levels: pod autoscaling (more replicas — the HPA) and node autoscaling (more machines). Key interaction: HPA wants more pods, but if nodes are full the pods sit Pending until a node autoscaler adds capacity — real elasticity needs both (HPA alone stalls when nodes fill).

Pod autoscaling (HPA): as in the Foundation — scale replicas to keep a metric near target; needs resource requests. EKS adds custom/external metrics and KEDA (event-driven, e.g. SQS depth, scale-to-zero).

Node autoscaling: Cluster Autoscaler (adds/removes nodes in fixed node groups when pods are Pending/ nodes idle) vs Karpenter (newer, recommended — launches right-sized, well-chosen instances to fit pending pods, prefers Spot/Graviton, faster and cheaper). Karpenter ties to the cost module: less, cheaper capacity.

Make it work: set resource requests (both autoscalers need them), spread across AZs, autoscale the burst on Spot with a stable baseline, and verify the full chain (load → pods → nodes → scheduled). HPA scales pods, Karpenter scales nodes, requests make both work, Spot + right-sizing keep it cheap.

Practice

  1. Explain the difference between pod autoscaling and node autoscaling and why both are needed.
  2. Explain what happens when the HPA scales pods but there is no room on the nodes and no node autoscaler.
  3. Explain how the HPA scales, and one EKS-specific way to scale on a non-CPU signal.
  4. Compare Cluster Autoscaler and Karpenter, and why Karpenter is often better for cost.
  5. Explain why resource requests are essential for both levels of autoscaling.
  6. Design cost-aware autoscaling for a stateless service: baseline, burst, instance choice, AZ spread.

Official documentation

Next: requests, limits and why pods get evicted.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship