RizTech Academy logo
RizTech Academy
Kubernetes on EKSLesson 5 of 635 min

Requests, limits and why pods get evicted

"My pods keep getting killed and I don't know why" is one of the most common EKS support questions, and the answer is almost always resource requests and limits — the CPU and memory settings on your containers. Get them wrong and pods get evicted, throttled, or OOM-killed unpredictably; get them right and the cluster is stable, schedulable and cost-efficient. This lesson is requests, limits, and why pods get evicted.

Requests and limits: two different numbers

Every container can declare two numbers each for CPU and memory, and they do very different jobs:

  • Requests — the amount a container is guaranteed. The scheduler uses requests to decide which node a pod goes on: it places the pod only where the sum of requests fits. Requests are about scheduling and guaranteed capacity.
  • Limits — the maximum a container may use. Exceed the limit and the container is throttled (CPU) or killed (memory). Limits are about capping a container so one pod cannot starve the others.
        resources:
          requests:      # guaranteed, used for scheduling
            cpu: "250m"
            memory: "256Mi"
          limits:        # maximum before throttle (CPU) / kill (memory)
            cpu: "500m"
            memory: "512Mi"

Read it: this container is guaranteed 0.25 CPU and 256 MB (and scheduled on that basis), and may burst to 0.5 CPU and 512 MB before CPU throttling or a memory kill. The gap between request and limit is burst room. Understanding that requests are for scheduling and limits are for capping is the key to everything below.

CPU vs memory: throttled vs killed

CPU and memory behave differently when a container hits its limit, and this difference explains most pod problems:

  • CPU is compressible. If a container exceeds its CPU limit, it is throttled — slowed down, but not killed. Your app gets less CPU and runs slower; nothing dies. So a too-low CPU limit causes latency, not crashes.
  • Memory is not compressible. If a container exceeds its memory limit, it is killed — an OOMKill (out of memory). There is no "run slower on less memory"; the kernel terminates the process. So a too-low memory limit causes crashes (you see OOMKilled in kubectl describe).

This is the answer to "why does my pod keep dying?": it is almost always memory — the container's memory limit is lower than what it actually uses under load, so it gets OOMKilled, restarts, and eventually CrashLoopBackOffs. The fix is to measure real memory usage and set a limit above the true peak. "Pod restarts mysteriously under load" → check memory limits first.

Why pods get evicted

Beyond hitting their own limits, pods can be evicted — removed from a node — for node-level reasons, and knowing why prevents surprises:

  • Node under memory pressure. If a node runs low on memory (because pods are using more than their requests, in the burst room), the kubelet evicts pods to reclaim memory — starting with pods using more than they requested. So a pod that requested little but uses a lot is a prime eviction target. This is why setting requests close to real usage matters: a pod whose request reflects its actual need is protected; one that lowballs its request gets evicted first under pressure.
  • Node scaling down. When a node autoscaler removes an underused node (the scaling lesson), it evicts and reschedules the pods elsewhere — normal, and handled gracefully by Deployments.
  • Spot reclamation. A Spot node (the compute module) being reclaimed evicts its pods with ~2 minutes' notice — which is fine for the stateless, replicated workloads you put on Spot.

The recurring lesson: eviction and OOMKills come from a mismatch between what pods request and what they actually use. Requests too low means eviction under pressure; limits too low (especially memory) means OOMKills. Right-sizing these numbers is what makes an EKS cluster stable.

Getting the numbers right — and the cost angle

Setting requests and limits well is a balance, and it ties straight to cost:

  • Set requests to real usage plus modest headroom. Measure actual CPU/memory (CloudWatch, metrics-server, or the Vertical Pod Autoscaler's recommendations) and request roughly that. Too high wastes capacity (the scheduler reserves it, so you pay for nodes you don't need — the cost angle); too low invites eviction. This is right-sizing (the cost module) at the pod level.
  • Set memory limits at or slightly above the real peak to avoid OOMKills, and be careful lowballing them.
  • Be cautious with CPU limits — because CPU throttling hurts latency, many teams set CPU requests but loose or no CPU limits, letting containers burst on spare CPU. Memory limits, by contrast, should always be set (an unbounded memory leak can take down a node).
  • Right-sized requests improve packing and cost. Because the scheduler packs by requests, accurate requests let more pods fit per node, so you run fewer nodes — directly cheaper. Over-inflated requests waste node capacity you pay for.

So requests and limits are not just a stability setting — they are a cost lever. Accurate requests mean stable scheduling, protection from eviction, tight packing and a smaller bill; sensible memory limits prevent OOMKills. Measuring real usage and setting these deliberately (not guessing high "to be safe") is a core EKS operational skill and a cost-optimisation one at once.

Check your work

Requests vs limits: requests = guaranteed amount, used by the scheduler to place pods (scheduling / guaranteed capacity); limits = maximum before throttle/kill (capping). The gap is burst room.

CPU vs memory at the limit: CPU is compressible → exceeding the limit throttles (slower, not killed) → too-low CPU limit = latency. Memory is not → exceeding = OOMKilled → too-low memory limit = crashes/CrashLoopBackOff. "Pod keeps dying under load" → check memory limits first.

Eviction: under node memory pressure the kubelet evicts pods using more than they requested first (so lowballing requests → evicted first); also on scale-down (rescheduled) and Spot reclamation (~2 min, fine for stateless). Evictions/OOMKills come from request/usage mismatch.

Getting it right (+ cost): set requests to real usage + modest headroom (VPA/metrics; too high wastes paid capacity, too low invites eviction); set memory limits at/above real peak (always set memory limits); be cautious with CPU limits (throttling hurts latency). Accurate requests pack tighter → fewer nodes → cheaper. Measure, don't guess high.

Practice

  1. Explain the different jobs of requests and limits, and what the gap between them represents.
  2. Explain why exceeding a CPU limit throttles but exceeding a memory limit kills, and what each implies.
  3. Diagnose "my pod keeps restarting under load" — what do you check first and why?
  4. Explain why a pod that requests little but uses a lot is evicted first under node memory pressure.
  5. Explain how accurate resource requests reduce your node count and therefore cost.
  6. Give sensible guidance on setting CPU limits vs memory limits, and why memory limits must always be set.

Official documentation

Next: debugging a broken deployment.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship