Debugging a broken deployment
Something is broken on EKS: a deployment will not come up, traffic will not reach a pod, or a pod keeps
restarting. The Foundation's debugging lesson gave you the core kubectl routine; this lesson adds the
EKS-specific failure modes — the ones that come from Kubernetes being wired into AWS — so you can diagnose
the problems that a plain-Kubernetes course never mentions. It closes the EKS module.
Start with the Foundation routine
The debugging method from the DevOps Foundation applies unchanged, and you always start there:
kubectl get pods— the status (Pending,ImagePullBackOff,CrashLoopBackOff,0/1ready).kubectl describe pod <name>— read the Events for the reason.kubectl logs <name>(--previousif crashing) — the app's actual error.kubectl exec/port-forward— get inside or test directly.
That routine solves most problems. What follows is the AWS-specific layer: when the status or Events point at something that is really an AWS integration issue, not a plain Kubernetes one.
EKS-specific failure modes
These are the ones unique to running Kubernetes on AWS, mapped to what you would see:
-
Pod stuck
Pending— no nodes / no IPs. In the Foundation,Pendingmeant "no room on nodes." On EKS it can also mean:- No node capacity and node autoscaling not adding nodes (the scaling lesson) — check Karpenter/Cluster Autoscaler is working and the pending pod's requests can be satisfied.
- No available IP addresses. Because the VPC CNI gives each pod a real VPC IP (the eks-basics lesson),
pods can fail to schedule if the subnet runs out of IPs — a distinctly EKS problem.
describeshows a CNI/IP-allocation error. Fix: bigger subnets (the CIDR planning from the VPC lesson) or CNI IP-management settings.
-
ImagePullBackOff— cannot pull from ECR. Often the node/pod lacks permission to pull from your ECR registry, or the image tag is wrong. Check the node role / IRSA has ECR pull permission and the image URI is correct (an ECR URI like123456789012.dkr.ecr.eu-west-1.amazonaws.com/app:tag). -
Pod runs but cannot call AWS — IRSA misconfigured. The app gets
AccessDeniedcalling S3/SQS/Secrets Manager. Almost always the IRSA setup (the workloads lesson): the service account annotation is missing or wrong, the IAM role's trust policy does not allow the service account, or the role lacks the needed permission. Check the pod uses the right service account and the role's policy and trust relationship. -
Ingress created but no load balancer / 503s. The AWS Load Balancer Controller (the ingress lesson) is not running, lacks permissions, or the Ingress annotations are wrong — so no ALB is created, or it has no healthy targets. Check the controller's logs, the Ingress events, and that the target pods pass health checks.
-
OOMKilled/ evictions. Memory limits too low, or requests too low under node pressure (the resources-and-limits lesson).describeshowsOOMKilledor an eviction event; fix the resource numbers.
The pattern: the symptom is a normal Kubernetes status, but the cause is an AWS integration — IP exhaustion
(CNI), ECR permissions, IRSA, the load balancer controller. Knowing these EKS-specific causes turns a baffling
Pending or AccessDenied into a quick fix.
Where to look: kubectl and AWS together
Debugging EKS means looking in both Kubernetes and AWS:
- In Kubernetes:
kubectl describe(Events),kubectl logs,kubectl get events --sort-by=.lastTimestamp(recent cluster events across objects), and the logs of the relevant controller (the AWS Load Balancer Controller, Karpenter) when the problem is theirs. - In AWS: CloudWatch Logs for the app and control-plane logs (EKS can send control-plane logs to
CloudWatch); the EC2 / Auto Scaling console for node issues; IAM for role/permission problems;
CloudTrail to see a denied AWS API call and why (the IAM lesson). When a pod's
AccessDeniedis the issue, CloudTrail shows the exact call and the role that made it.
The habit is to move between the two: the Kubernetes symptom tells you where the problem is (scheduling, image, AWS call, ingress), and then you look in the corresponding AWS place (subnet IPs, ECR, IAM/CloudTrail, the load balancer controller) for the cause. And, as always on Kubernetes, you fix the manifest or the AWS configuration and re-apply — you do not hand-patch a live pod, because the reconciler will replace it.
Check your work
Start with the Foundation routine: get pods (status) → describe (Events) → logs/--previous →
exec/port-forward. Solves most; the EKS layer is when the cause is an AWS integration.
EKS-specific causes (symptom → cause): Pending → no node capacity or subnet out of IPs (VPC CNI —
bigger subnets); ImagePullBackOff → ECR permission/URI; pod AccessDenied calling AWS → IRSA
misconfigured (service-account annotation, role trust, or missing permission); Ingress but no LB / 503s → AWS
Load Balancer Controller not running / wrong annotations / unhealthy targets; OOMKilled/evictions →
resource requests/limits (previous lesson).
Look in both places: Kubernetes (describe Events, logs, cluster events, the relevant controller's
logs) and AWS (CloudWatch app + control-plane logs, EC2/ASG for nodes, IAM/CloudTrail for denied calls
— CloudTrail shows the exact call and role). Symptom tells you where; look in the matching AWS place for the
cause. Fix the manifest/AWS config and re-apply — never hand-patch a live pod.
Practice
- State the core
kubectldebugging routine and when it is enough versus when you look at AWS. - Give two EKS-specific reasons a pod can be
Pendingand how you'd confirm each. - Diagnose a pod getting
AccessDeniedcalling S3 — what is almost always wrong, and what do you check? - Diagnose an Ingress that created no load balancer — where do you look?
- Explain how you'd use CloudTrail to diagnose a denied AWS API call from a pod.
- Explain why you fix the manifest/AWS config and re-apply rather than editing a live pod.
Official documentation
- AWS — EKS troubleshooting — Common EKS problems and fixes.
- AWS — EKS control plane logging (CloudWatch) — Sending cluster logs to CloudWatch.
- Kubernetes — Debug running pods — The core debugging routine.
Next: the Terraform at Scale module.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship