RizTech Academy logo
RizTech Academy
Kubernetes on EKSLesson 6 of 640 min

Debugging a broken deployment

Something is broken on EKS: a deployment will not come up, traffic will not reach a pod, or a pod keeps restarting. The Foundation's debugging lesson gave you the core kubectl routine; this lesson adds the EKS-specific failure modes — the ones that come from Kubernetes being wired into AWS — so you can diagnose the problems that a plain-Kubernetes course never mentions. It closes the EKS module.

Start with the Foundation routine

The debugging method from the DevOps Foundation applies unchanged, and you always start there:

  1. kubectl get pods — the status (Pending, ImagePullBackOff, CrashLoopBackOff, 0/1 ready).
  2. kubectl describe pod <name> — read the Events for the reason.
  3. kubectl logs <name> (--previous if crashing) — the app's actual error.
  4. kubectl exec / port-forward — get inside or test directly.

That routine solves most problems. What follows is the AWS-specific layer: when the status or Events point at something that is really an AWS integration issue, not a plain Kubernetes one.

EKS-specific failure modes

These are the ones unique to running Kubernetes on AWS, mapped to what you would see:

  • Pod stuck Pending — no nodes / no IPs. In the Foundation, Pending meant "no room on nodes." On EKS it can also mean:

    • No node capacity and node autoscaling not adding nodes (the scaling lesson) — check Karpenter/Cluster Autoscaler is working and the pending pod's requests can be satisfied.
    • No available IP addresses. Because the VPC CNI gives each pod a real VPC IP (the eks-basics lesson), pods can fail to schedule if the subnet runs out of IPs — a distinctly EKS problem. describe shows a CNI/IP-allocation error. Fix: bigger subnets (the CIDR planning from the VPC lesson) or CNI IP-management settings.
  • ImagePullBackOff — cannot pull from ECR. Often the node/pod lacks permission to pull from your ECR registry, or the image tag is wrong. Check the node role / IRSA has ECR pull permission and the image URI is correct (an ECR URI like 123456789012.dkr.ecr.eu-west-1.amazonaws.com/app:tag).

  • Pod runs but cannot call AWS — IRSA misconfigured. The app gets AccessDenied calling S3/SQS/Secrets Manager. Almost always the IRSA setup (the workloads lesson): the service account annotation is missing or wrong, the IAM role's trust policy does not allow the service account, or the role lacks the needed permission. Check the pod uses the right service account and the role's policy and trust relationship.

  • Ingress created but no load balancer / 503s. The AWS Load Balancer Controller (the ingress lesson) is not running, lacks permissions, or the Ingress annotations are wrong — so no ALB is created, or it has no healthy targets. Check the controller's logs, the Ingress events, and that the target pods pass health checks.

  • OOMKilled / evictions. Memory limits too low, or requests too low under node pressure (the resources-and-limits lesson). describe shows OOMKilled or an eviction event; fix the resource numbers.

The pattern: the symptom is a normal Kubernetes status, but the cause is an AWS integration — IP exhaustion (CNI), ECR permissions, IRSA, the load balancer controller. Knowing these EKS-specific causes turns a baffling Pending or AccessDenied into a quick fix.

Where to look: kubectl and AWS together

Debugging EKS means looking in both Kubernetes and AWS:

  • In Kubernetes: kubectl describe (Events), kubectl logs, kubectl get events --sort-by=.lastTimestamp (recent cluster events across objects), and the logs of the relevant controller (the AWS Load Balancer Controller, Karpenter) when the problem is theirs.
  • In AWS: CloudWatch Logs for the app and control-plane logs (EKS can send control-plane logs to CloudWatch); the EC2 / Auto Scaling console for node issues; IAM for role/permission problems; CloudTrail to see a denied AWS API call and why (the IAM lesson). When a pod's AccessDenied is the issue, CloudTrail shows the exact call and the role that made it.

The habit is to move between the two: the Kubernetes symptom tells you where the problem is (scheduling, image, AWS call, ingress), and then you look in the corresponding AWS place (subnet IPs, ECR, IAM/CloudTrail, the load balancer controller) for the cause. And, as always on Kubernetes, you fix the manifest or the AWS configuration and re-apply — you do not hand-patch a live pod, because the reconciler will replace it.

Check your work

Start with the Foundation routine: get pods (status) → describe (Events) → logs/--previous → exec/port-forward. Solves most; the EKS layer is when the cause is an AWS integration.

EKS-specific causes (symptom → cause): Pending → no node capacity or subnet out of IPs (VPC CNI — bigger subnets); ImagePullBackOff → ECR permission/URI; pod AccessDenied calling AWS → IRSA misconfigured (service-account annotation, role trust, or missing permission); Ingress but no LB / 503s → AWS Load Balancer Controller not running / wrong annotations / unhealthy targets; OOMKilled/evictions → resource requests/limits (previous lesson).

Look in both places: Kubernetes (describe Events, logs, cluster events, the relevant controller's logs) and AWS (CloudWatch app + control-plane logs, EC2/ASG for nodes, IAM/CloudTrail for denied calls — CloudTrail shows the exact call and role). Symptom tells you where; look in the matching AWS place for the cause. Fix the manifest/AWS config and re-apply — never hand-patch a live pod.

Practice

  1. State the core kubectl debugging routine and when it is enough versus when you look at AWS.
  2. Give two EKS-specific reasons a pod can be Pending and how you'd confirm each.
  3. Diagnose a pod getting AccessDenied calling S3 — what is almost always wrong, and what do you check?
  4. Diagnose an Ingress that created no load balancer — where do you look?
  5. Explain how you'd use CloudTrail to diagnose a denied AWS API call from a pod.
  6. Explain why you fix the manifest/AWS config and re-apply rather than editing a live pod.

Official documentation

Next: the Terraform at Scale module.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship