kubectl and debugging a broken pod
Something is wrong: a pod will not start, keeps restarting, or the app is not reachable. On Kubernetes, the
skill that makes you useful is not writing YAML — it is diagnosing what the cluster is actually doing when it
does not match what you wanted. kubectl gives you everything you need to find the problem, if you know the
handful of commands and how to read them. This lesson is kubectl for debugging, and it closes the Kubernetes
module.
The everyday kubectl commands
You drive the cluster with kubectl, and a small set of commands covers almost everything:
kubectl get pods # list pods and their status
kubectl get pods -o wide # ...with node and IP
kubectl get all # deployments, replicasets, pods, services at once
kubectl apply -f k8s/ # apply all manifests in a directory (declare desired state)
kubectl delete -f k8s/ # remove them
kubectl scale deployment/x --replicas=3
kubectl rollout status deployment/x
kubectl rollout undo deployment/x
kubectl get is your constant "what is the state?" check. kubectl apply -f is the declarative verb you use
to create and update everything. But when something is wrong, three commands do the diagnosis: describe,
logs, and exec.
kubectl describe: why is this pod not happy?
kubectl describe pod <name> is the first thing you run on a misbehaving pod. It shows the pod's full
state — its containers, its probes, its resource requests — and, most valuably, the Events at the bottom: a
timeline of what Kubernetes tried and what happened. The Events section usually tells you the problem directly:
kubectl describe pod greetings-b5d7b9fd9-4m8kq
- "Failed to pull image" → the image name is wrong or not available (an
ImagePullBackOff). - "Liveness probe failed" → the app is not responding on its health check.
- "Insufficient cpu/memory" → the scheduler cannot place the pod because no node has room (relates to your resource requests).
- "Back-off restarting failed container" → the container starts and immediately crashes (a
CrashLoopBackOff).
Reading the Events is the single highest-value Kubernetes debugging skill. Most "why won't this pod run?" questions are answered in that list.
Reading pod status: the common failure states
kubectl get pods shows a STATUS, and a few recurring ones each point at a specific cause:
Running— good (though checkREADY 1/1, not0/1:0/1means the readiness probe is failing, so it is running but getting no traffic).Pending— not scheduled yet; usually no node has the resources, or a volume cannot be attached.describesays which.ImagePullBackOff/ErrImagePull— cannot pull the image: wrong name/tag, or no access to the registry. Check the image reference.CrashLoopBackOff— the container starts, crashes, and Kubernetes keeps restarting it with growing back-off. The app itself is failing on startup — go to the logs.
Recognising these statuses tells you where to look next, which is half the battle.
kubectl logs: what did the app say?
When a container is crashing or misbehaving, its logs are where the real error is (a containerised app logs to stdout, so Kubernetes captures it — the containers module):
kubectl logs greetings-b5d7b9fd9-4m8kq # the container's logs
kubectl logs -f greetings-b5d7b9fd9-4m8kq # follow live
kubectl logs greetings-b5d7b9fd9-4m8kq --previous # logs of the PREVIOUS (crashed) container
That last one is essential for CrashLoopBackOff: the current container may be too new to have logged
anything, but --previous shows the output of the instance that just crashed — which usually contains the
actual error (a missing env var, a failed DB connection, a config typo). "The pod is CrashLoopBackOff" →
kubectl logs --previous → read the exception → fix it.
kubectl exec and port-forward: get inside
Two more tools for when logs are not enough:
kubectl exec -it <pod> -- sh— a shell inside a running pod (likedocker exec), to check env variables, files, network from the pod's point of view. Useful for "is the config actually there?" and "can this pod reach that service?".kubectl port-forward service/greetings 8080:80— forward a local port to a Service or pod, so you cancurl localhost:8080and test the app directly without exposing it publicly. This is how you verify "is the app itself responding?" (verified for greetings: it returns the JSON with the greeting). It separates "the app is broken" from "the routing/ingress is broken".
A debugging routine
Put it together into a routine for "something is wrong on Kubernetes":
kubectl get pods— what is the status? (Pending?ImagePullBackOff?CrashLoopBackOff?0/1ready?) — this points you at the cause.kubectl describe pod <name>— read the Events for the reason (bad image, failed probe, no resources).kubectl logs <name>(add--previousif crashing) — read the app's actual error.kubectl exec/port-forward— get inside or test the app directly when you need to.- Fix the desired state (the YAML) and
kubectl apply— because Kubernetes is declarative, you fix the manifest, not the live pod (a change to a live pod is lost when it is replaced).
That last point is the Kubernetes mindset: you do not hand-fix running pods; you correct the declared desired state and re-apply, and the reconciler makes it so. Master this routine and Kubernetes stops being a black box — a broken pod becomes a status to read, an Events list to check, and a log to inspect, ending in a one-line fix to a manifest. That diagnostic fluency, more than any amount of YAML, is what makes you effective on a cluster.
Check your work
Everyday: kubectl get pods/get all (state), kubectl apply -f (declare/update — the workhorse),
scale, rollout status/undo. When wrong, three commands diagnose: describe, logs, exec.
describe pod — read the Events (bottom): "Failed to pull image", "Liveness probe failed",
"Insufficient cpu/memory", "Back-off restarting". Highest-value debugging skill.
Statuses: Running (check READY 1/1, not 0/1 = readiness failing), Pending (unscheduled — no
resources/volume), ImagePullBackOff/ErrImagePull (bad image/registry access), CrashLoopBackOff (starts
then crashes — go to logs). The status tells you where to look.
logs — the app's real error; --previous for a crashed container (essential for CrashLoopBackOff).
exec -it -- sh (inside a pod) and port-forward (test the app directly — separates "app broken" from
"routing broken").
Routine: get pods (status) → describe (Events) → logs (--previous if crashing) → exec/port-forward → fix
the manifest and apply (never hand-fix a live pod — declarative). Diagnostic fluency > YAML.
Practice
- Run
kubectl get allon the greetings app and identify the Deployment, ReplicaSet, pods and Service. - Break the image name in the Deployment, apply, and use
kubectl get pods+describeto identify theImagePullBackOff. - Cause a crash (e.g. a bad start command), and use
kubectl logs --previousto find the error. - Explain what
READY 0/1on a Running pod means and which probe is involved. - Use
kubectl port-forwardto test the greetings app directly and confirm it responds. - Explain why you fix the manifest and re-apply rather than editing a live pod, in terms of the desired-state model.
Official documentation
- Kubernetes — kubectl overview & cheat sheet — The essential commands.
- Kubernetes — Debug running pods — describe, logs, exec and interpreting status.
- Kubernetes — Determine the reason for pod failure — Reading crash and failure states.
Next: the CI/CD module.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship