Health checks, self-healing and scaling
Two of the reasons orchestration exists (the first lesson) are self-healing — keeping the app running when things fail — and scaling — running the right number of replicas for the load. Kubernetes does both, driven by the desired-state model and by health checks that tell it which pods are actually working. This lesson makes those concrete, verified against a live cluster: probes, self-healing, scaling, and zero-downtime rolling updates.
Health checks: how Kubernetes knows a pod is working
A pod being "Running" does not mean the app inside it works — it could be started but hung, or still warming up. Kubernetes finds out with probes — checks it performs against your container. Two matter most:
- Liveness probe — "is the app alive?" If it fails repeatedly, Kubernetes restarts the pod (the app has hung or is broken). This is automatic recovery from a wedged process.
- Readiness probe — "is the app ready to serve traffic?" If it fails, Kubernetes removes the pod from the Service (stops sending it traffic) until it passes again — without restarting it. This is how a pod that is starting up, or temporarily busy, is kept out of rotation until it can actually handle requests.
The greetings Deployment defines both, verified on a live cluster, hitting the app's /health endpoint (the
one the app exposes for exactly this):
livenessProbe:
httpGet:
path: /health
port: 3000
initialDelaySeconds: 3
periodSeconds: 10
readinessProbe:
httpGet:
path: /health
port: 3000
initialDelaySeconds: 2
periodSeconds: 5
The distinction is worth holding: liveness failing → restart the pod; readiness failing → stop sending it
traffic (but leave it running). Getting these right is what makes deployments and recovery smooth — a readiness
probe is why a new pod does not receive traffic until it is actually ready, so a rolling update causes no
errors. This is why the greetings app has a /health endpoint at all: it exists so Kubernetes can probe it.
Self-healing in action
Because a Deployment (via its ReplicaSet) continuously reconciles toward the desired replica count, it heals failures automatically. Delete a pod and watch:
kubectl delete pod greetings-b5d7b9fd9-4m8kq
kubectl get pods -l app=greetings
The verified result: the deleted pod goes to Terminating, and a brand-new pod appears immediately to
restore the count to 2 — you did nothing. Kubernetes saw actual (1) ≠ desired (2) and created one. The same
happens if a whole node dies: its pods are rescheduled onto healthy nodes. This is the self-healing that no
human could provide by hand, and it falls straight out of the desired-state model — you declared "2", so
Kubernetes maintains "2", forever.
Scaling: more replicas on demand
Scaling is just changing the desired replica count. Manually:
kubectl scale deployment/greetings --replicas=3 # verified: pod count goes to 3
Kubernetes sees desired (3) > actual (2) and starts one more pod; scale back to 2 and it removes one. The Service (the previous lesson) automatically includes the new pods in load-balancing, because it selects by label — so scaling up genuinely spreads traffic wider with no extra wiring.
For automatic scaling, a HorizontalPodAutoscaler (HPA) adjusts the replica count based on load (CPU,
memory, or custom metrics): "keep CPU around 60%, between 2 and 10 replicas". Under a traffic spike it adds
pods; when it subsides it removes them. This is the cloud-native answer to load — and it is why the greetings
Deployment sets resource requests (cpu: 50m, memory: 64Mi): the autoscaler and the scheduler both rely
on those requests to know how much a pod needs and how loaded it is. (Right-sizing requests, and preferring a
sensible fixed replica count where load is predictable, is also the cost-conscious choice — you do not want an
autoscaler spinning up expensive capacity you did not need.)
Rolling updates and rollbacks: deploying without downtime
The last piece: updating to a new version without taking the app down. When you change the image and apply, the Deployment performs a rolling update — it brings up new-version pods and takes down old ones gradually, so there are always healthy pods serving traffic:
kubectl set image deployment/greetings greetings=greetings:1.1 # or edit the YAML and apply
kubectl rollout status deployment/greetings # watch it roll (verified: "successfully rolled out")
Under the hood this is the Deployment → ReplicaSet mechanism (the pods-and-deployments lesson): a new ReplicaSet scales up while the old one scales down, and the readiness probe gates each step — a new pod only receives traffic once it is ready, so users see no errors. If the new version is bad, roll back instantly:
kubectl rollout undo deployment/greetings # return to the previous version
Zero-downtime deploys and instant rollback are among the biggest practical wins of Kubernetes, and they come from combining desired state, ReplicaSets and readiness probes. (This is the foundation the CI/CD module's deployment-strategies lesson builds on — rolling is the default; blue-green and canary are variations.)
Check your work
Probes: liveness ("alive?" — fail → restart the pod) and readiness ("ready for traffic?" — fail
→ remove from the Service, don't restart). The greetings app's /health exists so Kubernetes can probe it.
Readiness gating is why rolling updates cause no errors.
Self-healing: the Deployment/ReplicaSet reconciles to the desired count — delete a pod and a new one appears immediately (verified: back to 2); a dead node's pods reschedule elsewhere. Falls straight out of desired state.
Scaling: change the replica count (kubectl scale, verified 2→3→2); the Service load-balances the new pods
automatically (label selector). HPA autoscales on load (CPU/memory), relying on resource requests (why
the Deployment sets cpu: 50m etc.); right-size requests and prefer fixed counts where load is predictable
(cost).
Rolling update: change the image + apply → new ReplicaSet up, old down gradually, gated by readiness →
zero downtime (verified rollout). kubectl rollout undo = instant rollback. Foundation for the CI/CD
deployment-strategies lesson.
Practice
- Explain the difference between a liveness and a readiness probe and what Kubernetes does when each fails.
- Add liveness and readiness probes hitting
/healthto the greetings Deployment and apply it. - Delete a pod and observe Kubernetes recreate one; explain how desired state drives this.
- Scale the Deployment to 3 and back to 2 with
kubectl scale, and explain how the Service includes new pods. - Explain what an HPA does and why resource requests are needed for it (and the cost trade-off).
- Perform a rolling update (change the image) and watch
kubectl rollout status; thenkubectl rollout undoand explain how readiness makes it zero-downtime.
Official documentation
- Kubernetes — Liveness, readiness and startup probes — Health checks and what each does.
- Kubernetes — Horizontal Pod Autoscaling — Automatic scaling on load.
- Kubernetes — Rolling updates & rollbacks — Zero-downtime deploys and undo.
Next: kubectl and debugging a broken pod.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship