RizTech Academy logo
RizTech Academy
Deployment PipelinesLesson 4 of 530 min

Rollback strategies that actually work

Deployments go wrong — a bad build reaches production, an unforeseen bug appears under real traffic. When that happens, the single thing that determines whether it is a minor blip or a long outage is how fast you can roll back. A team that can return to the last known-good version in seconds can take small risks cheaply; a team that cannot is one bad deploy away from a lengthy incident. This lesson is rollback strategies that actually work — building on the DevOps Foundation's deployment-strategies lesson, made concrete for EKS.

Rollback is the safety net under everything

The deployment-strategies lesson said it and it bears repeating: make rolling forward safe, and make rolling back fast. Because deployments will occasionally break, the ability to roll back quickly is what makes frequent deployment safe. Every practice in this lesson serves one goal: when a deploy is bad, get back to the last good version fast, with confidence. Rollback is not an admission of failure; it is the mechanism that lets you deploy often without fear.

Kubernetes rollback: the previous version is still there

On EKS, rolling back a Deployment is built in, because Kubernetes keeps the history of a Deployment's rollouts (its previous ReplicaSets — the Foundation's Deployment → ReplicaSet chain):

kubectl rollout undo deployment/greetings              # roll back to the previous version
kubectl rollout undo deployment/greetings --to-revision=3   # or a specific earlier revision
kubectl rollout history deployment/greetings           # see the revisions available

kubectl rollout undo returns the deployment to the previous image and config, performing a rolling update back to it — the old ReplicaSet is scaled up, the bad one scaled down, gated by readiness, zero downtime. It is fast because Kubernetes did not throw the previous version away. If you use Helm, helm rollback greetings <revision> does the same at the release level, reverting all the chart's resources together.

So the primary EKS rollback is one command, and it works because the previous version's definition is retained. This is why the deploy should be a clean Kubernetes/Helm operation — it gives you instant, built-in rollback.

Why commit-tagged images make rollback real

Rollback only works if you can name and redeploy a specific known-good version — which is exactly why the containers and building-images lessons insisted on commit-SHA tags, not latest:

  • To roll back, you redeploy the previous image — greetings:previous-sha. That image still exists in ECR (kept by the lifecycle policy's retention — the building-images lesson), and its SHA tag immutably identifies the exact code. You know precisely what you are rolling back to.
  • With latest, you could not — "the previous latest" is ambiguous, and the old image may have been overwritten. Rollback becomes guesswork or impossible.

So the immutable, commit-tagged, retained-in-ECR image is the foundation of reliable rollback: a specific, known-good artifact you can always redeploy. The tagging discipline from earlier lessons pays off precisely here, in the incident, when you need to get back to safety fast.

Rollback and deployment strategy

How you roll back depends partly on your deployment strategy (the Foundation's deployment-strategies lesson):

  • Rolling deployment (the default): roll back with kubectl rollout undo / helm rollback — a rolling update back to the previous version. Simple and built-in.
  • Blue-green: roll back by switching traffic back to the old (blue) environment, which is still running and untouched — the fastest possible rollback, essentially instant, because you are just repointing the load balancer. Blue-green's main appeal is exactly this instant, clean rollback.
  • Canary: if the canary (the small percentage on the new version) shows problems, roll back by routing all traffic back to the stable version and removing the canary — having exposed only a small slice of users to the bad version. Canary limits the blast radius and rolls back cleanly.

Each strategy has a fast rollback; the common thread is that the previous good version is kept available (as a retained image, an old ReplicaSet, or a whole blue environment) so returning to it is quick.

Automating and practising rollback

Two practices turn "we can roll back" into "we reliably do":

  • Automate rollback on failed verification. If the pipeline's post-deploy smoke test or health check fails (the deploying lesson), the pipeline can automatically roll back rather than leaving a broken version live while a human investigates. Some setups (progressive delivery tools like Argo Rollouts or Flagger) automate canary analysis and auto-rollback on bad metrics. Automatic rollback on a failed check is the fastest, most reliable recovery.
  • Practise rollback. A rollback procedure you have never run is not a procedure you can trust at 3am. Rehearse it — deliberately roll back a deploy in staging — so that when you need it in production, it is routine. The teams that recover fastest are the ones for whom rollback is a well-worn path, not a panicked first attempt.

The whole picture: commit-tagged images retained in ECR give you known-good versions; Kubernetes/Helm give you one-command rollback to them; your deployment strategy determines how instant it is; and automating rollback on failed verification plus practising it make recovery fast and reliable. Combined, they mean a bad deploy is a two-minute recovery, not an outage — which is what makes deploying frequently safe, and is one of the highest-value capabilities a team can have.

Check your work

Rollback is the safety net: deployments will break, so fast rollback is what makes frequent deployment safe (roll forward safe, roll back fast). Not a failure — the mechanism to deploy without fear.

Kubernetes rollback (built in): kubectl rollout undo (previous or --to-revision), rollout history; Helm helm rollback (whole release). Fast because Kubernetes keeps the previous ReplicaSet/version — a rolling update back, readiness-gated, zero downtime.

Commit-tagged images make it real: roll back = redeploy the previous SHA image (still in ECR via retention, immutably identifying the code). latest makes rollback ambiguous/impossible. The tagging discipline pays off in the incident.

By strategy: rolling → rollout undo/helm rollback; blue-green → switch traffic back to the untouched old environment (instant); canary → route all traffic back to stable, small slice exposed. All keep the previous good version available.

Automate + practise: auto-roll-back on failed post-deploy verification (fastest, most reliable; Argo Rollouts/Flagger for canary auto-rollback); rehearse rollback in staging so it's routine, not a 3am first attempt. Result: a bad deploy = 2-minute recovery, not an outage.

Practice

  1. Explain why fast rollback is what makes frequent deployment safe.
  2. Roll back an EKS Deployment with kubectl, and explain why it is fast (what Kubernetes retained).
  3. Explain how commit-SHA image tags and ECR retention make rollback to a specific known-good version possible.
  4. Describe how rollback works under rolling, blue-green and canary strategies, and which is most instant.
  5. Explain how a pipeline can automatically roll back on a failed post-deploy check, and why that helps.
  6. Explain why practising rollback in staging matters for real incidents.

Official documentation

Next: secrets — Parameter Store and Secrets Manager.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship