RizTech Academy logo
RizTech Academy
CI/CDLesson 5 of 630 min

Deployment strategies: rolling, blue-green, canary and rollback

You can deploy automatically — but how you roll a new version out to users decides whether a bad release causes a brief blip or a full outage. Deployment strategies are the patterns for releasing a new version safely: rolling, blue-green, and canary, each trading off risk, cost and complexity, plus the safety net that matters most — rollback. This lesson is choosing and understanding a deployment strategy. It was added to this course deliberately, because knowing how to release without downtime is core DevOps.

The problem: releasing without breaking things

Naively, deploying means "stop the old version, start the new one" — but that causes downtime (users hit nothing while it switches) and is all-or-nothing (if the new version is broken, everyone gets the broken version at once). Real deployment strategies avoid both: they keep the app available during the release, and they limit the blast radius if the new version is bad. The strategies differ in how they do that.

Rolling deployment: replace gradually

A rolling deployment replaces old instances with new ones a few at a time, so there are always healthy instances serving traffic. This is Kubernetes' default (the scaling-and-health lesson): it brings up new-version pods, waits for them to be ready (the readiness probe), shifts traffic to them, and removes old pods — gradually, until all are new.

  • Pro: zero downtime, no extra infrastructure (you reuse the same capacity), simple — it is the default.
  • Con: during the rollout, old and new versions run simultaneously, so both must be compatible (e.g. with the same database schema). And if the new version is subtly bad, it still reaches everyone by the end, just gradually.

Rolling is the sensible default for most services, and you already have it: kubectl set image triggers a rolling update. For many teams, rolling + good tests + fast rollback is all they need.

Blue-green: switch all at once, instantly reversible

A blue-green deployment runs two complete environments: "blue" (the current version, live) and "green" (the new version, deployed alongside but not yet receiving traffic). You deploy and test green fully while blue still serves everyone; then you switch all traffic from blue to green at once (by repointing the load balancer or Service). If green turns out bad, you switch straight back to blue — an instant rollback.

  • Pro: instant switch and instant rollback (blue is still there, untouched); you test the new version in a production-identical environment before any user hits it.
  • Con: you run two full environments at once, so it costs roughly double the infrastructure during the release; and database/state shared between them needs care.

Blue-green shines where a fast, clean, all-at-once switch with a guaranteed instant rollback is worth the extra cost — but that doubled capacity is a real cost to weigh (the cost-conscious default is to use it where the risk justifies it, not everywhere).

Canary: test on a small slice first

A canary deployment releases the new version to a small percentage of users first (say 5%), watches how it behaves — errors, latency, metrics (the observability module) — and, if it looks healthy, gradually increases the share (25%, 50%, 100%). If the canary shows problems, you roll it back having exposed only a tiny fraction of users.

  • Pro: the safest for risk — a bad release hits only a small slice, and you catch problems on real traffic before a full rollout. Best when you cannot fully trust tests to catch everything (and no tests catch everything).
  • Con: the most complex — it needs traffic-splitting, and, to be done well, automated monitoring of the canary's health to decide whether to proceed or roll back. More moving parts.

The name comes from the "canary in a coal mine": the small group is the early warning. Canary is the gold standard for high-stakes services, at the cost of the machinery to run it.

Rollback: the safety net under all of them

Whatever strategy you use, the single most important capability is fast rollback — getting back to the last known-good version quickly when a release goes wrong. Because things will go wrong, and the difference between a minor incident and a major outage is often just how fast you can roll back:

  • On Kubernetes, kubectl rollout undo deployment/greetings returns to the previous version (the scaling-and-health lesson) — one command, because the old ReplicaSet is still there.
  • Blue-green rolls back by switching traffic back to blue — instant.
  • This is exactly why images are tagged by commit/version, not latest (the build lesson): to roll back, you must be able to name and redeploy a specific known-good prior version.

The guiding principle: make rolling forward safe, and make rolling back fast. A team that can deploy frequently and roll back in seconds can take small risks cheaply, which is the whole point of CI/CD. A team that cannot roll back is one bad deploy away from a long outage.

Choosing a strategy

There is no single right answer; you match the strategy to the risk and budget:

  • Rolling — the default for most services: zero downtime, no extra cost, simple. Start here.
  • Blue-green — when you want an instant, all-at-once switch with guaranteed instant rollback, and can afford the temporary double capacity.
  • Canary — for high-risk, high-traffic services where limiting a bad release to a small slice is worth the extra complexity and monitoring.

And under all of them: fast rollback, always. Most teams run rolling by default and reach for canary or blue-green on the riskiest services — which is the cost-optimised, risk-matched approach. Knowing the trade-offs lets you pick deliberately rather than cargo-culting whatever the last blog post recommended.

Check your work

The problem: naive "stop old, start new" causes downtime and is all-or-nothing (a bad version hits everyone at once). Strategies keep the app available during release and limit the blast radius.

Rolling — replace instances a few at a time (Kubernetes default, gated by readiness). Pro: zero downtime, no extra cost, simple. Con: old+new run together (must be compatible); a bad version still reaches all, gradually. The sensible default.

Blue-green — two full environments; deploy/test green while blue serves, then switch all traffic at once; switch back for instant rollback. Pro: instant switch/rollback, prod-identical test. Con: ~double infrastructure during release.

Canary — release to a small % first, watch metrics, ramp up if healthy. Pro: safest — a bad release hits a tiny slice, caught on real traffic. Con: most complex (traffic-splitting + monitoring).

Rollback (the safety net under all): fast return to last known-good — kubectl rollout undo, or blue-green switch-back; why images are tagged by commit/version, not latest. Principle: roll forward safe, roll back fast.

Choose by risk/budget: rolling default; blue-green for instant-switch-with-instant-rollback (double cost); canary for high-risk/high-traffic (complexity). Always fast rollback.

Practice

  1. Explain why "stop the old, start the new" is a bad deployment approach.
  2. Describe a rolling deployment and its main pro and con; relate it to Kubernetes' default.
  3. Describe blue-green: how the switch and rollback work, and the cost trade-off.
  4. Describe canary: how it limits risk, and what machinery it needs to do well.
  5. Explain why fast rollback is the most important capability, and how commit-tagged images enable it.
  6. For three services (an internal tool, a payments API, a high-traffic site), choose a strategy and justify it by risk and cost.

Official documentation

Next: secrets in pipelines, handled safely.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship