What to do when something breaks
Eventually something will break in production — a deploy goes wrong, a dependency fails, traffic spikes, a certificate expires. When it does, what matters is not that it happened (it always will) but how you respond: calmly, methodically, and in a way the team learns from. This final lesson of the course is incident response basics — what to do when something breaks — and it ties together nearly everything you have learned. It closes the DevOps Foundation course.
An incident is normal — the response is what counts
First, the mindset: incidents are inevitable. Every real system fails sometimes; the best engineers and companies have outages. So the measure of a good operation is not zero incidents (impossible), but a fast, calm, blameless response and a system that gets more reliable over time because of what each incident teaches. Panic and blame make incidents worse and longer; process and learning make them shorter and rarer. Going in with that frame — "this is normal, we have a process" — is half the battle.
The response, step by step
When an incident hits (an alert fires, or someone reports the site is down), work through a clear sequence rather than flailing:
- Acknowledge and assess. Confirm it is real (not a flaky alert), and gauge severity — how many users, how badly affected? A total outage of payments is critical; a slow non-critical page is minor. Severity sets how urgently and widely you respond.
- Communicate. Tell the people who need to know — a status channel, affected teams, and (for user-facing outages) users via a status page. Saying "we are aware and investigating" early is far better than silence; it stops duplicate reports and builds trust. Keep updating as you learn more.
- Mitigate first, diagnose second. The priority is stopping the bleeding — restoring service — not understanding the root cause. If a bad deploy caused it, roll back (the deployment-strategies lesson) — do not debug the new version while users suffer. Restore service by the fastest safe means (roll back, scale up, fail over, disable the broken feature), then investigate at leisure. This is the single most important incident instinct: mitigate now, understand later.
- Diagnose with your observability. Now use everything from this module: metrics to see what changed and when, traces to find where, logs to find the specific error (the pillars lesson). "What changed just before this started?" — very often a recent deploy or config change — is the highest-yield question.
- Fix and verify. Apply the real fix, and confirm via the metrics and the live system that the incident is actually resolved (the deploy-verification habit) — not assumed resolved.
- Resolve and communicate the all-clear. Tell everyone it is over.
Notice how much of this course feeds into step 4 and 5: the networking debugging chain, kubectl for a broken
pod, the logs and metrics, the rollback. Incident response is where all the skills combine under pressure — and
the calm sequence above is what keeps that pressure manageable.
The blameless postmortem: where reliability actually comes from
After the incident is resolved comes the most valuable part, and the one teams most often skip: the postmortem (or incident review). You write up what happened — the timeline, the impact, the root cause, and above all the action items to prevent recurrence — and share it.
The essential quality is that it is blameless. The goal is never "who broke it?" but "what about our system allowed this, and how do we prevent it?" Because:
- Blame makes people hide problems. If incidents lead to punishment, people conceal mistakes and near-misses, and the organisation stops learning — which makes things less safe.
- The real causes are systemic. "A human ran the wrong command" is not a root cause; the root cause is "our system let a single wrong command take down production, with no confirmation and no easy rollback." The fix is systemic (add a safeguard), not "tell people to be careful."
A good blameless postmortem asks why until it reaches something fixable, and produces concrete action items — "add an alert for this condition", "make this deploy reversible", "add a check that would have caught this", "write a runbook for this failure". Those action items are how a system gets more reliable: each incident, properly reviewed, closes a gap. Teams that do blameless postmortems get steadily more reliable; teams that blame and move on hit the same problems repeatedly. This is the loop that turns failures into resilience — and it is a fitting end to the course, because it is the essence of DevOps: not avoiding failure, but building systems and habits that respond to it well and learn from it.
What you have learned
Step back and see the whole course. You started at the Linux command line and built up: navigating and operating a server, the networking that connects them, Git for collaborating on the code, containers to package applications, Kubernetes to run them at scale, CI/CD to ship them automatically and safely, Infrastructure as Code to manage the infrastructure itself, and now observability and incident response to operate it all in production. That is the full arc of getting software from a developer's machine to running reliably for users — which is what DevOps is. You will not be an expert in any one of these yet; nobody is after one course. But you now have the whole map, hands-on with the standard tools, and — as the AWS DevOps course builds on this — you are ready to go deeper and to do real infrastructure work. That is exactly what this foundation set out to give you.
Check your work
Mindset: incidents are inevitable; the measure is a fast, calm, blameless response and a system that learns. Panic/blame lengthen incidents; process/learning shorten them.
Response sequence: (1) acknowledge + assess severity (how many users, how bad); (2) communicate early ("aware and investigating") and keep updating; (3) mitigate first, diagnose second — stop the bleeding (roll back a bad deploy, scale, fail over, disable the feature) before understanding; (4) diagnose with metrics (what/when) → traces (where) → logs (what/why), asking "what changed?"; (5) fix and verify it is actually resolved; (6) all-clear. The whole course feeds steps 4-5.
Blameless postmortem (where reliability comes from): write timeline, impact, root cause, and action items; blameless because blame makes people hide problems and real causes are systemic ("the system let one wrong command take prod down", not "a human erred"). Ask why to something fixable; each reviewed incident closes a gap. Learn-don't-blame is the essence of DevOps.
The arc: Linux → networking → Git → containers → Kubernetes → CI/CD → IaC → observability/incidents — the full path from a dev's machine to reliably serving users. You have the whole map, hands-on, ready to go deeper.
Practice
- Explain why "zero incidents" is the wrong goal and what the right measure of an operation is.
- Put the incident-response steps in order and explain why "mitigate before diagnose" matters most.
- For a bad deploy causing errors, describe your first action and why (not debugging the new version).
- Describe how you would use metrics, traces and logs together to diagnose a live incident.
- Explain what makes a postmortem "blameless" and why blame makes systems less reliable.
- Turn a shallow root cause ("someone ran the wrong command") into a systemic one with a concrete action item.
Official documentation
- Google SRE Book — Managing incidents — Running an incident response.
- Google SRE Book — Postmortem culture — Blameless postmortems and learning from failure.
- Atlassian — Incident management handbook — Severity, communication and response, practically.
You have completed the DevOps Complete Foundation. Next: AWS DevOps in Practice builds on all of it with real cloud infrastructure.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship