Running an incident, and the report afterwards
An alert fired: production is broken. This lesson is what you do next — running the incident calmly and effectively, using all the observability you built, and writing the report afterwards that makes the system more reliable. It builds on the DevOps Foundation's incident-response lesson, adding the AWS specifics, and it closes the observability module. Incidents are where everything in this course combines under pressure.
The response, on AWS
The incident-response sequence from the Foundation applies unchanged; here it is with AWS specifics:
- Acknowledge and assess severity. Confirm the alert is real (not noise — the alerting lesson) and gauge impact: how many users, how badly? A payments outage is critical; a slow non-critical page is minor. Severity sets urgency and who is involved.
- Communicate early. Post to a status channel, tell affected teams, and for user-facing outages update a status page. "We are aware and investigating" beats silence. Keep updating.
- Mitigate before you diagnose. The priority is restoring service, not understanding root cause. On AWS,
the fastest mitigations are usually:
- Roll back if a recent deploy caused it —
kubectl rollout undo/helm rollback(the rollback lesson). "What changed just before this started?" is the highest-yield question, and it is very often a deploy. - Scale up if it is load — more replicas / nodes (the scaling lesson).
- Fail over / disable a broken dependency or feature. Stop the bleeding by the fastest safe means, then investigate.
- Roll back if a recent deploy caused it —
- Diagnose with your observability. Now everything you built pays off: metrics (CloudWatch/Grafana) to see what changed and when, traces to find where, logs (CloudWatch Logs Insights) to find the specific error, and CloudTrail to see any AWS API changes or denials (the IAM lesson). Metrics → traces → logs, the Foundation's flow, on AWS tools.
- Fix and verify. Apply the real fix, and confirm via the metrics and the live system that it is actually resolved — not assumed resolved.
- Resolve and communicate the all-clear.
The whole course feeds steps 3-5: the rollback, the scaling, the networking-debug chain, the EKS debugging, the logs and metrics. Incident response is where these combine — and the calm, ordered sequence is what keeps the pressure manageable.
Diagnosing with AWS observability
Concretely, when diagnosing an AWS incident:
- Start at the symptom's metric. The alert points at errors or latency; the dashboard (the dashboards lesson) shows when it started and what correlates — did latency rise when CPU saturated, when a deploy happened, when traffic spiked? The shared-timeline dashboard makes the trigger visible.
- "What changed?" Check recent deployments (the pipeline) and infrastructure changes (Terraform applies) around the start time — most incidents follow a change. CloudTrail logs AWS changes; your pipeline history logs deploys.
- Drill into logs at the right time and service. CloudWatch Logs Insights, filtered by the affected service/pod and the incident time window, and by a correlation ID if you have one (the logs lesson), gets you to the actual error fast.
- Check the AWS-integration failure modes (the EKS-debugging lesson) — an IRSA/permission issue (CloudTrail shows the denied call), an image pull failure, a load balancer problem, resource exhaustion.
The efficiency comes from having built observability before the incident: metrics to localise in time, logs to find the error, CloudTrail to see changes and denials. A team with this can go from "alert" to "cause" in minutes; a team without it guesses for hours. That is the entire return on the observability work.
The blameless postmortem
After resolution comes the most valuable part (the Foundation): the blameless postmortem. Write up the timeline (when it started, when detected, when mitigated, when resolved — your metrics and logs give you these timestamps), the impact (who/what was affected, for how long), the root cause, and — above all — the action items to prevent recurrence.
The essentials (the Foundation's incident lesson):
- Blameless. The question is "what about our system allowed this?", never "who broke it?" Blame makes people hide problems and stops the organisation learning. "A human ran the wrong command" is not a root cause; "our system let one command take down production with no guardrail" is — and the fix is a guardrail, not "be careful."
- Ask why until you reach something fixable, and produce concrete action items: "add an alert for this condition," "make this deploy auto-rollback on a failed check," "add an SCP/IAM guardrail that would have prevented this," "raise the resource limit that caused the OOMKill," "write a runbook for this failure."
- Track the action items to completion. A postmortem whose actions are never done teaches nothing; the value is in closing the gaps each incident reveals.
Each incident, properly reviewed, makes the system more reliable — an alert added, a deploy made reversible, a guardrail put in place. Teams that do blameless postmortems get steadily more reliable; teams that blame and move on hit the same problems repeatedly. This is the loop that turns failures into resilience, and it is the culmination of the whole course: not avoiding failure (impossible), but building systems and habits that respond to it well and learn from it.
Check your work
Response (AWS): acknowledge + assess severity (real? how many users?); communicate early ("aware, investigating") and keep updating; mitigate before diagnose — fastest safe fix: roll back a recent deploy (top cause — "what changed?"), scale up, or fail over/disable; diagnose with metrics (CloudWatch/ Grafana: what/when) → traces (where) → logs (CloudWatch Insights: what/why) → CloudTrail (AWS changes/ denials); fix and verify it's actually resolved; all-clear. The whole course feeds steps 3-5.
Diagnose efficiently: start at the symptom's metric + dashboard (when it started, what correlates); ask "what changed?" (deploys, Terraform applies — most incidents follow a change; CloudTrail + pipeline history); drill into logs at the right time/service/correlation ID; check AWS-integration failure modes (IRSA/ permissions via CloudTrail, image pulls, LB, resource exhaustion). Built-beforehand observability = minutes to cause, not hours.
Blameless postmortem: timeline, impact, root cause, action items. Blameless ("what let this happen?" not "who?" — blame hides problems; a human error → a missing guardrail is the real cause). Ask why to something fixable; concrete actions (add alert, auto-rollback, guardrail, raise limit, runbook); track to completion. Each reviewed incident closes a gap → steadily more reliable. The course's culmination: respond well and learn.
Practice
- Put the incident-response steps in order and explain why "mitigate before diagnose" is the key instinct.
- For a spike of 5xx errors right after a deploy, describe your first action and how you'd confirm the cause.
- Describe how you would use CloudWatch metrics, Logs Insights and CloudTrail together to diagnose an incident.
- Explain why "what changed?" is the highest-yield diagnostic question and where you'd check.
- Explain what makes a postmortem blameless and why blame makes systems less reliable.
- Turn a shallow root cause ("someone deployed a bad build") into a systemic one with two concrete action items.
Official documentation
- Google SRE Book — Managing incidents — Running an incident response.
- Google SRE Book — Postmortem culture — Blameless postmortems and learning from failure.
- AWS — Incident response with CloudWatch & CloudTrail — Using AWS observability during incidents.
Next: the Cost Optimisation module.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship