Alerts worth waking up for
An alert wakes someone up. That framing decides everything about good alerting: an alert must be worth a human's interrupted sleep, which means it fires only when a human genuinely needs to act, now. Most alerting fails the opposite way — too many alerts, most of them noise — until people ignore all of them, including the one that mattered. This lesson is alerts worth waking up for, applying the Foundation's alerting principles to AWS. It is the most important observability lesson, because the others are worthless if nobody is told when things break.
The one rule: every alert must be actionable
The Foundation's central alerting principle, restated: every alert must mean "a human needs to act, now." An alert that fires when nothing needs doing — a transient blip, a threshold set too tight, an internal metric that does not matter — trains people to ignore alerts. This is alert fatigue, and it is fatal: once people learn that alerts are usually noise, they mute the channel or ignore the page, and then the real alert is missed too. A noisy alerting system is worse than none, exactly like a flaky test suite (the QA course) — it costs attention and provides false comfort.
So the discipline is ruthless: a small number of alerts, every one of which, when it fires, requires action. If an alert routinely fires and the response is "ignore it" or "it recovered on its own," delete or fix that alert. Quality over quantity, always.
Alert on symptoms, not causes
The second Foundation principle: alert on symptoms (what users experience), not causes (internal metrics).
- Symptom alerts — high error rate, high latency, the service down/unreachable. These mean users are being hurt right now, whatever the underlying cause, so they always warrant action. Alert on the 5xx rate, on p99 latency crossing a threshold, on a health check failing.
- Cause alerts are usually noise — "CPU at 80%" may be perfectly fine (the service is handling load); "disk 60% full" is not urgent. Alerting on every internal metric fires constantly for things that do not matter to users.
You investigate causes (with dashboards and logs — the previous lessons, and the incident lesson next) once a symptom alert fires. Alerting on user-facing symptoms catches real problems without the false alarms of cause-based alerting. A short list of symptom alerts — errors, latency, availability, and a few critical resource-exhaustion warnings (a disk about to fill, which will cause an outage) — covers what matters.
Alerting on AWS: CloudWatch Alarms and beyond
The AWS mechanics:
- CloudWatch Alarms watch a CloudWatch metric (the metrics lesson) and fire when it crosses a threshold for a duration — e.g. "ALB 5xx rate > 5% for 5 minutes", "target latency p99 > 1s for 5 minutes". Alarms notify via SNS, which routes to email, chat (Slack/Teams), or a pager.
- Composite alarms combine conditions to reduce noise (alert only if errors and latency are both bad), and alarms can require a condition to persist (the duration) so a momentary blip does not page anyone.
- Prometheus alerting (Alertmanager, or AMP's alerting) fires on PromQL conditions over your app/Kubernetes metrics, for the signals Prometheus holds — routed similarly.
- On-call tools — PagerDuty, Opsgenie (or AWS's incident tools) receive the alerts, handle escalation (if the first person does not ack, escalate), on-call schedules, and acknowledgement. For anything that pages a human out of hours, an on-call tool manages who gets woken and what happens if they do not respond.
So an alert flows: a metric crosses a threshold (CloudWatch Alarm / Prometheus rule) → SNS/Alertmanager → chat for warnings, or a pager (PagerDuty/Opsgenie) for wake-someone-up severities → a human acts. Setting the right thresholds and durations (tight enough to catch real problems, loose enough not to cry wolf) is the craft.
Severity: not everything is a page
A key practice for avoiding fatigue is severity levels — not every alert should wake someone at 3am:
- Page (wake someone up) — reserved for user-impacting problems that need immediate action: the service is down, error rate is high, latency is breaking the SLO. These go to the pager.
- Ticket / chat (handle in hours) — things that need attention but not immediately: a non-critical warning, a slowly-filling disk with days of runway, a degraded but working component. These go to a channel or a ticket, not a page.
- Informational — logged for context, no notification.
Matching severity to urgency is what keeps the pager quiet enough to be trusted. If everything pages, nothing is urgent; reserve the page for genuine, immediate, user-impacting problems, and route the rest to where it gets seen without waking anyone. This severity discipline, plus symptom-based and actionable-only alerting, is what produces a pager that is quiet until it genuinely needs you — the only kind anyone actually responds to.
Tie it to SLOs
The mature framing (the Foundation): define an SLO (e.g. "99.9% of requests succeed under 300ms") and alert when you are at risk of breaching it — a symptom alert grounded in a concrete, user-centred target. Alerting on SLO burn (you are consuming your error budget too fast) is precisely "a user-impacting problem needs action," and it avoids both over-alerting (a brief blip within budget does not page) and under-alerting (sustained breach does). You do not need SLOs on day one, but they are where good alerting leads: alert on the user experience against a defined target, actionably, at the right severity.
Check your work
The one rule: every alert = "a human must act, now." Non-actionable alerts cause alert fatigue (people learn alerts are noise → ignore all → miss the real one) — a noisy system is worse than none (like flaky tests). Few alerts, every one actionable; delete/fix any that routinely fire-and-recover.
Symptoms not causes: alert on user-facing symptoms (error rate, p99 latency, service down/unreachable — always warrant action) not internal causes (80% CPU may be fine — noise). Investigate causes after a symptom fires. Plus a few critical resource-exhaustion warnings (a disk that will fill).
On AWS: CloudWatch Alarms (threshold + duration on a metric → SNS → chat/pager); composite alarms and durations reduce noise; Prometheus/Alertmanager for app/K8s metrics; PagerDuty/Opsgenie for on-call escalation/ack. Craft = right thresholds + durations.
Severity: page only for user-impacting/immediate (down, high errors, SLO breaking); ticket/chat for attention-not-urgent (slow disk, degraded-but-working); informational logged. Match severity to urgency to keep the pager trusted.
SLOs: alert on being at risk of breaching a user-centred target (SLO burn / error budget) — actionable, avoids over- and under-alerting. Where good alerting leads.
Practice
- Explain why every alert must be actionable, and how alert fatigue mirrors the flaky-test problem.
- Give three good symptom alerts and three cause "alerts" that would be noise, and explain the difference.
- Describe how a CloudWatch Alarm fires and routes a notification, including the role of duration and SNS.
- Explain severity levels and which alerts should page versus go to a ticket, with examples.
- Explain how on-call tools like PagerDuty handle escalation and why that matters for wake-someone alerts.
- Write an SLO and an alert condition based on being at risk of breaching it.
Official documentation
- AWS — CloudWatch Alarms — Threshold alarms and SNS notifications.
- Google SRE Book — Alerting on SLOs — Actionable, symptom- and SLO-based alerting.
- Prometheus — Alerting rules — Alerting on PromQL conditions.
Next: running an incident, and the report afterwards.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship