RizTech Academy logo
RizTech Academy
ObservabilityLesson 3 of 430 min

Monitoring and alerting basics

Logs and metrics are only useful if someone looks at them — and nobody can watch dashboards all day. Monitoring and alerting close the loop: dashboards let you see the system's health at a glance, and alerts tell you when something needs attention, so a problem reaches a human before it reaches (too many) users. Done well, this is what lets a small team run a reliable service; done badly, it is a firehose of noise everyone ignores. This lesson is monitoring and alerting basics.

Monitoring: seeing the system's health

Monitoring is continuously collecting metrics (the pillars lesson) and presenting them so you can see how the system is doing — usually on dashboards of graphs: request rate, error rate, latency, resource usage over time. A good dashboard answers "is the system healthy right now, and how has it been trending?" at a glance.

What to put on a dashboard is guided by the four golden signals (from Google's SRE practice) — the handful of metrics that best capture the health of almost any service:

  • Latency — how long requests take (track p50, p95, p99 — the tail matters; p99 latency reveals the slow requests an average hides).
  • Traffic — how much demand there is (requests per second).
  • Errors — the rate of failing requests (the 5xx rate from the HTTP lesson).
  • Saturation — how full the system is (CPU, memory, disk, queue depth) — how close to its limits.

If you track just these four for each service, you have a strong picture of its health. They are a far better starting point than a hundred random metrics, because they map directly to what users feel (slow, failing, overloaded) — start with the golden signals and add more only when a real question needs them.

Alerting: being told when something is wrong

Dashboards are for looking; alerts are for not having to look. An alert is a rule — "if the error rate exceeds 5% for 5 minutes, notify someone" — that watches the metrics and fires a notification (to chat, email, or a pager) when a condition is met. Alerting is what lets you not stare at dashboards: the system watches itself and calls you when it needs you.

The purpose shapes what a good alert is: it should mean "a human needs to act, now." Which leads to the single most important principle of alerting.

Alert on symptoms, and only on what needs action

Two rules separate useful alerting from noise:

  • Alert on symptoms, not causes. Alert on what users experience — high error rate, high latency, the service being down — not on every internal metric. "CPU is at 80%" is not necessarily a problem (maybe the service is handling load fine); "error rate is 10% and rising" definitely is. Alerting on user-facing symptoms catches real problems (however caused) without firing on internal fluctuations that do not matter. You investigate causes (with dashboards and logs) once a symptom alert fires.
  • Every alert must be actionable. An alert that fires when nothing needs doing trains people to ignore alerts — alert fatigue — and then the one alert that matters gets ignored too. This is exactly the flaky-test problem from the QA course: a signal that cries wolf becomes noise, and a noisy alert is worse than no alert. So every alert should require a human to do something; if an alert routinely fires and the response is "ignore it", delete or fix that alert. A small set of trustworthy, actionable alerts beats a hundred noisy ones.

The discipline: few alerts, all actionable, on user-facing symptoms. That gives you a pager that is quiet until it genuinely needs you — which is the only kind anyone actually responds to.

SLIs, SLOs and error budgets (briefly)

A more mature framing you will hear, worth recognising:

  • An SLI (Service Level Indicator) is a metric that measures user experience — e.g. the percentage of requests served successfully in under 300ms.
  • An SLO (Service Level Objective) is a target for it — e.g. "99.9% of requests succeed under 300ms over 30 days." It defines what "healthy enough" means, concretely.
  • The gap between 100% and your SLO is an error budget — the amount of failure you can tolerate. If you are within budget, you can ship features; if you are burning it, you focus on reliability. This turns "how reliable should we be?" from a vague argument into a number, and you alert when the SLO is at risk.

You do not need SLOs on day one, but the idea — define a concrete reliability target from the user's perspective, measure it, and alert on it — is where good monitoring leads, and it keeps you honest about reliability instead of chasing 100% (which is impossibly expensive) or having no target at all.

The tools

The common stack (the pillars lesson named these): Prometheus collects metrics, Grafana builds dashboards and can alert, Alertmanager routes alerts, and pager tools (PagerDuty, Opsgenie) handle on-call escalation. Clouds have CloudWatch; hosted platforms (Datadog, Grafana Cloud) combine it all. The mechanics vary; the concepts — dashboards of golden signals, actionable symptom-based alerts, SLOs — carry across every tool.

Check your work

Monitoring = continuously collecting metrics and showing them on dashboards so you can see health at a glance. Focus on the four golden signals: latency (p50/p95/p99 — the tail), traffic (req/s), errors (5xx rate), saturation (CPU/mem/disk/queue). They map to what users feel; start there, not a hundred metrics.

Alerting = rules that watch metrics and notify when a condition is met, so you don't have to watch dashboards. A good alert means "a human must act, now."

Two rules: alert on symptoms (user-facing: error rate, latency, down) not causes (80% CPU may be fine) — investigate causes after; and every alert must be actionable — noisy alerts cause alert fatigue (the flaky-test problem: cry wolf → ignored → the real one missed). Few alerts, all actionable, on symptoms.

SLI/SLO/error budget: SLI = a user-experience metric; SLO = a target for it (99.9% under 300ms/30d); error budget = allowed failure (within → ship; burning → fix reliability). Define a concrete target and alert on it; avoids chasing 100% or having no target.

Tools: Prometheus (metrics), Grafana (dashboards/alerts), Alertmanager + PagerDuty/Opsgenie (routing/on-call); CloudWatch/Datadog. Concepts carry across.

Practice

  1. List the four golden signals and give a concrete metric and why it matters for each.
  2. Explain the difference between a dashboard and an alert, and what each is for.
  3. Explain "alert on symptoms, not causes" with an example of a good symptom alert and a bad cause alert.
  4. Explain alert fatigue, how it mirrors the flaky-test problem, and why a noisy alert is worse than none.
  5. Write an SLI and an SLO for a web service, and explain what its error budget would mean.
  6. For the greetings service, choose three alerts you would set and justify that each is actionable.

Official documentation

Next: what to do when something breaks.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship