Metrics with Prometheus and CloudWatch
The DevOps Foundation taught the three pillars — logs, metrics, traces — and the golden signals. This module applies them to AWS: how you actually collect metrics, centralise logs, build dashboards, and alert, using AWS's tools and the Prometheus/Grafana stack. This first lesson is metrics on AWS: CloudWatch, Prometheus, and how they fit together on EKS. It builds on the Foundation's observability module.
CloudWatch: AWS's built-in metrics
Amazon CloudWatch is AWS's native monitoring service, and much of it works with no setup: every AWS service publishes metrics to CloudWatch automatically. An EC2 instance publishes CPU and network metrics; an ALB publishes request counts, latency and HTTP error codes; RDS publishes database metrics; EKS publishes control-plane metrics. So out of the box, CloudWatch gives you the infrastructure-level golden signals (the Foundation) for your AWS resources — latency and errors from the load balancer, saturation from the instances — without instrumenting anything.
- AWS-service metrics are there automatically — a strong baseline.
- Custom metrics — your application can publish its own metrics to CloudWatch (business metrics, custom measurements) via the API, though this has a per-metric cost to watch.
- CloudWatch metrics power alarms (the alerting lesson) and dashboards (the dashboards lesson), and are retained over time so you can see trends.
For infrastructure and AWS-service monitoring, CloudWatch is the path of least resistance: it is already collecting the metrics; you just use them.
Prometheus: the Kubernetes-native metrics standard
For application and Kubernetes metrics, the standard is Prometheus (the Foundation named it). Prometheus
scrapes metrics from endpoints your applications and Kubernetes expose (a /metrics endpoint in the
Prometheus format), storing time-series data you query with PromQL. It is the de facto standard for
container and Kubernetes metrics because:
- Kubernetes and its components expose Prometheus metrics natively (the API server, kubelet, and via exporters, nodes and pods), so Prometheus gives you deep cluster and pod-level visibility CloudWatch does not.
- Applications instrument with Prometheus client libraries — you add a
/metricsendpoint exposing your app's metrics (request counts, latencies, queue depths — the golden signals at the app level), and Prometheus scrapes it. - PromQL is a powerful query language for aggregating and analysing these metrics (rate of errors, p99 latency, etc.).
So on EKS, Prometheus is how you get rich application and Kubernetes metrics, complementing CloudWatch's infrastructure and AWS-service metrics.
Managed Prometheus and Grafana on AWS
Running Prometheus yourself (and its long-term storage) is operational work. AWS offers managed versions, which most teams prefer:
- Amazon Managed Service for Prometheus (AMP) — a managed, scalable Prometheus-compatible backend. You scrape metrics (or use the managed collector) and they are stored in AMP, so you get Prometheus without running and scaling the servers and storage yourself.
- Amazon Managed Grafana (AMG) — managed Grafana (the dashboards lesson) for visualising metrics from AMP, CloudWatch, and other sources in one place.
The common EKS observability stack is therefore: Prometheus (or AMP) for application/Kubernetes metrics, CloudWatch for AWS-service/infrastructure metrics, and Grafana (or AMG) to visualise both together. Grafana querying both Prometheus and CloudWatch gives you one pane of glass over the whole system — cluster, app, and AWS services.
What to collect: the golden signals, at every layer
The Foundation's guidance holds — focus on the four golden signals (latency, traffic, errors, saturation) — and on AWS you gather them at multiple layers:
- At the load balancer (CloudWatch): request count (traffic), target response time (latency), HTTP 5xx count (errors), and target health — the user-facing golden signals, automatically.
- At the application (Prometheus): your app's request rate, latency (p50/p95/p99), error rate, and any
business metrics — instrumented via a
/metricsendpoint. - At the cluster/nodes (Prometheus/CloudWatch): CPU, memory, pod counts, node saturation — the resource golden signal, tying to the requests/limits and scaling lessons.
Start with these signals at these layers rather than collecting hundreds of metrics for their own sake (the Foundation's warning against vanity metrics). The load balancer's latency and 5xx rate, the app's error rate and p99, and node saturation give you a strong, actionable picture of health — and they feed the dashboards and alerts the rest of this module builds.
Check your work
CloudWatch = AWS's native metrics; every AWS service publishes automatically (EC2 CPU, ALB request/ latency/5xx, RDS, EKS control plane) — infrastructure and AWS-service golden signals with no setup. Custom app metrics can be published (per-metric cost). Powers alarms and dashboards.
Prometheus = the standard for app + Kubernetes metrics: scrapes /metrics endpoints, stores
time-series, queried with PromQL. Kubernetes/components expose Prometheus metrics natively; apps instrument
with client libraries. Deeper cluster/pod visibility than CloudWatch.
Managed on AWS: AMP (managed Prometheus backend — no servers/storage to run) and AMG (managed Grafana). Common stack: Prometheus/AMP (app+K8s) + CloudWatch (AWS/infra) + Grafana/AMG (visualise both) = one pane of glass.
What to collect: the four golden signals at each layer — LB (CloudWatch: traffic/latency/5xx), app (Prometheus: rate/latency p99/errors), cluster/nodes (saturation). Start with these, not hundreds of vanity metrics.
Practice
- Explain what CloudWatch gives you automatically and why it is the baseline for AWS-service metrics.
- Explain what Prometheus is, how it collects metrics, and why it is the standard for Kubernetes/app metrics.
- Explain how CloudWatch and Prometheus complement each other on EKS.
- Explain what AMP and AMG are and why teams prefer managed versions.
- List the golden signals and where on AWS you would collect each (LB, app, cluster).
- Explain why you focus on the golden signals rather than collecting every possible metric.
Official documentation
- AWS — Amazon CloudWatch — Native AWS metrics, alarms and dashboards.
- AWS — Managed Service for Prometheus — Managed Prometheus metrics.
- Prometheus — Overview — Scraping, storage and PromQL.
Next: centralised logging and finding things in it.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship