Grafana dashboards people actually read
A dashboard is where metrics become understanding — a glance that tells you whether the system is healthy. But most dashboards fail: they are walls of graphs nobody reads, or they look impressive and answer no useful question. A dashboard people actually read is designed around the questions its viewers need answered. This lesson is building dashboards that are used, primarily with Grafana, applying the golden signals from the Foundation.
Why most dashboards fail
The common dashboard failure is a screen crammed with dozens of graphs — every metric someone could collect, arranged without thought. It looks comprehensive and is useless: nobody can tell at a glance whether things are okay, the important signals are lost among the trivial, and so people stop looking. A dashboard that is not read is worthless, however pretty (the same "trust and usefulness" theme as the QA course's suites and the Foundation's alerts).
The fix is to design a dashboard around a question and an audience: who looks at this, and what do they need to know? A good dashboard answers a specific question — "is the service healthy right now?", "how is the cluster's capacity?" — at a glance, with the important things prominent and the noise removed.
Design around the golden signals
For a service-health dashboard — the most important kind — the content is largely settled by the four golden signals (the Foundation): put them front and centre, and you have a dashboard that tells you if the service is healthy:
- Latency — request duration, showing p50, p95 and p99 (not just the average — the tail is what users feel; the Foundation). A rising p99 is an early warning.
- Traffic — requests per second, so you see load and can correlate spikes with problems.
- Errors — the error rate (5xx rate — the HTTP lesson), ideally as a percentage of requests. This is often the single most important number on the dashboard.
- Saturation — how full the system is: CPU, memory, and for EKS, pod/node capacity (the resources-and-limits and scaling lessons).
These four, clearly displayed, answer "is this service healthy?" better than fifty assorted graphs. Order them by importance (errors and latency near the top), size the important ones larger, and you have a dashboard someone can read in seconds. The RED method (Rate, Errors, Duration) and USE method (Utilisation, Saturation, Errors) are named patterns for exactly this — a small, principled set of signals rather than everything.
Grafana: one pane over CloudWatch and Prometheus
Grafana (or Amazon Managed Grafana — the metrics lesson) is the standard dashboarding tool, and its strength on AWS is combining sources: a single Grafana dashboard can query both Prometheus/AMP (app and Kubernetes metrics) and CloudWatch (AWS-service and infrastructure metrics), so you see the whole system in one place:
- App latency and error rate (Prometheus) next to load-balancer metrics and RDS metrics (CloudWatch) next to cluster CPU/memory (Prometheus) — correlated on one screen.
- Templating/variables let one dashboard serve many services or environments (a dropdown to pick the service), so you build a good service dashboard once and reuse it.
- Community dashboards exist for common things (Kubernetes cluster health, node metrics) — import a well-made one rather than building from scratch.
This one-pane-of-glass over both metric systems is why Grafana is the common choice on EKS: your metrics live in two places (Prometheus and CloudWatch), and Grafana unifies them into dashboards designed around real questions.
Dashboards that get used
A few practices that make dashboards genuinely useful (not just present):
- One dashboard per question/audience. A service-health dashboard (golden signals for one service) for on-call; a cluster-capacity dashboard for platform; a business dashboard for product. Do not mix audiences on one screen — each viewer should find what they need.
- Make "healthy vs not" obvious at a glance. Use clear thresholds and colour (green/red) so a glance answers the question; a viewer should not have to interpret raw graphs to know if something is wrong.
- Correlate on a shared timeline. Line up graphs on the same time axis so you can see that latency rose when errors spiked when CPU saturated — the cause becomes visible (the Foundation's metrics→traces→logs flow starts here).
- Keep it focused. Show what matters for the dashboard's question; resist adding every metric. A focused dashboard is read; a cluttered one is ignored.
- Dashboards support, not replace, alerts. You do not stare at dashboards (the alerting lesson); alerts tell you when to look, and the dashboard is where you look to understand. Design them to work together.
The picture: a focused dashboard per question and audience, built on the golden signals (RED/USE), in Grafana unifying Prometheus and CloudWatch, with health obvious at a glance and graphs correlated on a shared timeline. That is a dashboard people actually read — and a dashboard that is read is the difference between seeing a problem developing and being surprised by it.
Check your work
Why most fail: a wall of every-metric graphs is comprehensive-looking and useless — nobody can tell health at a glance, so it goes unread (worthless however pretty). Fix: design around a question + audience (who reads this, what do they need?).
Golden signals: a service-health dashboard = latency (p50/p95/p99 — the tail), traffic (req/s), errors (5xx rate/%), saturation (CPU/mem, pod/node capacity), ordered by importance. RED (Rate/Errors/ Duration) and USE (Utilisation/Saturation/Errors) name this focused set.
Grafana (AMG): unifies Prometheus/AMP (app/K8s) and CloudWatch (AWS/infra) on one pane; templating/ variables reuse one dashboard across services; import community dashboards. Why it's the EKS choice — metrics live in two places, Grafana unifies them.
Get used: one dashboard per question/audience (service-health / cluster / business — don't mix); health obvious at a glance (thresholds, colour); correlate on a shared timeline (latency↔errors↔CPU); keep focused; dashboards support alerts (alerts say when to look, dashboards are where you look).
Practice
- Explain why a dashboard crammed with every metric fails, and what to design around instead.
- Build (describe) a service-health dashboard from the golden signals, and say why latency needs percentiles.
- Explain the RED and USE methods as focused signal sets.
- Explain how Grafana gives one pane of glass over Prometheus and CloudWatch, and why that matters on EKS.
- Explain why you separate dashboards by audience and make health obvious at a glance.
- Explain how dashboards and alerts work together rather than one replacing the other.
Official documentation
- Grafana — Dashboards & panels — Building dashboards over multiple data sources.
- AWS — Amazon Managed Grafana — Managed Grafana over CloudWatch and Prometheus.
- Grafana — RED method / observability — The Rate/Errors/Duration approach to dashboards.
Next: alerts worth waking up for.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship