RizTech Academy logo
RizTech Academy
ObservabilityLesson 1 of 425 min

Logs, metrics and traces

You have deployed your app. It is running in production, serving real users — and now something is slow, or throwing errors, or just behaving oddly, and you need to know what is actually happening inside it. Observability is how you find out: the practice of instrumenting your systems so you can understand their behaviour from the outside. Its three pillars are logs, metrics and traces, and knowing what each is for is the foundation of operating anything in production. This module is observability, and this first lesson is the three pillars.

Why observability matters

In development you can attach a debugger and step through code. In production you cannot — the system is live, distributed across many containers and machines, serving thousands of users. You cannot pause it. So you must have built in, ahead of time, the ability to see what it is doing: what it logged, how it is performing, where a request went. That is observability, and the difference it makes is stark:

  • Without it, a production problem is a guess. "The site is slow" — why? You have no data, so you restart things and hope. Outages last longer, and you often never learn the real cause.
  • With it, a problem is a question you can answer. "The site is slow" → the metrics show latency spiked at 2pm → the logs show database timeouts starting then → the traces show the slow query. You find the cause and fix it.

Observability is what turns operating a system from guesswork into diagnosis. It is not optional for anything real; it is how you keep a system healthy and fix it fast when it is not.

The three pillars

Observability rests on three kinds of data, each answering a different question:

Logs — what happened

Logs are timestamped records of discrete events: "user 42 logged in", "payment failed: insufficient funds", "GET /transactions returned 500". They are the detailed, event-by-event story of what your application did. When you need to know what happened around a specific event — the exact error, the sequence leading to a bug — logs are where you look. They are the most detailed pillar, and the first one most people use (docker logs, kubectl logs, the good-logging lesson).

Metrics — how much, how fast, how many

Metrics are numeric measurements over time: requests per second, error rate, response latency (p50, p95, p99), CPU and memory usage, queue depth. They are aggregated numbers you track continuously and graph, and they answer how the system is performing — is it healthy, is it getting slower, is the error rate climbing? Metrics are cheap to store (just numbers), so you keep lots of them over long periods, and they are what dashboards and alerts are built on (the monitoring lesson). When you ask "is the system healthy?" or "when did this start?", metrics answer.

Traces — where a request went

Traces follow a single request as it travels through a distributed system — the web app calls an API, which calls a database and a cache, maybe across several services. A trace stitches those steps together into one timeline, showing where the time went and where a failure occurred. In a microservices system, where one user request touches many services, traces answer which service is the slow/broken one — a question logs and metrics alone struggle with. Traces are the newest and most specialised pillar, and they earn their keep once a system is spread across multiple services.

Using the three together

The pillars are complementary — you use them in sequence to diagnose a real problem, and the earlier example shows the flow:

  1. Metrics tell you that something is wrong, and when. An alert fires or a dashboard shows error rate spiking at 2pm. Metrics are the smoke alarm.
  2. Traces tell you where. For a slow request, a trace shows which service or query consumed the time — narrowing a distributed system down to the culprit.
  3. Logs tell you what and why. At the identified service and time, the logs show the actual error, stack trace, or bad input — the specific cause.

So: metrics to detect and locate in time, traces to locate in the system, logs for the detail. A team with all three can go from "something's wrong" to "here's the exact cause" in minutes; a team with none is guessing for hours. You do not always need all three — a small single-service app may live on logs and a few metrics — but knowing what each is for lets you reach for the right one, and add the others as the system grows.

The tools, briefly

You will meet these, and it helps to recognise the landscape (the specifics are for later, but the names recur):

  • Logs — collected and searched with tools like the ELK/OpenSearch stack, Loki, or a cloud service (CloudWatch Logs). Containers log to stdout (the containers module) and a collector ships them somewhere searchable.
  • Metrics — Prometheus is the standard for collecting them, Grafana for dashboards; clouds have CloudWatch, and there are hosted options (Datadog, etc.).
  • Traces — OpenTelemetry is the emerging standard for instrumenting apps, with backends like Jaeger, Tempo or Datadog.

Many of these combine into one platform (Grafana's LGTM stack — Loki, Grafana, Tempo, Mimir — covers all three). You do not need to master the tools now; you need the mental model of the three pillars, so that when you meet a tool you know which pillar it serves and what question it helps answer.

Check your work

Observability = instrumenting systems so you can understand their behaviour from outside — because in production you cannot attach a debugger. Without it, problems are guesswork (restart and hope); with it, they are answerable questions (data → cause → fix). Turns operating from guessing into diagnosis.

Three pillars: Logs = timestamped records of discrete events — what happened (detailed, event-by-event; docker/kubectl logs). Metrics = numeric measurements over time (req/s, error rate, latency p50/p95/p99, CPU) — how the system is performing (cheap, long-retained, power dashboards/alerts). Traces = one request's path across services — where time went / a failure occurred (shines in microservices).

Used together: metrics detect that/when (the smoke alarm) → traces locate where (which service) → logs give what/why (the detail). All three = minutes to root cause; none = hours guessing. Small apps may use just logs + a few metrics.

Tools: logs (ELK/OpenSearch, Loki, CloudWatch), metrics (Prometheus + Grafana, CloudWatch), traces (OpenTelemetry + Jaeger/Tempo); combined stacks like Grafana LGTM. Know which pillar each serves.

Practice

  1. Define logs, metrics and traces, and give the distinct question each answers.
  2. For "the checkout page is slow", describe how you would use metrics, then traces, then logs to find the cause.
  3. Explain why you cannot debug a production system the way you debug locally, and what observability provides instead.
  4. Give an example metric for each of: throughput, errors, latency, and resource usage.
  5. Explain when traces become important, and why logs and metrics alone struggle in a microservices system.
  6. Match each tool (Prometheus, Loki, OpenTelemetry, Grafana) to the pillar it serves.

Official documentation

Next: logging that is actually useful.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship