RizTech Academy logo
RizTech Academy
ObservabilityLesson 2 of 535 min

Centralised logging and finding things in it

When something breaks across many containers on many nodes, you cannot SSH to each and read log files — especially with pods that are created and destroyed constantly. You need logs centralised: shipped to one place, searchable, from every container. This lesson is centralised logging on AWS — how logs get to CloudWatch Logs, how to search them, and the practices (from the Foundation's good-logging lesson) that make them useful at scale. It builds on the Foundation's observability module.

Why centralised logging

The good-logging lesson established what to log (structured, with context, at the right level). This lesson is about where the logs go. In a containerised AWS system:

  • Pods are ephemeral. A pod that crashed and was replaced is gone, and its local logs with it — but you need those logs to find out why it crashed (the Foundation's kubectl logs --previous only reaches so far).
  • There are many sources. Dozens of pods across many nodes each produce logs; you need them in one searchable place, not scattered.
  • You need to search across them. "Show me every payment_failed error across all pods in the last hour" is impossible if logs live on individual containers.

So logs are shipped off the containers to a central store as they are produced, where they persist (surviving the pod) and are searchable across the whole system. On AWS the common central store is CloudWatch Logs.

How logs reach CloudWatch on EKS

A containerised app logs to stdout (the containers module), and a log collector running on each node ships those logs to CloudWatch Logs:

  • Fluent Bit (or Fluentd) is the standard collector — it runs as a DaemonSet (one per node), reads the logs of every container on its node, and forwards them to CloudWatch Logs (or another destination like OpenSearch). AWS provides Fluent Bit as an EKS add-on / container-insights setup.
  • Logs land in CloudWatch Logs log groups (typically one per application or namespace), where they are stored, searchable, and retained.
  • The collector can enrich logs with Kubernetes metadata (pod name, namespace, labels), so you can filter by which pod or app a log came from — essential for finding the right logs among many.

So the flow is: app → stdout → Fluent Bit (per node) → CloudWatch Logs (central, searchable, enriched with pod metadata). The pod can die; its logs live on in CloudWatch. (Alternatives exist — some teams ship to the OpenSearch/ELK stack for richer search, or to a third-party like Datadog — but CloudWatch Logs is the native default and the pattern is identical: a per-node collector ships stdout to a central store.)

Finding things: searching logs at scale

Central logs are only useful if you can find the relevant lines among millions, which is exactly why the good-logging lesson pushed structured (JSON) logs: a log system can query structured fields precisely.

  • CloudWatch Logs Insights is the query tool — a purpose-built query language over your logs. With structured logs, you can write queries like "all logs where event = payment_failed and userId = 42 in the last hour", or "count errors by reason", or "the p99 of a duration field" — filtering and aggregating on the JSON fields your app logs.
  • Filter by context — because the collector added pod/namespace metadata and your app added a correlation/ request ID (the good-logging lesson), you can narrow to one pod, one app, or one request across services — turning scattered lines into a coherent story.
  • Metric filters — CloudWatch can turn a log pattern into a metric (e.g. count of ERROR lines), so a log condition can drive an alarm (the alerting lesson) — bridging logs and metrics.

The payoff of structured logging is realised here: at scale, you do not grep files, you query fields — and the difference between finding a production problem in two minutes and an hour is whether your logs are structured and searchable. Log in JSON, ship to a central store, query with Insights.

Practices for logging at scale (and cost)

A few AWS-specific practices, tying to the Foundation and to cost:

  • Set log retention. CloudWatch Logs retains logs indefinitely by default, which costs money and grows forever (the storage-piles-up theme again). Set a retention policy per log group — keep logs as long as you need (say 30-90 days for app logs, longer for audit), not forever. This is a standard, easy cost saving.
  • Do not log secrets or PII (the good-logging lesson) — central logs are widely accessible and long-lived, so a secret in a log is a serious, persistent leak. Redact.
  • Control log volume. Excessive logging (debug logs in production, logging every request in full) generates huge CloudWatch ingestion and storage costs — log meaningful events at appropriate levels, and use log levels to keep production volume manageable.
  • Consider where to search. For very high volumes or complex search needs, teams sometimes ship to OpenSearch or a specialised platform; for most, CloudWatch Logs Insights is sufficient and native. Match the tool to the scale.

The complete picture: structured logs to stdout, shipped by a per-node collector (Fluent Bit) to CloudWatch Logs, enriched with pod metadata, searchable with Insights on their fields, with retention set to control cost and no secrets ever logged. That is centralised logging on AWS — the good-logging discipline made real at cluster scale.

Check your work

Why centralise: pods are ephemeral (crashed pod's logs gone — but you need them), many sources (dozens of pods/nodes), and you must search across them. So ship logs off the containers to one searchable, persistent store (CloudWatch Logs).

How on EKS: app → stdout → Fluent Bit (DaemonSet, one per node, reads every container's logs) → CloudWatch Logs log groups, enriched with pod/namespace metadata. Pod dies; logs live on. (Alternatives: OpenSearch/ELK, Datadog — same per-node-collector pattern.)

Finding things: CloudWatch Logs Insights queries structured (JSON) logs by field ("event=payment_failed and userId=42", "count errors by reason"); filter by pod metadata + correlation/request ID (one request across services); metric filters turn a log pattern into a metric for alarms. Structured logging pays off here — query fields, don't grep.

Practices (+cost): set retention per log group (default is forever = cost); no secrets/PII (widely accessible, long-lived); control volume (levels; excessive logging = ingestion/storage cost); match search tool to scale.

Practice

  1. Explain why centralised logging is necessary in a containerised system with ephemeral pods.
  2. Describe how logs get from an app's stdout to CloudWatch Logs on EKS, naming the collector and its shape.
  3. Explain how CloudWatch Logs Insights lets you find specific logs, and why structured logs make it powerful.
  4. Explain how pod metadata and a correlation ID let you narrow to one request across services.
  5. Explain why setting log retention matters for cost, and give a sensible retention approach.
  6. Explain the risks of logging secrets/PII to a central store and of excessive log volume.

Official documentation

Next: Grafana dashboards people actually read.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship