RizTech Academy logo
RizTech Academy
Linux FundamentalsLesson 5 of 735 min

grep, sed, awk and pipes

On a server, almost everything is text — logs, config files, command output, lists of things. The Unix philosophy is a set of small tools that each do one text job well, joined together with pipes to answer questions like "how many errors are in this log?" or "which IPs hit us most?". Learning to combine these turns a wall of log output into a precise answer in seconds. This lesson is text processing: grep, sed, awk, and the pipe that connects them.

The pipe: joining commands together

The idea that makes the shell powerful is the pipe, |. It takes the output of one command and feeds it as the input of the next, so you build a chain where each step transforms the text a little more:

cat access.log | grep "500" | wc -l

Read left to right: print the log, keep only lines containing "500", count them. Each tool does one thing; the pipe composes them. This is the whole game — small tools, piped together, to slice text down to what you want. Some pieces you will use in every pipeline:

wc -l file      # count lines (-w words, -c bytes)
sort            # sort lines
uniq -c         # collapse duplicate adjacent lines and count them (usually after sort)
head -n 5       # first 5 lines
cut -d',' -f2   # cut out field 2, using comma as the delimiter

grep: find lines that match

grep prints lines matching a pattern — the tool you reach for most, because "find the lines about X" is the commonest question:

grep "ERROR" app.log            # lines containing ERROR
grep -i "error" app.log         # case-insensitive
grep -v "healthcheck" app.log   # invert: lines that do NOT contain healthcheck
grep -c "ERROR" app.log         # count matching lines
grep -r "TODO" src/             # recursively search a directory tree
grep -n "ERROR" app.log         # show line numbers

grep also understands regular expressions — patterns like grep "^ERROR" (lines starting with ERROR) or grep "5[0-9][0-9]" (a 5xx status). You do not need to master regex to be useful, but knowing that grep can match patterns, not just fixed text, unlocks a lot. Filtering a log with grep is probably the single most common thing you will do on a server.

sed: find and replace in a stream

sed (stream editor) transforms text as it flows through — most often find-and-replace:

sed 's/old/new/' file.txt        # replace the FIRST 'old' on each line with 'new'
sed 's/old/new/g' file.txt       # g = global: replace ALL occurrences on each line
cat app.log | sed 's/[0-9]//g'   # delete all digits from the stream
sed -n '10,20p' file.txt         # print only lines 10 to 20

By default sed prints the transformed text to the screen without changing the file. To edit the file in place you add -i (sed -i 's/old/new/g' file) — but be careful, that changes the file for real, so test without -i first. sed is how you make a bulk change across a file or a stream without opening an editor.

awk: work with columns

awk shines when your text is in columns — which logs and command output usually are. It splits each line into fields ($1, $2, …, $0 is the whole line) and lets you pick, compute and filter:

awk '{print $1}' access.log            # print the first field (e.g. the IP) of each line
awk '{print $1, $9}' access.log        # first and ninth fields (IP and status)
awk '$9 == "500"' access.log           # only lines whose 9th field is 500
awk -F',' '{print $2}' data.csv        # use comma as the field separator (-F)
awk '{sum += $5} END {print sum}' file # sum the 5th column and print the total

That last one — accumulating a total across lines — shows awk is a small programming language, not just a filter. For everyday use, "print column N" and "keep rows where column N equals X" cover most needs, and they are enormously handy for turning structured output into exactly the fields you care about.

Putting it together: real one-liners

The power is in the combination. These are the kinds of one-liners you will actually write on a server:

# The top 5 IP addresses by number of requests in a web log
awk '{print $1}' access.log | sort | uniq -c | sort -rn | head -5

# How many 500 errors happened today
grep "$(date +%Y-%m-%d)" app.log | grep -c "500"

# Every unique error message, with a count of how often each occurred
grep "ERROR" app.log | sort | uniq -c | sort -rn

Read the first one: extract the IP (awk), sort them so identical ones are adjacent (sort), count each group (uniq -c), sort by count descending (sort -rn), take the top 5 (head). Five small tools, one precise answer. You are not expected to write these fluently yet — but recognise the pattern, because assembling pipelines like this is a defining Unix skill and it makes you fast at diagnosing real problems.

Check your work

The pipe | feeds one command's output into the next, composing small single-purpose tools into a chain. Building blocks: wc -l (count lines), sort, uniq -c (count adjacent duplicates — after sort), head, cut.

grep finds matching lines (the most-used): -i (case-insensitive), -v (invert), -c (count), -r (recursive), -n (line numbers); understands regex (^ERROR, 5[0-9][0-9]). Filtering a log with grep is the commonest server task.

sed transforms a stream — mainly find/replace: s/old/new/ (first per line), s/old/new/g (all); prints without changing the file unless -i (edit in place — test without it first).

awk works with columns: fields $1,$2,…,$0; print columns, filter rows ($9 == "500"), set the separator (-F','), and compute (sum a column). "Print column N" and "keep rows where N equals X" cover most needs.

Combine them: e.g. awk '{print $1}' | sort | uniq -c | sort -rn | head -5 = top IPs. Assembling pipelines is the defining Unix skill and makes you fast at diagnosis.

Practice

  1. Use grep -c to count lines containing "ERROR" in a log, and grep -v to exclude noisy healthcheck lines.
  2. Use sed 's/.../.../g' to replace a word throughout a file (without -i), then verify the file is unchanged.
  3. Use awk '{print $1}' to extract the first column of some columnar output, and awk '$N == "..."' to filter rows.
  4. Build the pipeline ... | sort | uniq -c | sort -rn | head to find the most frequent value in a column, and explain each stage.
  5. Write a one-liner that counts how many unique error messages appear in a log, with their frequencies.
  6. Explain the pipe in your own words and why "small tools joined together" is more flexible than one big tool.

Official documentation

Next: shell scripting basics.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship