RizTech Academy logo
RizTech Academy
Networking EssentialsLesson 5 of 530 min

Debugging with curl, dig and traceroute

"The site is down" is the start of a diagnosis, not a conclusion. The value of the last four lessons is that they give you a chain to check — DNS, connection, TLS, HTTP — and the command-line tools to probe each link until you find the break. A DevOps person who can methodically locate a network problem is worth a great deal; one who guesses and restarts things at random is not. This lesson ties the module together into a debugging method, with the tools you use at each step. It closes the networking module.

Debug the chain in order

Recall the journey a request takes (the first lesson): name → DNS → connect to IP:port → TLS → HTTP → response. When something is broken, the failure is at one of those links, so you check them in order, and the first one that fails tells you where the problem is. Guessing wastes time; walking the chain finds the break every time. The rest of this lesson is the tool for each link.

Is it DNS? — dig and nslookup

First: does the name resolve to the address you expect?

dig riztechacademy.com +short      # what IP does it resolve to right now?
dig @8.8.8.8 riztechacademy.com +short   # ask a public resolver — has the change propagated?

If dig returns nothing, or the wrong IP, the problem is DNS (a missing record, a typo, or a change that has not propagated — the DNS lesson). No point debugging the server if the name does not even point at it. A huge share of "it's down" turns out to be DNS: the record was wrong, or a recent change had not propagated. Rule out DNS first.

Can I reach the machine and port? — ping, nc, telnet

If DNS is right, can you actually reach the machine, and is the service listening on its port?

ping 203.0.113.10           # is the host reachable at all? (some hosts block ping, so absence isn't proof)
nc -zv 203.0.113.10 443     # is something listening on port 443? (netcat: connect test)
  • ping tests basic reachability of the host. If it fails and the host does not block ping, the machine may be down or unreachable. (Many servers block ping, so a failed ping is not conclusive — do not over-read it.)
  • nc -zv host port (netcat) tests whether a port is open and something is listening. "Connection refused" means the machine is up but nothing is listening on that port — the service is down or on a different port (the ports lesson). "Timed out" often means a firewall is dropping the traffic. This distinction — refused versus timeout — is genuinely useful: refused points at the service, timeout points at a firewall/network.

Is TLS okay? — openssl and curl -v

If the port is open but HTTPS fails, check the certificate (the TLS lesson):

curl -vI https://riztechacademy.com    # shows the TLS handshake and any cert error, plus the status
openssl s_client -connect riztechacademy.com:443 -servername riztechacademy.com </dev/null 2>/dev/null | openssl x509 -noout -dates

This tells you if the certificate is expired, for the wrong name, or otherwise rejected — the difference between "the app is broken" and "the cert lapsed".

What is the app saying? — curl

If DNS, connection and TLS are fine, the problem is at the HTTP/application layer, and curl reads it directly (the HTTP lesson):

curl -i https://riztechacademy.com/health   # status line, headers and body
curl -s -o /dev/null -w "%{http_code}\n" https://riztechacademy.com   # just the status code

Now the status code tells the story: a 5xx means the app or something behind it is failing (go to the app logs — the observability module); a 4xx means the request itself is being rejected; a 200 means the endpoint is actually fine and the problem is elsewhere (maybe a specific page, maybe the client). curl gives you the server's real answer, unmediated by a browser or a UI.

Where does the traffic go? — traceroute and mtr

Occasionally the problem is between you and the server — a routing issue somewhere on the path. traceroute shows the hops:

traceroute riztechacademy.com   # each network hop between you and the destination
mtr riztechacademy.com          # traceroute + ping, live — better for spotting loss/latency

You will reach for these rarely, but when a service is reachable from some places and not others, or is slow in a way that is not the app, a traceroute showing where packets stop or slow down is what tells you the problem is the network path, not your server.

The method, in one place

Put it together — this is the routine to run when something is "down":

  1. DNS — dig +short: does the name resolve to the right IP? (and has a recent change propagated?)
  2. Reachability/port — nc -zv host port: is the machine up and something listening? (refused = service; timeout = firewall)
  3. TLS — curl -vI / openssl: is the certificate valid and for the right name?
  4. HTTP/app — curl -i: what status and body does the app actually return? (5xx → app logs)
  5. Path (rarely) — traceroute/mtr: is the network route itself the problem?

The first check that fails is your answer. This ordered method is the difference between "I restarted things until it worked, I don't know why" and "DNS was pointing at the old server; I fixed the A record" — and the second is what makes you trusted with production.

Check your work

Debug the chain in order — name → DNS → connect IP:port → TLS → HTTP; the first link that fails is the problem. Walk it; don't guess.

DNS: dig +short (right IP? propagated? @8.8.8.8). A huge share of "down" is DNS — rule it out first.

Reachability/port: ping (host reachable — but many block ping, so not conclusive), nc -zv host port (is a service listening?). Connection refused = service down/wrong port; timeout = firewall.

TLS: curl -vI / openssl s_client (expired? wrong name?) — distinguishes "app broken" from "cert lapsed".

HTTP/app: curl -i / status-only curl. 5xx → app failing (go to logs); 4xx → request rejected; 200 → endpoint fine, look elsewhere.

Path (rare): traceroute/mtr — where packets stop/slow, when the network route itself is at fault.

The routine: DNS → port → TLS → HTTP → path; first failure is the answer. Turns "restarted randomly" into "found and fixed the actual cause".

Practice

  1. Write out the five-step chain you check, in order, and the tool for each step.
  2. Use dig +short and dig @8.8.8.8 on a domain and explain what a mismatch would mean.
  3. Use nc -zv host 443 on a server; explain what "connection refused" vs "timed out" each imply.
  4. Use curl -vI on an HTTPS site and locate the TLS handshake and the status code in the output.
  5. Use a status-only curl to get just the HTTP code from an endpoint, and say where a 5xx sends you next.
  6. Given "the site is down for some users but not others", describe which tools you would use and in what order.

Official documentation

Next: the Git and Collaboration module.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship