Reading and writing text files
Everything your programs have produced so far vanished when they stopped. Files are how work survives — and reading and writing them is most of what a great deal of real software does.
Reading a whole file
with open("notes.txt") as file:
contents = file.read()
print(contents)
open() gives you a file object. .read() returns the entire contents as one
string. The with closes it for you when the block ends — the next lesson
explains exactly what that does and why it matters. For now, always use
with; it is not optional style.
Note the indentation. The file is open inside the block and closed outside it, so reading must happen within.
Reading line by line
Loading a 2 GB log file into a string is a bad afternoon. Loop over the file instead:
with open("notes.txt") as file:
for line in file:
print(line.strip())
That reads one line at a time and never holds more than one in memory. It works identically on a four-line file and a four-million-line one, which is why it is the default way to read.
.strip() is there because each line keeps its newline character:
with open("notes.txt") as file:
for line in file:
print(repr(line))
'First line\n'
'Second line\n'
'Third line'
That \n causes the classic "why is there a blank line between everything"
problem, and it is why comparisons fail on lines that look correct. Strip unless
you specifically want it.
The last line often has no \n, which is a difference worth remembering when
processing files.
If you genuinely want every line at once:
with open("notes.txt") as file:
lines = file.readlines()
A list of strings, newlines included. Fine for small files.
Writing
with open("output.txt", "w") as file:
file.write("First line\n")
file.write("Second line\n")
Two things that catch people.
"w" destroys the file's existing contents immediately, the moment it is
opened — before you write anything. Opening a file for writing and then hitting
an error still leaves you with an empty file. There is no undo.
write() does not add a newline. Unlike print, it writes exactly what you
give it. Without the \n above, both strings land on one line.
To add rather than replace, use "a" for append:
with open("log.txt", "a") as file:
file.write("Another entry\n")
Appending creates the file if it does not exist, and adds to the end if it does. This is what you want for logs.
Writing many lines:
lines = ["first", "second", "third"]
with open("output.txt", "w") as file:
file.write("\n".join(lines))
"\n".join(lines) from the strings lesson is cleaner than a loop, and it avoids
a trailing blank line. There is also file.writelines(lines), which — despite
the name — adds no newlines at all, so you would have to include them yourself.
The modes
| Mode | Does | If the file does not exist |
|---|---|---|
"r" |
read (the default) | FileNotFoundError |
"w" |
write, replacing everything | creates it |
"a" |
append to the end | creates it |
"x" |
write, but only if new | creates it; errors if it exists |
"r+" |
read and write | FileNotFoundError |
"x" is worth knowing. When you must not overwrite something, it refuses at the
door rather than relying on you checking first.
Add "b" for binary — "rb", "wb" — when handling images, PDFs or anything
that is not text. You get bytes rather than str.
Always specify the encoding
with open("notes.txt", encoding="utf-8") as file:
Without it, Python uses the system default, which differs between machines. Code
that works on your Linux laptop then fails on a colleague's Windows machine with
UnicodeDecodeError the moment a file contains a rupee sign, an accent, or a
name in Devanagari.
Write encoding="utf-8" every time you open a text file. It costs nothing
and prevents a genuinely miserable class of bug. UTF-8 is what essentially
everything uses.
Handling the file not being there
try:
with open("config.txt", encoding="utf-8") as file:
config = file.read()
except FileNotFoundError:
print("No config file found, using defaults.")
config = ""
FileNotFoundError is the one you will meet most. Others worth knowing:
PermissionError— the file exists, you are not allowedIsADirectoryError— you gave it a folderUnicodeDecodeError— wrong encoding, or it is not a text file at all
Note the try wraps the with. The file still closes correctly — with
handles that regardless of how the block exits.
A worked example
Counting word frequency across a file, using the dictionary pattern from module 4:
counts = {}
with open("article.txt", encoding="utf-8") as file:
for line in file:
for word in line.lower().split():
word = word.strip(".,!?\"'")
if word:
counts[word] = counts.get(word, 0) + 1
top = sorted(counts.items(), key=lambda pair: pair[1], reverse=True)
with open("frequencies.txt", "w", encoding="utf-8") as output:
for word, count in top[:20]:
output.write(f"{count:>5} {word}\n")
That reads a file of any size, counts, sorts, and writes a report — using the
dictionary counting pattern, sorted with a lambda key, f-string alignment
and line-by-line reading. Everything from the last four modules, doing
something real.
Note .strip(".,!?\"'") — strip with an argument removes any of those
characters from both ends, not just whitespace.
Reading and writing the same file
Do not open a file for reading and writing simultaneously and expect it to go well. The safe pattern is read, transform, write:
with open("data.txt", encoding="utf-8") as file:
lines = file.readlines()
cleaned = [line.strip().title() + "\n" for line in lines if line.strip()]
with open("data.txt", "w", encoding="utf-8") as file:
file.writelines(cleaned)
Read fully, close, then reopen for writing. Safer still is writing to a new file and replacing the original only once it has succeeded — so a crash halfway through does not leave you with half a file.
Check your work
Each line keeps its newline.
'First line\n'
'Second line\n'
'Third line'
The last line usually has none. That \n is the "why is there a blank line
between everything" problem, and it is why comparisons fail on lines that look
right.
"w" run twice does not grow the file — it replaces the contents every
time, the moment the file is opened. "a" appends, and creates the file if it
is missing.
The FileNotFoundError, handled.
try:
with open("missing.txt", encoding="utf-8") as file:
contents = file.read()
except FileNotFoundError:
contents = ""
Line numbers.
with open("in.txt", encoding="utf-8") as source, \
open("out.txt", "w", encoding="utf-8") as target:
for number, line in enumerate(source, start=1):
target.write(f"{number:>4} {line.rstrip()}\n")
"x" twice gives FileExistsError on the second run — which is the point
of it. When you must not overwrite something, "x" refuses at the door rather
than relying on you checking first.
The encoding. Whether omitting it breaks depends on your system's default.
If it worked for you, note that it would fail on a colleague's machine with a
different default — which is exactly the bug encoding="utf-8" prevents, and
why it should be written every time.
Practice
- Create
notes.txtby hand with three lines. Read and print it whole. - Read it line by line, printing each with
repr(). Find the newlines. - Strip them and print cleanly.
- Write a new file with five lines using
"w". Run it twice and confirm it does not grow. - Change it to
"a"and run twice. Note the difference. - Open a file that does not exist. Read the
FileNotFoundError, then handle it withtry/except. - Write a program that reads a file and writes a copy with line numbers added.
- Write the word-frequency example against any text file you have.
- Use
"x"mode twice in a row and read the error on the second run. - Save a line containing
₹without specifying an encoding, then withencoding="utf-8". On some systems the first fails — if it does not, note that it would on a colleague's.
Next: what with is actually doing, and why leaving it out is worse than it
looks.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship