The N+1 problem and how to see it
The N+1 problem is the most common performance bug in JPA applications, and — like Django's version — it is
caused by the very convenience that makes the ORM pleasant: accessing a lazy relationship
(parcel.getHub()) looks like a field access but is secretly a database query. This lesson shows the bug
with real query counts, then fixes it, and — most importantly — teaches you to see it, which is the whole
"understand the SQL" theme of this course made concrete.
The bug, with real numbers
You list parcels and, for each, show its hub's name — an utterly ordinary requirement:
List<Parcel> parcels = parcelRepository.findAll(); // query 1: SELECT * FROM parcel
for (Parcel p : parcels) {
System.out.println(p.getHub().getName()); // each getHub() → another SELECT
}
This looks like one query. It is not. Because hub is LAZY, p.getHub() triggers a SELECT the first
time it is touched — once per parcel. So for N parcels you run 1 + N queries: one for the list, plus
one per parcel for its hub. Verified against a running Dakiya with Hibernate's query statistics: 5 parcels
with distinct hubs ran 6 queries (1 + 5). With 5,000 parcels that is 5,001 database round-trips for one
list page — the page crawls.
The bug is insidious because the query is hidden behind getHub() — there is no visible SQL in the loop.
You cannot spot N+1 by reading for a query call; you spot it by knowing that touching a lazy relationship
inside a loop is where it lives.
(A subtlety verified along the way: with only 2 distinct hubs shared among the 5 parcels, the naive count was 3, not 6 — because Hibernate's first-level cache reuses a hub already loaded in the same transaction. So N+1 is really "1 + the number of distinct related rows not already loaded". The persistence-context lesson explains that cache; the danger is unchanged — it grows with your data.)
The fix: fetch the relationship in one query
The fix is to load the parcels and their hubs together, in a single query, with a JPQL fetch join:
@Query("select p from Parcel p join fetch p.hub")
List<Parcel> findAllWithHub();
join fetch p.hub tells Hibernate to JOIN the hub into the same query and populate each parcel's hub
eagerly — so accessing p.getHub() afterwards needs no further query. Verified: the same loop over
findAllWithHub() ran 1 query instead of 6. That is the N+1 fix for a @ManyToOne: a fetch join,
loading the related data up front, in one round trip.
The other tools for the same job:
@EntityGraph— a declarative alternative to a fetch join, put on a repository method:@EntityGraph(attributePaths = "hub") List<Parcel> findAll(); // loads hub eagerly for this call, one query- A batch size (
@BatchSize, orspring.jpa.properties.hibernate.default_batch_fetch_size) — turns N separate hub queries into a fewIN (...)queries, mitigating N+1 for collections where a fetch join is awkward.
For a @ManyToOne like parcel→hub, a fetch join (or @EntityGraph) is the clean answer. Do not solve
it by making the relationship EAGER — that reintroduces the "loads the hub on every parcel query, even
when you do not need it" problem the relationships lesson warned about. Keep the relationship LAZY, and
fetch it eagerly only on the specific query that needs it.
Seeing it: this is the point of the course
You will not always reason out N+1 in advance, so you must be able to measure it. Three ways, in order of everyday usefulness:
- Log the SQL.
spring.jpa.show-sql=trueprints every statement Hibernate runs. If a single logical operation prints a wall of near-identicalSELECTs, that is N+1, visible. - Count queries with Hibernate statistics.
spring.jpa.properties.hibernate.generate_statistics=truemakes Hibernate report the query count (this is exactly how the 6-vs-1 numbers here were measured). In a test you can assert the count. - Assert the count in a test. A test that fails if an operation runs more than k queries is the regression guard — someone reintroducing N+1 breaks the test, not production.
The habit that separates a competent Spring developer from a beginner is the same as in every ORM: you know how many queries your endpoint runs, and why. Not a guess — a number you have looked at, in the SQL log or the statistics. An engineer who never looks at the generated SQL ships endpoints that work in a demo with ten rows and collapse on production data — exactly the "adds an annotation and hopes" developer this course exists to prevent.
The over-correction to avoid
The opposite mistake is real too: fetch-joining or @EntityGraph-ing relationships you do not use, or making
everything EAGER "to be safe". That loads data the request never touches — wasted queries and memory,
sometimes cartesian-product explosions when you fetch-join multiple collections at once. The goal is not
"always eager-load"; it is fetch exactly what this operation uses, in as few queries as sensible. LAZY by
default, a fetch join where you have seen (in the SQL log) that a relationship is accessed in a loop —
measure, then fix, not fix on a hunch.
Check your work
What N+1 is. One query for a list plus one per row for its lazy relationship — because getHub() (a
lazy access) is a hidden query. Verified: 5 parcels → 6 queries (1 + 5).
Why it hides. The per-row query is behind an innocent-looking getHub(); you find it by knowing that a
lazy relationship touched in a loop is where it lives — not by reading for query calls.
The first-level-cache nuance. N+1 is "1 + distinct related rows not already loaded" (verified 3 with 2 shared hubs), because Hibernate reuses loaded entities in the transaction — the danger still grows with data.
The fix. A JPQL fetch join (join fetch p.hub) or @EntityGraph loads the relationship in one
query. Verified: 6 → 1. Batch fetching mitigates collection cases. Do not switch to EAGER — keep LAZY,
fetch on the query that needs it.
Seeing it. show-sql (log the SQL), generate_statistics (count queries — how the numbers here were
measured), and an assertNumQueries-style test as a regression guard. Know your endpoint's query count.
The over-correction. Do not eager-load everything — fetch what the operation uses, in as few queries as sensible; measure, then fix.
Practice
- Create several parcels each with a distinct hub; loop, touch each
getHub().getName(), and withgenerate_statistics=trueconfirm the count is 1 + N (reproduce 6 for 5 parcels). - Add a
@Query("... join fetch p.hub")method and confirm the same loop now runs 1 query. - Try
@EntityGraph(attributePaths="hub")on a finder and confirm it also collapses to one query. - Repeat step 1 but with only 2 shared hubs; observe the count is 3, not 6, and explain it via the first-level cache.
- "Fix" the N+1 by making
@ManyToOneEAGER instead; then load parcels in a context that does not need the hub and observe the wasted hub query — reason about why LAZY + fetch-join is better. - Write a test asserting an endpoint runs at most one query for the list, and watch it fail if you remove the fetch join.
Official documentation
- Hibernate — Fetching (join fetch, batch) — Fetch strategies and N+1.
- Spring Data JPA — @EntityGraph — Declarative eager fetching per query.
- Spring Boot — SQL logging and statistics —
show-sqland Hibernate statistics.
Next: transactions and @Transactional.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship