The persistence context: the first-level cache and entity lifecycle
To use JPA well — and to debug it when it surprises you — you have to understand the persistence context:
the in-memory area, tied to a transaction, where Hibernate tracks the entities you are working with. Almost
every JPA behaviour that confuses people — why a query count is lower than expected, why a change is saved
without a save() call, why a lazy access sometimes works and sometimes throws — comes from the persistence
context. This lesson is what it is and how it behaves, verified against a running Dakiya.
What the persistence context is
The persistence context (Hibernate calls its holder the Session, JPA calls it the EntityManager) is
a workspace, scoped to a transaction, that holds the entities Hibernate is managing for you. When you load
or save an entity, it becomes managed — Hibernate keeps a reference to it in the context and watches it.
Think of it as a per-transaction cache-and-change-tracker sitting between your code and the database. Almost
everything JPA does that seems magical is the persistence context doing its two jobs: the first-level
cache and dirty checking.
The first-level cache: one instance per id per context
Within a single persistence context, Hibernate guarantees one object instance per entity id. Load the same row twice and you get the same object, from one query:
Parcel a = em.find(Parcel.class, id);
Parcel b = em.find(Parcel.class, id); // same id, same context
// a == b, and only ONE SELECT ran
Verified against the running app: loading the same id twice in one context ran 1 query and returned the
same instance (a == b was true). The first find hit the database and put the parcel in the context;
the second find found it already there and returned it without a query. This first-level cache is
automatic, always on, and scoped to the transaction (it is discarded when the context closes).
It also explains the N+1 nuance from the data-access module: iterating parcels that share only 2 hubs ran 3 queries, not 6, because once a hub was loaded into the context, the next parcel pointing at it got the cached instance — no second query. So N+1 is really "1 + distinct related rows not yet in the context". The first-level cache is why. Understanding it turns that surprising number from mysterious into predictable.
Dirty checking: changes save themselves
The second job surprises newcomers most: you do not always call save() to persist a change. A managed
entity's modifications are detected and written automatically when the context flushes:
@Transactional
public void markDelivered(Long id) {
Parcel p = repository.findById(id).orElseThrow(); // now MANAGED
p.setStatus(Status.DELIVERED); // just change the field...
// ...no save() call — Hibernate detects the change and UPDATEs at flush/commit
}
Because p is managed, Hibernate remembers its original state and, at flush (typically at transaction
commit), compares the current state to the original — dirty checking — and issues an UPDATE for what
changed. This is powerful and a common source of "why did the database change when I never called save?"
Inside a transaction, mutating a managed entity is persisting it. (Conversely, mutating a detached
entity — one from a closed context — does nothing until you re-attach it with save()/merge().)
Entity states: the lifecycle
An entity is always in one of four states, and knowing which explains its behaviour:
- Transient — a
newobject Hibernate knows nothing about (new Parcel(...)before saving). Not tracked; changes do nothing. - Managed — loaded or saved within an open context. Tracked; changes are dirty-checked and flushed.
- Detached — was managed, but the context closed (or you
clear()ed it). No longer tracked; changes are not persisted until re-attached. - Removed — marked for deletion, will be
DELETEd at flush.
The transitions: new → transient; save/persist or find → managed; transaction ends / clear →
detached; delete → removed. Most bugs are a state confusion: expecting a change on a detached entity to
save (it will not), or hitting LazyInitializationException because you touched a lazy field on a detached
entity outside the context.
Flush, and why lazy loading has a boundary
Flushing is when Hibernate synchronises the context to the database — running the accumulated
inserts/updates/deletes. It happens automatically before a query that might need the pending changes, and at
commit. You rarely flush manually (em.flush()), but knowing flush exists explains ordering: your changes
are not sent to the database the instant you make them, but batched and flushed.
And this ties together the whole lesson with the LAZY relationship behaviour from before: a lazy association
can only be loaded while the persistence context is open — because loading it needs the context's
Session to run the query. Access a lazy field on a managed entity (inside the transaction) and
Hibernate runs the query fine. Access it on a detached entity (after the transaction closed — e.g. when
Jackson serialises an entity in the controller, outside any transaction) and you get
LazyInitializationException. So the fix (map to DTOs inside the transaction) is now fully explained: do
the lazy access while the entity is still managed and the context still open. The persistence context is
the boundary within which managed entities live, are cached, are change-tracked, and can lazy-load — cross
that boundary and they become detached, inert, and unable to lazy-load.
Check your work
What it is. A transaction-scoped workspace (EntityManager/Hibernate Session) holding managed
entities; the source of most JPA behaviour, via its two jobs: first-level cache and dirty checking.
First-level cache. One instance per id per context — loading the same id twice returns the same object
from one query (verified: 1 query, a == b true). Explains the N+1 "distinct rows" nuance.
Dirty checking. Mutating a managed entity persists automatically at flush — no save() needed;
Hibernate compares to the original and UPDATEs the difference. (Detached entities do not.)
Entity states. Transient (new, untracked), managed (tracked, dirty-checked), detached (context closed, inert), removed (to be deleted) — most bugs are a state confusion.
Flush and the lazy boundary. Flush synchronises the context to the DB (auto before queries and at
commit). Lazy loading works only while the context is open (managed entity); a detached entity throws
LazyInitializationException — which is why you map to DTOs inside the transaction.
Practice
- Load the same id twice in one
@Transactionalmethod (with statistics on) and confirm 1 query and the same instance (reproduce the verified result). - In a
@Transactionalmethod,findByIda parcel, change a field, and call nosave(); confirm (show-sql) anUPDATEruns at commit — dirty checking. - Do the same change on a detached entity (fetched in a different transaction) and confirm it is not
persisted until you
save(). - Label the state of an entity at four points: after
new, aftersave, after the transaction ends, afterdelete. - Touch a lazy relationship inside the transaction (works) and after it closes (throws
LazyInitializationException); connect this to why DTO mapping happens inside the transaction. - Call
em.flush()manually and observe the SQL is sent at that point rather than at commit.
Official documentation
- Hibernate — The persistence context — First-level cache, dirty checking, flushing.
- Jakarta Persistence — entity lifecycle — Managed, detached, removed states.
- Vlad Mihalcea — A beginner's guide to the persistence context — Entity state transitions in depth.
Next: queries — derived, JPQL, native and @Query.
Stuck on this lesson?
Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.
About the internship