RizTech Academy logo
RizTech Academy
TestingLesson 4 of 530 min

Testing the data layer with @DataJpaTest

Your repositories contain real logic worth testing — derived query methods, custom @Querys, entity mappings — and Spring gives you @DataJpaTest to test them against a real database, fast, with each test rolled back. This lesson tests Dakiya's ParcelRepository, verified by running.

@DataJpaTest: the persistence slice

@DataJpaTest boots a slice containing just the JPA infrastructure — entities, repositories, an EntityManager, a transaction manager — and not the web layer, services, or security. It is focused on persistence and much faster than a full boot. Two behaviours make it pleasant:

  • Each test runs in a transaction that is rolled back at the end, so tests do not pollute each other or leave data behind — you can freely save rows in one test without affecting the next.
  • By default it uses an in-memory database (H2) auto-configured for the test, so you need no running database to run these tests. (You can point it at a real database — see the note below.)
@DataJpaTest
class ParcelRepositoryTest {
    @Autowired ParcelRepository repository;

    @Test
    void findByStatus_returnsMatching() {
        repository.save(new Parcel("PKG-1", "Asha", "Pune"));
        repository.save(new Parcel("PKG-2", "Bhavna", "Mumbai"));

        assertThat(repository.findByStatus(Parcel.Status.BOOKED)).hasSize(2);
        assertThat(repository.findByDestinationCity("Pune")).hasSize(1);
    }
}

Verified by running: this passes — both saved parcels are BOOKED (so findByStatus returns 2), and only one is in Pune (so findByDestinationCity returns 1). The repository is autowired, real, and backed by a real (in-memory) database — so the test exercises the actual generated SQL, not a mock.

What to test at the data layer

@DataJpaTest is for testing your persistence logic, not Spring Data's:

  • Custom query methods — a derived method or @Query returns the right rows for given data. This is the main use: does findByStatusAndDestinationCity, or that hand-written JPQL, actually select what you intend?
  • Entity mapping — a save-then-load round trip preserves the fields, the enum maps correctly (@Enumerated(STRING)), relationships persist.
  • Constraints — a @Column(unique=true) or a check constraint rejects a bad row (expect a DataIntegrityViolationException).
  • Fetch behaviour — a fetch-join query loads the relationship in one query (the N+1 fix), verifiable with query counting.

Do not test that save saves or findById finds — that is Spring Data's own code. Test the queries and mappings you wrote, where you can make a mistake (a wrong @Query, a missed mappedBy, an ordinal enum). Those are the data-layer bugs a test catches.

The in-memory-versus-real-database question

By default @DataJpaTest uses H2 in-memory, which is fast and needs nothing installed — great for the tests' speed. But it carries the same caveat as developing on H2 and deploying on PostgreSQL: H2 is not PostgreSQL, and a query or constraint can behave differently between them. A test that passes on H2 can miss a bug that only appears on the real database (a Postgres-specific type, a subtle SQL difference, a constraint enforced differently).

The options:

  • H2 in-memory (default) — fast, zero setup; fine for straightforward queries and mappings. The pragmatic default for most data-layer tests.
  • A real PostgreSQL via Testcontainers — the next lesson. For repositories using Postgres-specific features, or when you want the highest fidelity, run @DataJpaTest against a real Postgres in a container (@AutoConfigureTestDatabase(replace = NONE) plus a Testcontainers Postgres). Slower, but tests the actual database.

The judgement: H2 for fast, ordinary query tests; a real database (Testcontainers) when fidelity matters — Postgres-specific SQL, or the critical queries you cannot afford to have behave differently in production. Do not blindly trust H2 for queries that lean on database specifics.

Check your work

@DataJpaTest. Boots only the JPA slice (entities, repositories, EntityManager) — not web/services/ security; fast. Each test runs in a rolled-back transaction (no pollution); uses an in-memory H2 by default (no DB to install).

What it tests. Your custom queries (derived/@Query return the right rows — verified findByStatus=2, findByDestinationCity=1), entity mappings (round trips, enum mapping, relationships), constraints, and fetch behaviour — against a real (in-memory) database, exercising the actual SQL. Not save/findById themselves (Spring Data's code).

H2 vs real DB. H2 default is fast but not PostgreSQL — queries/constraints can differ. Use H2 for ordinary tests; a real Postgres via Testcontainers (next lesson) for Postgres-specific or critical queries where fidelity matters.

Practice

  1. Write a @DataJpaTest for ParcelRepository; save two parcels and assert findByStatus and findByDestinationCity counts (reproduce the verified 2 and 1).
  2. Confirm test isolation: save rows in one test method and verify a second method does not see them (rollback).
  3. Round-trip an entity with the status enum and assert it maps correctly (proving @Enumerated(STRING)).
  4. Add a unique constraint and test that a duplicate insert throws DataIntegrityViolationException.
  5. Test a fetch-join query and, with query counting, confirm it loads the relationship in one query.
  6. Identify a query that uses a Postgres-specific feature and reason about why H2 might not catch a bug in it — motivating Testcontainers.

Official documentation

Next: integration tests with Testcontainers.

Stuck on this lesson?

Being stuck is part of it — but being stuck alone for three days is not. Our internship programme pairs this curriculum with code review and one-to-one help from working developers, and it is free.

About the internship