# Hibernate & JPA From Zero

> Learn the ORM that sits under most Java backends: what an ORM really is, JPA vs Hibernate, entities and mapping, the EntityManager and persistence context, transactions and dirty checking, relationships, lazy-vs-eager fetching and the N+1 trap, JPQL and the Criteria API, inheritance, and caching. The magic Spring Data JPA hides, made plain.


---

# Hibernate & JPA From Zero

Almost every Java backend that talks to a relational database is, somewhere underneath, running
Hibernate. If you came from [Spring Boot](/guides/spring-boot-from-zero), you met it without meeting it:
Spring Data JPA's "declare an interface and queries appear" magic is Hibernate doing the work two layers
down. This guide pulls that layer into the light. Learning it directly is doubly worth it - it's a
hugely employable skill on its own *and* it turns Spring Data JPA from magic into something you can
reason about and debug.

We go mental-model-first the whole way. An ORM feels like magic until you understand the two ideas that
run it - the **persistence context** and **dirty checking** - and then the whole thing becomes
predictable. By the end you'll map objects to tables, navigate relationships without triggering a
thousand queries, write real JPQL, and know exactly what SQL Hibernate sends and why.

> 📝 This assumes you know **Java** (classes, generics, annotations, collections) and the basics of
> **relational databases** (tables, rows, keys, joins). If either is shaky, do
> [Java From Zero](/guides/java-from-zero) and [What a Database Is](/guides/what-a-database-is) first.
> New to frameworks generally? [What a Framework Even Is](/guides/what-a-framework-even-is) sets the stage.

## How to read this

Read in order - it builds one example domain (authors, books, reviews) and adds a layer each phase.
Crucially, **keep `show_sql` on** as you go: the whole skill is connecting the Java you write to the SQL
Hibernate emits. Phases carry difficulty badges.

## The phases

**Part 1 - The core model (🟢 Basic → 🟡)**
1. **[What an ORM Is & Why Hibernate Exists](01-what-an-orm-is.md)** 🟢 - the object/relational mismatch, JPA (the spec) vs Hibernate (the implementation).
2. **[Entities & Basic Mapping](02-entities-and-mapping.md)** 🟢 - `@Entity`, `@Id`, generated keys, columns, and the table they imply.
3. **[The EntityManager & Persistence Context](03-entitymanager-and-persistence-context.md)** 🟡 - the heart of it all: managed objects, the identity map, and the four entity states.
4. **[Transactions & the Unit of Work](04-transactions-and-unit-of-work.md)** 🟡 - commit/rollback, flushing, and **dirty checking** (why changing a field updates the row with no `save` call).

**Part 2 - Relationships & queries (🔴 Advanced → 🟡)**
5. **[Mapping Relationships](05-mapping-relationships.md)** 🔴 - `@ManyToOne`/`@OneToMany`/`@ManyToMany`, owning vs inverse side, join columns.
6. **[Lazy vs Eager Fetching & the N+1 Problem](06-fetching-and-n-plus-1.md)** 🔴 - the single most important ORM performance lesson, and how to fix it.
7. **[Querying: JPQL, Criteria & Native SQL](07-querying-jpql-criteria.md)** 🟡 - query objects not tables, parameters, projections, and the escape hatch to raw SQL.
8. **[Inheritance & Embeddables](08-inheritance-and-embeddables.md)** 🟡 - mapping class hierarchies and value objects.

**Part 3 - Making it fast & real (🔴 Advanced → 🟢)**
9. **[Caching & Performance](09-caching-and-performance.md)** 🔴 - first- vs second-level cache, batching, and reading the SQL Hibernate emits.
10. **[Hibernate in the Real World & Where to Go Next](10-where-to-go-next.md)** 🟢 - JPA inside Spring, schema migrations, when to drop to SQL, and what to build.

> The payoff: after this, [Spring Boot's persistence phase](/guides/spring-boot-from-zero) stops being
> magic - you'll see exactly what Spring Data JPA generates and how to make it fast.


---

# What an ORM Is & Why Hibernate Exists

Here's the friction that started it all. In Java, you think in **objects**: an `Author` has a name, holds
a `List<Book>`, maybe inherits from some base class. References point from one object to another, and you
navigate them with a dot. But the database your data actually *lives* in thinks in **tables**: rows, columns,
and foreign keys - numbers in one table pointing at numbers in another. Two worlds, two completely different
shapes for the same information.

Hibernate exists because translating between those two shapes by hand - over and over, in every method that
touches the database - is some of the most tedious, error-prone code you'll ever write. This phase covers
*why* that translation is painful, what an ORM does to make it disappear, and the one habit that keeps an
ORM from quietly betraying you. We'll write almost no real code yet - the goal here is the mental model.
The actual domain (authors, books, reviews) starts in [Phase 2](02-entities-and-mapping.md).

## The problem: the object/relational impedance mismatch

📝 **Object/relational impedance mismatch** - the structural gap between the **object** world (classes,
references, inheritance, collections) and the **relational** world (tables, rows, columns, foreign keys).
They model the same data in fundamentally different ways, so moving information between them always takes
translation work.

The phrase sounds intimidating; the idea is concrete. In your Java program an author *contains* a list of
books - you write `author.getBooks()` and you have them. In the database there's no "contains." There's an
`authors` table and a `books` table, and each book row holds an `author_id` column - a number pointing back at
its author. To turn that row-with-a-number into a real `Author` object holding a real `List<Book>`, *someone*
has to do the stitching.

Before ORMs, that someone was you, using raw **JDBC** (Java's low-level database API). Here's the kind of code
you wrote for *every single query*:

```java
// Raw JDBC: turn one ResultSet row into one Author object - by hand.
String sql = "SELECT id, name, born_year FROM authors WHERE id = ?";
try (PreparedStatement ps = connection.prepareStatement(sql)) {
    ps.setLong(1, authorId);                 // bind the parameter by position
    try (ResultSet rs = ps.executeQuery()) {
        if (rs.next()) {
            Author a = new Author();
            a.setId(rs.getLong("id"));           // column -> field, by hand
            a.setName(rs.getString("name"));     // column -> field, by hand
            a.setBornYear(rs.getInt("born_year")); // column -> field, by hand
            return a;
        }
        return null;
    }
}
```

*What just happened:* you wrote the SQL string yourself, bound the `?` parameter by its position (get the
index wrong and you bind the wrong value - with no compiler to stop you), then pulled each column out of the
`ResultSet` *one at a time* and copied it into the matching field. That's the whole translation, done by hand.
Now imagine the author also has books and reviews, imagine doing the *reverse* mapping to save an object, and
imagine repeating all of it across forty different queries. Add one column to the table and you're hunting
through every method that mentions it.

⚠️ Every line of that hand-mapping is a place a typo, a wrong column name, or a mismatched type slips through
silently. The bug doesn't show up at compile time - it shows up at runtime, in production, as a field that's
mysteriously `null` or a `ClassCastException` deep in a stack trace. This boilerplate isn't just boring; it's
where data bugs are *born*. That repetitive, fragile copying is exactly what an ORM removes.

## What an ORM does

📝 **ORM (Object-Relational Mapper)** - a library that maps your **classes to tables** and your **objects to
rows** automatically. You declare the correspondence once ("this `Author` class maps to the `authors`
table"), then work with plain objects: ask for an author, get an `Author`; save an author, and the ORM writes
and runs the `INSERT` for you. It generates the SQL so you don't hand-write it.

With an ORM, that entire JDBC block from above collapses to roughly:

```java
Author a = entityManager.find(Author.class, authorId);  // SQL generated & run for you
```

*What just happened:* you told the ORM "fetch the `Author` with this id," and it figured out the `SELECT`,
ran it, read the `ResultSet`, and built the `Author` object with all its fields populated - the work that took
fifteen lines of JDBC, gone. You stayed entirely in the object world; the SQL happened underneath.

```mermaid
flowchart LR
  O[Java objects<br/>Author, Book] <--> ORM[ORM<br/>Hibernate]
  ORM <--> SQL[(SQL<br/>SELECT/INSERT)]
  SQL <--> DB[(Database<br/>tables &amp; rows)]
```

*What just happened:* the diagram is the whole pitch. Your code talks to objects on the left; the database
stores rows on the right; the ORM sits in the middle and translates *both directions* - objects into SQL when
you save, rows into objects when you load.

The trade is real and worth naming up front. You write far less boilerplate, but you've added a layer of
behavior between you and the database - a layer that makes decisions (when to run a query, what SQL to
generate, when to cache) that you don't see unless you go looking. This whole guide is about understanding
that layer so it stays a helper and never becomes a mystery. An ORM doesn't excuse you from knowing SQL - 
it sits *on top* of that knowledge.

## JPA vs Hibernate: the distinction everyone trips on

This is the one piece of vocabulary beginners get tangled in, so let's untangle it cleanly. **JPA and
Hibernate are not competitors, and they're not the same thing.** One is a standard; the other follows it.

📝 **JPA (Jakarta Persistence API)** - a *specification*. It's a written standard: a set of annotations and
interfaces (the `jakarta.persistence.*` package) that *describe* how Java objects should map to database
tables. JPA is a rulebook. By itself, it doesn't *do* anything - it defines the contract.

📝 **Hibernate** - an *implementation* of that specification. It's the actual library that reads your
JPA annotations, generates the SQL, and runs it. Hibernate is the engine that does the work the rulebook
describes. It's the most popular JPA implementation by a wide margin; **EclipseLink** is another.

The relationship is the same one you've already met elsewhere: a *standard* and the *products that implement
it*. SQL is a standard; PostgreSQL and MySQL implement it. JPA is the standard; Hibernate implements it.

```java
import jakarta.persistence.Entity;   // <- a JPA annotation (the SPECIFICATION)
import jakarta.persistence.Id;       // <- also JPA

@Entity                              // "map this class to a table" - defined by JPA
public class Author {
    @Id                              // "this field is the primary key" - defined by JPA
    private Long id;
    private String name;
}
```

*What just happened:* every annotation here (`@Entity`, `@Id`) comes from `jakarta.persistence` - the JPA
specification. Your class names the standard, not the engine. At runtime, **Hibernate** reads these
annotations and produces the `CREATE TABLE` and `SELECT`/`INSERT` statements that actually hit the database.
You wrote to the rulebook; Hibernate did the work.

💡 **Code to JPA, run on Hibernate.** This is the practical payoff of the split. If you write your mappings
and queries using only the JPA API, your code isn't welded to Hibernate - in principle you could swap in
EclipseLink and your annotations still mean the same thing. In practice almost everyone runs Hibernate, and
Hibernate offers extra features beyond the standard. The discipline worth keeping: prefer the JPA way when
JPA covers it, and reach for Hibernate-specific features knowingly, aware you're stepping outside the standard.

## Where it sits in the stack

ORMs rarely sit alone. A typical Spring application has a small tower of layers, each a thinner convenience
wrapper over the one below it:

```mermaid
flowchart TD
  A[Your repository code] --> B[Spring Data JPA<br/>generates repository methods]
  B --> C[JPA<br/>the specification / API]
  C --> D[Hibernate<br/>the implementation]
  D --> E[JDBC<br/>low-level DB driver]
  E --> F[(Database)]
```

*What just happened:* a `findById` call enters at the top through **Spring Data JPA** (which generated that
method for you), which calls the **JPA** API, which is implemented by **Hibernate**, which builds SQL and
hands it to **JDBC** - the same low-level API from our first example - which finally talks to the database.
Each layer removes boilerplate from the one above it, all the way down to the raw `PreparedStatement` you saw
at the start.

💡 If you've been through [Spring Boot from zero](/guides/spring-boot-from-zero), this is the machinery that
was running *under* your repositories the whole time. Those tidy `findById` and `save` methods didn't talk to
the database directly - they sat on JPA, which sat on Hibernate. This guide is you opening that box and
learning what's inside. (And the JDBC layer at the bottom is the raw approach we'll keep [Java](/guides/java-from-zero)
developers grateful to never touch by hand again.)

## Minimal setup

You don't need much to start, and we'll keep it light - the real domain begins next phase. Four things have
to be in place:

1. **The dependency** - pull Hibernate (and the JPA API) into your project.
2. **Configuration** - either a `persistence.xml` file (plain JPA) or, in Spring, a few properties.
3. **A datasource** - the database URL, username, and password so Hibernate knows where to connect.
4. **`show_sql` turned on** - so you can *see* every SQL statement Hibernate generates.

In a Spring Boot project, the configuration is a handful of lines in `application.properties`:

```properties
# Where the database lives (the datasource)
spring.datasource.url=jdbc:postgresql://localhost:5432/library
spring.datasource.username=app
spring.datasource.password=secret

# THE most important line in this whole guide:
spring.jpa.show-sql=true
spring.jpa.properties.hibernate.format_sql=true
```

*What just happened:* the first three lines tell Hibernate which database to talk to. The last two are the
ones that matter for your sanity: `show-sql` prints every generated SQL statement to your console, and
`format_sql` lays it out readably instead of as one long line. With these on, the queries Hibernate writes
stop being invisible. When you load an author and watch this scroll past, the magic becomes legible:

```console
Hibernate:
    select
        a1_0.id,
        a1_0.name,
        a1_0.born_year
    from
        authors a1_0
    where
        a1_0.id=?
```

*What just happened:* that's Hibernate *showing you its work* - the exact `SELECT` it generated for a
`find(Author.class, id)` call, the same query you'd have hand-written in JDBC, now produced for you. Seeing it
confirms what ran, in what shape, against which tables.

⚠️ **Turn on `show_sql` now and never fly blind.** This is the single most important habit for everything
that follows. An ORM's whole job is to hide SQL from you - which is wonderful right up until it generates
something wasteful (like the dreaded hundreds-of-queries "N+1" problem we'll meet later) and you have no idea,
because you never looked. Reading the generated SQL is how you catch those problems early instead of in a
production incident. Every gotcha in this guide is one you'll *see coming* if `show_sql` is on. Leave it on
while you learn; the noise is the point.

## Recap

1. The **object/relational impedance mismatch** is the structural gap between Java's world (objects,
   references, collections) and the database's world (tables, rows, foreign keys) - bridging it by hand with
   raw **JDBC** is tedious and a breeding ground for silent bugs.
2. An **ORM** maps classes to tables and objects to rows automatically: you work with objects, it generates
   and runs the SQL. The trade is less boilerplate in exchange for a behavior layer you must understand.
3. **JPA is the specification** (the `jakarta.persistence.*` annotations and interfaces - a rulebook);
   **Hibernate is the most popular implementation** (the engine that does the work). EclipseLink is another.
4. **Code to JPA, run on Hibernate** - write to the standard, and Hibernate does the heavy lifting underneath.
5. The stack runs **Spring Data JPA → JPA → Hibernate → JDBC → database** - this is exactly what was powering
   your Spring Boot repositories all along.
6. ⚠️ Minimal setup is a dependency, config, and a datasource - and crucially, turn on **`show_sql`** so you
   *see* every query. Never fly blind; it's the habit that catches every gotcha ahead.

## Quick check

Three questions on the ideas that have to stick before Phase 2:

```quiz
[
  {
    "q": "What is the 'object/relational impedance mismatch'?",
    "choices": [
      "The structural gap between Java's object world (classes, references, collections) and the database's relational world (tables, rows, foreign keys)",
      "A performance problem caused by slow database hardware",
      "A bug that happens when two threads write to the same row at once",
      "The difference in speed between Hibernate and raw JDBC"
    ],
    "answer": 0,
    "explain": "Objects and relational tables model the same data in fundamentally different shapes - objects contain references and collections; tables use rows and foreign-key numbers. Translating between the two is the 'mismatch' an ORM exists to bridge."
  },
  {
    "q": "What is the relationship between JPA and Hibernate?",
    "choices": [
      "JPA is the specification (a standard set of annotations and interfaces); Hibernate is the most popular implementation of it",
      "They are two competing ORMs and you must pick one or the other",
      "Hibernate is the specification and JPA is one implementation of it",
      "They are different names for exactly the same library"
    ],
    "answer": 0,
    "explain": "JPA (Jakarta Persistence API) is the rulebook - the jakarta.persistence.* annotations and interfaces. Hibernate is the engine that reads those annotations and actually generates and runs the SQL. You code to JPA and run on Hibernate."
  },
  {
    "q": "Why is turning on `show_sql` called the single most important habit in this guide?",
    "choices": [
      "An ORM hides SQL by design, so seeing the generated queries is how you catch wasteful or unexpected SQL (like the N+1 problem) early instead of in production",
      "It makes Hibernate generate faster queries automatically",
      "It is required or Hibernate refuses to connect to the database",
      "It encrypts the SQL so the queries are more secure"
    ],
    "answer": 0,
    "explain": "An ORM's job is to hide SQL - which is great until it generates something wasteful and you never notice because you never looked. show_sql prints every generated statement so the ORM's behavior stays visible and you spot problems early."
  }
]
```


---

# Entities & Basic Mapping

In Phase 1 you saw the big idea: an ORM lets you work with Java objects and quietly keeps a database
table in sync underneath. This phase is where that promise gets concrete. We take an ordinary Java
class - a `Book` - and teach Hibernate to treat it as a row in a table, and see how a class becomes a
table, a field becomes a column, and how to nudge any of that when the defaults aren't what you want.

The mental model to hold the whole time: **the entity class is the map between two worlds.** On one
side, Java objects in memory. On the other, rows in a SQL table. Everything in this phase is you
drawing that map - and Hibernate following it in both directions.

We're keeping it deliberately small here. Our domain across this guide is `Author`, `Book`, and
`Review`, but relationships between them don't arrive until [Phase 5](05-mapping-relationships.md). For
now, `Book` stands alone - just its own data, mapped to its own table.

## `@Entity` - marking a class as a table

📝 **`@Entity`** - an annotation you put on a class to tell JPA "instances of this class correspond to
rows in a database table." That single annotation is what flips a normal class into something Hibernate
will load, save, and track.

Here's our `Book`, stripped to the bone:

```java
import jakarta.persistence.Entity;
import jakarta.persistence.Id;

@Entity
public class Book {

    @Id
    private Long id;

    private String title;
    private String isbn;
    private int publishedYear;

    // JPA needs this no-arg constructor
    public Book() {
    }

    public Book(String title, String isbn, int publishedYear) {
        this.title = title;
        this.isbn = isbn;
        this.publishedYear = publishedYear;
    }

    // getters and setters omitted for brevity
}
```

*What just happened:* `@Entity` told JPA that `Book` is mapped to a table. By default the table is named
after the class - `Book` → a table called `book` (or `Book`, depending on the database). Each field
becomes a column, also by name: `title`, `isbn`, `publishedYear`. We haven't written a line of SQL, yet
Hibernate now knows enough to read and write `Book` rows.

⚠️ **Entities are mutable classes, not records.** It's tempting to reach for a Java `record` here - it's
concise and immutable, perfect for a data holder. But JPA can't use records as entities. It needs a
**no-arg constructor** and a **non-final class with non-final fields**, and a record gives you none of
those. The reason is mechanical: Hibernate constructs a *blank* `Book` with `new Book()` and then fills
the fields in one by one as it reads a row - it can't do that if the only constructor demands all the
values up front. It also sometimes wraps your entity in a **proxy** (a generated subclass used for lazy
loading, which you'll meet in [Phase 6](06-fetching-and-n-plus-1.md)), and you can't subclass a final
class. So entities stay plain, mutable classes with a no-arg constructor. That's not Hibernate being
old-fashioned; it's the price of the magic.

> 💡 You can keep the no-arg constructor `protected` rather than `public` if you'd rather callers didn't
> use it directly - Hibernate only needs to be able to reach it via reflection, and `protected` is
> enough.

## `@Id` and key generation

Every row in a relational table needs a way to be uniquely identified - its **primary key**. JPA
mirrors that exactly: every entity needs a field marked `@Id`.

📝 **`@Id`** - marks the field that is the entity's primary key. It's the identity of the object in the
database; Hibernate uses it to know whether two rows are "the same record."

In the example above we marked `id` with `@Id` but left it to us to assign. That works, but it's a
chore - you'd have to invent a unique number every time you create a `Book`. The usual move is to let
the database (or Hibernate) generate the key for you:

```java
import jakarta.persistence.Entity;
import jakarta.persistence.GeneratedValue;
import jakarta.persistence.GenerationType;
import jakarta.persistence.Id;

@Entity
public class Book {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    private String title;
    private String isbn;
    private int publishedYear;

    // constructors, getters, setters...
}
```

*What just happened:* `@GeneratedValue` says "don't make me supply the id - generate it." The `strategy`
chooses *how*. With `IDENTITY`, you leave `id` null when you create a new `Book`; the database fills it
in on insert (using an auto-increment column), and Hibernate reads the generated value back out
afterward.

The strategy you pick changes the actual SQL Hibernate sends, so it's worth knowing the three you'll
meet:

- **`IDENTITY`** - the database owns the counter (an auto-increment / `IDENTITY` column). The id only
  exists *after* the `INSERT` runs, because the database assigns it. Simple and common (MySQL,
  PostgreSQL `SERIAL`), but it has a subtle cost: Hibernate can't batch inserts well, because it must
  run each insert immediately to learn the new id.
- **`SEQUENCE`** - the database has a separate **sequence** object whose only job is handing out
  numbers. Hibernate asks the sequence for the next id *before* inserting, so it knows the id up front
  and can batch inserts together. This is the preferred strategy on databases that support sequences
  (PostgreSQL, Oracle).
- **`AUTO`** - "you decide, Hibernate." It picks a strategy based on the database. Convenient, but
  you're handing the choice to a default you can't see, so once you care about performance, name the
  strategy explicitly.

💡 If you're on PostgreSQL or Oracle and have any volume of inserts, prefer `SEQUENCE` - the ability to
batch is a real performance win. On MySQL, `IDENTITY` is the natural fit. When in doubt early on, `AUTO`
is fine; just don't be surprised later when the generated SQL looks different than a teammate's on
another database.

## Column and table mapping

So far we've leaned entirely on defaults: class name → table name, field name → column name. Most of the
time that's exactly what you want. When it isn't, two annotations let you take over.

📝 **`@Column`** customizes how a single field maps to its column - the name, whether it can be null, its
length, whether it must be unique. **`@Table`** customizes the table the whole entity maps to - most
often just its name.

Here's `Book` with the mapping spelled out:

```java
import jakarta.persistence.Column;
import jakarta.persistence.Entity;
import jakarta.persistence.GeneratedValue;
import jakarta.persistence.GenerationType;
import jakarta.persistence.Id;
import jakarta.persistence.Table;

@Entity
@Table(name = "books")
public class Book {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    @Column(nullable = false, length = 200)
    private String title;

    @Column(name = "isbn_13", length = 13, unique = true)
    private String isbn;

    @Column(name = "published_year", nullable = false)
    private int publishedYear;

    // constructors, getters, setters...
}
```

*What just happened:* we renamed the table to `books` with `@Table`. We told Hibernate `title` can't be
null and caps at 200 characters. We mapped the `isbn` field to a column actually called `isbn_13`, made
it 13 characters, and marked it `unique` so no two books can share an ISBN. And `publishedYear` (Java's
camelCase) now lands in a column named `published_year` (SQL's usual snake_case). The Java names and the
SQL names no longer have to match - the annotations are the translation layer.

Those annotations aren't just runtime hints; they describe a real table. If you let Hibernate generate
the schema (next section), this is the DDL it would produce from the class above:

```sql
CREATE TABLE books (
    id            BIGINT       NOT NULL AUTO_INCREMENT,
    title         VARCHAR(200) NOT NULL,
    isbn_13       VARCHAR(13),
    published_year INTEGER     NOT NULL,
    PRIMARY KEY (id),
    UNIQUE (isbn_13)
);
```

*What just happened:* read it side by side with the entity and the mapping clicks into place.
`@Id @GeneratedValue(IDENTITY)` became the `BIGINT ... AUTO_INCREMENT` primary key. `nullable = false`
became `NOT NULL`. `length = 200` became `VARCHAR(200)`. The `@Column(name = ...)` values became the
real column names, and `unique = true` became a `UNIQUE` constraint. The entity class and this table are
two views of the same thing.

## Basic type mapping

You may have noticed we never told Hibernate that `title` is text and `publishedYear` is a number - it
just knew. That's because JPA has built-in rules for mapping common Java types to SQL types. The ones
you'll use constantly:

| Java type | SQL type (typical) |
|-----------|--------------------|
| `String` | `VARCHAR` |
| `int` / `Integer` | `INTEGER` |
| `long` / `Long` | `BIGINT` |
| `boolean` / `Boolean` | `BOOLEAN` (or a 0/1 column) |
| `double` / `BigDecimal` | `DOUBLE` / `NUMERIC` |
| `LocalDate` | `DATE` |
| `LocalDateTime` | `TIMESTAMP` |

For these, you write the field and Hibernate handles the rest. Two cases need a small hint from you:

**Enums** map either as a number or as text, and you choose with `@Enumerated`:

```java
import jakarta.persistence.Enumerated;
import jakarta.persistence.EnumType;

public enum Format { HARDCOVER, PAPERBACK, EBOOK }

@Enumerated(EnumType.STRING)   // store the name "PAPERBACK", not the position 1
private Format format;
```

*What just happened:* `@Enumerated(EnumType.STRING)` tells Hibernate to store the enum's *name* as text
in the column. The alternative, `EnumType.ORDINAL`, stores the enum's *position* (0, 1, 2…) as a number.
⚠️ Strongly prefer `STRING`. With `ORDINAL`, reordering your enum constants or inserting a new one in the
middle silently changes what every existing row means - a `1` that meant `PAPERBACK` yesterday might
mean something else tomorrow. `STRING` survives reordering because the name, not the position, is what's
stored.

**Skipping a field** entirely uses `@Transient`:

```java
import jakarta.persistence.Transient;

@Transient
private int cachedWordCount;   // computed in memory, never stored
```

*What just happened:* `@Transient` tells JPA "this field is not part of the mapping - don't give it a
column, don't load it, don't save it." Use it for values you compute on the fly or hold temporarily but
never want persisted. (Don't confuse it with Java's own `transient` keyword, which is about
serialization - `@Transient` is the JPA annotation, and that's the one Hibernate reads.)

## Schema generation with `hibernate.hbm2ddl.auto`

You've now seen entities turn into DDL twice. Hibernate can do that for you automatically - read your
entities at startup and create or update the matching tables.

📝 **`hibernate.hbm2ddl.auto`** - a configuration setting that controls whether (and how) Hibernate
generates database schema from your entities when the app starts. The values you'll see:

- **`create`** - drop all the mapped tables and recreate them from scratch on every startup. Clean slate
  each run. Everything in those tables is erased.
- **`create-drop`** - like `create`, and also drop them again when the app shuts down. Handy for tests.
- **`update`** - compare your entities to the existing tables and apply additive changes (e.g. add a new
  column for a new field). It never deletes or alters existing columns, so it drifts over time.
- **`validate`** - change nothing; just check that the tables match your entities and fail fast at
  startup if they don't. Great as a safety net.
- **`none`** - do nothing. You manage the schema yourself.

You set it in `persistence.xml` or your Spring config, for example:

```console
hibernate.hbm2ddl.auto = update
```

*What just happened:* with `update`, you can add a `pageCount` field to `Book`, restart, and Hibernate
adds the matching column for you - no hand-written `ALTER TABLE`. During development that feedback loop
is wonderful: change the class, restart, the schema follows.

💡 For local development and quick experiments, `create` or `update` is genuinely great - it removes
schema busywork while you're still shaping your model.

⚠️ **Never use `create` or `update` in production.** `create` would erase your real data on every
deploy. `update` is safer but still untrustworthy: it silently ignores column type changes, can't rename
or drop anything, and gives you no record of what changed - so your schema slowly drifts in ways nobody
reviewed. In production you control schema changes deliberately, with **migrations** (versioned,
reviewed SQL scripts via tools like Flyway or Liquibase). We cover that properly in
[Phase 10](10-where-to-go-next.md). A common, sane setup is `validate` in production: Hibernate changes
nothing but refuses to start if the live schema and your entities have drifted apart.

The throughline of this whole phase: **the entity class is the single source of truth for the mapping.**
Whether Hibernate generates the table for you or you write the migration by hand, the annotations on
your `Book` define what a `Book` *is* in both worlds. Get the entity right and everything downstream - 
queries, saves, the generated SQL - follows from it.

## Recap

1. **`@Entity`** marks a class as mapped to a table; by default the class name becomes the table name
   and each field becomes a column of the same name.
2. JPA entities must be **mutable, non-final classes with a no-arg constructor** - not records - 
   because Hibernate constructs blank instances and may wrap them in proxies.
3. **`@Id`** marks the primary key, and **`@GeneratedValue`** generates it for you: `IDENTITY` (database
   auto-increment, id known after insert), `SEQUENCE` (sequence object, id known before insert, allows
   batching), or `AUTO` (Hibernate decides).
4. **`@Column`** customizes a field's column (`name`, `nullable`, `length`, `unique`) and **`@Table`**
   renames the table - together they're the translation layer between Java names and SQL names.
5. Common Java types map to SQL automatically (`String`→`VARCHAR`, `int`→`INTEGER`, `LocalDate`→`DATE`);
   use **`@Enumerated(EnumType.STRING)`** for enums and **`@Transient`** to skip a field.
6. **`hibernate.hbm2ddl.auto`** can generate the schema (`create`, `update`, `validate`, `none`) - great
   in dev, but never `create`/`update` in production; use migrations there. The entity is the single
   source of truth for the mapping.

## Quick check

Test yourself on the ideas most likely to trip you up in real mapping code:

```quiz
[
  {
    "q": "Why can't you use a Java `record` as a JPA entity?",
    "choices": [
      "Records are too new for Hibernate to recognize",
      "JPA needs a no-arg constructor and a non-final, mutable class so it can build a blank instance and possibly proxy it - records provide none of those",
      "Records can't have an @Id field",
      "Records map to multiple tables, which Hibernate forbids"
    ],
    "answer": 1,
    "explain": "Hibernate constructs an empty object with `new Book()` and fills fields in as it reads a row, and it may wrap the entity in a generated subclass (a proxy). A record has no no-arg constructor and is effectively final/immutable, so neither is possible."
  },
  {
    "q": "Your @Id uses `@GeneratedValue(strategy = GenerationType.IDENTITY)`. When does the id value exist?",
    "choices": [
      "Before the INSERT - Hibernate asks a sequence object for it first",
      "After the INSERT - the database assigns it via an auto-increment column and Hibernate reads it back",
      "As soon as you call `new Book(...)`",
      "Only after you commit the transaction"
    ],
    "answer": 1,
    "explain": "With IDENTITY, the database owns the counter via an auto-increment column, so the id only exists once the row is inserted. That's also why IDENTITY can't batch inserts well - each insert must run immediately to learn its id. SEQUENCE, by contrast, gets the id before inserting."
  },
  {
    "q": "Which `hibernate.hbm2ddl.auto` setting is the safe choice for production?",
    "choices": [
      "`create` - guarantees a clean, correct schema every deploy",
      "`update` - applies just the changes you need automatically",
      "`validate` - changes nothing but fails startup if the live schema and your entities have drifted apart",
      "`none` is the only acceptable value; any other risks data loss"
    ],
    "answer": 2,
    "explain": "`create` erases data every deploy and `update` silently ignores type changes and drifts. `validate` makes no changes and just checks the schema matches your entities, failing fast if not - a good safety net while you manage real changes with versioned migrations."
  }
]
```


---

# The EntityManager & Persistence Context

In Phase 2 you mapped a `Book` to a table - `@Entity`, `@Id`, columns, the works. But a mapping just
sits there. Something has to actually *do* things with it: insert a new book, fetch one back, change a
title and have that change reach the database. That something is the **`EntityManager`**, and the place
it does its work is the **persistence context**.

This is the phase. If you take one idea from this whole guide, take this one. Almost every Hibernate
surprise you'll ever hit - a change that saved without you calling save, a lazy load that worked here
but blew up there, the dreaded N+1 - traces straight back to what's in this file. Once the persistence
context clicks, the rest of Hibernate stops being magic and starts being obvious.

## The mental model: a workbench, not a pipe

**What it actually is.** People picture an ORM as a pipe: Java object goes in one end, SQL comes out the
other, row lands in the table. That picture will mislead you for years. Hibernate isn't a pipe - it's a
**workbench**. When you load or save objects, Hibernate lays them out on a workbench it keeps for the
duration of your transaction, watches them, and only sends SQL to the database when it decides it's
time. The `EntityManager` is your handle to that workbench. The **persistence context** *is* the
workbench - the in-memory area holding the objects Hibernate is currently managing for you.

> 💡 **Key point.** Hibernate doesn't write to the database the instant you touch an object. It tracks a
> small set of objects in memory (the persistence context) and synchronizes them with the database on
> *its* schedule. Every "wait, when did *that* SQL run?" question is answered by understanding this one
> idea.

## The `EntityManager` - your handle to JPA

📝 **`EntityManager`** - the single object you use to talk to JPA. Nearly everything goes through it:

| Method | What it does |
|--------|--------------|
| `persist(entity)` | Make a brand-new object managed (schedules an `INSERT`) |
| `find(Class, id)` | Look up one entity by primary key |
| `merge(entity)` | Copy a detached object's state back into a managed one |
| `remove(entity)` | Mark a managed entity for deletion (schedules a `DELETE`) |
| `createQuery(...)` | Run JPQL - the topic of Phase 7 |

> 📝 Hibernate predates JPA, and its own native equivalent of `EntityManager` is called **`Session`**.
> They do the same job; `EntityManager` is the standard JPA name and the one you'll see in Spring. When
> an old Stack Overflow answer says "the Hibernate Session," mentally read "the EntityManager."

Let's save a `Book` and read it back. (We'll treat the transaction boilerplate as a given here - 
Phase 4 dissects it.)

```java
EntityManager em = emf.createEntityManager();
em.getTransaction().begin();

Book book = new Book("Dune", "Frank Herbert");  // a plain Java object, nothing special yet
em.persist(book);                                // hand it to the EntityManager

em.getTransaction().commit();
em.close();
```
```sql
insert into book (author, title, id) values ('Frank Herbert', 'Dune', 1)
```
*What just happened:* `new Book(...)` created an ordinary object - at that moment Hibernate knows nothing
about it. `em.persist(book)` placed it on the workbench: now it's *managed*, and Hibernate has scheduled
an `INSERT`. Notice the SQL didn't fire on the `persist` line - it fired at `commit`. That gap between
"I told Hibernate about this object" and "the SQL actually ran" is the whole story of this phase.

Now read it back:

```java
EntityManager em = emf.createEntityManager();

Book found = em.find(Book.class, 1L);   // SELECT by primary key
System.out.println(found.getTitle());

em.close();
```
```sql
select b.id, b.author, b.title from book b where b.id = 1
```
```console
Dune
```
*What just happened:* `find(Book.class, 1L)` asked the EntityManager for the book with primary key `1`.
Not finding it on the workbench, Hibernate ran a `SELECT`, built a `Book` object from the row, placed it
on the workbench (now *managed*), and handed it to you. Plain, predictable. The interesting behavior
shows up when you ask for the *same* book twice.

## The persistence context - an identity map and a first-level cache

This is the core idea. 📝 The **persistence context** is a per-transaction, in-memory area that holds
the entities the EntityManager is currently managing. It does two jobs at once, and both surprise people
the first time:

1. **Identity map** - within one persistence context, a given database row maps to exactly **one** Java
   object. Ask for book `1` ten times and you get the *same instance* back every time.
2. **First-level cache** - once an entity is in the context, looking it up again by id returns it from
   memory. **No second SQL query.**

Here's the proof. Watch how many `SELECT`s come out:

```java
EntityManager em = emf.createEntityManager();

Book first  = em.find(Book.class, 1L);   // hits the database
Book second = em.find(Book.class, 1L);   // same id, same context

System.out.println(first == second);     // not .equals - identity, ==
```
```sql
select b.id, b.author, b.title from book b where b.id = 1
```
```console
true
```
*What just happened:* Two `find` calls, but **only one `SELECT`**. The first call ran the query and put
the `Book` on the workbench. The second call found it already there and returned it straight from
memory - no database round trip. And `first == second` is `true`: not just equal *values*, the very
**same object** (remember `==` vs `.equals()` from
[Java's classes phase](/guides/java-from-zero) - this is `==`, raw identity). That's the identity map
guaranteeing one row, one object, per context.

> ⚠️ This cache is **per persistence context** - it lives and dies with one transaction. It is *not* a
> shared application-wide cache that survives across requests. Open a new `EntityManager` and you get a
> fresh, empty workbench; the next `find` hits the database again. (The shared, long-lived cache is the
> *second*-level cache, and it's a whole separate opt-in feature - Phase 9.)

Why does Hibernate work this way? Because the identity map is what makes the next idea - automatic change
tracking - even possible. If you and three other lines of code each loaded book `1` into a *different*
object, Hibernate couldn't know which one's changes to save. One row, one object means there's exactly
one source of truth on the workbench to watch.

## The four entity states

Every entity, from Hibernate's point of view, is always in exactly one of **four states**. Learn these
names cold - error messages, docs, and your own debugging all speak this language.

📝 The four states:

- **Transient** - a brand-new object you made with `new`. Hibernate has never heard of it; it's not on
  the workbench and has no database row. (`new Book(...)` before any `persist`.)
- **Managed** (also *persistent*) - on the workbench, tracked by the persistence context, tied to a
  database row. Hibernate watches it and will save changes to it. (After `persist`, or anything `find`
  returns.)
- **Detached** - *was* managed, but its persistence context has closed. It still holds data, but nobody's
  watching it anymore; changes to it go nowhere. (A `Book` you loaded, after `em.close()`.)
- **Removed** - a managed entity you've marked for deletion with `remove`. It's scheduled to disappear at
  the next flush.

Here's the lifecycle as a diagram:

```mermaid
stateDiagram-v2
    [*] --> Transient: new Book(...)
    Transient --> Managed: persist()
    [*] --> Managed: find() loads it
    Managed --> Detached: context closes
    Detached --> Managed: merge()
    Managed --> Removed: remove()
    Removed --> [*]: flush -> DELETE
```

Let's walk one object through three of those states:

```java
Book book = new Book("Dune", "Frank Herbert");  // TRANSIENT - Hibernate doesn't know it

EntityManager em = emf.createEntityManager();
em.getTransaction().begin();
em.persist(book);                                // now MANAGED - on the workbench
book.setTitle("Dune (Special Edition)");         // tracked: this change WILL be saved
em.getTransaction().commit();
em.close();

book.setTitle("ignored");                        // now DETACHED - this change goes nowhere
```
```sql
insert into book (author, title, id) values ('Frank Herbert', 'Dune (Special Edition)', 1)
```
*What just happened:* The object started **transient** - a plain object Hibernate ignored. `persist`
made it **managed**, so when we changed the title *before commit*, Hibernate noticed and the `INSERT`
used the new value. After `em.close()` the context was gone, leaving the object **detached** - so the
final `setTitle("ignored")` changed the in-memory object but emitted no SQL, because nothing was watching
it. Same object, three different relationships to the database, depending purely on state. (That
"changing a managed field updates the row with no save call" behavior is *dirty checking* - Phase 4
makes it the star.)

## `persist` vs `merge` - the classic confusion

⚠️ This trips up nearly everyone, so read it twice. `persist` and `merge` sound interchangeable. They are
not, and reaching for the wrong one causes some of the most baffling Hibernate bugs.

- **`persist`** is for a **transient** (brand-new) object. It takes the object you pass and makes *that
  object* managed.
- **`merge`** is for a **detached** object. It does **not** make your object managed. It copies your
  detached object's state onto a managed copy and **returns that managed copy** - and the object you
  passed in *stays detached*.

That return value is the trap. Watch:

```java
// 'book' was loaded in a previous context, which closed - so it's DETACHED.
book.setTitle("New Title");                      // change the detached object

EntityManager em = emf.createEntityManager();
em.getTransaction().begin();

Book managed = em.merge(book);                    // returns the MANAGED copy

managed.setAuthor("Updated Author");              // change THIS one - it's the tracked one
book.setAuthor("goes nowhere");                   // change to the detached one - ignored

em.getTransaction().commit();
em.close();
```
```sql
select b.id, b.author, b.title from book b where b.id = 1
update book set author='Updated Author', title='New Title' where b.id=1
```
*What just happened:* `merge` ran a `SELECT` to load the current managed instance, copied `book`'s state
onto it, and **returned that managed instance** as `managed`. The update saved "New Title" (merged from
`book`) and "Updated Author" (set on `managed`) - but **not** "goes nowhere," because `book` is still
detached and nobody's watching it. The rule to burn in: **after `merge`, work with the returned object,
never the one you passed in.** Calling `persist` on a detached entity instead would throw - `persist` is
strictly for transient objects.

## Why this is the lens for everything that follows

💡 Step back, because this is the payoff. Nearly every Hibernate behavior that feels like magic or
mystery is one of these ideas wearing a costume:

- *"I changed a field and it saved without calling save"* → it was **managed**, and dirty checking caught
  the change (Phase 4).
- *"The same query ran once instead of twice"* → the **first-level cache** served the second call.
- *"`==` returned true for two loads"* → the **identity map** gave you one object per row.
- *"Why did this update fail silently?"* → you changed a **detached** object, or worked with the wrong
  side of a `merge`.
- *The N+1 problem* (Phase 6) → entities loaded one-by-one into the context, each triggering its own
  query.

⚠️ One more forward-reference worth planting now: a **detached entity can't lazy-load**. If you load a
`Book`, close the context, and *then* try to walk to a relationship that wasn't fetched yet, Hibernate
has no open persistence context to run the query through - and you get the infamous
`LazyInitializationException`. We'll meet it properly in Phase 6, but you already understand *why* it
happens: no open context, no workbench, nothing to do the lazy load. That's the whole point of learning
states first.

## Recap

1. The **`EntityManager`** is your handle to JPA - `persist`, `find`, `merge`, `remove`, `createQuery`.
   (Hibernate's native equivalent is the **`Session`**.)
2. The **persistence context** is a per-transaction, in-memory workbench holding managed entities.
   Hibernate syncs it to the database on *its* schedule, not the instant you touch an object.
3. It's a **first-level cache + identity map**: within one context, `find` the same id twice → one
   `SELECT` and the **same** object instance (`==` is true).
4. Every entity is **transient** (new, unknown), **managed** (tracked, tied to a row), **detached**
   (context closed, unwatched), or **removed** (marked for delete).
5. ⚠️ **`persist`** is for transient objects; **`merge`** is for detached ones - and `merge` **returns**
   the managed copy while your original stays detached. Always use the returned object.
6. 💡 This is the lens for the rest of the guide: dirty checking, lazy loading, and N+1 all reduce to the
   persistence context and entity states. A detached entity can't lazy-load (forward-ref Phase 6).

## Quick check

The three ideas that explain the most future bugs:

```quiz
[
  {
    "q": "Inside one persistence context, you call `em.find(Book.class, 1L)` twice. How many SELECT queries does Hibernate run, and is the result the same object?",
    "choices": [
      "One SELECT; both calls return the same object instance (== is true) - the first-level cache and identity map serve the second call from memory",
      "Two SELECTs; you get two separate objects with equal data",
      "Two SELECTs, but Hibernate returns the same object both times",
      "Zero SELECTs; find never touches the database"
    ],
    "answer": 0,
    "explain": "The persistence context is a first-level cache plus an identity map. The first find runs the SELECT and stores the Book; the second find returns that same instance from memory with no new query, so == is true."
  },
  {
    "q": "You call `Book managed = em.merge(detachedBook);` and then change a title on `detachedBook` (not on `managed`). What happens to that change?",
    "choices": [
      "Nothing - `detachedBook` is still detached after merge; only changes to the returned `managed` object are tracked and saved",
      "It's saved, because merge makes `detachedBook` managed",
      "It throws a LazyInitializationException",
      "Both objects are now managed, so the change is saved"
    ],
    "answer": 0,
    "explain": "merge does not make the object you pass in managed. It copies its state onto a managed copy and returns that copy. The original stays detached, so changes to it go nowhere. Always work with the object merge returns."
  },
  {
    "q": "Which entity state describes a `Book` you loaded with `find`, after its EntityManager has been closed?",
    "choices": [
      "Detached - it was managed, but its persistence context is gone, so it's no longer tracked",
      "Transient - it has no connection to Hibernate",
      "Managed - find always returns managed entities",
      "Removed - closing the context schedules it for deletion"
    ],
    "answer": 0,
    "explain": "An entity that was managed but whose persistence context has closed is detached. It still holds its data, but nothing watches it, so changes won't be saved - and it can no longer lazy-load relationships."
  }
]
```


---

# Transactions & the Unit of Work

In [Phase 3](03-entitymanager-and-persistence-context.md) you met the persistence context - the in-memory workspace where the `EntityManager` keeps your managed entities, the identity map that guarantees one object per row, and the four states an entity can live in. That phase answered "where do my objects live while Hibernate is looking after them?" This one answers the question everyone asks next, usually in a panic: *"I changed a field and never called `save` - why did the database update?"*

Before any code, here's the whole phase in one sentence - paste it on your monitor:

> **You work inside a transaction; Hibernate watches the managed objects in the persistence context, and at the right moment it figures out the SQL and sends it as one batch.**

Everything below is that sentence unpacked. The "magic" save isn't magic - it's the persistence context from Phase 3 doing exactly the job it was built for. Once you see the mechanism, Hibernate stops surprising you and starts being predictable, which is the entire point of learning it directly.

Keep `show_sql` on, as the guide overview insists - this phase is *all* about connecting the Java you write to the SQL Hibernate emits, and you can only see the timing if Hibernate prints it.

## Everything happens in a transaction

📝 **A transaction** is a bracket around a group of database changes that either all happen or none do. You met the full story in [Transactions & ACID](/guides/transactions-and-acid) - atomicity, consistency, isolation, durability. JPA leans on that idea completely: essentially all your persistence work runs *inside* a transaction, and the persistence context itself is scoped to one. Open a transaction, do your work, commit. If anything goes wrong, roll back and it's as if none of it happened.

The raw JPA shape is three calls bracketing your work:

```java
EntityManager em = emf.createEntityManager();
EntityTransaction tx = em.getTransaction();
try {
    tx.begin();                          // open the bracket

    Author author = new Author("Ursula K. Le Guin");
    em.persist(author);                  // now managed

    tx.commit();                         // close it - changes become permanent
} catch (RuntimeException e) {
    if (tx.isActive()) tx.rollback();    // something broke - undo everything
    throw e;
} finally {
    em.close();
}
```
*What just happened:* `tx.begin()` opened a database transaction and the persistence context that rides along with it. We created an `Author` and called `persist`, which made it *managed* (Phase 3's term) - but note, no `INSERT` has run yet. `tx.commit()` is the moment Hibernate flushes the pending work to the database and the transaction makes it durable. The `try/catch/finally` is the same disciplined shape you saw in [Java's try-with-resources phase](/guides/java-from-zero) - on any failure we roll back so we never leave a half-finished change behind, and we always `close()` the `EntityManager`.

💡 You will rarely type those three lines in real life. In a Spring application you annotate a method with `@Transactional` and Spring writes the `begin`/`commit`/`rollback` for you - opening the transaction before your method runs, committing if it returns normally, rolling back if it throws. That's the same machinery you see above, hidden one layer down. We show the raw version so that when Spring's annotation does something surprising, you know exactly what it's doing on your behalf.

## The unit of work

Here's the mental shift that makes Hibernate click. A naive ORM might send one SQL statement every time you touch an object: set a field, fire an `UPDATE`; add to a list, fire an `INSERT`. Hibernate deliberately does *not* work that way.

📝 **The unit of work** is Hibernate's core operating model: across one transaction it *collects* all your intended changes in the persistence context and writes them to the database as a single coordinated batch at the right moment - not one statement per method call. You don't issue SQL. You change managed objects in memory - set a title, add a review, delete a book - and Hibernate works out the minimal set of `INSERT`, `UPDATE`, and `DELETE` statements needed to make the database match what you did, then sends them together.

This is why "I never called `save`" is the wrong question. You're not telling Hibernate *which statement* to run; you're telling it *what the world should look like*, and it reverse-engineers the SQL. Two consequences fall straight out of this and define the rest of the phase: Hibernate must somehow *detect* what you changed (dirty checking), and it must pick a *moment* to send the SQL (flush).

## Dirty checking - the "magic" save

This is the one that trips up every newcomer, so let's hit it head-on.

📝 **Dirty checking** is Hibernate noticing, all by itself, that a managed entity's fields have changed since you loaded it - and issuing an `UPDATE` to match, *with no `save`, `merge`, or `persist` call from you*. A "dirty" object is one whose current state differs from what was loaded.

Watch it happen. We load a `Book`, change one field, and commit:

```java
tx.begin();

Book book = em.find(Book.class, 1L);   // SELECT runs; book is now managed
book.setTitle("A Wizard of Earthsea (Revised)");   // just a setter - no em call

tx.commit();   // <-- an UPDATE appears here, out of nowhere
```
```sql
select b.id, b.title, b.author_id from book b where b.id=?
update book set title=?, author_id=? where id=?
```
*What just happened:* `em.find` ran the `SELECT` and handed back a *managed* `Book`. We called an ordinary Java setter - no `EntityManager` method anywhere near it. Yet at `tx.commit()` Hibernate emitted an `UPDATE`. It noticed the title differed from what it had loaded and synced the database to match. There is genuinely no `save` call because there is no `save` method in JPA's vocabulary at all - managed objects update themselves. (`em.persist` is for brand-new entities; `em.merge` is for *detached* ones, which we'll get to.)

So *why* does this work? Here's the mechanism, and it's pleasingly mundane. 💡 When Hibernate loads an entity into the persistence context, it keeps a private **snapshot** of every field value at load time. At flush, it walks each managed entity and compares the current field values against that snapshot, field by field. Any entity whose values drifted is dirty, and Hibernate generates an `UPDATE` for exactly the changed columns (or all of them, depending on configuration). No change, no snapshot mismatch, no SQL:

```java
tx.begin();

Book book = em.find(Book.class, 1L);   // loaded; snapshot taken
String title = book.getTitle();        // only reading - no field changed

tx.commit();   // no UPDATE - nothing differs from the snapshot
```
```sql
select b.id, b.title, b.author_id from book b where b.id=?
```
*What just happened:* Same `find`, same managed entity, but this time we only *read* a field. At commit the snapshot still matches the live object, so dirty checking finds nothing dirty and emits no `UPDATE`. This is the flip side of the magic: Hibernate won't write what didn't change, so an accidental no-op transaction costs you only the `SELECT`. The snapshot is the persistence context (Phase 3) earning its keep - the identity map gives you one object per row, and the snapshot beside it is how Hibernate knows when that object has drifted from the database.

⚠️ One trap that follows directly: dirty checking only works on **managed** entities. A `Book` you `new`ed up yourself but never loaded or `persist`ed is *transient* - Hibernate has no snapshot of it and no idea it exists, so changing its fields does nothing. The magic is a property of the persistence context, not of the object.

## Flush - sending the SQL vs making it permanent

We've said "at commit, the SQL appears." The precise term for that send is **flush**, and pulling it apart from **commit** clears up a whole category of confusion.

📝 **Flush** is the act of synchronizing the persistence context's pending changes *to the database* - translating your in-memory dirty objects into the actual `INSERT`/`UPDATE`/`DELETE` statements and sending them over the connection. Crucially, flush happens *inside* the transaction; the statements are now visible to your own session but not yet permanent.

⚠️ **Flush is not commit.** This is the distinction to nail:

- **Flush** sends the SQL within the current transaction. The database has executed your statements, but they're still inside the open transaction and can still be rolled back.
- **Commit** ends the transaction and makes everything durable and visible to *other* sessions. (Commit always flushes first - you can't commit changes you haven't sent.)

Think of flush as "push my changes to the database's working memory" and commit as "and now make them official."

You almost never call `flush()` yourself, because Hibernate flushes automatically at two moments: **at commit** (so nothing is lost), and **before a query runs** (so your query sees your own pending changes). That second one is the order-of-operations gotcha worth seeing:

```java
tx.begin();

Book book = em.find(Book.class, 1L);
book.setTitle("New Title");        // dirty in memory; no SQL sent yet

// This query forces a flush FIRST, so it doesn't read stale data
List<Book> hits = em.createQuery(
        "select b from Book b where b.title = 'New Title'", Book.class)
    .getResultList();

tx.commit();
```
```sql
select b.id, b.title, b.author_id from book b where b.id=?
update book set title=?, author_id=? where id=?     -- auto-flush before the query
select b.id, b.title, b.author_id from book b where b.title='New Title'
```
*What just happened:* We made the `Book` dirty in memory, then ran a JPQL query searching for the *new* title. Hibernate knows the query hits the `book` table and that you have a pending change to a `book`, so it auto-flushed the `UPDATE` *before* running the `SELECT` - otherwise the query would have read the old title from the database and missed your own change. The order in the SQL log tells the story: the `UPDATE` jumps ahead of the query. This is Hibernate keeping its promise that your queries see a consistent picture including your uncommitted work.

💡 You *can* force a flush early with `em.flush()` - useful when you need a database-generated ID immediately, or want a constraint violation to surface now rather than at commit. But reach for it rarely; trust the automatic flush points unless you have a concrete reason not to.

## Rollback, detachment, and the rule that saves you

The other half of "all or nothing" is rollback. If anything throws and you call `tx.rollback()`, the transaction unwinds: every `INSERT`, `UPDATE`, and `DELETE` it sent is undone, and the database is exactly as it was before `begin()`. The flushed SQL never becomes permanent - that's the whole power of flush-before-commit. Nothing you did inside that transaction survives.

But there's a subtler consequence, and it's the source of a classic bug. When the transaction ends and the `EntityManager` closes, every entity that was managed becomes **detached** (Phase 3's fourth state). A detached entity is a perfectly normal Java object holding data - but Hibernate is no longer watching it. There's no snapshot, no persistence context, no dirty checking. It's just a POJO now.

⚠️ Here's the bug everyone writes once. You load an entity, the transaction closes, and *then* you change a field, expecting the database to update because "dirty checking does that." It doesn't - the entity is detached, and your setter mutates an object Hibernate has forgotten:

```java
Book book;

tx.begin();
book = em.find(Book.class, 1L);   // managed
tx.commit();                      // transaction ends -> book is now DETACHED
em.close();

book.setTitle("This change goes nowhere");   // mutating a detached object
// No transaction, no persistence context, no dirty checking. The DB never hears about it.
```
*What just happened:* The `find` and the load happened inside the transaction, but the moment we committed and closed, `book` detached. The setter ran fine - it's just Java - but there was no managed context to notice the change and no transaction to flush it into. The title in the database is untouched. To actually persist a change to a detached entity you'd have to re-attach it (`em.merge(book)` inside a fresh transaction), which is a different, heavier operation than the effortless dirty checking you get on managed objects.

💡 The rule that prevents this entire class of bug, and the one habit to carry out of this phase: **load, modify, and commit all inside one transaction.** Open the transaction, `find` the entity (now managed), change its fields, commit - and let dirty checking do the rest while the object is still being watched. Spring's `@Transactional` makes this natural by wrapping a whole service method in one transaction, so your load and your mutations stay safely on the managed side of the line.

## Recap

1. **Everything runs in a transaction** - `begin`/`commit`/`rollback` bracket your work, and the persistence context is scoped to it. Spring's `@Transactional` writes those calls for you; ACID is the foundation underneath ([Transactions & ACID](/guides/transactions-and-acid)).
2. **Hibernate works as a unit of work** - you change managed objects in memory describing what the world should look like; Hibernate figures out the minimal `INSERT`/`UPDATE`/`DELETE` batch and sends it at the right moment, not one statement per call.
3. **Dirty checking is the "magic" save** - change a field on a *managed* entity and Hibernate emits an `UPDATE` at flush with no `save`/`merge`/`persist`, because it snapshots loaded state and diffs against it. JPA has no `save` method; managed objects update themselves.
4. **Flush ≠ commit** - flush sends the SQL *within* the transaction (auto-fired at commit and before queries); commit makes it permanent and visible to others. Use `em.flush()` only when you need IDs or errors early.
5. **Rollback undoes everything; closing detaches** - after the transaction/context ends, entities are detached and no longer dirty-checked. Mutating a detached object silently does nothing.
6. **The rule:** load → modify → commit, all inside one transaction, so your changes happen while the entity is still managed.

With the persistence context, the unit of work, and dirty checking in hand, you understand the engine room. Next we start connecting entities to each other - and that's where the SQL gets genuinely interesting.

## Quick check

Make sure the one idea that defines this phase stuck - why a field change becomes an `UPDATE` with no `save` call:

```quiz
[
  {
    "q": "You load a Book with em.find inside a transaction, call book.setTitle(\"New\"), and commit - never calling save, merge, or persist. What happens?",
    "choices": [
      "Hibernate detects the changed field via dirty checking and issues an UPDATE at commit",
      "Nothing - without a save call the change stays only in memory",
      "It throws an exception because you must call persist to write changes",
      "The change is written immediately when setTitle runs, before commit"
    ],
    "answer": 0,
    "explain": "The Book is managed, so Hibernate kept a snapshot at load time. At flush (which commit triggers) it diffs the live object against the snapshot, sees the title changed, and emits an UPDATE. There's no save method in JPA - managed entities update themselves."
  },
  {
    "q": "What's the difference between a flush and a commit?",
    "choices": [
      "Flush sends the pending SQL within the transaction; commit ends the transaction and makes the changes permanent",
      "They're the same thing - flush is just an older name for commit",
      "Flush makes changes permanent; commit only sends them to the database's cache",
      "Flush rolls back changes, while commit saves them"
    ],
    "answer": 0,
    "explain": "Flush translates your dirty in-memory objects into SQL and sends it inside the open transaction (it can still be rolled back). Commit ends the transaction, making everything durable and visible to other sessions. Commit always flushes first, but a flush alone is not permanent."
  },
  {
    "q": "After tx.commit() and em.close(), you call book.setTitle(\"Changed\"). Why doesn't the database update?",
    "choices": [
      "Once the transaction and EntityManager close, the entity is detached - no persistence context is watching it, so dirty checking doesn't apply",
      "The setter silently failed because the object is now read-only",
      "It does update - dirty checking works on any entity at any time",
      "Hibernate batches the change and applies it on the next find call"
    ],
    "answer": 0,
    "explain": "Closing the transaction/EntityManager detaches the entity. A detached entity is an ordinary Java object with no snapshot and no managing context, so the setter just mutates memory. The rule that avoids this: load, modify, and commit all inside one transaction."
  }
]
```


---

# Mapping Relationships

So far each entity has been an island: one class, one table, a handful of columns. Real data isn't like
that. An author writes many books; a book collects many reviews; a book wears many tags. The whole reason
you reach for a relational database is that things *relate* - and the whole reason ORM relationships feel
fiddly is that the database and your objects describe those relations in two completely different
languages. This phase is the translation guide.

The mental model to hold onto before any annotation: **a relationship is stored as a foreign key in the
database, but it shows up as a reference (or a collection) in your objects, and JPA's job is to keep those
two views in sync.** Get that one sentence into your bones and every annotation below is just spelling out
*which* table holds the key and *which* field points where.

## The mental model: two views of the same link

📝 In the database, "this book was written by that author" is a single column: `book.author_id` holds the
`id` of a row in the `author` table. That's it. A foreign key is just a column whose value matches a
primary key somewhere else. (If that's fuzzy, [Relationships &
Keys](/guides/relationships-and-keys) is the prerequisite, and [SQL Joins
Explained](/guides/sql-joins-explained) shows how you stitch the rows back together.)

In your Java objects, that *same* link looks different. A `Book` object holds a reference:
`book.getAuthor()` hands you the actual `Author` object. And an `Author` can hold a `List<Book>` - the
collection view of the same relationship, walked from the other end. One foreign key column, two object-side
shapes depending on which way you're looking.

Here's our domain for the rest of the guide:

```mermaid
erDiagram
    AUTHOR ||--o{ BOOK : writes
    BOOK ||--o{ REVIEW : receives
    BOOK }o--o{ TAG : "tagged with"
```

*What just happened:* one author writes many books (`1 - *`), one book receives many reviews (`1 - *`), and
books and tags form a many-to-many (`* - *`). Three relationship shapes, and JPA has an annotation for each.
We'll do them in order of how often you'll write them - most common first.

## `@ManyToOne`: the foreign-key side

Start here, because this is the side that actually holds the foreign key, and it's the one you'll write
most. Many `Book`s point to one `Author`, so on `Book` we add a reference to its author.

```java
@Entity
public class Book {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    private String title;

    @ManyToOne                          // many books → one author
    @JoinColumn(name = "author_id")     // the FK column lives in THIS table
    private Author author;

    // getters and setters
}
```

*What just happened:* `@ManyToOne` tells JPA "many of these belong to one of those." `@JoinColumn(name =
"author_id")` names the foreign-key column that JPA will manage in the `book` table - the column that stores
which author this book belongs to. The `author` field isn't an `id`; it's a whole `Author` object, and
Hibernate handles loading it from that `author_id` value behind the scenes.

The table this implies is exactly what you'd write by hand:

```sql
CREATE TABLE book (
    id        BIGINT GENERATED ALWAYS AS IDENTITY PRIMARY KEY,
    title     VARCHAR(255),
    author_id BIGINT REFERENCES author(id)   -- the foreign key
);
```

*What just happened:* `author_id` is an ordinary column on `book` holding the id of a row in `author`. When
you call `book.setAuthor(someAuthor)` and the transaction commits, Hibernate writes `someAuthor`'s id into
that column. The object reference *is* the foreign key - JPA just lets you work with the object instead of
the raw number.

💡 If you only ever need to go from a book to its author, **you can stop here.** A `@ManyToOne` on its own is
a complete, working relationship. You don't have to add the collection on the other side unless you actually
need to walk it that direction.

## `@OneToMany` and the bidirectional link

Sometimes you *do* want the other direction: given an `Author`, list their `Book`s. That's the inverse
view - the same foreign key, walked from the one-side back to the many-side.

```java
@Entity
public class Author {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    private String name;

    @OneToMany(mappedBy = "author")     // inverse side: "look at Book.author"
    private List<Book> books = new ArrayList<>();

    // getters and setters
}
```

*What just happened:* `@OneToMany(mappedBy = "author")` gives `Author` a collection of its books. The
crucial word is `mappedBy = "author"`: it tells JPA *this side does not own the foreign key - the
relationship is already mapped by the `author` field over on `Book`.* No new column is created for the
author side. There's still exactly one foreign key in the database (`book.author_id`); we've just added a
second way to read it.

📝 **Owning side vs inverse side - the concept that explains everything.** A bidirectional relationship has
two ends but only *one* foreign key column. The end that holds the FK is the **owning side** - for us,
`Book` with its `@ManyToOne` and `@JoinColumn`. The other end is the **inverse side**, marked with
`mappedBy`. Here's the rule that trips up everyone: **Hibernate only looks at the owning side when deciding
what to write to the database.** The inverse collection is read-only as far as persistence goes. Change the
owning side, the FK changes. Change *only* the inverse side, and nothing gets written.

⚠️ **The #1 relationship bug: setting only one side.** Because the owning side is what persists, this looks
right and silently does the wrong thing:

```java
Author author = new Author("Ursula K. Le Guin");
Book book = new Book("A Wizard of Earthsea");

author.getBooks().add(book);   // touched the INVERSE side only...
// book.setAuthor(author) was never called - the owning side is still null!

em.persist(author);
em.persist(book);
// commit → book.author_id is NULL. The link you "made" was never saved.
```

*What just happened:* you added the book to the author's in-memory list, but `book.author` - the side that
owns the foreign key - was never set. Hibernate persists based on the owning side, sees `null`, and writes
`null` into `author_id`. Your objects look connected in memory; the database disagrees. And the next time
you load the author fresh, the books are gone.

The fix is to **keep both sides in sync with a helper method**, so you can never set one without the other:

```java
@Entity
public class Author {
    // ... fields as above ...

    public void addBook(Book book) {
        books.add(book);            // update the inverse collection (for in-memory consistency)
        book.setAuthor(this);       // update the OWNING side (this is what gets persisted)
    }
}
```

*What just happened:* `addBook` does both halves of the link in one call. `book.setAuthor(this)` sets the
owning side - that's the line that actually saves the foreign key. `books.add(book)` keeps your in-memory
object graph accurate so the author's list reflects reality without a database round-trip. Calling
`author.addBook(book)` now does the right thing every time. 💡 Make the raw setters less tempting and route
all linking through helpers like this; it's the single highest-leverage habit for avoiding relationship
bugs.

## `@ManyToMany`: a link with no obvious home

Books and tags are many-to-many: a book has many tags, a tag labels many books. Neither table can hold the
foreign key (which row would it point to?), so the database uses a **join table** - a third table whose only
job is to pair ids.

```java
@Entity
public class Book {
    // ... id, title, author as before ...

    @ManyToMany
    @JoinTable(
        name = "book_tag",                              // the join table
        joinColumns = @JoinColumn(name = "book_id"),    // FK back to THIS entity (Book)
        inverseJoinColumns = @JoinColumn(name = "tag_id") // FK to the other entity (Tag)
    )
    private Set<Tag> tags = new HashSet<>();
}
```

*What just happened:* `@ManyToMany` plus `@JoinTable` tells Hibernate to manage a `book_tag` table with two
foreign keys - `book_id` and `tag_id` - where each row means "this book wears this tag." `Book` is the
**owning side** here (it declares the `@JoinTable`); `Tag` would carry `mappedBy = "tags"` if you wanted the
relationship navigable from that end too. The same owning-side rule applies: add to the owning collection to
persist the link.

The join table is exactly:

```sql
CREATE TABLE book_tag (
    book_id BIGINT REFERENCES book(id),
    tag_id  BIGINT REFERENCES tag(id),
    PRIMARY KEY (book_id, tag_id)
);
```

⚠️ **When the link itself has data, `@ManyToMany` is the wrong tool.** A plain `@ManyToMany` join table can
hold *only* the two foreign keys. The moment you need to record anything *about* the relationship - when the
tag was applied, who applied it, a relevance score - you can't, because there's nowhere to put it. The fix
is to promote the join table to a real entity (say `BookTag`) with its own `@Id` and two `@ManyToOne`s back
to `Book` and `Tag`. 💡 Rule of thumb: pure association → `@ManyToMany`; association *with attributes* → a
join entity.

## Cascade and orphan removal

One more pair of settings, because they decide what happens to the *related* rows when you save or delete.

📝 **Cascade** propagates an operation from a parent to its children. With `cascade = CascadeType.ALL` on
`Author`'s `books`, calling `em.persist(author)` also persists every `Book` in the collection, and removing
the author removes its books - you don't have to persist or delete each child by hand.

📝 **Orphan removal** (`orphanRemoval = true`) goes further: if you *remove a child from the collection*,
Hibernate deletes that child's row. The child became an "orphan" - no longer referenced by its parent - so
it's deleted rather than left dangling with a `null` foreign key.

```java
@Entity
public class Book {
    // ... id, title, author ...

    @OneToMany(
        mappedBy = "book",
        cascade = CascadeType.ALL,    // persist/remove a Book → persist/remove its Reviews
        orphanRemoval = true          // remove a Review from the list → DELETE that review
    )
    private List<Review> reviews = new ArrayList<>();
}
```

*What just happened:* a `Review` has no life of its own - it exists only as part of a `Book`, so cascading
all operations and removing orphans is exactly right here. Save the book, its reviews save with it. Delete
the book, its reviews go too. Pull a review out of `book.getReviews()`, and that review's row is deleted on
commit. This is the model where aggressive cascade *belongs*: a true parent-child where the child can't
outlive the parent.

⚠️ **Don't cascade like this on a `@ManyToOne`.** It's tempting to slap `cascade = CascadeType.ALL` on
`Book`'s `author` field, but think about what `REMOVE` would mean: deleting one book would delete its
author - and with the author gone, potentially every *other* book by that author. The `@ManyToOne` points to
something shared and longer-lived; you don't own the author, you just reference it. Cascade belongs on the
side that owns the children's lifecycle (the parent's `@OneToMany`), almost never on the `@ManyToOne` back to
a shared parent.

💡 The throughline of this entire phase: **model the owning side carefully, and be deliberate about
cascade.** Nearly every relationship bug you'll hit in practice is one of two mistakes - setting only the
inverse side so the FK never gets written, or a cascade that deletes more than you meant. Get those two
right and relationships stop being scary. (Fetching - when these related objects actually load from the
database - is its own large topic, and it's next.)

## Recap

1. A relationship is **one foreign key** in the database but **a reference or collection** in your objects;
   JPA's job is to translate between the two views.
2. **`@ManyToOne` + `@JoinColumn`** is the owning side - it holds the FK and is the simplest, most common
   mapping. On its own it's a complete relationship.
3. **`@OneToMany(mappedBy = "...")`** adds the inverse collection without a new column. Hibernate persists
   based on the **owning side only**, so a bidirectional link needs a **helper method that sets both sides**.
4. **`@ManyToMany` + `@JoinTable`** maps a pair of entities through a join table; promote it to a **join
   entity** the moment the link needs its own data.
5. **Cascade** propagates persist/remove to children and **`orphanRemoval`** deletes children pulled from
   the collection - right for true parent-child (Book→Review), wrong on a `@ManyToOne` to a shared parent.
6. Most relationship bugs are **owning-side** mistakes (only the inverse side set) or **cascade** mistakes
   (deleting more than intended). Model the owning side deliberately.

## Quick check

Lock in the two ideas that cause the most real-world relationship bugs:

```quiz
[
  {
    "q": "In a bidirectional Author–Book relationship (Book has @ManyToOne author, Author has @OneToMany(mappedBy=\"author\")), which side does Hibernate use to decide what foreign key to write?",
    "choices": [
      "The owning side - Book's @ManyToOne, which holds the @JoinColumn",
      "The inverse side - Author's @OneToMany collection",
      "Both sides equally; it merges them",
      "Whichever side you modified most recently"
    ],
    "answer": 0,
    "explain": "The owning side is the one with the foreign key (the @ManyToOne with @JoinColumn). Hibernate persists based on the owning side only; the mappedBy collection is the read-only inverse view. That's why setting only the collection silently fails to save the link."
  },
  {
    "q": "You write `author.getBooks().add(book)` but never call `book.setAuthor(author)`, then persist and commit. What ends up in book.author_id?",
    "choices": [
      "NULL - you only set the inverse side, so the owning FK was never written",
      "The author's id - adding to the collection is enough",
      "A constraint error prevents the commit",
      "The id of the most recently saved author"
    ],
    "answer": 0,
    "explain": "You updated the inverse collection but not the owning side (book.author). Hibernate persists from the owning side, sees null, and writes NULL into author_id. This is the #1 bidirectional bug - fix it with a helper method that sets both sides at once."
  },
  {
    "q": "Where is `cascade = CascadeType.ALL` (with orphanRemoval) appropriate, and where is it dangerous?",
    "choices": [
      "Appropriate on a parent's @OneToMany to children it owns (Book→Review); dangerous on a @ManyToOne to a shared parent (Book→Author)",
      "Appropriate everywhere - it just saves typing",
      "Appropriate on @ManyToOne; dangerous on @OneToMany",
      "Only ever appropriate on @ManyToMany join tables"
    ],
    "answer": 0,
    "explain": "Cascade belongs where one entity truly owns another's lifecycle - a Book owns its Reviews, so persist/remove should propagate. On a @ManyToOne to a shared Author, cascading REMOVE would delete the author (and potentially their other books) when you delete one book. You reference the author; you don't own it."
  }
]
```


---

# Lazy vs Eager Fetching & the N+1 Problem

You mapped the relationships in Phase 5: an `Author` has many `Book`s, a `Book` has many `Review`s. The
foreign keys are in place, the object graph navigates cleanly. Here's the question nobody asks until it's
too late: when you load one `Author`, **does Hibernate also load their books? Their books' reviews? The
whole tree?**

The answer to that question is the difference between an app that returns in 8 milliseconds and one that
returns in 8 seconds. This is the single most important performance lesson in the entire guide, and almost
every "Hibernate is slow" complaint on the internet traces back to getting it wrong without noticing. So
slow down here. We're going to make it visceral - you're going to *see* the flood of queries - because once
you've watched it happen, you'll never write a blind loop over a collection again.

## The mental model: a switch with two settings

📝 Every association in JPA has a **fetch strategy** - a setting that answers "when do you load this?" There
are exactly two settings:

- **EAGER** - load this association *immediately*, in the same breath as the parent. Load an `Author`, and
  their `Book`s come along for the ride whether you asked for them or not.
- **LAZY** - *don't* load it yet. Hibernate puts a stand-in object in the field - a **proxy** - and only runs
  the real query the moment you actually touch the data (call `getBooks()` and read it).

Think of it like ordering at a restaurant. EAGER is the waiter bringing the appetizer, main, dessert, and
coffee all at once before you've decided what you want. LAZY is bringing each course only when you ask for
it. Both can be wrong: EAGER hauls food you'll never eat; LAZY makes a separate trip to the kitchen for
every single item.

Here's the part that bites everyone, so commit it to memory. The **JPA defaults are not uniform**:

| Annotation | Default fetch |
|------------|---------------|
| `@ManyToOne` | **EAGER** |
| `@OneToOne` | **EAGER** |
| `@OneToMany` | LAZY |
| `@ManyToMany` | LAZY |

📝 The pattern: the *to-one* sides are eager by default, the *to-many* collections are lazy. This split is a
trap. Your `Book.author` (`@ManyToOne`) is silently eager - load 50 books and you may quietly load their
authors too, even on pages that never show the author.

⚠️ **Recommendation: make everything LAZY, then fetch what you need explicitly.** Eager-by-default is the
gift that keeps on taking - it loads data you didn't ask for, on code paths you forgot about, and you only
find out when production slows down. Turn the to-one defaults off:

```java
@Entity
public class Book {

    @ManyToOne(fetch = FetchType.LAZY)        // override the EAGER default
    @JoinColumn(name = "author_id")
    private Author author;

    @OneToMany(mappedBy = "book", fetch = FetchType.LAZY)   // already lazy, stated for clarity
    private List<Review> reviews = new ArrayList<>();
}
```

*What just happened:* we forced `Book.author` to LAZY, overriding JPA's eager default for `@ManyToOne`. Now
loading a `Book` loads exactly the `book` row - no surprise `SELECT` for the author tagging along. The
reviews were already lazy, but writing it out makes the intent obvious to the next person (and to you, six
months from now). The rule of thumb that will save you: **lazy everywhere, fetch deliberately.**

### A lazy collection is a proxy, not your data

When a collection is lazy, the field doesn't hold your `Book`s. It holds a Hibernate stand-in:

```java
EntityManager em = emf.createEntityManager();
Author author = em.find(Author.class, 1L);

System.out.println(author.getBooks().getClass().getName());
// org.hibernate.collection.spi.PersistentBag   ← not ArrayList!

author.getBooks().size();   // touching it NOW triggers the SELECT
```
```sql
select a.id, a.name from author a where a.id = 1
-- ...later, the moment you call size():
select b.id, b.title, b.author_id from book b where b.author_id = 1
```

*What just happened:* `find` ran *one* query for the author. The `books` field came back as a Hibernate
`PersistentBag` - a proxy that knows how to load the books but hasn't yet. Nothing hit the `book` table until
`size()` forced it. That second query is the proxy "waking up." Lazy isn't *if* you pay for the books - it's
*when*. And that timing is exactly what causes the next two problems.

## `LazyInitializationException`: the classic crash

📝 A lazy proxy can only wake up while its **persistence context is still open**. Remember from
[Phase 3](03-entitymanager-and-persistence-context.md): when the `EntityManager` closes, every entity it
loaded becomes **detached** - nobody's watching it, and there's no open context to run a query through. So if
you touch a lazy association *after* the context closed, the proxy has no way to load its data, and Hibernate
throws.

This is the most famous beginner crash in all of Hibernate. The shape is always the same - load in one
layer, touch in another:

```java
// --- service / repository layer: context opens and CLOSES here ---
public Author loadAuthor(Long id) {
    EntityManager em = emf.createEntityManager();
    Author author = em.find(Author.class, id);   // books NOT fetched (lazy)
    em.close();                                    // ← context closed; author is now DETACHED
    return author;
}

// --- controller / view layer: context is long gone ---
Author author = loadAuthor(1L);
for (Book book : author.getBooks()) {   // 💥 touching the lazy proxy now...
    System.out.println(book.getTitle());
}
```
```console
org.hibernate.LazyInitializationException: failed to lazily initialize a
collection of role: com.example.Author.books: could not initialize proxy - no Session
```

*What just happened:* `loadAuthor` opened a context, loaded the author *without* the books, then closed the
context - detaching the author. Back in the controller, `author.getBooks()` asks the lazy proxy to load, but
its context is gone. No open session, no query, no data - exception. 💡 The fix isn't "make it eager" (that
just trades this crash for the N+1 you're about to meet). The real fix is to **fetch the books while the
context is still open**, which is the whole rest of this phase. Tie this back to Phase 3's rule: *a detached
entity can't lazy-load.* This is that rule biting.

## The N+1 problem: the main event

This is the one. The performance killer that ships to production looking completely innocent. Watch closely.

You load all your authors - one clean query - and then loop over them to print each author's book count. The
collection is lazy, so each `getBooks()` call wakes up its proxy:

```java
List<Author> authors = em.createQuery("select a from Author a", Author.class)
                         .getResultList();           // query #1: load the authors

for (Author author : authors) {
    System.out.println(author.getName() + ": " + author.getBooks().size());
    //                                            ↑ each iteration triggers ANOTHER query
}
```

Looks harmless. It is a disaster. Here is the SQL Hibernate actually emits with, say, 100 authors:

```sql
select a.id, a.name from author a;                              -- the "1": one query for all authors

select b.id, b.title, b.author_id from book b where b.author_id = 1;    -- the "N" begins...
select b.id, b.title, b.author_id from book b where b.author_id = 2;
select b.id, b.title, b.author_id from book b where b.author_id = 3;
select b.id, b.title, b.author_id from book b where b.author_id = 4;
-- ... one more SELECT for every single author ...
select b.id, b.title, b.author_id from book b where b.author_id = 99;
select b.id, b.title, b.author_id from book b where b.author_id = 100;
```

*What just happened:* **1 query to load the authors, then N more - one per author - to load each one's
books.** That's `1 + N` queries. 100 authors = **101 queries**. A thousand authors = 1001. Every one is a
separate round trip to the database: network hop, parse, plan, execute, return. Individually they're fast;
multiplied by N they're a stampede. This is the **N+1 problem**, and it is the number-one reason ORMs get
blamed for being slow.

⚠️ The cruelty of N+1 is that it's *invisible in the code*. The Java reads like a normal loop. It works
perfectly with 3 authors in your test database. Then it meets 5,000 authors in production and falls over - 
and nobody changed a line. The query count grows with your data, not with your code, so it sails through
code review and load-tests-you-didn't-run. You have to *watch the SQL* to even know it's there. Which is
exactly why the discipline at the end of this phase matters.

## Fixing it: load the tree in one query

The cure for N+1 is to tell Hibernate up front: *I'm going to need the books, so fetch them together.* You
have three tools.

### 1. `JOIN FETCH` - the workhorse

In JPQL, `join fetch` says "join to this association and load it into the result, in the same query":

```java
List<Author> authors = em.createQuery(
        "select distinct a from Author a join fetch a.books", Author.class)
        .getResultList();

for (Author author : authors) {
    System.out.println(author.getName() + ": " + author.getBooks().size());
    // no extra queries - the books are already loaded
}
```
```sql
select distinct a.id, a.name, b.id, b.title, b.author_id
from author a
join book b on b.author_id = a.id;
```

*What just happened:* **one query** now does the whole job. `join fetch a.books` told Hibernate to join
`author` to `book` and hydrate the `books` collection right there in the result set. The loop runs without
emitting a single extra `SELECT` - the books arrived with their authors. 101 queries collapsed to 1. The
`distinct` keyword de-duplicates authors in the Java result (a join repeats the author row once per book);
it's almost always what you want with a `join fetch` on a collection.

⚠️ **Two `JOIN FETCH` caveats that bite hard.** First: **`JOIN FETCH` a collection + pagination don't mix.**
If you add `setMaxResults`/`setFirstResult` to a query that fetches a collection, Hibernate can't apply the
limit in SQL (the join multiplied your rows), so it pulls *everything* into memory and paginates there - 
quietly defeating the point and risking an OOM on a large table. Second: **you can't `JOIN FETCH` two
collections at once** (e.g. an author's books *and* each book's reviews in one query) - that's a cartesian
product, and Hibernate throws `MultipleBagFetchException`. Fetch one collection per query, or use batch
fetching (below) for the second.

### 2. `@EntityGraph` - the declarative alternative

If you'd rather not write JPQL - say you're using Spring Data repositories - an **entity graph** declares
which associations to fetch eagerly *for this one call*, without touching the entity's default mapping:

```java
@EntityGraph(attributePaths = "books")
List<Author> findAll();    // Spring Data: this findAll fetches books in one query
```

*What just happened:* `@EntityGraph(attributePaths = "books")` tells Hibernate "for this query, treat `books`
as eager - load it with the author." It produces the same single join query as `JOIN FETCH`, but you express
*what to load* declaratively instead of hand-writing the join. Same fix, different syntax - reach for
whichever fits your codebase. The key idea is identical: **the fetch decision belongs to the use-case, not
the mapping.**

### 3. Batch fetching - collapse N into a few

Sometimes you genuinely can't fetch up front - the books load lazily deep in some other code. Batch fetching
softens the blow: instead of one query *per* author, Hibernate loads the lazy collections in **batches**
using an `IN` clause.

```java
@Entity
public class Author {

    @OneToMany(mappedBy = "author")
    @BatchSize(size = 25)       // load up to 25 authors' books per query
    private List<Book> books = new ArrayList<>();
}
```
```sql
-- instead of 100 separate SELECTs, with batch size 25 you get 4:
select b.id, b.title, b.author_id from book b where b.author_id in (1,2,3, ... ,25);
select b.id, b.title, b.author_id from book b where b.author_id in (26,27, ... ,50);
select b.id, b.title, b.author_id from book b where b.author_id in (51,52, ... ,75);
select b.id, b.title, b.author_id from book b where b.author_id in (76,77, ... ,100);
```

*What just happened:* `@BatchSize(size = 25)` told Hibernate "when you have to wake up these lazy
collections, grab 25 at a time." The N+1's 100 follow-up queries became `ceil(100/25) = 4`. You can set this
globally with `hibernate.default_batch_fetch_size` instead of annotating each collection. 💡 Batch fetching
is the safety net for the lazy loads you can't restructure away - it won't beat a single `JOIN FETCH`, but
turning 101 queries into 5 is a massive win for one config line.

## The discipline: watch the SQL count

💡 Here's the throughline, and the habit that separates people who fight Hibernate from people who command
it: **default to lazy, fetch what each use-case needs explicitly, and always watch the number of queries you
emit.** N+1 doesn't announce itself. The only way to catch it is to *see* the SQL.

So make the SQL visible while you develop:

- Turn on `hibernate.show_sql=true` (and `format_sql=true`) and actually look at the console for a given
  page or endpoint. If one user action emits dozens of near-identical `SELECT`s, you've found an N+1.
- Better, add a **query counter** to your tests - a tool like Hibernate's `Statistics` or a library like
  datasource-proxy that asserts "this endpoint runs at most 3 queries." That turns N+1 from a thing you
  notice in prod into a test that fails in CI.

📝 The plain summary of this entire phase: **the #1 reason people say "Hibernate is slow" is really "N+1
that nobody noticed."** Hibernate isn't slow; a loop that secretly runs 500 queries is slow, and the ORM
just made it easy to write without seeing it. Counting your queries is how you stay on the fast side of that
line. (When a query you *did* write deliberately is the slow one, that's a different skill - measuring and
reading query plans - covered in [Why Is My Query Slow?](/guides/why-is-my-query-slow).)

## Recap

1. Every association has a **fetch strategy**: **EAGER** (load with the parent) or **LAZY** (load a proxy
   now, run the real query only when you touch it). JPA defaults: `@ManyToOne`/`@OneToOne` **EAGER**,
   `@OneToMany`/`@ManyToMany` LAZY.
2. ⚠️ Prefer **LAZY everywhere** and fetch explicitly per use-case - eager-by-default loads data you didn't
   ask for on code paths you forgot about.
3. **`LazyInitializationException`** happens when you touch a lazy association after the persistence context
   closed (the entity is detached). Fix it by fetching while the context is open - not by going eager.
4. The **N+1 problem**: load N parents in 1 query, then trigger 1 query per parent to load each one's lazy
   collection = `1 + N` queries. 100 authors → 101 queries. It's invisible in code and scales with your
   data, not your logic.
5. **Fixes:** `JOIN FETCH` (load the tree in one query), `@EntityGraph` (the declarative version), and
   `@BatchSize` / `hibernate.default_batch_fetch_size` (collapse N into a few `IN` queries). Mind the
   pagination and `MultipleBagFetchException` caveats with `JOIN FETCH`.
6. 💡 The discipline: **watch your query count.** Turn on `show_sql`, add a query counter to tests. Most
   "Hibernate is slow" is really an unnoticed N+1.

## Quick check

Lock in the one idea that wrecks more Hibernate apps than any other:

```quiz
[
  {
    "q": "You load 50 Authors with `select a from Author a`, then loop over them calling `author.getBooks().size()` on each (books is a LAZY @OneToMany). How many SQL queries does Hibernate run?",
    "choices": [
      "51 - one to load the authors, then one more per author to load each one's books (the N+1 problem)",
      "1 - Hibernate loads everything in a single query",
      "2 - one for authors, one for all books",
      "50 - one per author"
    ],
    "answer": 0,
    "explain": "This is the textbook N+1: 1 query for the authors, then N=50 lazy loads (one per author when you touch getBooks()) = 51 total. The loop looks innocent but each iteration wakes up a lazy proxy with its own SELECT."
  },
  {
    "q": "Your service loads an Author with `find` and closes the EntityManager, then a controller loops over `author.getBooks()` and crashes with LazyInitializationException. What's the correct fix?",
    "choices": [
      "Fetch the books while the context is still open (e.g. JOIN FETCH or an entity graph)",
      "Change the @OneToMany to FetchType.EAGER",
      "Catch the exception and ignore it",
      "Call em.close() later, in the controller"
    ],
    "answer": 0,
    "explain": "The crash happens because the author is detached (context closed) and a lazy proxy can't load with no open session. The right fix is to fetch the books deliberately while the context is open - via JOIN FETCH or @EntityGraph. Going EAGER 'fixes' the crash but reintroduces N+1 elsewhere and loads books even when you don't need them."
  },
  {
    "q": "Which statement about `JOIN FETCH` on a collection is a real caveat to watch out for?",
    "choices": [
      "Combining it with pagination (setMaxResults) forces Hibernate to paginate in memory, defeating the limit",
      "It always runs slower than lazy loading",
      "It can only be used with @ManyToOne, never collections",
      "It permanently changes the entity's default fetch type"
    ],
    "answer": 0,
    "explain": "JOIN FETCH on a collection plus setMaxResults can't apply the limit in SQL (the join multiplied the rows), so Hibernate loads everything and paginates in memory - slow and memory-risky. Separately, you can't JOIN FETCH two collections at once (MultipleBagFetchException). JOIN FETCH affects only the one query, not the mapping's default."
  }
]
```


---

# Querying: JPQL, Criteria & Native SQL

Up to now you've mostly found things by their id - `em.find(Book.class, 1L)`. That's fine when you already know exactly which row you want. But real applications ask open-ended questions: "all books by this author," "the five most-reviewed titles," "every book missing an ISBN." For those you need to *query*, and JPA hands you three tools for the job.

The mental model to carry through this whole phase: **JPA queries are layered, and you reach down a layer only when the one above can't reach.** JPQL covers the vast majority of what you'll write. The Criteria API exists for queries you have to *build* in code, piece by piece. Native SQL is the trapdoor for when you need something only your specific database can do. Same data, same entities - three levels of control, each more powerful and more verbose than the last.

We'll use the `Author`, `Book`, `Review` domain from [Phase 5](05-mapping-relationships.md): an author writes many books, a book collects many reviews.

## JPQL - query objects, not tables

📝 **JPQL** (Java Persistence Query Language) looks almost exactly like SQL, but with one profound difference: **it operates on your entities and their fields, not on tables and columns.** You write `Book`, not `books`. You write `b.author.name`, walking the object reference, not `JOIN author ON ...`. Hibernate reads your JPQL, looks at your entity mappings, and translates it into the real SQL for your database.

Here's "every book written by a given author":

```java
String jpql = "select b from Book b where b.author.name = :name";
List<Book> books = em.createQuery(jpql, Book.class)
                     .setParameter("name", "Ursula K. Le Guin")
                     .getResultList();
```

*What just happened:* `select b from Book b` says "give me `Book` entities, calling each one `b`." The `where b.author.name = :name` walks from a book *through its `author` reference* to the author's `name` field - no explicit join written, even though one is clearly needed. `createQuery(jpql, Book.class)` returns the books as fully-formed `Book` objects, ready to use.

Notice what you did *not* write: no table names, no `author_id`, no join condition. You described the question in terms of your objects, and Hibernate turned it into this:

```sql
select b.id, b.title, b.isbn_13, b.published_year, b.author_id
from book b
join author a on b.author_id = a.id
where a.name = ?
```

*What just happened:* Hibernate inferred the join from `b.author.name`. It knew, from your `@ManyToOne` mapping, that reaching `author.name` requires joining `book` to `author` on the foreign key - so it generated the `JOIN` for you. This is the whole point of JPQL: you think in objects, Hibernate writes the SQL.

💡 If you've never written raw SQL joins, the companion guide [SQL Joins, Finally Explained](/guides/sql-joins-explained) shows what Hibernate is doing under the hood here. JPQL hides the join, but the join is still happening - and understanding it is what lets you predict the SQL your queries generate (and why the next phase's N+1 problem bites).

## Parameters - and why you must use them

You saw `:name` above. That's a **named parameter** - a placeholder you fill in with `setParameter`. It is not a convenience. It is the single most important security habit in this entire guide.

⚠️ **Never, ever build a query by gluing user input into the string.** This looks innocent and is a gaping security hole:

```java
// DANGER - never do this
String userInput = request.getParameter("name");
String jpql = "select b from Book b where b.author.name = '" + userInput + "'";
List<Book> books = em.createQuery(jpql, Book.class).getResultList();
```

*What just happened:* you concatenated raw user input straight into the query text. If someone submits `' or '1'='1`, your `where` clause becomes always-true and leaks every book. Worse inputs can read or destroy data. This is **SQL injection**, and it's been at the top of the security-flaw lists for two decades. The string-building *is* the vulnerability.

The fix is parameters - and it's also less code:

```java
String jpql = "select b from Book b where b.author.name = :name";
TypedQuery<Book> query = em.createQuery(jpql, Book.class);
query.setParameter("name", userInput);
List<Book> books = query.getResultList();
```

*What just happened:* `:name` is a typed placeholder. `setParameter("name", userInput)` hands the value to Hibernate *separately* from the query text, so the database treats it strictly as data - never as part of the query structure. Even the `' or '1'='1` string is just searched for literally and finds nothing. Parameters aren't only safer; they also let the database reuse a query plan across calls.

💡 The `TypedQuery<Book>` is worth calling out. By passing `Book.class` you get a `TypedQuery<Book>`, so `getResultList()` returns `List<Book>` with no casting - the compiler checks the type for you. The untyped `createQuery(jpql)` hands back a raw `Query` and `List` of `Object`, which you then have to cast by hand. Prefer the typed form everywhere.

## Joins and projections - fetching only what you need

JPQL joins across relationships when you need to filter or select through them. And often you *don't* want whole entities back - you want a few fields. Pulling the full `Book` (and triggering loads of its reviews, its author) just to show a title and an author name is wasteful.

📝 A **projection** is a query that selects specific values instead of entire entities. The cleanest form maps those values straight into a small read-only object - a **DTO** (Data Transfer Object):

```java
public class BookSummary {
    private final String title;
    private final String authorName;

    public BookSummary(String title, String authorName) {
        this.title = title;
        this.authorName = authorName;
    }
    // getters
}
```

```java
String jpql = """
    select new com.example.BookSummary(b.title, b.author.name)
    from Book b
    join b.author a
    order by b.title""";
List<BookSummary> summaries = em.createQuery(jpql, BookSummary.class)
                                .getResultList();
```

*What just happened:* `select new com.example.BookSummary(...)` is JPQL's **constructor expression** - for each matching row, Hibernate calls that constructor with the two selected values and hands you a `BookSummary`. `join b.author a` joins through the `author` reference explicitly (here it reads clearly and lets you alias `a`). You get back exactly the two fields you asked for, in lightweight objects.

The generated SQL selects only those columns:

```sql
select b.title, a.name
from book b
join author a on b.author_id = a.id
order by b.title
```

*What just happened:* two columns, one join, nothing else. Compare that to loading full `Book` entities - every column, plus whatever lazy collections you might accidentally trip later. 💡 Projections are a real, measurable performance win for read-heavy screens: list views, dashboards, search results. When you only need a handful of fields to *display* something, project into a DTO instead of loading managed entities you'll never modify. (Loading full entities you don't need is closely tied to the N+1 problem in [Phase 6](06-fetching-and-n-plus-1.md).)

## The Criteria API - queries you build in code

JPQL is a string, and a string is great until the query is *dynamic*. Picture a search screen with optional filters: maybe an author name, maybe a minimum review count, maybe a published-year range - any combination, depending on what the user filled in. Building that JPQL by hand means concatenating fragments and tracking commas and `and`s, which drags you right back toward the injection-prone string-gluing you just learned to avoid.

📝 The **Criteria API** lets you build a query *programmatically* - as Java objects, method call by method call - instead of as a string. Because it's code, you can add a `where` clause inside an `if`, type-safely, with no string surgery.

```java
CriteriaBuilder cb = em.getCriteriaBuilder();
CriteriaQuery<Book> cq = cb.createQuery(Book.class);
Root<Book> book = cq.from(Book.class);

List<Predicate> filters = new ArrayList<>();
if (authorName != null) {
    filters.add(cb.equal(book.get("author").get("name"), authorName));
}
if (minYear != null) {
    filters.add(cb.greaterThanOrEqualTo(book.get("publishedYear"), minYear));
}

cq.select(book).where(cb.and(filters.toArray(new Predicate[0])));
List<Book> books = em.createQuery(cq).getResultList();
```

*What just happened:* `Root<Book> book` is the Criteria stand-in for `from Book b` - the thing you build expressions off. Each optional filter becomes a `Predicate` (a `where` condition) only if its input is present, collected into a list. `cb.and(...)` combines whatever predicates you gathered, and `cq.where(...)` applies them. Add an `if`, get a clause - no fragile string assembly, and user values still flow through as bound parameters automatically, so it stays injection-safe.

⚠️ **The Criteria API is verbose, and that verbosity is a real cost.** The same query as a one-line JPQL string is far easier to read and review. So use Criteria for what it's *good* at - queries genuinely assembled at runtime from optional pieces. For a fixed query whose shape never changes, JPQL reads better, and you should prefer it. Reaching for Criteria everywhere out of habit makes a codebase harder, not safer.

## Native SQL - the escape hatch

JPQL is deliberately portable: it speaks "entities," and Hibernate translates to whatever database you're on. But that portability means JPQL can only express what's common across databases. Sooner or later you'll need something it cannot say - a window function, a vendor-specific function, a hand-tuned query the optimizer needs.

📝 For that, there's **native SQL** via `createNativeQuery` - you write the real SQL, Hibernate runs it as-is, and can still map the results back onto entities or projections.

```java
String sql = """
    select * from book
    where id in (
        select book_id from review
        group by book_id
        having count(*) >= :minReviews
    )""";
List<Book> popular = em.createNativeQuery(sql, Book.class)
                       .setParameter("minReviews", 10)
                       .getResultList();
```

*What just happened:* this is genuine SQL against the `book` and `review` tables - table names, not entity names. Passing `Book.class` tells Hibernate to map each returned row into a managed `Book` entity, so even though you dropped to raw SQL, you get back the same objects JPQL would have given you. Parameters still work (`:minReviews`), so it stays injection-safe down here too. When you need a database feature JPQL doesn't expose, this is how you reach it without abandoning JPA.

💡 Step back and look at the three layers together:

- **JPQL** for the overwhelming majority of queries - object-oriented, portable, concise.
- **Criteria API** when the query must be *built* dynamically from optional parts.
- **Native SQL** when you need something only the database can express.

That layering is the real lesson of this phase. A good ORM gives you a comfortable high-level language for everyday work but **never traps you away from the real SQL underneath**. You're never stuck - you just move down a layer, trading portability for power, exactly as far as the problem demands.

## Recap

1. **JPQL operates on entities, not tables** - `select b from Book b where b.author.name = :name` walks object references, and Hibernate translates it into SQL (inferring the join) for your database.
2. **Always use parameters** (`:name` + `setParameter`); never concatenate user input into a query string - that's the SQL injection hole. **`TypedQuery<Book>`** also gives you compile-time type safety and no casting.
3. **Projections** select specific fields, often into a **DTO** via JPQL's `select new ...` constructor expression - a real performance win when you only need a few fields to display, instead of loading full entities.
4. **The Criteria API** builds queries programmatically in type-safe Java, ideal for **dynamic queries** assembled from optional filters - but it's verbose, so prefer JPQL for fixed queries.
5. **Native SQL** (`createNativeQuery`) is the escape hatch for database-specific features JPQL can't express, still mapping results to entities or projections and still using parameters.
6. The layering - **JPQL → Criteria → native SQL** - means a good ORM never traps you away from real SQL; you drop down a level only when the level above can't reach.

## Quick check

Lock in when to use which layer, and the one habit that keeps your queries safe:

```quiz
[
  {
    "q": "What is the key difference between JPQL and SQL?",
    "choices": [
      "JPQL operates on entities and their fields (Book, b.author.name); SQL operates on tables and columns",
      "JPQL is faster because it skips the database",
      "JPQL can only query by primary key, unlike SQL",
      "There is no difference; JPQL is just an alias for SQL"
    ],
    "answer": 0,
    "explain": "JPQL looks like SQL but works against your object model - entities and fields, walking references like b.author.name. Hibernate translates it (including inferring joins) into real SQL for your specific database."
  },
  {
    "q": "Why must you use `setParameter(\"name\", value)` instead of concatenating the value into the query string?",
    "choices": [
      "Concatenation opens a SQL injection hole; parameters send the value as data the database never treats as query structure",
      "Concatenation is slower to type",
      "setParameter is the only way to query strings at all",
      "Parameters are required only for numbers, not text"
    ],
    "answer": 0,
    "explain": "Gluing user input into the query text lets a crafted value like `' or '1'='1` change the query's meaning - classic SQL injection. Parameters pass the value separately, so it's always treated as data, never as part of the query."
  },
  {
    "q": "You have a search screen with several optional filters that combine in any way the user chooses. Which querying tool fits best?",
    "choices": [
      "The Criteria API - it builds the query programmatically from whichever filters are present",
      "Native SQL - only raw SQL can handle optional filters",
      "A single fixed JPQL string with all filters always applied",
      "em.find by id, called once per filter"
    ],
    "answer": 0,
    "explain": "Dynamic queries assembled from optional parts are exactly what the Criteria API is for: add a Predicate inside an if, type-safely, with no fragile string-building. For fixed-shape queries, JPQL still reads better."
  }
]
```


---

# Inheritance & Embeddables

In Phase 6 of [Java from zero](/guides/java-from-zero) you learned how Java classes share behavior:
`extends`, overriding, polymorphism. Java leans on inheritance happily. Relational databases, on the
other hand, have never heard of it. A table is a flat grid of rows and columns - there is no "this table
is a kind of that table." So the moment your `Book` and `Magazine` both want to live under a common
`Publication` parent in Java, you hit a wall: how does a hierarchy of *classes* become a layout of
*tables*?

The mental model to hold the whole time: **JPA can't change the database, so it offers you a few
different ways to flatten a class tree into tables - and each one trades query speed against
normalization.** There's no single right answer. The strategy you pick changes the actual tables
Hibernate creates and the SQL it runs, so the choice is about *how you'll query*, not about taste.

Inheritance isn't the only "compound" shape you'll want. Sometimes you have a clump of fields - street,
city, zip - that belong together but don't deserve their own table or their own identity. That's what
**embeddables** are for, and they're the tool you'll actually reach for far more often than inheritance.
We'll get there at the end, because it's the most important idea in this phase.

## The challenge: a hierarchy that has to land somewhere

Let's set up a small, straightforward "is-a" hierarchy. A library holds **publications**. A `Book` is a
publication, and so is a `Magazine`. They share a title and a year, but each has its own extra field:

```java
import jakarta.persistence.Entity;
import jakarta.persistence.Id;
import jakarta.persistence.GeneratedValue;
import jakarta.persistence.GenerationType;
import jakarta.persistence.Inheritance;
import jakarta.persistence.InheritanceType;

@Entity
@Inheritance(strategy = InheritanceType.SINGLE_TABLE)
public abstract class Publication {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    private String title;
    private int year;

    // constructors, getters, setters...
}
```
```java
@Entity
public class Book extends Publication {
    private String isbn;          // only Books have this
    // ...
}

@Entity
public class Magazine extends Publication {
    private int issueNumber;      // only Magazines have this
    // ...
}
```

*What just happened:* `Publication` is marked `@Entity` *and* `@Inheritance`, which tells JPA "this is the
root of a mapped hierarchy - expect subclasses." Both `Book` and `Magazine` are entities that `extend` it,
so they inherit the `id`, `title`, and `year` mapping for free, exactly like ordinary Java inheritance.
The one new knob is `@Inheritance(strategy = ...)`. That single attribute decides how all of this lands in
the database - and that's the whole topic of this phase.

📝 **`@Inheritance`** - the annotation on the root entity that selects a mapping strategy for the whole
hierarchy. There are three: `SINGLE_TABLE` (the default), `JOINED`, and `TABLE_PER_CLASS`. You set it once,
on the parent.

## `SINGLE_TABLE` - one table for the whole family

📝 **`SINGLE_TABLE`** - every class in the hierarchy shares **one** table. The columns are the *union* of
all fields from the parent and every subclass, plus one extra **discriminator column** whose value records
which subclass each row actually is.

This is the default, and it's the default for a reason: it's the fastest. Here's the table Hibernate
generates for our `Publication` hierarchy:

```sql
CREATE TABLE publication (
    dtype         VARCHAR(31)  NOT NULL,   -- the discriminator
    id            BIGINT       NOT NULL AUTO_INCREMENT,
    title         VARCHAR(255),
    year          INTEGER,
    isbn          VARCHAR(255),            -- only used by Book rows
    issue_number  INTEGER,                 -- only used by Magazine rows
    PRIMARY KEY (id)
);
```

*What just happened:* one table, `publication`, holds *everything*. The `dtype` column ("discriminator
type") gets the value `Book` or `Magazine` so Hibernate knows which class to rebuild when it reads a row. A
`Book` row fills in `isbn` and leaves `issue_number` null; a `Magazine` row does the reverse. Loading any
publication is a single-row read with no joins - which is why queries against this layout fly.

⚠️ **The tradeoff: nullable columns.** Look at that table again. A `Book` row *cannot* fill in
`issue_number`, and a `Magazine` row *cannot* fill in `isbn` - those columns sit null for half the rows.
And here's the sharp edge: **a subclass field can never be `NOT NULL`** in a single table, because the
*other* subclass's rows have nothing to put there. So `SINGLE_TABLE` quietly costs you database-level "this
field is required" constraints on anything specific to a subclass. The more subclasses you add, the wider
and sparser the table gets.

💡 Reach for `SINGLE_TABLE` (or just accept the default) when read performance matters, the subclasses
don't differ by much, and you don't need hard `NOT NULL` constraints on subclass-only fields. For a great
many hierarchies, that's the right call.

## `JOINED` - a table per class, stitched by key

📝 **`JOINED`** - the parent gets a table for the *shared* fields, and **each subclass gets its own table**
for *its* extra fields. The subclass tables share the parent's primary key, and Hibernate `JOIN`s them back
together when it loads a row.

```java
@Entity
@Inheritance(strategy = InheritanceType.JOINED)
public abstract class Publication {
    @Id @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;
    private String title;
    private int year;
}
```

That produces three clean tables instead of one wide one:

```sql
CREATE TABLE publication (
    id     BIGINT       NOT NULL AUTO_INCREMENT,
    title  VARCHAR(255),
    year   INTEGER,
    PRIMARY KEY (id)
);

CREATE TABLE book (
    id    BIGINT       NOT NULL,           -- same id as the parent row
    isbn  VARCHAR(255) NOT NULL,           -- now this CAN be NOT NULL
    PRIMARY KEY (id),
    FOREIGN KEY (id) REFERENCES publication (id)
);

CREATE TABLE magazine (
    id            BIGINT  NOT NULL,
    issue_number  INTEGER NOT NULL,
    PRIMARY KEY (id),
    FOREIGN KEY (id) REFERENCES publication (id)
);
```

*What just happened:* the shared `title`/`year` live once in `publication`, and each subclass's extra field
lives in its own slim table keyed by the *same* `id`. A `Book` is really two rows that share an id - one in
`publication`, one in `book` - joined on read. Notice `isbn` and `issue_number` are now `NOT NULL`: because
each subclass owns its table, a required field can actually be required. No wasted nulls, fully normalized.

The cost is right there in the name: **reading a `Book` means a join** (`publication` ⋈ `book`), and a
polymorphic query like "all publications" joins to *every* subclass table. More tables, more joins, slower
reads than `SINGLE_TABLE`.

💡 Prefer `JOINED` when the data model matters more than raw speed: subclasses have many distinct fields,
you want real `NOT NULL` constraints and a clean normalized schema, and the join cost is acceptable. It's
the database purist's choice.

## `TABLE_PER_CLASS` - a full table per concrete class

📝 **`TABLE_PER_CLASS`** - each *concrete* class gets a complete, standalone table holding both inherited
and own fields, with no shared parent table; a query across the hierarchy becomes a `UNION` of all those
tables, which is the downside that makes it the least-used of the three.

## Embeddables - folding a value object into the table

Now the idea you'll use constantly. Step away from inheritance entirely.

📝 **Embeddable** - a value object with *no identity of its own*. It's a small class whose fields become
**columns of the owning entity's table** - not a separate row, not a separate table. You mark the class
`@Embeddable` and the field that holds one `@Embedded`. Think `Address`, `Money`, `Dimensions`: things that
are *part of* an entity, not entities in their own right.

Here's an `Address` value object embedded into an `Author`:

```java
import jakarta.persistence.Embeddable;

@Embeddable
public class Address {
    private String street;
    private String city;
    private String zip;

    // constructors, getters, setters...
}
```
```java
import jakarta.persistence.Embedded;
import jakarta.persistence.Entity;
import jakarta.persistence.Id;
import jakarta.persistence.GeneratedValue;
import jakarta.persistence.GenerationType;

@Entity
public class Author {

    @Id
    @GeneratedValue(strategy = GenerationType.IDENTITY)
    private Long id;

    private String name;

    @Embedded
    private Address address;      // not a separate table - its fields land here

    // constructors, getters, setters...
}
```

*What just happened:* `Address` is `@Embeddable`, so it has no `@Id` and no table of its own. When `Author`
holds one via `@Embedded`, Hibernate *inlines* the address's three fields straight into the `author` table:

```sql
CREATE TABLE author (
    id      BIGINT       NOT NULL AUTO_INCREMENT,
    name    VARCHAR(255),
    street  VARCHAR(255),         -- from Address
    city    VARCHAR(255),         -- from Address
    zip     VARCHAR(255),         -- from Address
    PRIMARY KEY (id)
);
```

One table, one row per author, with the address columns sitting right alongside `name`. In Java you get a
tidy `author.getAddress().getCity()`; in SQL it's all flat. You got the grouping *for free* - no join, no
extra table.

**Contrast with `@Entity`.** An entity *has identity* (an `@Id`), lives in its own table, and can be
referenced and shared. An embeddable has *none* of that: it's owned wholly by its parent row, lives and
dies with it, and two authors with the "same" address have two separate copies of those column values. If you
ever need to ask "give me that address by id" or share one address across rows, you've outgrown an
embeddable and want a real entity with a relationship (Phase 5).

💡 **`@ElementCollection` for a *collection* of values.** When you want many simple values or many
embeddables attached to one entity - say, a set of an author's `phoneNumbers`, or a list of `Address`es - 
mark the field `@ElementCollection`. Hibernate puts them in a small **side table** keyed back to the owner,
without making them full entities. It's the embeddable idea, one-to-many.

Two pieces of plain guidance to take away:

- 💡 **Reach for embeddables to group related fields** without spinning up a separate table. They keep your
  Java model expressive (`Money`, `Address`, `Dimensions`) while the database stays flat and fast.
- 💡 **Choose the inheritance strategy by your query patterns:** `SINGLE_TABLE` when you read a lot and want
  speed, `JOINED` when normalization and real constraints matter more than join cost.

⚠️ **Inheritance is over-used - prefer composition and embeddables.** Just as in plain Java (Phase 6), the
classic mistake is building a class hierarchy where you didn't need one. Mapped inheritance adds real
complexity to every query and migration. Only model an `@Inheritance` hierarchy when there's a genuine,
stable **"is-a"** relationship you'll actually query polymorphically. The rest of the time, group fields
with an embeddable or model a relationship between entities - those are almost always the simpler, sturdier
choice.

## Recap

1. Relational tables have no inheritance, so **`@Inheritance`** on the root entity picks how a Java class
   hierarchy is flattened into tables - the choice drives the SQL, so decide by query patterns.
2. **`SINGLE_TABLE`** (the default) puts the whole hierarchy in one table with a discriminator column;
   fastest to read, but subclass-only fields must be nullable, so you lose `NOT NULL` on them.
3. **`JOINED`** gives the parent and each subclass its own table joined by shared primary key - normalized,
   allows real `NOT NULL` constraints, but every read costs a join.
4. **`TABLE_PER_CLASS`** gives each concrete class a full standalone table; polymorphic queries become a
   `UNION`, which is why it's the least-used strategy.
5. An **`@Embeddable`** is a value object with no identity whose fields become **columns of the owning
   entity's table** (`@Embedded`) - contrast with `@Entity`, which has its own id and table.
   `@ElementCollection` stores a collection of such values in a side table.
6. ⚠️ Inheritance is often over-used; prefer **composition/embeddables** unless a real "is-a" hierarchy
   exists.

## Quick check

Test yourself on the distinctions most likely to bite you in real mapping code:

```quiz
[
  {
    "q": "With `InheritanceType.SINGLE_TABLE`, why can't a field that only exists on a subclass be `NOT NULL` in the database?",
    "choices": [
      "Because the whole hierarchy shares one table, so rows of other subclasses have nothing to put in that column and would violate NOT NULL",
      "Because Hibernate forbids NOT NULL on any inherited field",
      "Because the discriminator column already enforces nullability for you",
      "Because SINGLE_TABLE stores subclass fields in a separate side table"
    ],
    "answer": 0,
    "explain": "SINGLE_TABLE merges every class into one table whose columns are the union of all fields. A subclass-only column is irrelevant to other subclasses' rows, which leave it null - so it cannot be NOT NULL. That lost constraint is the main tradeoff for the strategy's read speed."
  },
  {
    "q": "What is the key cost of `InheritanceType.JOINED` compared to `SINGLE_TABLE`?",
    "choices": [
      "It can't generate a schema automatically",
      "Reading an entity requires joining the parent table to the subclass table(s), so reads are slower",
      "It erases the discriminator so you can't tell subclasses apart",
      "It forces every subclass field to be nullable"
    ],
    "answer": 1,
    "explain": "JOINED stores shared fields in the parent table and each subclass's fields in its own table keyed by the same id, so loading a row means a join (and polymorphic queries join across all subclass tables). You gain normalization and real NOT NULL constraints; you pay in join cost on reads."
  },
  {
    "q": "How does an `@Embeddable` value object like `Address` differ from a separate `@Entity`?",
    "choices": [
      "An embeddable gets its own table and primary key; an entity does not",
      "They're identical - @Embeddable is just an alias for @Entity",
      "An embeddable has no identity of its own and its fields become columns of the owning entity's table; an entity has an @Id and its own table",
      "An embeddable can be shared across many rows by reference, but an entity cannot"
    ],
    "answer": 2,
    "explain": "An embeddable is a value object with no @Id and no table of its own - its fields are inlined as columns into the owner's table. An entity has identity (an @Id), lives in its own table, and can be referenced and shared. Reach for an embeddable to group related fields without a separate table."
  }
]
```


---

# Caching & Performance

Here's the thing nobody tells you when you pick up an ORM: it will happily make your app slow, and it
will do it quietly. Not with errors - with extra round-trips. A page that should fire two queries fires
two hundred, and everything still *works*, so nothing screams. Performance with Hibernate isn't about
clever tricks - it's one habit (looking at the SQL it actually emits), a couple of levers (caching,
batching), and the wisdom to know when to put the ORM down.

The mental model for this whole phase: **Hibernate trades round-trips for memory, and it trades
convenience for control.** Caching keeps data in memory so you skip round-trips. Batching bundles
round-trips so you make fewer of them. And the convenience that hides SQL from you is exactly what you
have to switch off when speed matters. Hold those three ideas and the rest is detail.

## First-level cache - you already have it, and it's free

📝 You met this in [the persistence context phase](03-entitymanager-and-persistence-context.md), but it
belongs here too, because it's the cheapest performance win you'll ever get: the **persistence context
itself is a cache**. Within one transaction, asking for the same entity by id twice runs **one** query.
The second lookup comes straight from the workbench in memory.

```java
EntityManager em = emf.createEntityManager();

Author first  = em.find(Author.class, 1L);   // hits the database
Author second = em.find(Author.class, 1L);   // same id, same context

System.out.println(first == second);          // identity, not equality
```
```sql
select a.id, a.name from author a where a.id = 1
```
```console
true
```
*What just happened:* Two `find` calls, **one** `SELECT`. The first ran the query and parked the
`Author` on the workbench; the second found it already there and handed back the *same instance*. This is
the **first-level cache**, and it's always on, scoped to a single transaction, and impossible to turn
off. Its whole reach is one persistence context - open a new `EntityManager` and the next `find` hits the
database again.

> 💡 **Key point.** The first-level cache is free and automatic, but its lifespan is one transaction.
> Don't reach for anything fancier until you've confirmed you actually need cross-transaction caching - 
> most apps don't.

## Second-level cache - optional, shared, and dangerous if you're careless

📝 The **second-level cache (L2)** is a separate, *opt-in* cache that lives **across** transactions and
sessions, shared by the whole application (and configured per entity type). Where the first-level cache
forgets everything at `commit`, the L2 cache holds onto entity data so the *next* transaction's `find`
can skip the database entirely.

Hibernate doesn't store the cache itself - it delegates to a provider you plug in: **EhCache**,
**Caffeine**, **Hazelcast**, **Infinispan**. You enable it, point Hibernate at a provider, and mark which
entities are cacheable:

```java
@Entity
@Cacheable
@org.hibernate.annotations.Cache(usage = CacheConcurrencyStrategy.READ_WRITE)
public class Author {
    @Id @GeneratedValue
    private Long id;
    private String name;
    // ...
}
```
```sql
-- First transaction, first request ever:
select a.id, a.name from author a where a.id = 1
-- Second transaction, brand-new EntityManager, same id:
-- (no SQL - served from the second-level cache)
```
*What just happened:* `@Cacheable` opts `Author` into the L2 cache; the `@Cache` annotation tells
Hibernate the concurrency strategy (`READ_WRITE` is the safe default for data that changes occasionally).
The first time *anyone* loads author `1`, Hibernate runs the `SELECT` and stashes the row's data in the
shared cache. A later transaction - a different `EntityManager`, a different web request - asks for the
same id and gets it with **no query at all**. That's the payoff: read-mostly reference data served from
memory across the whole app.

⚠️ Now the catch, and it's a big one. **Cache invalidation is one of the genuinely hard problems in
computing**, and the second-level cache hands it to you. The moment data lives in two places - the
database and the cache - they can disagree:

- **Staleness.** If a row changes *outside* Hibernate (a raw SQL update, another service, a DBA fixing
  data by hand), the cache doesn't know. It keeps serving the old value until it expires.
- **Clustering.** Run several app instances and each has its own cache. Instance A updates an author;
  instance B's cache still holds the old one until they're wired to talk to each other (distributed
  caches like Hazelcast/Infinispan exist for exactly this, and they add real complexity).
- **Volatile data.** Caching a row that changes every few seconds buys you almost nothing and risks
  serving stale data constantly. The cost of invalidation swamps the benefit.

> ⚠️ **The rule.** Cache **read-mostly reference data** - country lists, product categories, config that
> changes daily not hourly. Never cache volatile, frequently-written data. And before you enable L2 at
> all, prove with the SQL count (below) that repeated reads are actually your bottleneck. A cache you
> didn't need is just a stale-data bug waiting to happen.

### The query cache - a sharper edge still

📝 There's a cousin: the **query cache**, which caches the *result of a query* (specifically, the list of
entity ids a query returned) rather than entities by primary key. It only helps if you run the *exact
same query with the exact same parameters* repeatedly. ⚠️ It's notoriously finicky: it needs the L2 cache
turned on to be useful, it gets invalidated whenever *any* row of *any* touched table changes (so a
frequently-written table makes it nearly worthless), and a naive setup can end up *slower* than no cache
at all. Treat it as a specialist tool for a measured, repeated, read-only query - not a default.

## Batching writes - stop dribbling out one INSERT at a time

⚠️ Here's a slow path you'll write without noticing. Loop over a thousand new `Review` objects, `persist`
each one, commit:

```java
em.getTransaction().begin();
for (int i = 0; i < 1000; i++) {
    Review r = new Review("Review #" + i, 5, book);
    em.persist(r);
}
em.getTransaction().commit();
```
```sql
insert into review (book_id, rating, text, id) values (1, 5, 'Review #0', 1)
insert into review (book_id, rating, text, id) values (1, 5, 'Review #1', 2)
insert into review (book_id, rating, text, id) values (1, 5, 'Review #2', 3)
-- ... 997 more, one statement per round-trip ...
```
*What just happened:* a thousand separate `INSERT` statements, each its own network round-trip to the
database. The compute is trivial; the *latency* is the killer - a millisecond per round-trip is a full
second of nothing but waiting. The database could swallow these in one gulp, but Hibernate is sending
them one spoonful at a time.

The fix is **JDBC batching**: tell Hibernate to bundle statements and send them in groups. One config
property turns it on:

```java
// persistence.xml / application.properties
hibernate.jdbc.batch_size = 50
// helps the batch stay tight when inserts/updates interleave:
hibernate.order_inserts = true
hibernate.order_updates = true
```

⚠️ But there's a second half nobody mentions: the **persistence context keeps growing**. Every `persist`
adds an entity to the workbench, and Hibernate tracks all of them for dirty checking. Insert a million
rows in one context and you'll run the JVM out of memory long before you finish. So in big loops you
**flush and clear** periodically:

```java
em.getTransaction().begin();
for (int i = 0; i < 1000; i++) {
    Review r = new Review("Review #" + i, 5, book);
    em.persist(r);
    if (i % 50 == 0) {     // every batch_size rows
        em.flush();        // push this batch's INSERTs to the DB
        em.clear();        // empty the workbench so memory stays flat
    }
}
em.getTransaction().commit();
```
```sql
-- batched: roughly 1000 / 50 = 20 round-trips instead of 1000
insert into review (book_id, rating, text, id) values (1, 5, 'Review #0', 1), (1, 5, 'Review #1', 2), ... (50 rows)
insert into review (book_id, rating, text, id) values (1, 5, 'Review #50', 51), ... (50 more)
-- ... ~18 more batches ...
```
*What just happened:* with `batch_size = 50`, Hibernate accumulates 50 inserts and ships them as one
round-trip - turning ~1000 trips into ~20. The `flush()` sends the pending batch and the `clear()` empties
the persistence context so it doesn't balloon. One caution worth planting: this works cleanly with a
manually-assigned or sequence id; the old `GenerationType.IDENTITY` strategy forces Hibernate to insert
rows one at a time to read back each auto-increment id, which **silently disables batching** - another
reason to prefer sequence-based ids for bulk work.

## Reading the SQL is the real skill

💡 This is the throughline of the entire guide, so let it land: an ORM's job is to hide SQL, and your job
is to **un-hide it when it matters**. Every performance problem in this phase - N+1, accidental eager
loads, missing indexes, un-batched writes - is invisible until you look at the queries Hibernate emits.
Once you *can* see them, the problems become obvious. So make them visible.

```java
// the blunt instrument - log every statement (dev only):
hibernate.show_sql = true
hibernate.format_sql = true

// the better instrument - counts, not noise:
hibernate.generate_statistics = true
```
```console
Session Metrics {
    1247 jdbc statements executed   <-- this number is the whole game
}
```
*What just happened:* `show_sql` dumps every statement to the log - fine for eyeballing a single request,
useless once volume is high. The real tool is **statistics**: it tells you *how many* statements one
operation fired. That single count is your performance dashboard. Render a page, glance at the count: 3
queries? Healthy. 300? You just found your N+1 - entities loaded one-by-one, each triggering its own
`SELECT`, exactly the trap covered in the fetching phase. The count doesn't lie and it doesn't theorize.

> 💡 **Make this a habit, not a fire drill.** Watch the query count during normal development, not only
> when something's already on fire. An N+1 caught the day you write the loop is a one-line `JOIN FETCH`;
> the same N+1 found in production three months later is an incident.

When the count is high and the *why* isn't obvious, that's where deeper diagnosis comes in - reading
query plans, checking indexes, profiling the slow statement itself. Those skills live in their own
guides: [Why Is My Query Slow?](/guides/why-is-my-query-slow) for hunting down the expensive query and
the missing index behind it, and [Profiling 101](/guides/profiling-101) for measuring where time
actually goes instead of guessing.

## When the ORM is the wrong tool

⚠️ Hibernate is built for one thing brilliantly: loading objects, letting you change them, and saving
them back - the object-graph, domain-logic, CRUD world. There are jobs where forcing that shape on the
problem makes it slower *and* uglier. Recognize them and step around the ORM on purpose.

**Bulk updates and deletes.** You need to mark every review older than a year as archived. The ORM-shaped
instinct - load them all, set a flag on each, save - drags thousands of rows into memory and dirty-checks
every one. Don't. Issue a single bulk statement:

```java
// load-then-save loop: thousands of SELECTs + UPDATEs, huge memory footprint
// DON'T do this for bulk changes.

// JPQL bulk update - one statement, runs in the database, touches no workbench:
int updated = em.createQuery(
        "update Review r set r.archived = true where r.createdAt < :cutoff")
    .setParameter("cutoff", oneYearAgo)
    .executeUpdate();
```
```sql
update review set archived = true where created_at < '2025-06-22'
```
*What just happened:* the JPQL `update`/`delete` runs **directly in the database** as one statement - no
entities loaded, no dirty checking, no memory bloat. The one trade-off to know: bulk operations bypass
the persistence context, so any entities already on your workbench won't reflect the change. Run bulk ops
in their own transaction (or `clear()` afterward) and you're fine.

**Heavy reporting and analytics.** "Total revenue per author per quarter" is not an object-graph problem;
it's aggregation. Mapping it through entities is wasteful. Drop to a **projection** (select only the
columns you need into a DTO) or raw SQL:

```java
List<AuthorSales> rows = em.createQuery(
        "select new com.example.AuthorSales(a.name, sum(b.price)) " +
        "from Author a join a.books b group by a.name", AuthorSales.class)
    .getResultList();
```
*What just happened:* instead of loading full `Author` and `Book` entities only to add up a number, the
query selects exactly two values straight into a lightweight DTO. The database does the grouping; you
move a handful of columns instead of whole object graphs. For genuinely complex reports, plain native SQL
through `createNativeQuery` is often the clearest, fastest choice - and that's not a failure of the ORM,
it's using the right tool.

💡 The clear framing: use Hibernate for the **95%** - CRUD, domain logic, the everyday loading and
saving of objects, where its convenience is a genuine gift. Drop to SQL for the **5%** - bulk operations,
analytics, the rare white-hot path - where that convenience costs more than it's worth. Don't fight the
ORM, and don't worship it. Know which 5% you're in, and step around it without guilt.

## Recap

1. The **first-level cache** is the persistence context: free, always on, one transaction wide. Same id
   twice → one query. It's your cheapest win and you already have it.
2. The **second-level cache** is optional, shared across transactions, backed by a provider
   (EhCache/Caffeine/Hazelcast), and opted into per entity with `@Cacheable`. Great for **read-mostly
   reference data**.
3. ⚠️ L2's price is **invalidation**: stale data on out-of-band writes, divergence across clustered
   instances, and near-zero value on volatile rows. Never cache frequently-written data; prove the need
   with SQL counts first. The **query cache** is sharper still - a specialist tool, not a default.
4. **Batch writes** with `hibernate.jdbc.batch_size`, and **flush + clear** periodically in big loops so
   the persistence context doesn't run you out of memory. (`IDENTITY` ids silently disable batching.)
5. 💡 **Reading the emitted SQL is the real skill.** Turn on `generate_statistics`, watch the query
   count, and N+1, accidental eager loads, and missing indexes stop hiding. Make it a daily habit, not a
   fire drill.
6. ⚠️ Step around the ORM for the 5% it's wrong for: **bulk `update`/`delete` JPQL**, **projections/raw
   SQL for reporting**, and very hot paths. Use Hibernate for the 95% it's brilliant at.

## Quick check

Three ideas that decide whether your Hibernate app is fast or quietly slow:

```quiz
[
  {
    "q": "You enable the second-level cache on an Author entity that's read constantly but updated by a nightly batch job running raw SQL outside Hibernate. What's the main risk?",
    "choices": [
      "Stale data - Hibernate doesn't see the out-of-band SQL update, so the cache keeps serving the old author until it expires",
      "Nothing - the second-level cache automatically detects all database changes",
      "The first-level cache will conflict with the second-level cache and throw an exception",
      "Reads will become slower because every read now checks two caches"
    ],
    "answer": 0,
    "explain": "The L2 cache only knows about changes made through Hibernate. A raw SQL update (another job, a DBA, another service) bypasses it, so the cache holds stale data until it expires. Read-mostly data changed out-of-band is exactly where invalidation bites."
  },
  {
    "q": "You're inserting 100,000 Review rows in a loop, persisting each. You set hibernate.jdbc.batch_size=50 but the loop still runs out of memory. What did you forget?",
    "choices": [
      "To flush() and clear() the persistence context periodically - every persisted entity stays on the workbench for dirty checking, so the context grows without bound",
      "To set batch_size higher; 50 is too small to matter",
      "To wrap the loop in a transaction",
      "To enable the second-level cache, which would hold the entities instead"
    ],
    "answer": 0,
    "explain": "batch_size controls how many statements are bundled per round-trip, but every persisted entity still lives in the persistence context. Without periodic flush() + clear(), the context keeps growing and exhausts memory. Both halves are needed for bulk inserts."
  },
  {
    "q": "A page that should be fast is slow. You turn on hibernate.generate_statistics and see one request fired 312 JDBC statements. What's the most likely cause and the right next move?",
    "choices": [
      "An N+1 problem - a collection or association is being loaded one row at a time; fix it with a JOIN FETCH or entity graph, found by reading the query count",
      "The database is missing RAM; restart it",
      "Hibernate is broken; switch to raw JDBC for the whole app",
      "The second-level cache is too small; increase its size"
    ],
    "answer": 0,
    "explain": "Hundreds of statements for one logical operation is the classic N+1 signature: entities loaded individually, each firing its own SELECT. The statistics count is what makes it visible, and the fix is to fetch the association in one query (JOIN FETCH / entity graph)."
  }
]
```


---

# Hibernate in the Real World & Where to Go Next

Stop for a second and look at how far you've come. You started this guide thinking of an ORM as a black box that turned objects into rows somehow. Now you can name the gears inside it. You understand the **persistence context** - the managed identity map that makes an entity feel like a live object instead of a dead snapshot. You know **dirty checking** is what updates a row when you never called `save`. You can map a `@ManyToOne`, reason about owning versus inverse sides, and - this is the big one - you can *see* the **N+1 problem** coming and reach for a `JOIN FETCH` before it ever hits production.

Most of all, you can read the SQL. With `show_sql` on, Hibernate stopped being magic and started being a tool whose output you can predict and debug. That's the whole game - a data layer is no longer something that happens *to* you; it's something you reason about.

This last phase isn't new mechanics - it's about where all of that lives in real teams, and there's one revelation waiting that ties the whole guide together.

## The magic, revealed

💡 Here's the moment everything clicks. Remember [Spring Boot](/guides/spring-boot-from-zero) and its uncanny trick - you declare an *interface*, never write a line of implementation, and somehow `findByLastName` returns rows from the database? You now know exactly what's happening down there. It's this. It's Hibernate.

```mermaid
flowchart TD
  A[You declare an interface] --> B[Spring Data JPA]
  B --> C[JpaRepository wraps EntityManager]
  C --> D[Method name to JPQL to SQL]
  D --> E[Hibernate persistence context]
```

Pull it apart and every piece is something you've already met:

- **`JpaRepository`** is a thin wrapper around the `EntityManager` you spent Phase 3 inside. `save`, `findById`, `delete` - those are `persist`, `find`, and `remove` with a friendlier name.
- **Derived query methods** (`findByTitleAndPublishedYear`) are parsed from the method name into **JPQL** - the exact query language you wrote by hand in Phase 7. Spring just generates it for you.
- **`@Transactional`** opens and commits the unit of work from Phase 4. The flush, the dirty checking, the commit-or-rollback - same machinery, declared instead of written.

This matters because **most real Java doesn't use Hibernate directly - it uses Hibernate via Spring Data JPA.** That's the normal, expected path, and it's a good one. The point of this guide was never to make you write `EntityManager` code forever. It was to make sure that when the generated query is slow, or the lazy-loading exception fires, or the SQL looks wrong, you're not staring at a sealed box. You can open it.

## Schema migrations - the one thing you must get right for production

⚠️ Back in Phase 2 you saw `hbm2ddl.auto` quietly create and alter tables to match your entities. It's a wonderful convenience while you're learning. Let me be the friend who says the hard thing once, plainly: **never let `hbm2ddl.auto=update` manage a production schema.** Not ever. It's fine for a throwaway dev database; it is a foot-gun pointed at your real data.

Why? Because `update` mode makes schema changes *implicit*. Hibernate looks at your classes, looks at the tables, and silently does what it thinks is the difference. It won't drop a renamed column (so it leaks). It can't review itself. It can't be rolled back. And two developers' machines can end up with two different schemas and no record of how either got there.

Real teams make schema changes the way they make code changes - **deliberate, versioned, and reviewed.** That's what **Flyway** and **Liquibase** are for:

- You write each change as a small migration file (`V3__add_review_rating.sql`) and check it into git, right next to the code that needs it.
- The tool keeps a table of which migrations have run, applies new ones in order on startup, and never re-runs an old one.
- Changes go through pull-request review like everything else, and they're repeatable: the same sequence builds the same schema on a laptop, in CI, and in production.

📝 The mental shift is small but total: your **entities describe how Java sees the data; your migrations are the source of truth for the schema itself.** Set `hbm2ddl.auto` to `validate` in production - it'll check that your entities and the real schema agree and refuse to start if they've drifted, without ever touching a table. This isn't optional polish. In a real team, it's table stakes.

## When to reach past the ORM

Hibernate is excellent, and it is not the answer to everything. Part of the maturity you've earned is knowing *which tool for which job* - so here's the clear-eyed map:

- **Hibernate for CRUD and domain logic - the 95%.** Loading an order, updating a user, saving a book with its reviews, navigating relationships. This is exactly what an ORM is built for, and it's where most of your code lives.
- **jOOQ or raw SQL for gnarly queries and reporting.** When you need a seven-way join, window functions, a recursive CTE, or a reporting query tuned to the bone, an ORM fights you. Don't fight back. Drop to SQL (or a SQL-first library like jOOQ) and let the database do what it's great at.
- **Bulk operations via JPQL or native queries.** Updating ten thousand rows by loading each entity, mutating it, and flushing is the slow way. A single `UPDATE` statement - JPQL bulk update or native SQL - is the right way.

💡 Notice that you can make every one of these calls now. Knowing *when the ORM is the wrong tool* is itself a skill the ORM can't teach you - you only get it by understanding what Hibernate does and what it costs. You have that.

## What to build, and a last word

Reading got you here. Building is what makes it stick. The model from this guide - **authors, books, reviews** - is a perfect sandbox because it has every relationship shape and the N+1 trap baked in. A few no-nonsense projects:

- **Build the model into a small standalone app.** Wire up the entities, write a query that loads books with their reviews using a `JOIN FETCH` (and watch the query count drop from N+1 to one), and add a **Flyway migration** to create the schema instead of `hbm2ddl`. Now you've practiced the two things that separate a tutorial from production: fetch strategy and real migrations.
- **Or drop it into a Spring Boot app.** Turn the model into a `JpaRepository`, expose a couple of REST endpoints, keep `show_sql` on, and *watch the SQL Spring Data generates.* This is the most satisfying exercise in the whole guide - you'll see, line by line, the thing you just learned to read being written for you.

Whichever you pick, **finish one.** A small, finished app that you debugged teaches more than three half-built ones. And when you want the canonical reference, bookmark the **Hibernate User Guide** and the **Jakarta Persistence (JPA) specification** - between them they answer almost any question you'll have, and you can now actually read them.

You came in seeing an ORM as a magic trick. You're leaving able to map objects to tables, dodge the N+1 trap, write real JPQL, choose when *not* to use the ORM at all, and read the SQL underneath every bit of it. The ORM was never magic - it's the persistence context doing exactly what you now understand. Go build the small thing.

## Recap

1. **Spring Data JPA *is* Hibernate, generated.** `JpaRepository` wraps the `EntityManager`, derived method names become JPQL, and `@Transactional` runs the unit of work - all machinery you've already met. Most real Java uses Hibernate this way.
2. **Never let `hbm2ddl.auto=update` run production.** Schema changes must be deliberate, versioned, and reviewed.
3. **Use Flyway or Liquibase** for migrations: small SQL files checked into git, applied in order, repeatable across every environment. Set `hbm2ddl.auto=validate` in prod.
4. **Reach past the ORM on purpose:** Hibernate for CRUD and domain logic, jOOQ or raw SQL for complex reporting, JPQL or native queries for bulk operations.
5. **Build the authors/books/reviews model** with a `JOIN FETCH` and a Flyway migration, or wire it into Spring Boot and watch the generated SQL. Finish one, and keep the Hibernate User Guide and JPA spec handy.

## Quick check

One last check - on how Hibernate actually shows up in the real world:

```quiz
[
  {
    "q": "In a typical Spring Boot app, what is Spring Data JPA's JpaRepository actually doing?",
    "choices": [
      "Wrapping the EntityManager - derived method names become JPQL and @Transactional runs the unit of work, all the Hibernate machinery you already learned",
      "Replacing Hibernate entirely with its own brand-new ORM engine",
      "Talking to the database with hand-written JDBC and no ORM involved",
      "Caching every table in memory so the database is never queried"
    ],
    "answer": 0,
    "explain": "Spring Data JPA is Hibernate, generated for you: JpaRepository wraps the EntityManager, method names are parsed into JPQL, and @Transactional opens and commits the same unit of work you wrote by hand in earlier phases."
  },
  {
    "q": "Why should you not use hbm2ddl.auto=update to manage a production schema?",
    "choices": [
      "Schema changes become implicit, unreviewable, and non-repeatable - use versioned Flyway or Liquibase migrations instead",
      "It only works on MySQL and fails silently on every other database",
      "It is too slow to run on a database with more than a few rows",
      "It deletes the entire database every time the application restarts"
    ],
    "answer": 0,
    "explain": "update mode silently guesses the difference between your entities and the tables: it won't drop renamed columns, can't be reviewed, and can't be rolled back. Real teams use versioned, reviewed migrations (Flyway/Liquibase) and set hbm2ddl.auto=validate in production."
  },
  {
    "q": "You need a heavy reporting query with several joins and window functions. What's the mature call?",
    "choices": [
      "Drop to raw SQL or a SQL-first tool like jOOQ - the ORM is great for CRUD and domain logic, but complex reporting is a job for SQL",
      "Force it through entity navigation, loading every related object into memory first",
      "Avoid the query entirely because Hibernate cannot coexist with raw SQL",
      "Rewrite all your entities so the query becomes a simple findById call"
    ],
    "answer": 0,
    "explain": "Hibernate handles the ~95% that's CRUD and domain logic. For complex reporting and gnarly joins, reaching for raw SQL or jOOQ is the right instinct - and knowing when the ORM is the wrong tool is itself a sign you understand it."
  }
]
```
