# PyTorch From Zero

> Learn the deep-learning framework that runs modern AI: tensors and the GPU, autograd (automatic differentiation), building models with nn.Module, loss functions and optimizers, the training loop, Dataset/DataLoader, training a real classifier, saving and inference, performance pitfalls, and where to go. The three ideas under all of it - tensors, autograd, the loop - made plain.


---

# PyTorch From Zero

PyTorch is the framework most modern AI is actually built in - the research papers, the image models, and
the large language models you've heard of were, in huge part, trained with it. It has a reputation for
being deep and mathematical, and the math is real, but the framework itself rests on just **three ideas**:
a **tensor** (a multi-dimensional array you do math on, fast, on a GPU), **autograd** (PyTorch
automatically computes the derivatives needed to learn), and **the training loop** (a short, repeating
ritual that nudges a model toward being right). Understand those three and the rest is detail.

This guide builds those ideas first, in plain language, then assembles them into a real model you train
and run. We connect it to what you may already know: a tensor is NumPy's array with superpowers; "learning"
is the gradient descent from [How a Model Learns](/guides/how-a-model-learns), made concrete; a model is a
Python class. By the end you'll have trained a working classifier and understand every line of the loop
that did it.

> 📝 This teaches the **framework**, not the math from scratch. It assumes **Python**
> ([Python From Zero](/guides/python-from-zero)) and is far richer if you've met the concepts in
> [What AI & ML Are](/guides/what-ai-and-ml-are) and especially [How a Model Learns](/guides/how-a-model-learns)
> (gradients, loss, training). Helpful too: [pandas From Zero](/guides/pandas-from-zero) for data prep.
>
> ⚠️ PyTorch needs a native install (and ideally a GPU), so examples here are shown with their output
> rather than run on the page - follow along in a notebook or Google Colab (free GPUs).

## How to read this

Read in order - it builds from a single tensor up to a trained, saved classifier, one idea per phase.
Phases carry difficulty badges; the 🔴 ones (autograd, the loop, performance) are the conceptual core.

## The phases

**Part 1 - The three core ideas (🟢 → 🔴)**
1. **[What PyTorch Is & Tensors](01-what-pytorch-is-and-tensors.md)** 🟢 - the tensor: a GPU-ready, autograd-aware array.
2. **[Tensor Operations & the GPU](02-tensor-operations-and-gpu.md)** 🟢 - math, broadcasting, reshaping, and moving work to the GPU.
3. **[Autograd: Automatic Differentiation](03-autograd.md)** 🔴 - how PyTorch computes the gradients that make learning possible.

**Part 2 - Building & training a model (🟡 → 🔴)**
4. **[Building Models with nn.Module](04-building-models-with-nn-module.md)** 🟡 - layers, `forward()`, and a model as a Python class.
5. **[Loss Functions & Optimizers](05-loss-and-optimizers.md)** 🟡 - measuring wrongness and the algorithm that fixes it.
6. **[The Training Loop](06-the-training-loop.md)** 🔴 - forward → loss → backward → step: the ritual that trains everything.
7. **[Data: Dataset & DataLoader](07-datasets-and-dataloaders.md)** 🟡 - batching, shuffling, and feeding data efficiently.
8. **[Training a Real Classifier](08-training-a-classifier.md)** 🟡 - putting it all together on a real dataset, with evaluation.

**Part 3 - Using & shipping models (🟡 → 🟢)**
9. **[Saving, Loading & Inference](09-saving-loading-inference.md)** 🟡 - `state_dict`, `eval()` mode, `no_grad`, and running a trained model.
10. **[GPUs, Performance & Common Pitfalls](10-gpus-performance-pitfalls.md)** 🔴 - devices, speed, and the bugs that bite every beginner.
11. **[Where to Go Next](11-where-to-go-next.md)** 🟢 - pretrained models, transfer learning, the ecosystem, and LLMs.

> Tensors, autograd, the loop. Everything in deep learning - from a 3-line model to a giant LLM - is those
> three ideas at scale. This guide makes them yours.


---

# What PyTorch Is & Tensors

If you've ever wondered what the big AI models are actually *built in* - the chatbots, the image
generators, the recommendation engines - the straight answer, most of the time, is PyTorch. It's the tool
researchers reach for the moment "teach a computer from data" enters the conversation. Before you train a
single model, there's one idea that makes the whole library fall into place. Get it, and PyTorch stops
being a wall of unfamiliar functions and starts feeling like something you already half-know.

That idea is the **tensor**. If you've touched NumPy or pandas, you already understand 90% of it - a tensor
is a grid of numbers you do math on. PyTorch adds two superpowers to that grid, and those two powers are
the entire reason deep learning uses tensors instead of plain Python lists. We'll build up to them slowly.

> 💡 PyTorch isn't available in this guide's in-browser runtime, so the code blocks here aren't runnable. To
> follow along for real, type these into a Jupyter notebook or a Python REPL where PyTorch is installed
> (`pip install torch`). Each example shows its output so you can check yourself either way.

## What PyTorch actually is

📝 **PyTorch** - an open-source deep-learning framework. It gives you three things: **tensors** (fast
multi-dimensional arrays), **automatic differentiation** (it can compute the math derivatives a model needs
to learn - more in Phase 3), and **neural-network building blocks** (ready-made layers and training tools).
And it runs all of that fast on **GPUs**.

If those words are fuzzy, that's fine - this is the toolkit machine learning is *done* in. When a team
trains a model to recognize images or generate text ([What AI & ML Are](/guides/what-ai-and-ml-are)),
PyTorch is very often what's underneath. It won the research world by being deeply **Pythonic**: tensors
behave like the arrays you already know, and code runs line by line so you can poke at it the way you'd poke
at any Python.

By near-universal convention, you import it as `torch`:

```python
import torch
```

*What just happened:* You pulled in the library under its real package name, `torch`. (The project is called
PyTorch, but the thing you `import` is `torch` - a quirk worth knowing so the import line doesn't confuse
you.) Every PyTorch tutorial and codebase writes `torch.something`; follow the convention.

⚠️ **One plain aside about TensorFlow.** PyTorch isn't the only major deep-learning framework - Google's
**TensorFlow** is the other big one, and you'll see it in plenty of production systems. Both can train the
same kinds of models. PyTorch dominates research and is widely considered the more Pythonic, friendlier
place to learn. You don't need both; learn one well, and the concepts carry over.

## The tensor - a grid of numbers

Everything in PyTorch is made of tensors. So let's pin down exactly what one is.

📝 **Tensor** - a multi-dimensional array of numbers, and PyTorch's fundamental data type. The number of
dimensions is its **rank**: a single number is a **scalar** (0-D), a list of numbers is a **vector** (1-D),
a grid is a **matrix** (2-D), and you can keep stacking into 3-D, 4-D, and beyond.

That's the same shape-family you've met before. A tensor is essentially NumPy's `ndarray`
([pandas From Zero](/guides/pandas-from-zero) is built on NumPy, so a pandas column is a close cousin too) - 
a typed, rectangular block of numbers. What makes it a *tensor* rather than just an array is the two
superpowers we keep teasing:

1. **It can live on a GPU.** Move it to graphics hardware and the same math runs hundreds of times faster
   (Phase 2).
2. **It can track gradients.** A tensor can record every operation done to it, so PyTorch can later work
   backward and figure out how to nudge a model's numbers to make it better - the engine of learning
   (Phase 3, **autograd**).

Hold those two in mind. For now, just think: *tensor = array of numbers, with a turbo button.*

```mermaid
flowchart LR
  A["scalar (0-D)<br/>3.14"] --> B["vector (1-D)<br/>[1, 2, 3]"]
  B --> C["matrix (2-D)<br/>[[1, 2],<br/> [3, 4]]"]
  C --> D["3-D and beyond<br/>stacks of matrices"]
```

## Creating tensors

There are a handful of ways to make a tensor, and you'll use all of them constantly. Start with the most
direct one - handing PyTorch a Python list:

```python
import torch

scalar = torch.tensor(3.14)
vector = torch.tensor([1, 2, 3])
matrix = torch.tensor([[1, 2], [3, 4]])

print(scalar)
print(vector)
print(matrix)
```
```console
tensor(3.1400)
tensor([1, 2, 3])
tensor([[1, 2],
        [3, 4]])
```

*What just happened:* `torch.tensor(...)` took ordinary Python numbers and lists and turned each into a
tensor. Notice the **rank** showing through in the brackets: `scalar` has none (it's a lone number),
`vector` has one level, and `matrix` has two - a list of lists becomes a 2-D grid. The repr always wraps the
values in `tensor(...)` so you can tell at a glance you're holding a PyTorch object, not a plain list.

Typing out values by hand only goes so far. Most real tensors start filled with a pattern - all zeros, all
ones, or random numbers - and you say *what shape* you want with a tuple of dimensions:

```python
import torch

zeros = torch.zeros(2, 3)        # a 2-row, 3-column grid of 0.0
ones  = torch.ones(2, 3)         # same shape, all 1.0
rand  = torch.randn(2, 3)        # same shape, random numbers
steps = torch.arange(0, 10, 2)   # like Python's range: start, stop, step

print(zeros)
print(rand)
print(steps)
```
```console
tensor([[0., 0., 0.],
        [0., 0., 0.]])
tensor([[ 0.4967, -0.1383,  0.6477],
        [ 1.5230, -0.2342, -0.2341]])
tensor([0, 2, 4, 6, 8])
```

*What just happened:* `torch.zeros(2, 3)` and `torch.ones(2, 3)` built 2×3 grids pre-filled with a constant
 - the workhorses for setting up a tensor of a known size before you fill it. `torch.randn(2, 3)` filled the
same shape with random values drawn from a normal distribution (your numbers will differ - that's the point
of random). `torch.arange(0, 10, 2)` mirrors Python's `range`, producing a 1-D tensor of evenly spaced
values. Random init like `randn` is how neural-network weights start life before training shapes them.

You can also convert straight from a NumPy array, which matters because so much of the data world speaks
NumPy:

```python
import torch
import numpy as np

arr = np.array([1.0, 2.0, 3.0])
t = torch.from_numpy(arr)
print(t)
```
```console
tensor([1., 2., 3.], dtype=torch.float64)
```

*What just happened:* `torch.from_numpy(arr)` wrapped an existing NumPy array as a tensor - no manual
copying of values. This is the bridge between the NumPy/pandas data-prep world and the PyTorch
model-training world: load and clean data with the tools you know, then hand it to PyTorch as tensors. (Note
the `dtype=torch.float64` it picked up from NumPy - more on dtype right now.)

## Tensor attributes - shape, dtype, device

Every tensor carries three pieces of metadata you'll check over and over. They're how you reason about what
you're holding, and - fair warning - how you'll diagnose most of your bugs.

```python
import torch

t = torch.randn(2, 3)
print(t.shape)    # how many in each dimension
print(t.dtype)    # what kind of number
print(t.device)   # where it lives (CPU or GPU)
```
```console
torch.Size([2, 3])
torch.float32
cpu
```

*What just happened:* `.shape` reported `[2, 3]` - two rows, three columns - the size along each dimension.
`.dtype` reported `torch.float32`, the kind of number stored (here, 32-bit floats). `.device` reported
`cpu`, meaning this tensor lives in regular memory, not on a GPU (Phase 2 covers moving it). Three
attributes, and together they tell you everything about *how* a tensor is laid out.

A word on each, because they pay rent:

- **`.shape`** - the single most important attribute. Deep-learning math is mostly about matching shapes:
  to multiply or add tensors, their dimensions have to line up. ⚠️ **The vast majority of bugs you'll hit in
  PyTorch are shape mismatches** - a tensor that's `[32, 10]` where the next step expected `[10, 32]`, or a
  stray extra dimension. When something breaks, your first move is almost always to print `.shape` and see
  what you've actually got. We'll lean on this hard in Phase 2.
- **`.dtype`** - the number type. `float32` is the default for anything a model learns, because gradient math
  needs decimals (you can't nudge a weight by 0.0001 if it's stored as an integer). Integer dtypes show up
  for things like labels and indices. Mixing dtypes that don't match is another classic source of errors.
- **`.device`** - `cpu` or something like `cuda:0` (a GPU). A tensor can only do math with other tensors on
  the *same* device, which is the whole subject of the next phase.

## Why tensors, not plain Python lists

You might reasonably ask: a list of lists already holds a grid of numbers - why invent a whole new type?
Here's the payoff, and it's the same trio of reasons every time.

💡 **Key point.** Tensors beat lists for deep learning on three counts, and you need all three:

1. **Vectorized math.** A tensor does arithmetic on the *whole grid at once*, in fast compiled code - 
   `a + b` adds every element in one sweep, no Python loop. This is the exact same "don't write a `for` loop
   over the rows" habit you'd use in NumPy or [pandas](/guides/pandas-from-zero); PyTorch rewards it just as
   much. Looping element-by-element in Python is slow and un-PyTorch-ish.
2. **GPU parallelism.** A GPU has thousands of tiny cores that do the same operation on different numbers
   simultaneously. Tensors can run on it; plain lists can't. For the huge matrix multiplies that models are
   made of, this is the difference between minutes and days.
3. **Gradients.** A tensor can remember the operations performed on it and let PyTorch compute derivatives
   automatically (autograd). That's literally how a model learns from its mistakes - and a Python list has no
   idea what was ever done to it.

That trio - vectorized, GPU-able, gradient-tracking - is the entire reason deep learning is built on
tensors. Vectorized math and the GPU are where we go next, in **Phase 2: Tensor Operations & the GPU**.
Gradients get their own spotlight in Phase 3.

## Recap

1. **PyTorch** (imported as `torch`) is an open-source deep-learning framework: tensors + automatic
   differentiation + neural-network building blocks, run fast on GPUs. It's what much of modern AI is built
   in, and it's the dominant, very Pythonic choice in research. **TensorFlow** is the other major framework.
2. A **tensor** is a multi-dimensional array of numbers - scalar (0-D), vector (1-D), matrix (2-D), and up.
   It's NumPy's array with two superpowers: it can live on a **GPU** and it can **track gradients**.
3. Create tensors with `torch.tensor([...])` from a list, `torch.zeros` / `torch.ones` / `torch.randn` from
   a shape, `torch.arange` like `range`, or `torch.from_numpy` from a NumPy array.
4. Every tensor has a **`.shape`** (size per dimension), a **`.dtype`** (number type - `float32` by default
   for learning), and a **`.device`** (CPU or GPU). ⚠️ Most PyTorch bugs are shape mismatches - print
   `.shape` first when things break.
5. Tensors beat plain Python lists for three reasons that all matter: **vectorized math** (whole-grid at
   once, no loops), **GPU parallelism**, and **gradient tracking** for learning.
6. The "think in whole arrays, not loops" habit from NumPy and pandas carries straight over - it's the core
   PyTorch instinct too.

## Quick check

Test yourself on the one idea this whole guide builds on - what a tensor is and why it's special:

```quiz
[
  {
    "q": "What is the best mental model for a PyTorch tensor?",
    "choices": [
      "A NumPy-style multi-dimensional array of numbers that can also run on a GPU and track gradients",
      "A special kind of Python for-loop optimized for numbers",
      "A file format PyTorch uses to save trained models to disk",
      "A function that automatically downloads pre-trained AI models"
    ],
    "answer": 0,
    "explain": "A tensor is a multi-dimensional array of numbers - like a NumPy array - with two extra superpowers: it can live on a GPU for fast parallel math, and it can track gradients so models can learn."
  },
  {
    "q": "You print `t.shape` and see `torch.Size([2, 3])`. What does that tell you?",
    "choices": [
      "The tensor has 2 rows and 3 columns - 2 along the first dimension, 3 along the second",
      "The tensor contains exactly the numbers 2 and 3",
      "The tensor uses 2.3 bytes of memory per value",
      "The tensor is stored on GPU number 23"
    ],
    "answer": 0,
    "explain": "`.shape` reports the size along each dimension. `[2, 3]` means a 2-D grid with 2 along the first axis (rows) and 3 along the second (columns). Matching shapes is most of what PyTorch math is about - which is why shape mismatches cause most bugs."
  },
  {
    "q": "Why does deep learning use tensors instead of plain Python lists?",
    "choices": [
      "Tensors do vectorized whole-array math, can run on GPUs, and can track gradients for learning - lists can do none of these",
      "Tensors take up less disk space than lists",
      "Python lists cannot store decimal numbers, only tensors can",
      "Tensors are the only way to print numbers to the console in Python"
    ],
    "answer": 0,
    "explain": "That trio is the whole reason: vectorized math (the whole grid at once, no loops), GPU parallelism for massive speed, and automatic gradient tracking so models can learn. A plain list offers none of these."
  }
]
```


---

# Tensor Operations & the GPU

In [Phase 1](01-what-pytorch-is-and-tensors.md) you met the tensor: a multi-dimensional array that's
GPU-ready and autograd-aware - that was the noun. This phase is the verbs, the things you *do* to tensors.

Here's the mental model to carry through everything below: **a tensor operation acts on the whole tensor
at once, not one number at a time.** When you write `a + b`, you're not asking for a loop over elements - 
you're asking the underlying engine (highly optimized C++, and on a GPU, thousands of cores) to do the
whole thing in one shot. Your job is to line the shapes up correctly; the engine does the grunt work. Get
comfortable here and the rest of PyTorch is mostly arranging these operations in the right order.

We'll cover five things: elementwise math, matrix multiply (the operation neural nets are built from),
broadcasting (combining different-but-compatible shapes), reshaping and indexing (shuffling data into the
shape a layer expects), and finally moving the work onto a GPU.

## 1. Elementwise math

The everyday operators - `+`, `-`, `*`, `/`, `**` - work on whole tensors, position by position. No loop
required. This is called **vectorized** math, and it's the first reason PyTorch is fast.

```python
import torch

a = torch.tensor([1.0, 2.0, 3.0])
b = torch.tensor([10.0, 20.0, 30.0])

print(a + b)      # add, element by element
print(a * b)      # multiply, element by element
print(a ** 2)     # square each element
```

```console
tensor([11., 22., 33.])
tensor([10., 40., 90.])
tensor([1., 4., 9.])
```

*What just happened:* Each operation lined up `a` and `b` by position and combined them. `a + b` added
the first elements (1 + 10), then the second (2 + 20), and so on - all at once, no `for` loop in your
code. `a ** 2` squared every element independently. To you it reads like ordinary arithmetic; underneath,
it's a single fast pass over the data.

PyTorch also ships the math functions you'd expect, plus reductions that collapse a tensor down to a
summary number:

```python
x = torch.tensor([1.0, 4.0, 9.0])

print(torch.sqrt(x))   # square root of each element
print(torch.exp(x))    # e^x for each element
print(x.sum())         # add everything up -> one number
print(x.mean())        # average -> one number
```

```console
tensor([1., 2., 3.])
tensor([2.7183e+00, 5.4598e+01, 8.1031e+03])
tensor(14.)
tensor(4.6667)
```

*What just happened:* `torch.sqrt` and `torch.exp` are elementwise - same shape out as in. `x.sum()` and
`x.mean()` are **reductions**: they squash the whole tensor into a single scalar tensor. You'll lean on
`sum` and `mean` constantly later - a loss value, for instance, is usually the *mean* error over a batch.

💡 Notice `x.sum()` returns `tensor(14.)`, not `14.` - it's still a tensor (a zero-dimensional one). If
you need a plain Python number out of it, call `.item()`.

## 2. Matrix multiply - the heart of a neural net

Elementwise math is the warm-up. The operation deep learning is actually *built* from is **matrix
multiplication**, written `@` (or `torch.matmul`).

Why does this one matter so much? Because a neural-network layer, stripped of mystique, is a matrix
multiply plus a bias: `output = input @ weights + bias`. That's it. Stack a few of those with some
non-linear functions between them and you have a model. So when people say training a model is
"expensive," they mostly mean: *a staggering number of matrix multiplies.*

```python
m = torch.tensor([[1.0, 2.0],
                  [3.0, 4.0]])          # shape (2, 2)
v = torch.tensor([[1.0],
                  [1.0]])               # shape (2, 1)

print(m @ v)                            # matrix-vector multiply
```

```console
tensor([[3.],
        [7.]])
```

*What just happened:* `m @ v` is true matrix multiplication, not elementwise. Row 1 of `m` (`[1, 2]`)
dotted with `v` (`[1, 1]`) gives `1*1 + 2*1 = 3`; row 2 (`[3, 4]`) gives `3*1 + 4*1 = 7`. The result has
shape `(2, 1)`. This dot-product-of-rows-and-columns pattern is the single most-run computation in all of
deep learning.

⚠️ **Shapes must align.** To multiply `(m, k) @ (k, n)` the inner dimensions must match - the number of
columns on the left must equal the number of rows on the right. Mismatch them and PyTorch stops you cold:

```python
left = torch.randn(2, 3)    # shape (2, 3)
right = torch.randn(2, 3)   # shape (2, 3) -- inner dims 3 and 2 don't match

print(left @ right)
```

```console
RuntimeError: mat1 and mat2 shapes cannot be multiplied (2x3 and 2x3)
```

*What just happened:* `left` is `(2, 3)` and `right` is `(2, 3)`. For `@` to work the inner dimensions
have to agree: `(2, 3) @ (3, n)` is fine, but here the left's `3` columns meet the right's `2` rows - 
no match, so PyTorch refuses rather than guessing. This is one of the most common errors you'll hit, and
the message tells you exactly which two shapes collided. Read it, fix the shapes, move on.

💡 The vast majority of the math inside a neural net is matrix multiplies. If you internalize the
`(m, k) @ (k, n) -> (m, n)` rule, you've internalized the shape-checking skill that prevents most layer
bugs before they happen.

## 3. Broadcasting

📝 **Broadcasting** is how PyTorch combines tensors of *different but compatible* shapes without you
manually copying data around. The classic case: you have a matrix and you want to add the same bias
vector to every row.

```python
matrix = torch.tensor([[1.0, 2.0, 3.0],
                       [4.0, 5.0, 6.0]])   # shape (2, 3)
bias = torch.tensor([10.0, 20.0, 30.0])    # shape (3,)

print(matrix + bias)
```

```console
tensor([[11., 22., 33.],
        [14., 25., 36.]])
```

*What just happened:* `matrix` is `(2, 3)` and `bias` is just `(3,)` - a single row. Instead of erroring,
PyTorch **broadcast** the bias: it acted as if `bias` were stretched to `(2, 3)` (copied down both rows)
and then added elementwise. `[10, 20, 30]` got added to row 1 *and* row 2. No actual copy happens in
memory - it's a view-level trick - which is why it's fast and memory-cheap. This is exactly the
`input + bias` step inside a layer.

The rule (identical to NumPy's): line the shapes up from the **right**. Two dimensions are compatible if
they're equal, or one of them is `1` (or missing). Here `(2, 3)` and `(3,)` align as `(2, 3)` and
`(1, 3)` - the `1` stretches to `2`. Done.

⚠️ Broadcasting is powerful but it's *silent* - it won't always error when you make a mistake; sometimes
it produces a valid-but-wrong shape. A `(3, 1)` and a `(1, 3)` will broadcast to `(3, 3)`, which is
occasionally what you wanted and occasionally a bug you won't notice until your loss looks insane.
**Always sanity-check the shape of a broadcast result** with `.shape` if you're unsure:

```python
col = torch.tensor([[1.0], [2.0], [3.0]])   # shape (3, 1)
row = torch.tensor([[10.0, 20.0, 30.0]])    # shape (1, 3)

print((col + row).shape)                     # not (3,) or (3,1) -- it's (3,3)!
```

```console
torch.Size([3, 3])
```

*What just happened:* A column `(3, 1)` plus a row `(1, 3)` broadcast both directions and produced a full
`(3, 3)` grid - every column value added to every row value. Perfectly legal PyTorch, and a real surprise
if you expected a length-3 result. The lesson isn't "avoid broadcasting" - it's "print the shape when the
result matters."

## 4. Reshaping & indexing

Most of practical PyTorch is getting data into the *shape* a layer expects, then pulling pieces back out.
PyTorch gives you sharp tools for both.

**Reshape** rearranges the same data into a new shape (the total number of elements must stay the same).
`.reshape()` and `.view()` do the same thing for our purposes - `.reshape()` is the safe default.

```python
x = torch.arange(6)          # tensor([0, 1, 2, 3, 4, 5]), shape (6,)

print(x.reshape(2, 3))       # rearrange into 2 rows, 3 cols
print(x.reshape(3, 2))       # or 3 rows, 2 cols
```

```console
tensor([[0, 1, 2],
        [3, 4, 5]])
tensor([[0, 1],
        [2, 3],
        [4, 5]])
```

*What just happened:* The six numbers never changed - only how they're laid out. `(6,)` became `(2, 3)`
and then `(3, 2)`. The element count (6) is identical each time, which is the one rule reshape enforces.

**`.squeeze()` and `.unsqueeze()`** remove or add a dimension of size 1. This sounds fussy, but it's
everywhere - a model often expects a batch dimension, so you wrap a single example with `.unsqueeze(0)`,
and you peel an extra dimension back off the output with `.squeeze()`.

```python
single = torch.tensor([1.0, 2.0, 3.0])   # shape (3,)

batched = single.unsqueeze(0)            # add a dim at position 0
print(batched.shape)                     # (1, 3) -- now it's "a batch of 1"

print(batched.squeeze().shape)           # remove the size-1 dim -> back to (3,)
```

```console
torch.Size([1, 3])
torch.Size([3])
```

*What just happened:* `unsqueeze(0)` inserted a new axis at the front, turning a lone `(3,)` vector into a
`(1, 3)` "batch of one" - the shape many models demand. `squeeze()` then dropped that size-1 axis, getting
us back to `(3,)`. You'll do this dance constantly when feeding single examples to a model built for
batches.

**Indexing and slicing** work just like NumPy (and Python lists), including for multiple dimensions:

```python
g = torch.tensor([[10, 11, 12],
                  [20, 21, 22],
                  [30, 31, 32]])

print(g[0])        # first row
print(g[:, 1])     # second column (all rows, column 1)
print(g[1, 2])     # single element: row 1, col 2
```

```console
tensor([10, 11, 12])
tensor([11, 21, 31])
tensor(22)
```

*What just happened:* `g[0]` grabbed the whole first row. `g[:, 1]` used `:` to mean "all rows" and `1`
to pick column 1 - that's how you slice a column. `g[1, 2]` indexed both dimensions at once for a single
element. Same syntax you already know from NumPy, no relearning required.

💡 Shape-wrangling is a genuinely large part of day-to-day PyTorch, and here's the plain truth most
tutorials skip: **the majority of PyTorch bugs are shape bugs.** A model that "doesn't work" is far more
often a `(batch, features)` that should've been `(features, batch)` than a deep conceptual error. When
something breaks, print `.shape` first.

## 5. The GPU

Now the payoff. 📝 Everything above runs fine on your CPU - but deep learning runs on **GPUs**, and it's
worth understanding *why.*

A CPU has a handful of very fast, very general cores (see
[CPU, RAM & Storage](/guides/cpu-ram-and-storage) for what a core actually is). It's brilliant at doing
one complicated thing after another. A GPU is the opposite: **thousands of simpler cores** that all do
math in parallel. And what is a matrix multiply? Thousands of independent multiply-and-add operations that
can all happen at once. That's a perfect match. The same training that takes a CPU hours can take a GPU
minutes, purely because the GPU does the parallel arithmetic of deep learning far faster.

In PyTorch, a tensor lives on a **device** - either `"cpu"` or `"cuda"` (NVIDIA GPU). You move tensors
between devices with `.to(device)`. The standard, do-it-once-at-the-top pattern looks like this:

```python
# Pick the device ONCE, near the top of your program
device = "cuda" if torch.cuda.is_available() else "cpu"
print(f"Using device: {device}")

x = torch.tensor([1.0, 2.0, 3.0])   # created on the CPU by default
x = x.to(device)                    # move it to the chosen device

print(x.device)
```

```console
Using device: cuda
tensor([1., 2., 3.], device='cuda:0')
```

*What just happened:* `torch.cuda.is_available()` checks whether a usable GPU exists; if so we pick
`"cuda"`, otherwise we fall back to `"cpu"`. Then `x.to(device)` moved the tensor onto that device - note
the output now shows `device='cuda:0'` (GPU number 0). The exact same code runs unchanged on a laptop with
no GPU (it'd print `cpu`) and on a Colab machine with one. That's the goal: **device-agnostic code.**

⚠️ Here's THE classic GPU error, the one that bites everyone exactly once: doing math across tensors on
**different devices.** If your model is on the GPU but your data is still on the CPU, PyTorch refuses to
guess where the work should happen:

```python
a = torch.tensor([1.0, 2.0]).to("cuda")   # on the GPU
b = torch.tensor([3.0, 4.0])              # still on the CPU

print(a + b)
```

```console
RuntimeError: Expected all tensors to be on the same device, but found at least
two devices, cuda:0 and cpu!
```

*What just happened:* `a` lives on `cuda:0` and `b` lives on `cpu`. An operation can't span two devices - 
the data would have to physically move, and PyTorch won't do that silently behind your back. The error
names both devices so you can see the mismatch. The fix is to put `b` on the GPU too (`b = b.to("cuda")`).

💡 The whole headache disappears if you adopt one habit: **pick `device` once, then `.to(device)`
everything** - your model and every batch of data. Because you reuse the same `device` variable, your
tensors can't drift onto different devices. This one discipline prevents the most common GPU bug in
PyTorch, and it's the pattern you'll see in every training loop from Phase 6 onward.

## Recap

- **Elementwise math** (`+`, `*`, `torch.sqrt`, `torch.exp`, …) acts on the whole tensor at once, no
  loops; reductions like `.sum()` and `.mean()` collapse a tensor to a summary value.
- **Matrix multiply** (`@` / `torch.matmul`) is the operation neural nets are built from - a layer is a
  matmul plus a bias. Inner dimensions must align: `(m, k) @ (k, n) -> (m, n)`.
- **Broadcasting** combines different-but-compatible shapes (like adding a bias vector to every row)
  without copying - but it's silent, so check the result's `.shape` when in doubt.
- **Reshaping & indexing** (`.reshape`/`.view`, `.squeeze`/`.unsqueeze`, NumPy-style slicing) get data
  into the shape a layer expects. Most PyTorch bugs are shape bugs - print `.shape` first.
- **GPUs** do the parallel matmuls of deep learning far faster than CPUs. Write device-agnostic code: pick
  `device` once, `.to(device)` everything, and never mix tensors across devices.

## Quick check

```quiz
[
  {
    "q": "You write a @ b where a has shape (4, 3) and b has shape (4, 3). What happens?",
    "choices": ["It returns a (4, 3) tensor, multiplied elementwise", "RuntimeError: the inner dimensions (3 and 4) don't match", "It returns a (4, 4) tensor"],
    "answer": 1,
    "explain": "Matrix multiply needs inner dims to match: (m, k) @ (k, n). Here it's (4, 3) @ (4, 3), so the left's 3 columns meet the right's 4 rows -- mismatch, RuntimeError. For elementwise multiply you'd use *, not @."
  },
  {
    "q": "You add a tensor of shape (2, 3) to one of shape (3,). What does broadcasting do?",
    "choices": ["Errors, because the shapes are different", "Stretches the (3,) vector across both rows, giving a (2, 3) result", "Sums everything into a single number"],
    "answer": 1,
    "explain": "Aligning from the right, (3,) is treated as (1, 3) and stretched to (2, 3) -- the same vector is added to every row. No data is actually copied in memory."
  },
  {
    "q": "What's the reliable way to avoid the 'tensors on different devices' RuntimeError?",
    "choices": ["Always use the CPU and never touch the GPU", "Pick device once at the top, then .to(device) both the model and every batch of data", "Call torch.cuda.is_available() before every operation"],
    "answer": 1,
    "explain": "Device-agnostic code: choose one device variable up front and move everything to it. Reusing the same variable means your tensors can't drift onto different devices."
  }
]
```


---

# Autograd: Automatic Differentiation

This is the phase where PyTorch stops looking like a fancy NumPy and starts looking like a machine that
*learns*. Everything in [Phase 2](02-tensor-operations-and-gpu.md) - elementwise math, matmul,
broadcasting - was you doing arithmetic. Autograd is PyTorch quietly watching that arithmetic and, on
request, working out the calculus of it for you. If you've ever wondered what "training" actually *does*
under the hood, this is the engine room.

## 1. Why gradients matter at all

📝 First, a quick refresher from [How a Model Learns](/guides/how-a-model-learns). A model is a big bundle
of adjustable numbers (its **parameters**, or weights). Training is a loop: the model makes a guess, you
measure how wrong it was (the **loss**), and then you nudge every parameter a tiny bit in the direction
that makes the loss smaller. Repeat millions of times and the guesses get good. That nudging process is
called **gradient descent**.

The word doing all the work there is *direction*. For each parameter, which way should it move - up or
down - to reduce the loss, and by how much? That answer is the **gradient**: the derivative of the loss
with respect to that parameter. A positive gradient means "increasing this parameter increases the loss,
so go the other way"; a large gradient means "this parameter has a big effect, take a bigger step."

Here's the problem. A real model has *millions* of parameters, all tangled together through layers of
matrix multiplies and non-linear functions. Computing the derivative of the loss with respect to every
single one of them, by hand, with the chain rule, is not hard - it's *impossible* at that scale. Nobody
does it. Instead:

💡 **Autograd does the calculus for you.** You write the forward math (the guess and the loss). PyTorch
figures out all the derivatives. That's the whole deal, and it's why PyTorch exists.

## 2. `requires_grad` - telling PyTorch to pay attention

📝 By default, a tensor is just data - PyTorch does the arithmetic and forgets how it got there. But set
`requires_grad=True` on a tensor and PyTorch flips into record mode: from then on, **every operation you
do with that tensor gets logged**, building up a behind-the-scenes map of the computation. That map is
what makes the calculus possible later.

```python
import torch

x = torch.tensor(3.0, requires_grad=True)
print(x)

y = x ** 2 + 1      # do some math with x
print(y)
```

```console
tensor(3., requires_grad=True)
tensor(10., grad_fn=<AddBackward0>)
```

*What just happened:* We created `x` with `requires_grad=True`, marking it as something we'll want
gradients for. When we computed `y = x ** 2 + 1`, PyTorch didn't just store the answer (`10.`) - look at
`y`'s printout: `grad_fn=<AddBackward0>`. That `grad_fn` is PyTorch's note-to-self: "this tensor was
produced by an addition, and here's how to differentiate back through it." Every operation downstream of a
`requires_grad` tensor carries one of these. That trail of `grad_fn`s *is* the recording.

⚠️ Only **floating-point** tensors can require gradients. Calculus needs continuous numbers - you can't
take the derivative with respect to an integer. `torch.tensor(3, requires_grad=True)` (note: no decimal
point) will raise an error. Use `3.0`.

## 3. `.backward()` and `.grad` - collecting the answers

So PyTorch has recorded the forward computation. Now you ask for the gradients.

📝 You call `.backward()` on the final scalar result (usually the loss). PyTorch walks the recorded graph
**backwards** - from the result, back through every operation, applying the chain rule at each step - and
deposits the gradient for each input tensor into its `.grad` attribute. One call, all the derivatives.

Let's verify it on something we can check by hand. We know from basic calculus that if `y = x²`, then
`dy/dx = 2x`. So at `x = 3`, the gradient should be `2 × 3 = 6`.

```python
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2

y.backward()        # walk the graph backwards, compute dy/dx

print(x.grad)       # should be 2 * x = 6
```

```console
tensor(6.)
```

*What just happened:* `y.backward()` triggered the reverse pass. PyTorch knew (from the recorded `grad_fn`)
that `y` came from squaring `x`, applied the rule `d(x²)/dx = 2x`, plugged in `x = 3`, and stored the
result - `6.` - into `x.grad`. We never wrote the derivative ourselves; we only wrote the forward math
`x ** 2`. That number in `x.grad` is exactly what gradient descent needs to know which way to nudge `x`.

Here's the same computation as a picture. The forward pass (left to right) builds the graph; `.backward()`
runs the reverse pass (right to left), accumulating the gradient back to `x`:

```mermaid
flowchart LR
  x["x = 3.0<br/>requires_grad"] -->|"square"| y["y = 9.0<br/>grad_fn"]
  y -.->|"backward: dy/dx = 2x = 6"| x
```

💡 `.backward()` expects a **scalar** (a single number) by default - which is exactly what a loss is. If you
try to call it on a multi-element tensor, PyTorch won't know how to collapse those into one number to
differentiate and will ask you for guidance. In practice you almost always call it on a loss, so this
rarely bites you. If it does, the fix is usually `.mean()` or `.sum()` to reduce to a scalar first (you
saw reductions in [Phase 2](02-tensor-operations-and-gpu.md)).

## 4. The computation graph (and why it's *dynamic*)

That map PyTorch records has a name: the **computation graph**. The crucial thing about PyTorch's version - 
and the reason it feels so natural to write - is that it's built **dynamically, as your code runs.** This is
called **define-by-run**.

📝 There's no separate "define the model, then compile it" step. The graph is constructed on the fly, one
operation at a time, exactly as Python executes your lines. The graph for a given forward pass only exists
because you *ran* that forward pass.

This has a wonderful consequence: **ordinary Python control flow just works inside a model.** An `if` that
takes different branches depending on the data, a `for` loop whose length varies, a `while` - all fine.
Whatever path the code actually took, that's the graph that got recorded, so that's the graph `.backward()`
walks back through.

```python
def wobble(x):
    y = x * 2
    if y.sum() > 0:        # a real Python branch, decided at runtime
        y = y * 3
    else:
        y = y - 100
    return y.sum()

x = torch.tensor([1.0, 2.0], requires_grad=True)
loss = wobble(x)
loss.backward()
print(x.grad)
```

```console
tensor([6., 6.])
```

*What just happened:* The function ran like normal Python. `y.sum()` was positive, so the `if` branch
fired and `y` got multiplied by 3. The graph PyTorch recorded reflects the path actually taken:
`x → ×2 → ×3 → sum`. Differentiating that chain gives `2 × 3 = 6` for each element, which is what landed in
`x.grad`. With a different input that took the `else` branch, a *different* graph would have been built and
a different gradient computed. You didn't have to declare those branches to any framework in advance - you
just wrote Python.

💡 One more detail with big practical impact: **the graph is freed after `.backward()`.** Once the reverse
pass finishes, PyTorch throws the graph away to save memory (it assumes you're done with it). Call
`.backward()` a second time on the same graph and you'll get a "backward through the graph a second time"
error. That's by design - each training step builds a fresh graph from a fresh forward pass.

## 5. Turning autograd off, and the one gotcha that gets everyone

Recording every operation costs time and memory. When you're **not** training - say you're just running the
finished model to make predictions (**inference**) - you don't need gradients, and you should switch the
machinery off.

The tool is `torch.no_grad()`, a context manager that says "don't record anything inside this block":

```python
x = torch.tensor(3.0, requires_grad=True)

with torch.no_grad():
    y = x ** 2          # computed, but NOT recorded
print(y.requires_grad)  # the result is a plain tensor

with torch.no_grad():
    pass
print(x.requires_grad)  # x itself is unchanged
```

```console
False
True
```

*What just happened:* Inside the `torch.no_grad()` block, PyTorch did the math but skipped the bookkeeping - 
notice `y.requires_grad` is `False`, so `y` has no `grad_fn` and can't be backpropagated through. That's
faster and lighter, which is exactly what you want for inference. Outside the block, `x` is untouched and
still tracks gradients. You'll wrap your evaluation and prediction code in `with torch.no_grad():` as a
matter of habit.

There's also `.detach()`, which gives you a copy of a tensor that's cut loose from the graph - same numbers,
no history. Reach for it when you want to pull a value *out* of a computation to log it, store it, or feed
it somewhere that shouldn't backpropagate.

```python
x = torch.tensor(3.0, requires_grad=True)
y = x ** 2

clean = y.detach()      # same value, no graph attached
print(y.requires_grad, clean.requires_grad)
```

```console
True False
```

*What just happened:* `y` is still part of the graph (`requires_grad=True`), but `clean` is `y`'s value
snapped off from its history - a plain `tensor(9.)` you can use freely without dragging the whole graph
along. This is the safe way to say "I just want the number, not the calculus."

Now the big one. ⚠️ **Gradients accumulate.** When you call `.backward()`, PyTorch *adds* the new gradients
into `.grad` - it does not overwrite them. So if you compute backward twice without clearing in between, the
gradients pile up and your update is wrong.

```python
x = torch.tensor(3.0, requires_grad=True)

y = x ** 2
y.backward()
print(x.grad)           # 6, as expected

z = x ** 2
z.backward()
print(x.grad)           # NOT 6 -- it's 12, the two added together!
```

```console
tensor(6.)
tensor(12.)
```

*What just happened:* The first `backward()` put `6.` in `x.grad`. The second one computed another `6.` and
**added** it to what was already there, giving `12.`. PyTorch did exactly what it always does - accumulate - 
but if you expected each step to start clean, this is a silent, brutal bug: your model trains on garbage
gradients and you stare at a loss that won't go down.

💡 This is the **single most common training bug** in PyTorch, and the fix is one line you'll run at the top
of every training step: zero out the gradients before each backward pass. In a real loop that's
`optimizer.zero_grad()` (manually, it's `x.grad.zero_()`). We'll wire this into the full training loop in
[Phase 6](06-the-training-loop.md) - for now, just burn into memory: **gradients accumulate; you
must clear them every step.**

Step back and admire the deal you're getting: you write the forward math - the guess and the loss, in plain
Python - and autograd derives the entire backward pass for free. That division of labor is what the rest of
this guide is built on. Layers (Phase 4), optimizers (Phase 5), and the training loop (Phase 6) are all
just convenient ways to organize the forward math and let autograd handle the rest.

## Recap

- A model learns by **gradient descent**: nudge each parameter toward lower loss. The **gradient** is the
  derivative of the loss w.r.t. a parameter - it tells you which way and how far to nudge. At scale, no one
  computes these by hand; autograd does it.
- Set **`requires_grad=True`** (on float tensors) and PyTorch records every operation, attaching a
  `grad_fn` to each result - that trail is the **computation graph**.
- Call **`.backward()`** on the final scalar (the loss); PyTorch walks the graph in reverse with the chain
  rule and fills each input's **`.grad`**. Verified: `y = x²` gives `x.grad == 2x`.
- The graph is **define-by-run** - built dynamically as your code executes, so normal Python `if`/`for`
  control flow works inside models. It's freed after `.backward()`.
- Turn autograd off for inference with **`torch.no_grad()`**; use **`.detach()`** to pull a value out of the
  graph. ⚠️ Gradients **accumulate** in `.grad`, so you must **zero them every training step** - the #1
  PyTorch training bug.

## Quick check

```quiz
[
  {
    "q": "You create x = torch.tensor(2.0, requires_grad=True), compute y = x ** 3, then call y.backward(). What ends up in x.grad?",
    "choices": ["8.0, the value of y", "12.0, the derivative 3x² evaluated at x=2", "Nothing -- you must compute the gradient by hand", "2.0, the value of x"],
    "answer": 1,
    "explain": "backward() applies the chain rule: d(x³)/dx = 3x², and at x=2 that's 3·4 = 12. Autograd derives and stores the gradient in x.grad; you only wrote the forward math."
  },
  {
    "q": "Why does PyTorch's define-by-run computation graph let you use ordinary Python if/for control flow inside a model?",
    "choices": ["Because PyTorch compiles your whole model ahead of time and analyzes every branch", "Because the graph is built dynamically as the code runs, so it records exactly the path actually taken", "Because control flow is automatically converted to matrix operations", "It doesn't -- branches and loops are forbidden inside models"],
    "answer": 1,
    "explain": "The graph is constructed on the fly as Python executes. Whatever branch or loop iterations actually ran are what gets recorded, so backward() walks back through that exact path."
  },
  {
    "q": "You call loss.backward() twice in a row without clearing anything in between. What goes wrong?",
    "choices": ["Nothing -- the second call overwrites the first", "The gradients accumulate (add together), so .grad is now wrong -- you must zero gradients each step", "PyTorch automatically averages the two gradients for you", "The model trains twice as fast"],
    "answer": 1,
    "explain": "backward() ADDS into .grad rather than overwriting. Without zeroing gradients between steps they pile up and corrupt the update -- the most common PyTorch training bug. Fix: optimizer.zero_grad() (or x.grad.zero_()) each step."
  }
]
```


---

# Building Models with nn.Module

In [Phase 3](03-autograd.md) you saw autograd quietly record every operation on a tensor that has
`requires_grad=True`, then hand you the gradients on demand. That's the engine of learning; this phase is
about the thing autograd runs *inside* - the model.

Here's the mental model to hold onto, and it's one you already know from Python: **a model is a class.**
Specifically, a class that subclasses `nn.Module`. You met classes in
[Objects & Classes](/guides/python-from-zero) - data bundled with the behavior that acts on it. A PyTorch
model is exactly that: the *data* is the layers (each holding learnable weights), and the *behavior* is the
forward pass (how an input flows through those layers to an output). Nothing mystical. If you can write a
`Dog` class, you can write a neural network.

`nn.Module` is the parent class you inherit from, and inheriting from it buys you a lot for free - 
parameter tracking, device moves, train/eval switching. We'll build up from the smallest possible model to
a real two-layer network, and end by looking at the parameters the optimizer will update in
[Phase 5](05-loss-and-optimizers.md).

## 1. A model is a class

📝 To define a model you **subclass `nn.Module`**, define your layers in `__init__`, and define the forward
pass in a method called `forward`. That's the whole pattern. Here is the smallest model that does anything:

```python
import torch
import torch.nn as nn

class TinyModel(nn.Module):
    def __init__(self):
        super().__init__()              # let nn.Module do its setup FIRST
        self.layer = nn.Linear(3, 1)    # one layer, defined as an attribute

    def forward(self, x):
        return self.layer(x)            # the forward pass: input -> output

model = TinyModel()
print(model)
```

```console
TinyModel(
  (layer): Linear(in_features=3, out_features=1, bias=True)
)
```

*What just happened:* `TinyModel(nn.Module)` means "a TinyModel *is an* `nn.Module`" - the same `is-a`
inheritance from the Python guide. The `super().__init__()` call runs `nn.Module`'s own constructor, which
sets up the bookkeeping that tracks your layers. Then `self.layer = nn.Linear(3, 1)` stored a layer *on
this model*, exactly like storing `self.name` on a dog. Printing the model shows PyTorch already knows about
that layer - because `nn.Module` was watching when you assigned it.

⚠️ **Always call `super().__init__()` first**, before assigning any layers. `nn.Module`'s constructor sets
up the internal machinery that records your layers and their parameters. Skip it (or assign layers before
it) and you'll get a confusing `AttributeError` like *"cannot assign module before Module.__init__() call"*.

> 💡 **Key point.** What `nn.Module` gives you for inheriting from it: it **tracks every parameter** in
> every layer you assign (so the optimizer can find them), it **moves the whole model to a device** with one
> `model.to(device)` call, and it **toggles train/eval mode** with `model.train()` / `model.eval()`. You get
> all of that by writing `class MyModel(nn.Module)` and calling `super().__init__()`. That's the payoff.

## 2. nn.Linear - a layer is a matmul plus a bias

📝 A **linear layer** (also called *fully-connected* or *dense*) computes `output = input @ W + b`. That's
the exact matrix-multiply-plus-bias from [Phase 2](02-tensor-operations-and-gpu.md) - `nn.Linear` is just
that operation wrapped up with its weights bundled inside.

`nn.Linear(in_features, out_features)` creates two tensors for you: a weight matrix `W` and a bias vector
`b`. Crucially, it creates them already marked as **learnable** - their `requires_grad` is `True`
automatically (tying back to [Phase 3](03-autograd.md)), so autograd will track them and produce gradients.
You don't set that up by hand.

```python
layer = nn.Linear(3, 2)        # 3 inputs in, 2 outputs out

print(layer.weight.shape)      # the W matrix
print(layer.bias.shape)        # the b vector
print(layer.weight.requires_grad)
```

```console
torch.Size([2, 3])
torch.Size([2])
True
```

*What just happened:* `nn.Linear(3, 2)` built a weight of shape `(2, 3)` and a bias of shape `(2,)` - sized
so that an input with 3 features maps to 2 outputs. (PyTorch stores `W` as `(out, in)` and computes
`input @ W.T + b` under the hood, which is why it's `(2, 3)` and not `(3, 2)` - you rarely need to think
about the transpose.) Both were initialized to small random values and, as the last line shows, both already
have `requires_grad=True`. These are the numbers training will adjust.

## 3. forward() and calling the model

📝 You define the forward pass in a method named `forward(self, x)` - but you **call the model directly**,
as `model(x)`, *not* `model.forward(x)`. Writing `model(x)` triggers `nn.Module`'s `__call__`, which runs
some setup (like hooks and train/eval handling) and *then* calls your `forward`. This is the same dunder-method
trick you saw with `__init__` in the Python guide: PyTorch defines `__call__` so that `model(x)` "just works."

```python
model = TinyModel()             # has one nn.Linear(3, 1) inside

x = torch.randn(4, 3)           # a batch of 4 examples, each with 3 features
output = model(x)               # call the model -> runs forward()

print(output.shape)
```

```console
torch.Size([4, 1])
```

*What just happened:* `model(x)` invoked `__call__`, which ran your `forward`, which pushed `x` through the
linear layer. The input was `(4, 3)` - 4 examples of 3 features each - and the layer mapped each example's
3 features to 1 output, giving `(4, 1)`. Notice the batch dimension (4) flows straight through untouched;
layers operate per-example. This is the shape-tracking habit from Phase 2 paying off.

⚠️ **Call `model(x)`, never `model.forward(x)` directly.** Calling `forward` yourself skips the wrapper
work `__call__` does (hooks, and the train/eval bookkeeping that layers like dropout and batch-norm rely on).
Most days it'll *seem* to work, then silently misbehave the one time it matters. Build the `model(x)` habit
now and you'll never get bitten.

## 4. Activations and stacking layers

Here's a subtle, important truth: 📝 **stacking linear layers with nothing between them gains you nothing.**
Two matrix multiplies in a row are mathematically just *one* matrix multiply (the product of the two weight
matrices). So a 10-layer all-linear network has exactly the same expressive power as a single linear layer - 
it can only draw straight lines.

The fix is a **nonlinearity** (an *activation function*) between the linear layers. The most common is
**ReLU** (`nn.ReLU` or `torch.relu`), which is dead simple: it turns negatives into zero and leaves
positives alone. That tiny kink is enough to break the "stacked linears collapse to one" trap and let the
network learn curved, complicated boundaries. The pattern is **Linear → activation → Linear**.

Let's build a real two-layer **MLP** (multi-layer perceptron):

```python
class MLP(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(3, 8)      # 3 features -> 8 hidden units
        self.relu = nn.ReLU()           # the nonlinearity
        self.fc2 = nn.Linear(8, 1)      # 8 hidden units -> 1 output

    def forward(self, x):
        x = self.fc1(x)                 # first linear layer
        x = self.relu(x)                # nonlinearity in between
        x = self.fc2(x)                 # second linear layer
        return x

model = MLP()
out = model(torch.randn(4, 3))
print(out.shape)
```

```console
torch.Size([4, 1])
```

*What just happened:* The input `(4, 3)` flowed through `fc1` to become `(4, 8)` - 8 hidden features per
example. `relu` then zeroed out the negatives (same shape, `(4, 8)`), and `fc2` mapped those 8 hidden
features down to 1 output, giving `(4, 1)`. The `forward` method reads top-to-bottom like a recipe: that's
the whole point of defining it yourself - you control exactly how data flows.

For a plain stack like this, `nn.Sequential` is shorthand that chains layers in order, so you don't write
the `forward` by hand at all:

```python
model = nn.Sequential(
    nn.Linear(3, 8),
    nn.ReLU(),
    nn.Linear(8, 1),
)

out = model(torch.randn(4, 3))
print(out.shape)
```

```console
torch.Size([4, 1])
```

*What just happened:* `nn.Sequential` built a module that runs each layer in the order listed, feeding each
one's output into the next - the same Linear → ReLU → Linear pipeline as the `MLP` class, in fewer lines. It
produced the identical `(4, 1)` output. 💡 Use `nn.Sequential` when your model is a straight chain; write a
full `nn.Module` subclass with a custom `forward` when you need branches, skip connections, or any logic that
isn't a simple line. Most real models start as `Sequential` and grow into a custom class.

## 5. Parameters - what gets learned

Every layer you defined holds learnable tensors (the `W`s and `b`s). `nn.Module` collects them all so you
never have to round them up yourself. Two methods matter:

- **`model.parameters()`** - yields every learnable tensor in the model. This is exactly what you'll hand
  to the optimizer in [Phase 5](05-loss-and-optimizers.md) so it knows what to update.
- **`model.state_dict()`** - a dictionary mapping each layer's name to its current values. This is what you
  save to disk in [Phase 9](09-saving-loading-inference.md).

A common sanity check is counting how many learnable numbers a model has:

```python
model = MLP()       # Linear(3,8) + Linear(8,1)

total = sum(p.numel() for p in model.parameters())
print(f"Trainable parameters: {total}")
```

```console
Trainable parameters: 41
```

*What just happened:* `model.parameters()` walked every layer and yielded each weight and bias tensor;
`p.numel()` counted the elements in each. `fc1` has a `(8, 3)` weight (24) plus an `(8,)` bias (8) = 32, and
`fc2` has a `(1, 8)` weight (8) plus a `(1,)` bias (1) = 9, for 41 total. Those 41 numbers *are* the model - 
training is the process of nudging exactly these values until the outputs are good.

💡 **The big picture for this phase.** A model is a class made of learnable layers. You define what's in it
(`__init__`) and how data flows through it (`forward`), call it as `model(x)`, and `nn.Module` keeps track of
every parameter inside. Autograd (Phase 3) tracks those parameters; the optimizer (Phase 5) updates them.
That's the division of labor - and you've now got the middle piece.

## Recap

1. **A model is a class** that subclasses `nn.Module`. Define layers in `__init__` (after
   `super().__init__()`), define the forward pass in `forward(self, x)`.
2. **`nn.Linear(in, out)`** is a layer computing `input @ W + b`. It creates `W` and `b` for you, already
   marked learnable (`requires_grad=True`).
3. **Call the model as `model(x)`** - this runs `__call__`, which runs your `forward`. Never call
   `model.forward(x)` directly.
4. **Activations** (`nn.ReLU` / `torch.relu`) between linear layers add nonlinearity; without them, stacked
   linear layers collapse into a single linear layer. `nn.Sequential` is shorthand for a straight chain.
5. **`model.parameters()`** yields what the optimizer updates; **`model.state_dict()`** is what you save.
   The parameters *are* the model.

With the model built, the next two pieces complete the picture: a way to measure how wrong it is, and an
algorithm that uses the gradients to fix it. That's loss functions and optimizers.

## Quick check

```quiz
[
  {
    "q": "How should you run a forward pass on a model named `model` for input `x`?",
    "choices": ["model.forward(x)", "model(x)", "model.run(x)"],
    "answer": 1,
    "explain": "Call the model directly as model(x). That triggers nn.Module's __call__, which does setup (hooks, train/eval handling) and then runs your forward(). Calling model.forward(x) skips that wrapper and can silently misbehave."
  },
  {
    "q": "Why put a nn.ReLU() between two nn.Linear layers?",
    "choices": ["To make the model run faster", "Without a nonlinearity, the two linear layers collapse into a single linear layer", "ReLU is required for the model to compile"],
    "answer": 1,
    "explain": "Two matrix multiplies in a row equal one matrix multiply, so stacked linears have no more power than one. A nonlinearity like ReLU breaks that, letting the network learn curved, complex patterns."
  },
  {
    "q": "What does nn.Linear(in, out) set up for you automatically?",
    "choices": ["A weight and bias, both with requires_grad=True so autograd tracks them", "Just a weight matrix, with no bias", "A weight and bias that must be manually marked as learnable"],
    "answer": 0,
    "explain": "nn.Linear creates the weight W and bias b for you, already initialized and already marked learnable (requires_grad=True), so autograd tracks them and the optimizer can update them."
  }
]
```


---

# Loss Functions & Optimizers

In [Phase 4](04-building-models-with-nn-module.md) you built a model: a `nn.Module` with layers and a
`forward()` that turns an input into a prediction. But a fresh model is *random* - its weights are
nonsense, so its predictions are nonsense. Training fixes that, and fixing something takes two things: a
way to measure how wrong you are, and a way to act on that measurement.

Here's the mental model for this whole phase, and it's short: **the loss function tells you how wrong the
model is, and the optimizer is the thing that does something about it.** Loss is the score. The optimizer
is the player trying to lower the score. Everything below is just the PyTorch names for those two roles - 
and the three magic lines that connect them. This is the missing half of the training loop you'll
assemble in Phase 6.

## 1. Loss = how wrong, in one number

📝 A **loss function** takes the model's predictions and the true answers and boils the gap between them
down to a single number. Lower is better. A loss of zero means the predictions matched the truth exactly;
a big loss means the model is badly off. That's the entire idea - *training is the act of making that one
number smaller.*

This is the same picture from [How a Model Learns](/guides/how-a-model-learns): a model learns by being
wrong, measuring *how* wrong, and nudging its numbers to be a little less wrong next time. The loss
function is the "how wrong" part, made concrete. PyTorch ships the common ones in `torch.nn`, ready to use.

💡 One number is the point, not a limitation. The optimizer needs a single value to push downhill - you
can't minimize ten numbers at once. The loss function's job is to be the impartial scorekeeper that compresses
"how did the whole batch do?" into one comparable score.

## 2. The two losses you'll reach for most

📝 Two loss functions cover the overwhelming majority of beginner work, and which one you pick is decided
by *what kind of problem you have*:

- **`nn.MSELoss`** - for **regression** (predicting a number: a price, a temperature). It's the mean
  squared error: average of `(prediction − target)²`.
- **`nn.CrossEntropyLoss`** - for **classification** (predicting a category: cat vs. dog, digit 0–9).

Let's compute a regression loss. You create the loss object once, then call it like a function with
`(predictions, targets)`:

```python
import torch
import torch.nn as nn

loss_fn = nn.MSELoss()

predictions = torch.tensor([2.5, 0.0, 2.1])   # what the model guessed
targets     = torch.tensor([3.0, 0.0, 2.0])   # the true values

loss = loss_fn(predictions, targets)
print(loss)
```

```console
tensor(0.0867)
```

*What just happened:* `nn.MSELoss()` built a loss object; calling `loss_fn(predictions, targets)` measured
the gap. Element by element the errors are `-0.5`, `0.0`, `0.1`; squared they're `0.25`, `0.0`, `0.01`;
their mean is `0.0867`. One small number, because the guesses were close. If a prediction had been wildly
off, squaring would have blown that error up and the loss would be large - that's MSE punishing big misses
hard.

Now classification, where there's a notorious trap. ⚠️ **`nn.CrossEntropyLoss` expects RAW logits - the
plain, un-softmaxed numbers straight out of your model's last layer - together with the true class labels
as plain integers.** Applying a softmax yourself before passing predictions in is the classic
CrossEntropyLoss bug: it double-applies the math and quietly wrecks your training.

```python
loss_fn = nn.CrossEntropyLoss()

# Raw scores (logits) for 2 examples over 3 classes -- NO softmax applied
logits = torch.tensor([[2.0, 0.5, 0.1],    # example 1: model leans toward class 0
                       [0.1, 0.2, 3.0]])   # example 2: model leans toward class 2

targets = torch.tensor([0, 2])              # true classes, as integers (not one-hot)

loss = loss_fn(logits, targets)
print(loss)
```

```console
tensor(0.2559)
```

*What just happened:* We passed `logits` (raw, unnormalized scores) and `targets` as a tensor of integer
class indices - `0` means "example 1's correct answer is class 0," `2` means "example 2's is class 2."
`CrossEntropyLoss` internally does the softmax *for* us and then measures how much probability the model
put on the right class. Both examples leaned toward the correct class, so the loss is low. Pass it
pre-softmaxed numbers or one-hot labels and you'll either get an error or, worse, silently wrong training.

💡 Remember the contract: **raw logits in, integer labels in, softmax stays out of your hands.** If you
ever catch yourself writing `softmax(...)` right before a `CrossEntropyLoss`, delete it.

## 3. The optimizer - the thing that updates the weights

📝 The loss tells you *how wrong*. Autograd (Phase 3) tells you *which direction* each weight should move
to reduce that wrongness - the gradients. The **optimizer** is what actually takes those gradients and
*adjusts the weights*. It's the mechanism of learning: no optimizer, no improvement.

Optimizers live in `torch.optim`. You create one by handing it two things: the parameters it's allowed to
change, and a learning rate. Remember `model.parameters()` from Phase 4 - that's the bundle of every
weight and bias in your model. You pass it in so the optimizer knows exactly *what* it's responsible for
updating:

```python
import torch.optim as optim

model = nn.Linear(4, 2)     # a tiny model from Phase 4: 4 inputs -> 2 outputs

optimizer = optim.SGD(model.parameters(), lr=0.01)
print(optimizer)
```

```console
SGD (
Parameter Group 0
    dampening: 0
    lr: 0.01
    ...
)
```

*What just happened:* `optim.SGD(model.parameters(), lr=0.01)` created an optimizer wired directly to this
model's weights. By passing `model.parameters()`, we told it "these are the numbers you may change." From
now on, when we ask the optimizer to take a step, it walks through exactly those parameters and nudges each
one. The `lr=0.01` is the learning rate - coming up next.

## 4. SGD vs. Adam, and the learning rate

📝 You'll meet two optimizers early. **SGD** (Stochastic Gradient Descent) is the textbook one: for each
weight, step a little bit in the downhill direction - `new_weight = old_weight − (gradient × learning
rate)`. Simple and predictable. **Adam** is the smarter default: it adapts the step size per-parameter as it
goes, which usually means it learns faster and needs less hand-tuning. Swapping between them is a one-line
change:

```python
sgd  = optim.SGD(model.parameters(), lr=0.01)
adam = optim.Adam(model.parameters(), lr=1e-3)   # 1e-3 = 0.001

print(type(sgd).__name__, type(adam).__name__)
```

```console
SGD Adam
```

*What just happened:* Same `model.parameters()`, two different update strategies. SGD will take steps of a
fixed size scaled by the gradient; Adam will quietly tune each parameter's step on the fly. The API is
identical - you'll use the exact same three lines (next section) regardless of which one you chose.

📝 That `lr` - the **learning rate** - is the size of each step downhill, and ⚠️ **it's the single most
important hyperparameter you'll touch.** Set it too high and the model overshoots the bottom on every step,
bouncing around or blowing up (loss goes to `nan`). Set it too low and learning crawls - technically
correct, but it might take a thousand times longer than it should. Most "my model won't learn" problems
trace back to the learning rate.

💡 When in doubt, start with **Adam and `lr=1e-3` (0.001)**. It's the closest thing PyTorch has to a safe
default, and it's where the majority of real projects begin before any tuning. Get something training
first, fiddle with the learning rate second.

## 5. The three-line update

Here's where loss and optimizer finally meet. Every PyTorch training step - for the simplest linear model
and for a giant language model alike - runs these three lines after computing the loss. Learn them once and
you've learned the engine of all of deep learning:

```python
optimizer.zero_grad()   # 1. clear the old gradients
loss.backward()         # 2. autograd fills in fresh gradients
optimizer.step()        # 3. apply the update to every weight
```

```console
(no output -- this is the work itself)
```

*What just happened:* Three jobs, in order. **`optimizer.zero_grad()`** wipes the gradients from the last
step - ⚠️ this matters because PyTorch *accumulates* gradients by default (you saw this in Phase 3); skip
this line and old and new gradients pile up, corrupting the update. **`loss.backward()`** runs autograd
backward from the loss, computing a fresh gradient for every parameter - the "which way is downhill"
answer. **`optimizer.step()`** then reads those gradients and actually moves each weight, using whatever
strategy (SGD, Adam) you chose. Old grads cleared, new grads computed, step taken.

💡 The clean way to hold this in your head: **the loss says how wrong you are, autograd (`backward`) says
which way to go, and the optimizer (`step`) takes the step.** Three roles, three lines, in that exact
order. That ordering - clear, backward, step - is non-negotiable, and getting it wrong (especially
forgetting `zero_grad`) is one of the most common training bugs.

This is the heart of training. In [Phase 6](06-the-training-loop.md) we wrap these three lines inside a
loop that runs them over and over, batch after batch, epoch after epoch - and you'll watch the loss
actually fall.

## Recap

- A **loss function** measures how wrong the model is in one number; lower is better, and training is the
  act of shrinking it. It makes the "learn by being wrong" idea concrete.
- **`nn.MSELoss`** is for regression (predicting a number); **`nn.CrossEntropyLoss`** is for
  classification. ⚠️ CrossEntropyLoss wants **raw logits and integer labels** - never pre-apply softmax.
- The **optimizer** (`torch.optim.SGD`, `torch.optim.Adam`) takes autograd's gradients and updates the
  weights. You pass it `model.parameters()` so it knows what to change.
- **SGD** steps by gradient × learning rate; **Adam** adapts and is the usual default. The **learning
  rate** is the most important hyperparameter - too high diverges, too low crawls. Start with Adam + `1e-3`.
- The update is three lines, in order: **`optimizer.zero_grad()`** (clear old grads - they accumulate),
  **`loss.backward()`** (autograd fills grads), **`optimizer.step()`** (apply the update).

## Quick check

```quiz
[
  {
    "q": "What does a loss function compute?",
    "choices": ["The model's prediction for a new input", "One number measuring how far the predictions are from the true answers", "The learning rate for the optimizer"],
    "answer": 1,
    "explain": "A loss function turns predictions-vs-truth into a single number, lower is better. Training is the process of making that number smaller."
  },
  {
    "q": "You're doing classification with nn.CrossEntropyLoss. What should you feed it?",
    "choices": ["Softmax probabilities and one-hot labels", "Raw logits and integer class labels", "Raw logits and softmax probabilities"],
    "answer": 1,
    "explain": "CrossEntropyLoss expects raw logits (it applies softmax internally) plus integer class indices. Pre-applying softmax yourself is the classic bug that quietly breaks training."
  },
  {
    "q": "What is the correct order of the three update lines, and why call zero_grad() first?",
    "choices": ["step(), backward(), zero_grad() -- to apply before measuring", "zero_grad(), backward(), step() -- because PyTorch accumulates gradients, so old ones must be cleared first", "backward(), zero_grad(), step() -- to compute then reset before stepping"],
    "answer": 1,
    "explain": "Clear, backward, step. PyTorch adds new gradients onto existing ones by default, so zero_grad() wipes the previous step's grads before backward() computes fresh ones and step() applies them."
  }
]
```


---

# The Training Loop

You've now met all the pieces. In [Phase 4](04-building-models-with-nn-module.md) you built a model
that turns inputs into predictions. In [Phase 5](05-loss-and-optimizers.md) you got a **loss function**
that measures how wrong those predictions are, and an **optimizer** that knows how to adjust the model's
weights. This phase is where those pieces start moving - the ritual that actually trains a model, and it
never really changes, from a three-line toy to a giant language model.

So let's get the mental model dead clear before a single line of code, because this is THE thing to
internalize about PyTorch.

## 1. The mental model: a loop you write yourself

📝 Training is one short cycle, repeated thousands of times:

1. **Show the model some data** → it makes predictions (the *forward pass*).
2. **Measure how wrong it was** → the loss, one number.
3. **Compute which way to nudge each weight** to make the loss smaller → the *backward pass*.
4. **Take a small step** in that direction → the optimizer updates the weights.

Then do it again. And again. Each pass, the model is a little less wrong than the last. That slow,
patient nudging *is* learning - it's the gradient descent from
[How a Model Learns](/guides/how-a-model-learns), now made concrete in code.

Here's the part that surprises people coming from other libraries: **PyTorch has no `model.fit()`.**
There's no magic "train this for me" button. *You* write the loop. That sounds like more work, and the
first time it's a little intimidating - but it's a gift. Nothing is hidden. You can see and change every
step, which is exactly why researchers reach for PyTorch. And the loop is short and always the same shape,
so once you've written it once, you've written it forever.

```mermaid
flowchart LR
  A[Data X, y] --> B[pred = model X]
  B --> C[loss = loss_fn pred, y]
  C --> D[optimizer.zero_grad]
  D --> E[loss.backward]
  E --> F[optimizer.step]
  F -->|next epoch| A
```

That diagram is the whole phase. Everything below is just making each box concrete.

## 2. The canonical loop

Here it is - the most important code in this entire guide. Read it slowly. We'll dissect every line
afterward, but first take in the *shape* of it: setup, then a loop that repeats five steps.

```python
import torch
import torch.nn as nn

# Some toy data: learn y = 2x. X is (4, 1), y is (4, 1).
X = torch.tensor([[1.0], [2.0], [3.0], [4.0]])
y = torch.tensor([[2.0], [4.0], [6.0], [8.0]])

model = nn.Linear(1, 1)                              # one input, one output
loss_fn = nn.MSELoss()                               # mean squared error
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)

for epoch in range(100):                             # repeat the ritual 100 times
    pred = model(X)                                  # 1. forward: predictions
    loss = loss_fn(pred, y)                          # 2. measure wrongness

    optimizer.zero_grad()                            # 3. clear old gradients
    loss.backward()                                  # 4. backward: compute gradients
    optimizer.step()                                 # 5. nudge the weights

    if epoch % 20 == 0:
        print(f"epoch {epoch:3d} | loss {loss.item():.4f}")
```

```console
epoch   0 | loss 28.4631
epoch  20 | loss 0.6118
epoch  40 | loss 0.1370
epoch  60 | loss 0.0312
epoch  80 | loss 0.0072
```

*What just happened:* The setup ran once - a model, a loss function, an optimizer wired to the model's
parameters. Then the loop ran 100 times, and each pass did the same five steps: predict, measure loss,
clear gradients, compute new gradients, step. The thing to *feel* is the loss column: it starts at 28.46
(the untrained model is very wrong) and falls toward zero. That falling number is the model learning
`y = 2x`. Nothing here is special to this toy - swap in a deep network and a real dataset and the loop
body is identical.

💡 Notice `loss.item()` in the print. `loss` is a tensor (a zero-dim one, from
[Phase 2](02-tensor-operations-and-gpu.md)); `.item()` pulls out the plain Python number for printing.
Get in the habit - logging the raw tensor every step also quietly holds onto its computation graph and
wastes memory.

## 3. Epochs and batches

Two words you'll see everywhere, and they're simpler than they sound.

📝 An **epoch** is one full pass over your entire dataset. The loop above ran 100 epochs - it showed the
model all four examples, 100 times over. More epochs means more chances to learn (up to a point - past
that, the model starts memorizing instead of learning, which is [overfitting](/guides/how-a-model-learns)).

📝 A **batch** is a chunk of the data processed together in one forward/backward pass. In the toy loop
we fed all four examples at once, so the whole dataset *was* one batch. Real datasets are far too big for
that, so each epoch is split into many batches, and you loop over them *inside* the epoch:

```python
for epoch in range(num_epochs):
    for X_batch, y_batch in data_loader:            # inner loop: one batch at a time
        pred = model(X_batch)
        loss = loss_fn(pred, y_batch)

        optimizer.zero_grad()
        loss.backward()
        optimizer.step()
```

*What just happened:* The five steps didn't change at all - they just moved one level deeper, inside an
inner loop over batches. Each `X_batch` is a slice of the data; one trip through the inner loop is one
weight update. One trip through the *outer* loop is one epoch (every batch seen once). That `data_loader`
is the piece we haven't built yet - it's the [Dataset & DataLoader](07-datasets-and-dataloaders.md) of
Phase 7, which hands you batches automatically.

Why bother with batches instead of the whole dataset at once? Two reasons. **Memory:** a million images
won't fit on your GPU all at once, but a batch of 64 will. **Better learning:** updating the weights
after every small batch - rather than once per full pass - gives many more, slightly-noisy steps, and that
noise actually helps the model find better solutions. So batching isn't a compromise; it's how modern
training is *supposed* to work.

## 4. Order matters - the #1 beginner bug

Look back at the three middle lines:

```python
optimizer.zero_grad()    # 3. clear old gradients
loss.backward()          # 4. compute new gradients
optimizer.step()         # 5. apply them
```

⚠️ **That order is not optional.** `zero_grad` → `backward` → `step`, every single time. Getting it wrong
is the most common way a beginner's training silently breaks. Let's spell out exactly what each line does
and what happens without it.

- **`loss.backward()`** runs the backward pass from [Phase 3](03-autograd.md). It walks the computation
  graph and computes, for every weight, the gradient - the direction that would make the loss bigger.
  After this line, each parameter has its `.grad` filled in. **Without it,** there are no gradients at
  all, so the next line has nothing to act on and the model never changes.

- **`optimizer.step()`** uses those `.grad` values to nudge each weight a small amount in the
  loss-reducing direction. This is the actual learning. **Without it,** you compute perfect gradients and
  then ignore them - the model stays frozen.

- **`optimizer.zero_grad()`** is the one beginners forget, and it's the subtle one. Here's the trap: in
  PyTorch, `backward()` *adds* the new gradients to whatever is already in `.grad` - it **accumulates**,
  it doesn't overwrite (this is the gradient-accumulation behavior from [Phase 3](03-autograd.md)). So if
  you don't reset to zero before each `backward()`, this step's gradients pile on top of last step's, and
  last step's, and so on. Your weight updates get computed from a stale, ever-growing sum of gradients,
  the steps go haywire, and training quietly falls apart - no error, just a loss that refuses to fall or
  explodes. **Without `zero_grad()`,** the loop runs fine and the result is garbage, which is the worst
  kind of bug.

Here's what forgetting it looks like:

```python
# BROKEN: no optimizer.zero_grad()
for epoch in range(100):
    pred = model(X)
    loss = loss_fn(pred, y)
    loss.backward()          # gradients ACCUMULATE across every epoch
    optimizer.step()
```

```console
epoch   0 | loss 28.4631
epoch  20 | loss 1453.8079
epoch  40 | loss 98211.4453
epoch  60 | loss nan
epoch  80 | loss nan
```

*What just happened:* With no `zero_grad()`, each `backward()` added its gradients to the leftover pile
from all previous epochs. The "step" the optimizer took kept growing, overshooting wildly, until the
numbers blew up to `nan` (not-a-number). Same model, same data, same learning rate as the working loop in
section 2 - the *only* difference is the missing reset line. That's how much one line matters. When your
loss explodes to `nan`, "did I forget `zero_grad()`?" should be your first thought.

💡 The order is a tiny story: **clear the slate (`zero_grad`), figure out which way to go
(`backward`), take the step (`step`).** Say it that way once and you'll never reorder it.

## 5. Tracking progress, and train vs. eval

The loss printout isn't decoration - it's your only window into whether training is working. **A healthy
run shows the loss falling** and roughly leveling off. If it *doesn't* fall, something is wrong, and it's
almost always one of three things: the learning rate is off (too high → it explodes; too low → it barely
moves - see [Phase 5](05-loss-and-optimizers.md)), there's a bug in the loop (often the `zero_grad` one),
or the data is bad. Print the loss every epoch from day one; it's the cheapest diagnostic you have.

There's also a part of training the toy loop skipped: checking how the model does on data it *didn't*
train on. A model that aces its training data but flops on new data hasn't learned - it's memorized
([overfitting](/guides/how-a-model-learns) again). So you hold out some data and evaluate on it. Two
PyTorch habits make that correct:

📝 **`model.train()` and `model.eval()`** flip the model between two modes. Some layers behave
differently while training versus while being evaluated (dropout, batch-norm - you'll meet them later),
so you tell the model which phase it's in. Call `model.train()` before the training loop and
`model.eval()` before you evaluate.

📝 **`torch.no_grad()`** wraps your evaluation code so PyTorch *doesn't* build the computation graph or
track gradients. You're only measuring here, not learning - there's nothing to update - so skipping the
gradient bookkeeping makes evaluation faster and lighter.

```python
model.train()                              # training mode
for epoch in range(num_epochs):
    for X_batch, y_batch in train_loader:
        pred = model(X_batch)
        loss = loss_fn(pred, y_batch)
        optimizer.zero_grad()
        loss.backward()
        optimizer.step()

model.eval()                               # evaluation mode
with torch.no_grad():                      # no gradients needed -- just measuring
    for X_batch, y_batch in test_loader:
        pred = model(X_batch)
        # ... compute accuracy or test loss on held-out data ...
```

*What just happened:* The training loop is exactly the five-step ritual you already know, bracketed by
`model.train()`. Then `model.eval()` switches modes, and `torch.no_grad()` turns off gradient tracking
for the measurement pass over held-out test data - no `backward()`, no `step()`, because we're judging
the model, not training it. This train-then-evaluate shape is the skeleton of the real classifier you'll
build in [Phase 8](08-training-a-classifier.md).

💡 Here's the payoff to sit with: **every model is trained by a scaled-up version of this exact loop.**
The image recognizers, the recommendation engines, the giant LLMs behind the chatbots you use - under the
hood, they all run forward → loss → `zero_grad` → backward → step, over and over, across mountains of
data and hardware. The scale is staggering; the ritual is the one you just learned. Master these five
lines and you understand how all of deep learning actually trains.

## Recap

- **Training is a loop you write yourself** - PyTorch has no `model.fit()`. The body is always the same
  five steps: forward → loss → `zero_grad` → `backward` → `step`, repeated many times.
- **An epoch** is one full pass over the data; **a batch** is a chunk processed in one step. Real
  training loops over batches inside each epoch - for memory and for better, noisier learning.
- **Order is non-negotiable:** `optimizer.zero_grad()` → `loss.backward()` → `optimizer.step()`.
  Forgetting `zero_grad()` lets gradients accumulate across steps and silently wrecks training (loss
  explodes to `nan`).
- **Watch the loss fall** - it's your main diagnostic. If it doesn't drop, suspect the learning rate, a
  loop bug, or the data.
- **`model.train()` vs. `model.eval()`** sets the model's mode; evaluate held-out data under
  `torch.no_grad()` since you're measuring, not learning.
- **Every model - up to giant LLMs - trains with a scaled-up version of this exact loop.**

## Quick check

```quiz
[
  {
    "q": "What is the correct order of the three middle steps in a PyTorch training loop?",
    "choices": ["loss.backward() → optimizer.zero_grad() → optimizer.step()", "optimizer.zero_grad() → loss.backward() → optimizer.step()", "optimizer.step() → loss.backward() → optimizer.zero_grad()"],
    "answer": 1,
    "explain": "Clear the slate (zero_grad), compute which way to go (backward), then take the step (step). Any other order breaks training."
  },
  {
    "q": "You remove optimizer.zero_grad() from your loop. What's the most likely result?",
    "choices": ["A clear RuntimeError that stops the program immediately", "Nothing changes - zero_grad() is optional cleanup", "The loss climbs or explodes to nan, because gradients accumulate across steps"],
    "answer": 2,
    "explain": "backward() ADDS to .grad rather than overwriting it. Without zero_grad(), gradients pile up every step, the updates overshoot, and the loss blows up - usually with no error at all."
  },
  {
    "q": "Why wrap evaluation code in `with torch.no_grad()`?",
    "choices": ["You're only measuring, not learning, so there's no need to track gradients - it saves memory and time", "It makes the model more accurate on the test set", "It is required or backward() will throw an error"],
    "answer": 0,
    "explain": "During evaluation there's nothing to update, so building the computation graph is wasted work. no_grad() skips gradient tracking, making the pass faster and lighter."
  }
]
```


---

# Data: Dataset & DataLoader

In [Phase 6](06-the-training-loop.md) you wrote the loop that trains every model - and in the middle of
it sat a line we waved at and moved past:

```python
for X_batch, y_batch in data_loader:
    ...
```

That `data_loader` is the missing piece. The loop *consumes* batches; something has to *produce* them.
This phase is that something. Before any code, let's get the mental model clear, because it's two ideas
and a clean division of labor.

## 1. The mental model: two jobs, two objects

📝 Feeding a model isn't "pass in the data." Real training needs the data delivered in **batches** (a
handful of samples at a time, for the memory and learning reasons from
[Phase 6](06-the-training-loop.md)), **shuffled** fresh each epoch (so the model doesn't learn the order
instead of the pattern), and delivered **fast** (so your expensive GPU isn't sitting idle waiting for the
next batch).

You could hand-roll all of that - slice arrays, track indices, reshuffle every epoch, maybe spin up
threads to load ahead. It's fiddly, repetitive, and exactly the kind of code that hides off-by-one bugs.
So PyTorch splits the work into two objects, each with one job:

- **`Dataset`** - knows how to get *one* sample. "Give me item 37" → returns the features and label for
  example 37. That's its entire responsibility.
- **`DataLoader`** - wraps a `Dataset` and handles everything *around* the samples: grouping them into
  batches, shuffling the order, loading in parallel, handing you batch after batch in a `for` loop.

💡 The one-line version to keep in your head: **Dataset = "how to get one item." DataLoader = "batch them,
shuffle them, feed them fast."** Get that split and the rest of this phase is just syntax.

```mermaid
flowchart LR
  A[Raw data] --> B[Dataset: one sample at a time]
  B --> C[DataLoader: batch + shuffle + parallel]
  C --> D[Training loop: for X, y in loader]
```

## 2. The Dataset - how to get one sample

📝 A `Dataset` is a class with exactly two methods you must implement:

- **`__len__(self)`** - returns how many samples there are.
- **`__getitem__(self, i)`** - returns sample `i`, as tensors (typically a `(features, label)` pair).

That's the whole contract. If your class can answer "how many?" and "give me number `i`," PyTorch knows
how to use it. Here's a small custom Dataset wrapping two arrays - features and labels:

```python
import torch
from torch.utils.data import Dataset

class PointsDataset(Dataset):
    def __init__(self, features, labels):
        # store the raw data as tensors
        self.features = torch.tensor(features, dtype=torch.float32)
        self.labels = torch.tensor(labels, dtype=torch.float32)

    def __len__(self):
        return len(self.features)            # how many samples

    def __getitem__(self, i):
        return self.features[i], self.labels[i]   # one (X, y) pair

ds = PointsDataset([[1.0], [2.0], [3.0], [4.0]], [2.0, 4.0, 6.0, 8.0])
print(len(ds))        # uses __len__
print(ds[0])          # uses __getitem__
```

```console
4
(tensor([1.]), tensor(2.))
```

*What just happened:* You defined a class that subclasses `Dataset` and filled in the two required
methods. `len(ds)` quietly called your `__len__` and got back 4; `ds[0]` quietly called your
`__getitem__(0)` and got back the first feature/label pair as tensors. Notice you *never call these
methods by name* - Python's `len()` and `[]` syntax route to them, and (next section) the DataLoader will
call `__getitem__` for you, over and over, to assemble batches. This tiny class is the standard way to
wrap *any* data source: arrays in memory, rows in a CSV, image files on disk. The body of `__getitem__`
changes; the shape of the class doesn't.

💡 For the common case of "I already have my data as tensors," you don't even need a custom class.
`TensorDataset` does the wrapping for you:

```python
from torch.utils.data import TensorDataset

X = torch.tensor([[1.0], [2.0], [3.0], [4.0]])
y = torch.tensor([[2.0], [4.0], [6.0], [8.0]])

ds = TensorDataset(X, y)
print(len(ds))
print(ds[0])
```

```console
4
(tensor([1.]), tensor([2.]))
```

*What just happened:* `TensorDataset(X, y)` built a ready-made Dataset out of two tensors - same `__len__`
and `__getitem__` behavior as the class above, zero boilerplate. Reach for `TensorDataset` when your data
already fits in memory as tensors; write a custom `Dataset` when fetching a sample takes real work (read a
file, decode an image, look up a row).

## 3. The DataLoader - batch, shuffle, feed

📝 A `DataLoader` takes a `Dataset` and turns it into something you iterate over to get **batches**. The
two arguments you'll set constantly:

- **`batch_size`** - how many samples per batch.
- **`shuffle`** - whether to reorder the samples each epoch.

You hand it a Dataset, then loop:

```python
from torch.utils.data import DataLoader

loader = DataLoader(ds, batch_size=2, shuffle=True)

for X_batch, y_batch in loader:
    print(X_batch.shape, y_batch.shape)
    print(X_batch)
```

```console
torch.Size([2, 1]) torch.Size([2, 1])
tensor([[3.],
        [1.]])
torch.Size([2, 1]) torch.Size([2, 1])
tensor([[4.],
        [2.]])
```

*What just happened:* The DataLoader called the Dataset's `__getitem__` behind the scenes, **stacked** the
individual samples into batches of 2, and handed them to you one batch at a time. Look at the shapes: each
sample was `(1,)`, and the loader added a **batch dimension** in front, giving `(2, 1)` - two samples,
one feature each. Four samples at `batch_size=2` means two batches, which is exactly what the loop
printed. And because `shuffle=True`, the rows came out reordered (`3, 1` then `4, 2`, not `1, 2, 3, 4`) - 
a *different* order next epoch. That batch dimension is why your model from
[Phase 4](04-building-models-with-nn-module.md) always expects a leading batch axis: the DataLoader is
where it comes from.

💡 **`shuffle=True` for training, `shuffle=False` for evaluation.** Shuffling during training breaks any
order bias in your data (imagine a file sorted by label - without shuffling, the model would see all the
0s, then all the 1s, and learn the order rather than the content). During evaluation there's nothing to
learn, so shuffling buys you nothing - leave it off so results are reproducible.

## 4. The real training loop

Now put it together. Here is the Phase 6 ritual - unchanged in its five steps - but fed by a DataLoader
instead of a single hand-built batch. *This* is what real training actually looks like:

```python
import torch.nn as nn

train_loader = DataLoader(ds, batch_size=2, shuffle=True)

model = nn.Linear(1, 1)
loss_fn = nn.MSELoss()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)

model.train()
for epoch in range(100):                         # outer loop: epochs
    for X_batch, y_batch in train_loader:        # inner loop: batches
        pred = model(X_batch)                     # 1. forward
        loss = loss_fn(pred, y_batch)             # 2. measure
        optimizer.zero_grad()                     # 3. clear gradients
        loss.backward()                           # 4. backward
        optimizer.step()                          # 5. step

    if epoch % 20 == 0:
        print(f"epoch {epoch:3d} | loss {loss.item():.4f}")
```

```console
epoch   0 | loss 11.9032
epoch  20 | loss 0.2451
epoch  40 | loss 0.0517
epoch  60 | loss 0.0109
epoch  80 | loss 0.0023
```

*What just happened:* The five-step body is *byte-for-byte the same* as Phase 6 - forward, loss,
`zero_grad`, `backward`, `step`. The only structural change is the **inner loop**: instead of feeding the
whole dataset once per epoch, you now loop over `train_loader` and take one optimizer step *per batch*. So
one epoch is now several weight updates (one per batch), not one - which is precisely the "many small,
slightly-noisy steps" that Phase 6 said helps the model learn. The loss still falls toward zero; the model
still learns `y = 2x`. You've just swapped the hand-fed batch for the real machinery. This exact
two-level loop - epochs outside, batches inside - is the skeleton of the classifier in
[Phase 8](08-training-a-classifier.md).

## 5. Transforms and performance

Two more pieces you'll meet constantly in real code.

📝 **Transforms** preprocess each sample as it's fetched - normalize numbers into a sensible range, or for
images, resize, crop, convert to a tensor, and (during training) randomly flip or rotate to make the data
more varied. Transforms are usually handed to the Dataset, which applies them inside `__getitem__` so
every sample comes out ready to use:

```python
from torchvision import datasets, transforms

transform = transforms.Compose([
    transforms.ToTensor(),                          # image -> tensor, scaled to [0, 1]
    transforms.Normalize((0.5,), (0.5,)),           # shift to roughly [-1, 1]
])

train_ds = datasets.MNIST(root="data", train=True, download=True, transform=transform)
train_loader = DataLoader(train_ds, batch_size=64, shuffle=True)
```

*What just happened:* `transforms.Compose` chained two preprocessing steps into one pipeline, and you
passed it to the Dataset via `transform=`. Now every time the DataLoader pulls a sample, the Dataset
applies that pipeline first - the raw image becomes a normalized tensor automatically, before it ever
reaches your model. (This is `torchvision`, PyTorch's companion library for image data; `datasets.MNIST`
is a ready-made Dataset of handwritten digits - the exact one you'll train on in Phase 8.) You write the
preprocessing once, declaratively, and it runs on every sample for free.

📝 Two DataLoader arguments tune *speed*:

- **`num_workers`** - how many background processes load and preprocess data in parallel. With
  `num_workers=0` (the default), loading happens in your main process, blocking everything else. Bumping
  it to, say, 4 lets the loader prepare upcoming batches while the GPU is busy with the current one.
- **`pin_memory=True`** - speeds up the copy of each batch from CPU to GPU. Worth turning on when you
  train on a GPU.

```python
train_loader = DataLoader(
    train_ds, batch_size=64, shuffle=True,
    num_workers=4, pin_memory=True,
)
```

*What just happened:* Same loader, now with parallel loading and faster GPU transfers. Functionally
identical - same batches, same order behavior - but the data pipeline can keep ahead of the model instead
of making it wait.

⚠️ **A slow data pipeline starves the GPU.** This is one of the most common real-world performance traps:
your GPU can crunch a batch in milliseconds, but if loading and preprocessing the *next* batch is slow
(reading thousands of files, heavy transforms, `num_workers=0`), the GPU finishes and then sits idle,
waiting. You bought a fast engine and feed it through a straw. When training feels mysteriously slow even
though the model is small, the culprit is usually the data pipeline, not the math - check `num_workers`
first.

💡 Step back and see the shape of it: **Dataset** says how to get one item, **transforms** make that item
model-ready, and **DataLoader** batches, shuffles, and feeds it fast enough to keep the GPU busy. That's
the entire input side of training - and it's the last piece you needed. In [Phase 8](08-training-a-classifier.md)
you'll wire all of it - real data, real transforms, a real loop - into a working image classifier.

## Recap

- **Two objects, two jobs.** `Dataset` knows how to fetch *one* sample; `DataLoader` batches, shuffles,
  and parallel-loads them. The training loop just consumes the batches the DataLoader produces.
- **A `Dataset`** implements `__len__` (how many) and `__getitem__(i)` (return sample `i` as tensors).
  For data that's already tensors, `TensorDataset(X, y)` skips the boilerplate.
- **A `DataLoader`** wraps a Dataset: `DataLoader(ds, batch_size=32, shuffle=True)`. Iterate it to get
  batches; it adds the leading **batch dimension** your model expects. Use `shuffle=True` for training,
  `False` for evaluation.
- **The real loop** is the Phase 6 five-step ritual with an inner loop over the DataLoader - one optimizer
  step per batch, epochs on the outside.
- **Transforms** preprocess each sample (normalize, augment) and are passed to the Dataset.
  `num_workers` and `pin_memory` make loading fast.
- ⚠️ **A slow data pipeline starves the GPU** - it finishes a batch and idles waiting for the next. When
  training is mysteriously slow, suspect the data side first.

## Quick check

Make sure the division of labor and the loop shape stuck:

```quiz
[
  {
    "q": "What is each object's job?",
    "choices": ["Dataset batches and shuffles; DataLoader fetches one sample", "Dataset knows how to get one sample; DataLoader batches, shuffles, and parallel-loads them", "Both do the same thing; DataLoader is just a faster Dataset"],
    "answer": 1,
    "explain": "Dataset = 'how to get one item' (__len__ + __getitem__). DataLoader wraps it and handles everything around the samples: batching, shuffling, and parallel loading."
  },
  {
    "q": "Which two methods must a custom Dataset implement?",
    "choices": ["__init__ and forward", "__batch__ and __shuffle__", "__len__ and __getitem__"],
    "answer": 2,
    "explain": "__len__ returns how many samples there are; __getitem__(i) returns sample i as tensors. PyTorch calls __getitem__ repeatedly to assemble batches."
  },
  {
    "q": "Your tiny model trains far slower than expected and the GPU usage is mostly idle. What's the most likely cause?",
    "choices": ["The data pipeline is starving the GPU - slow loading/preprocessing, often num_workers=0", "The learning rate is too low", "The model needs more layers"],
    "answer": 0,
    "explain": "A slow data pipeline (heavy transforms, many file reads, no parallel workers) means the GPU finishes a batch and waits for the next. Bumping num_workers and enabling pin_memory usually fixes it."
  }
]
```


---

# Training a Real Classifier

This is the phase you've been building toward. Every piece you've met so far has been one corner of a
picture, and now we snap them together into something that actually *works* - a neural network that looks
at a handwritten digit and tells you which one it is.

Here's the mental model to hold onto before any code. **A real training program is always the same five-part
skeleton**, and you already know all five parts:

1. **Data** - load it and hand it out in batches ([Phase 7](07-datasets-and-dataloaders.md)).
2. **Model** - an `nn.Module` that turns inputs into predictions ([Phase 4](04-building-models-with-nn-module.md)).
3. **Loss + optimizer** - one to measure wrongness, one to fix it ([Phase 5](05-loss-and-optimizers.md)).
4. **Training loop** - the forward → loss → `zero_grad` → backward → step ritual ([Phase 6](06-the-training-loop.md)).
5. **Evaluation** - check it on data it never saw.

That skeleton doesn't change whether you're classifying digits or training a model with billions of
parameters. We'll build it once, end to end, on the classic "hello world" of deep learning: MNIST.

## 1. The task and the data

📝 **MNIST** is 70,000 grayscale images of handwritten digits, each 28×28 pixels, each labeled with the
digit it shows (0 through 9). The job is **classification**: given a 28×28 image, predict which of the 10
classes it belongs to. It's small, it's clean, and a simple network gets very good at it - which makes it
perfect for seeing the whole pipeline without drowning in detail.

The data comes pre-split into two parts, and that split is the single most important idea in this phase:

- **Training set** (60,000 images) - the model *learns* from these.
- **Test set** (10,000 images) - the model is *judged* on these, and it never trains on them.

Why hold data back? Because a model that has seen an image can just memorize the answer - that tells you
nothing about whether it actually *learned* to recognize digits. The only real measure of learning is
performance on examples it has never encountered. That's the [overfitting](/guides/how-a-model-learns)
problem made concrete: train great, test poorly, and you've memorized, not learned. We hold out the test
set so we can catch exactly that.

`torchvision` (PyTorch's companion library for images) downloads MNIST for us and gives us a `Dataset`,
which we wrap in the `DataLoader` from [Phase 7](07-datasets-and-dataloaders.md):

```python
import torch
from torch.utils.data import DataLoader
from torchvision import datasets, transforms

# ToTensor() turns each 28x28 image into a float tensor with values in [0, 1].
transform = transforms.ToTensor()

train_data = datasets.MNIST(root="data", train=True,  download=True, transform=transform)
test_data  = datasets.MNIST(root="data", train=False, download=True, transform=transform)

train_loader = DataLoader(train_data, batch_size=64, shuffle=True)
test_loader  = DataLoader(test_data,  batch_size=1000, shuffle=False)

print(f"train: {len(train_data)} images | test: {len(test_data)} images")
images, labels = next(iter(train_loader))
print(f"one batch: images {tuple(images.shape)}, labels {tuple(labels.shape)}")
```

```console
train: 60000 images | test: 10000 images
one batch: images (64, 1, 28, 28), labels (64,)
```

*What just happened:* `datasets.MNIST` downloaded the data (the first time only) and gave us two
`Dataset` objects, one per split. `transforms.ToTensor()` is the recipe that converts each PIL image into a
tensor with pixel values scaled to the 0–1 range PyTorch likes. We wrapped each `Dataset` in a
`DataLoader`: the training one **shuffles** every epoch (so the model doesn't learn the order) and hands out
batches of 64; the test one doesn't shuffle (order is irrelevant when you're only measuring) and uses big
batches of 1000 for speed. The batch shape `(64, 1, 28, 28)` reads as *64 images, 1 color channel,
28 tall, 28 wide* - and the 64 matching labels are just the digit for each one.

💡 The `1` in `(64, 1, 28, 28)` is the channel dimension - grayscale has one channel. Color images would
have 3 (red, green, blue). It's there even for grayscale because image layers in PyTorch always expect a
channel dimension; you'll be glad of the consistency later.

## 2. The model

For digits, a small **multilayer perceptron** (MLP) does the job: flatten the image into a flat row of
numbers, push it through one hidden layer with a ReLU in between, then out to 10 numbers - one score per
digit class. We define it as an `nn.Module`, exactly the pattern from
[Phase 4](04-building-models-with-nn-module.md).

```python
import torch.nn as nn

class DigitClassifier(nn.Module):
    def __init__(self):
        super().__init__()
        self.flatten = nn.Flatten()          # (B, 1, 28, 28) -> (B, 784)
        self.fc1 = nn.Linear(28 * 28, 128)   # 784 inputs -> 128 hidden
        self.relu = nn.ReLU()                # nonlinearity
        self.fc2 = nn.Linear(128, 10)        # 128 hidden -> 10 class scores

    def forward(self, x):
        x = self.flatten(x)
        x = self.relu(self.fc1(x))
        return self.fc2(x)                   # raw logits, one per class

device = "cuda" if torch.cuda.is_available() else "cpu"
model = DigitClassifier().to(device)
print(model)
print("device:", device)
```

```console
DigitClassifier(
  (flatten): Flatten(start_dim=1, end_dim=-1)
  (fc1): Linear(in_features=784, out_features=128, bias=True)
  (relu): ReLU()
  (fc2): Linear(in_features=128, out_features=10, bias=True)
)
device: cpu
```

*What just happened:* We described the network as layers in `__init__` and the data's path through them in
`forward`. `Flatten` squashes each 28×28 image into a flat vector of 784 numbers (a `Linear` layer wants a
flat row, not a grid). Then 784 → 128 → 10, with a `ReLU` in the middle so the network can learn curves
rather than just straight lines. The final layer spits out **10 logits** - raw, unbounded scores where the
biggest one is the model's guess. We never squeeze them into probabilities ourselves; the loss function does
that for us in the next step. Finally, `.to(device)` moves the model's weights onto the GPU if there is one
([Phase 2](02-tensor-operations-and-gpu.md)) - and the iron rule is that the model and its data must live on
the *same* device, which is why we'll move each batch too.

## 3. Loss, optimizer, and the training loop

Now the engine. For multi-class classification the standard loss is **`CrossEntropyLoss`**, and a reliable
default optimizer is **`Adam`** - both from [Phase 5](05-loss-and-optimizers.md).

```python
loss_fn = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.001)

epochs = 5
model.train()                                       # training mode
for epoch in range(epochs):
    running_loss = 0.0
    for images, labels in train_loader:             # one batch at a time
        images, labels = images.to(device), labels.to(device)

        loss = loss_fn(model(images), labels)       # forward + measure
        optimizer.zero_grad()                       # clear old gradients
        loss.backward()                             # compute new gradients
        optimizer.step()                            # nudge the weights

        running_loss += loss.item()

    avg = running_loss / len(train_loader)
    print(f"epoch {epoch + 1}/{epochs} | avg loss {avg:.4f}")
```

```console
epoch 1/5 | avg loss 0.3372
epoch 2/5 | avg loss 0.1502
epoch 3/5 | avg loss 0.1064
epoch 4/5 | avg loss 0.0823
epoch 5/5 | avg loss 0.0674
```

*What just happened:* This is the exact five-step ritual from [Phase 6](06-the-training-loop.md), now with
real data flowing through it. The outer loop counts epochs (full passes over all 60,000 images); the inner
loop walks the batches the `DataLoader` hands us. For each batch we move it to the device, run the forward
pass and measure the loss, then `zero_grad` → `backward` → `step` in that exact, non-negotiable order. We
accumulate `loss.item()` to print an average per epoch. The thing to *feel* is the loss column falling from
0.34 to 0.07 - that downward march is the network learning to read digits. A few important details hide in
plain sight:

- **`CrossEntropyLoss` takes raw logits**, not probabilities. It applies the softmax internally. A classic
  beginner bug is adding your own softmax to the model and feeding it in - that double-counts and quietly
  hurts training. Hand it the raw logits.
- **The labels are plain integers** (`3`, `7`, ...), not one-hot vectors. `CrossEntropyLoss` expects exactly
  that. It just works.

💡 If your loss starts high and *doesn't* fall, it's almost always one of the three suspects from Phase 6:
the learning rate, a loop-order bug, or the data. Print the loss every epoch from the start - it's your
cheapest diagnostic.

## 4. Evaluation: how good is it, really?

The falling loss is encouraging, but it's measured on the *training* data - the data the model is allowed to
study. The real question is: **how does it do on the 10,000 test images it has never seen?** For
classification, the natural metric is **accuracy**: out of all test images, what fraction did it label
correctly?

📝 The evaluation pass has its own ritual, and it's deliberately different from training:

- **`model.eval()`** flips the model into evaluation mode (some layers behave differently when training vs.
  measuring - Phase 6).
- **`torch.no_grad()`** turns off gradient tracking. We're only measuring, not learning, so there's nothing
  to update and no reason to build the computation graph - it's faster and lighter.
- **No `backward()`, no `step()`.** We never adjust the weights here. We're judging, not teaching.

To turn logits into a prediction, we take the **argmax** - the index of the biggest score is the model's
guessed digit. Compare that to the true label, count the matches, divide by the total.

```python
model.eval()                                        # evaluation mode
correct = 0
total = 0
with torch.no_grad():                               # no gradients -- just measuring
    for images, labels in test_loader:
        images, labels = images.to(device), labels.to(device)
        logits = model(images)
        predicted = logits.argmax(dim=1)            # index of the biggest score
        correct += (predicted == labels).sum().item()
        total += labels.size(0)

accuracy = correct / total
print(f"test accuracy: {accuracy:.4f}  ({correct}/{total})")
```

```console
test accuracy: 0.9743  (9743/10000)
```

*What just happened:* We switched to `eval()` mode, wrapped everything in `no_grad()`, and ran the whole
test set through the model - no training, just measuring. For each batch, `logits.argmax(dim=1)` picked the
highest-scoring class per image. `(predicted == labels)` is a tensor of `True`/`False`; `.sum().item()`
counts the `True`s as a plain number, and we tallied them across all batches. The result: **97.4% of digits
it had never seen were classified correctly.** That's the real score - the number that means it *learned*,
not memorized.

⚠️ **Never tune your model against the test set.** It's tempting to peek at the test accuracy, tweak a
setting, re-check, tweak again - but the moment you make decisions based on the test set, you've started
leaking its information into your model, and your "test" score stops being trustworthy. The test set is for one
thing: a final, untouched verdict. (When you need a set to tune against, you hold out a *third* split - a
validation set - and leave the test set sealed until the very end.) Report the number you actually got, not
the one you wish you got.

## 5. Reading the results

So you've got a falling training loss and a 97.4% test accuracy. 💡 Here's how to read those two numbers
together, because the *relationship* between them tells you almost everything:

- **Training loss falling AND test accuracy high** → the model genuinely learned. This is the win. The
  patterns it found on the training data transfer to data it's never seen, which is the whole point.
- **Training loss low (or accuracy near-perfect on training) BUT test accuracy poor** → classic
  [overfitting](/guides/ml-basics-for-data-people). The model memorized the training set instead of learning
  general patterns. The gap between train and test performance is the alarm bell - when training looks great
  but test lags badly, suspect overfitting first.
- **Both poor** → the model hasn't learned enough yet. Train longer, give it more capacity, or check the
  learning rate and data.

💡 And here's the payoff to carry out of this whole guide: **this exact skeleton scales to anything.** Data →
model → loss + optimizer → training loop → evaluation. Swap MNIST for medical scans, swap the MLP for a
giant network, swap the digit labels for any target you can measure a loss against - the shape of the
program is *identical* to what you just wrote. You didn't just train a digit classifier. You learned the
template that every supervised deep-learning project on Earth is built from. From here on, you're not
learning *whether* you can train a model - you're just changing what goes in the five boxes.

## Recap

- **MNIST** is 70,000 labeled 28×28 digit images, pre-split into a 60,000-image training set and a
  10,000-image test set. `torchvision.datasets.MNIST` + `transforms.ToTensor()` + a `DataLoader` give you
  batches ready to train on.
- **The train/test split exists so you can evaluate on unseen data.** Performance on data the model never
  trained on is the only real measure of learning, and the way you catch overfitting.
- **The model** is a small MLP (`Flatten` → `Linear` → `ReLU` → `Linear` → 10 logits) defined as an
  `nn.Module`, moved to `device` along with every batch.
- **The training loop is the same Phase 6 ritual:** `CrossEntropyLoss` (fed raw logits + integer labels) +
  `Adam`, looping `zero_grad` → `backward` → `step` over batches and epochs while the loss falls.
- **Evaluation** runs under `model.eval()` and `torch.no_grad()`: take `argmax` of the logits, compare to
  the labels, and report accuracy (~97%). Never tune on the test set; report it straight.
- **The five-part skeleton - data → model → loss/optimizer → train loop → eval - is universal.** You now have
  it end to end, and it scales to any supervised task.

## Quick check

```quiz
[
  {
    "q": "Why does MNIST come split into a training set and a separate test set?",
    "choices": ["To make the download smaller", "So the model can be evaluated on data it never trained on - the only real measure of learning", "Because PyTorch requires exactly two DataLoaders"],
    "answer": 1,
    "explain": "A model can memorize data it has seen, which proves nothing. Performance on the held-out test set shows whether it actually learned to generalize - and reveals overfitting."
  },
  {
    "q": "What should you feed into nn.CrossEntropyLoss as the model's predictions?",
    "choices": ["Softmax probabilities you computed yourself", "The raw logits straight from the final Linear layer", "The argmax (predicted class index)"],
    "answer": 1,
    "explain": "CrossEntropyLoss applies softmax internally and expects raw logits plus integer labels. Adding your own softmax double-counts and hurts training."
  },
  {
    "q": "Training loss is very low but test accuracy is poor. What does this most likely indicate?",
    "choices": ["The model has learned well and is ready to ship", "Overfitting - the model memorized the training data instead of learning general patterns", "The learning rate is too low"],
    "answer": 1,
    "explain": "A big gap between strong training performance and weak test performance is the classic signature of overfitting: the model fit the training set rather than learning patterns that generalize."
  }
]
```


---

# Saving, Loading & Inference

In [Phase 8](08-training-a-classifier.md) you trained a real classifier - you watched the loss fall, the
accuracy climb, and ended up with a model that actually works. Then your Python process exits, and it's
all gone. The weights lived in RAM; closing the program threw them away.

So here's the mental model for this whole phase, and it's the one that makes everything else fall into
place: **the value of training isn't the running program - it's the numbers it produced.** A trained model
is, at the end of the day, a bag of learned tensors (the weights and biases the optimizer nudged into shape
over Phase 8). Saving a model means writing those numbers to disk. Loading means recreating the model and
pouring the numbers back in. And *using* the model - inference - means flipping two switches that tell
PyTorch "we're done learning, just give me an answer."

Train once, save the numbers, then load and run them anywhere - your laptop, a server, someone else's
machine. That's the loop this phase closes.

## 1. Save the `state_dict`, not the whole model

📝 The recommended way to save a PyTorch model is to save its **`state_dict`** - the dictionary of learned
parameters you met back in [Phase 4](04-building-models-with-nn-module.md). It maps each layer's name to
its current weight and bias tensors. Those tensors *are* what training produced. Save them and you've saved
everything that matters.

```python
import torch

# `model` is the classifier you trained in Phase 8
torch.save(model.state_dict(), "model.pt")

# peek at what's inside that dictionary
for name, tensor in model.state_dict().items():
    print(name, tuple(tensor.shape))
```

```console
fc1.weight (16, 4)
fc1.bias (16,)
fc2.weight (3, 16)
fc2.bias (3,)
```

*What just happened:* `model.state_dict()` handed back a plain dictionary - keys like `fc1.weight` naming
each parameter, values being the actual tensors of learned numbers. `torch.save(...)` pickled that
dictionary to a file called `model.pt`. Notice what is *not* in there: no Python class, no `forward` method,
no architecture. Just named tensors. The shapes are exactly the layers you defined - that's the proof that
the file is your learning and nothing more.

⚠️ **Don't save the whole model object** (`torch.save(model, "model.pt")`). It works, and it's tempting
because loading looks like one line - but it pickles your *Python class along with the weights*. That ties
the file to your exact code, file layout, and library versions at save time. Rename the class, move the
file, or bump your PyTorch version, and the load can break in confusing ways. The `state_dict` is just
numbers, so it survives all of that. Save the `state_dict`.

> 💡 **Why `.pt` files are portable but not magic.** A `state_dict` file is a snapshot of numbers, not a
> program. It doesn't know what shape the model is, only what shape its own tensors are. That's the whole
> reason the next section needs your model *code* - the file can't rebuild the architecture, only refill it.

## 2. Load: recreate the architecture, then pour the numbers in

📝 Loading is a two-step dance, and the order matters:

1. **Recreate the model** - instantiate the same class with the same architecture you trained.
2. **Load the weights into it** - `model.load_state_dict(torch.load("model.pt"))`.

⚠️ You need the model **code** to load weights. The file is just numbers; it has no idea what a `fc1` layer
is until you build an object that *has* an `fc1`. The keys and shapes in the file have to line up with the
keys and shapes of the model you create. Build the wrong architecture and the load fails loudly (which is
better than failing silently).

```python
import torch
import torch.nn as nn

# 1. Recreate the SAME architecture you trained (same class, same sizes)
class IrisNet(nn.Module):
    def __init__(self):
        super().__init__()
        self.fc1 = nn.Linear(4, 16)
        self.relu = nn.ReLU()
        self.fc2 = nn.Linear(16, 3)

    def forward(self, x):
        return self.fc2(self.relu(self.fc1(x)))

model = IrisNet()                                   # fresh, random weights

# 2. Load the saved numbers into it
model.load_state_dict(torch.load("model.pt"))
print("Loaded.")
```

```console
Loaded.
```

*What just happened:* `IrisNet()` built a brand-new model with *random* weights - useless on its own.
`torch.load("model.pt")` read the dictionary of trained tensors back off disk, and `load_state_dict(...)`
copied each one into the matching layer by name. The freshly-built random model is now the trained model
again. The architecture came from your code; the intelligence came from the file. Both halves are required - 
that's why "the file is just numbers" is the line to remember.

## 3. Inference mode: two switches before you predict

You've loaded a trained model. Before you ask it for predictions, you flip **two switches**. Forgetting
either is one of the most common beginner mistakes, so let's make the mental model crisp: 📝 **training mode
and inference mode are different**, and PyTorch doesn't switch automatically.

- **`model.eval()`** - puts the model into *evaluation* behavior. Some layers act differently when training
  versus predicting: dropout randomly zeros activations during training but must pass everything through at
  inference; batch-norm uses batch statistics during training but running averages at inference.
  ⚠️ Forget `model.eval()` and those layers stay in training mode, giving you wrong, unstable predictions.
- **`torch.no_grad()`** - tells PyTorch *not* to track gradients for these operations. You're not learning
  anymore, so there's no backward pass coming. Skipping gradient tracking is faster and uses less memory.

Here's a single prediction done correctly - load, `eval()`, `no_grad()`, forward, `argmax`:

```python
model.eval()                                # switch 1: inference behavior

# one flower: [sepal_len, sepal_width, petal_len, petal_width]
sample = torch.tensor([[5.1, 3.5, 1.4, 0.2]])

with torch.no_grad():                       # switch 2: no gradient tracking
    logits = model(sample)                  # forward pass -> raw scores
    predicted_class = logits.argmax(dim=1)  # index of the highest score

print("Logits:", logits)
print("Predicted class index:", predicted_class.item())
```

```console
Logits: tensor([[ 4.21, -0.88, -3.05]])
Predicted class index: 0
```

*What just happened:* `model.eval()` flipped every layer into inference behavior. The `with torch.no_grad():`
block ran the forward pass without building the autograd graph - faster and lighter, because we threw away
the machinery we'd only need for training. `model(sample)` returned three raw scores (one per class), and
`argmax(dim=1)` picked the index of the largest. Class `0` won. 💡 Build the habit now: **load → `eval()` →
`no_grad()` → forward → `argmax`** is the inference ritual, the same shape every time, just as the training
loop was.

## 4. From logits to an answer

The numbers the model spits out are **logits** - raw, unbounded scores, not probabilities. They're enough
to pick a winner with `argmax`, but they don't tell a human *how confident* the model is. To get that:

- **`softmax`** turns logits into probabilities that sum to 1 - a confidence for each class.
- **`argmax`** picks the predicted class (the highest one).

Let's turn those raw scores into a labeled prediction with a confidence:

```python
import torch.nn.functional as F

labels = ["setosa", "versicolor", "virginica"]

with torch.no_grad():
    logits = model(sample)
    probs = F.softmax(logits, dim=1)        # logits -> probabilities
    confidence, idx = probs.max(dim=1)      # best probability + its index

print(f"Prediction: {labels[idx.item()]}")
print(f"Confidence: {confidence.item():.1%}")
```

```console
Prediction: setosa
Confidence: 98.7%
```

*What just happened:* `F.softmax` squashed the three logits into probabilities that add up to 1, so they
read as confidences. `probs.max(dim=1)` returned both the largest probability *and* its index in one call.
We used the index to look up a human-readable label and formatted the probability as a percent. Now the
output is something a person - or an API caller - can actually use: *"setosa, 98.7% confident,"* instead of
`tensor([[4.21, -0.88, -3.05]])`. ⚠️ Don't apply `softmax` during *training* with a loss like
`CrossEntropyLoss` - that loss expects raw logits and applies its own softmax internally. Softmax is for
*reading* predictions, not for feeding the loss.

## 5. Toward deployment

You now have the complete cycle: train, save, load, predict. Real deployment builds on exactly that - here's
the lay of the land so you know what to reach for next.

- **Serve it behind an API.** The common pattern is to load the model once at startup (`eval()` + the
  weights), then call it inside a request handler. A small web framework like
  [FastAPI](/guides/fastapi-from-zero) is the usual home for this: receive JSON, run the same inference
  ritual from sections 3–4, return the labeled prediction.
- **Export for production and portability.** Three options, one line each: **TorchScript**
  (`torch.jit.script(model)`) serializes the model so it can run without Python; **`torch.compile(model)`**
  speeds up inference by compiling the forward pass; **ONNX** (`torch.onnx.export(...)`) converts the model
  to a framework-neutral format other runtimes can load.
- ⚠️ **Checkpoint during long training.** Don't wait until the end to save. Periodically write the
  `state_dict` (say, every few epochs) so a crash, a power cut, or a killed job doesn't vaporize hours of
  training. A checkpoint is just a `state_dict` saved mid-run.

💡 The throughline of this whole guide: a model is learnable layers (Phase 4), a loss and optimizer teach
them (Phases 5–6), training produces good weights (Phase 8), and you save that `state_dict` and reload it
with `eval()` + `no_grad()` to use the model anywhere. Train once, run forever.

## Recap

1. **Save the `state_dict`, not the whole model** - `torch.save(model.state_dict(), "model.pt")` writes the
   learned tensors. Pickling the whole object ties the file to your code and versions; the `state_dict` is
   just numbers and stays portable.
2. **Loading is two steps** - recreate the same architecture (you need the model *code*), then
   `model.load_state_dict(torch.load("model.pt"))`. The file refills the model; it can't rebuild it.
3. **Flip two switches to predict** - `model.eval()` for correct inference behavior (dropout, batch-norm),
   and `torch.no_grad()` for faster, lighter forward passes with no gradient tracking.
4. **Logits aren't answers** - the model outputs raw scores. `argmax` picks the class; `softmax` turns
   logits into probabilities you can report as a confidence.
5. **Deployment is this cycle, scaled up** - serve it behind an API (e.g. FastAPI), export with
   TorchScript / `torch.compile` / ONNX, and checkpoint periodically during long training.

That closes the loop from random weights to a model you can ship. The last phase looks outward: how to run
all of this fast on a GPU, and the pitfalls that trip people up along the way.

## Quick check

```quiz
[
  {
    "q": "Why is saving model.state_dict() preferred over saving the whole model object?",
    "choices": ["The state_dict trains faster", "The state_dict is just the learned tensors, so it stays portable across code and version changes; pickling the whole object ties the file to your exact class and library versions", "You can only load a state_dict on a GPU"],
    "answer": 1,
    "explain": "state_dict() saves only the parameters (named tensors). Saving the whole model pickles the Python class too, which can break when your code, file layout, or PyTorch version changes."
  },
  {
    "q": "What do you need in order to load weights from a saved state_dict file?",
    "choices": ["Nothing - the file rebuilds the model itself", "The model CODE, to recreate the same architecture before calling load_state_dict()", "A GPU to deserialize the tensors"],
    "answer": 1,
    "explain": "The file is just numbers with no architecture. You recreate the same model class first, then load_state_dict() pours the saved tensors into the matching layers by name."
  },
  {
    "q": "Before running predictions on a trained model, which two things should you do?",
    "choices": ["Call model.train() and enable gradients", "Call model.eval() and wrap the forward pass in torch.no_grad()", "Apply softmax and call loss.backward()"],
    "answer": 1,
    "explain": "model.eval() puts layers like dropout and batch-norm into inference behavior, and torch.no_grad() skips gradient tracking for a faster, lighter forward pass since you're not training."
  }
]
```


---

# GPUs, Performance & Common Pitfalls

Here's the plain truth nobody tells you when you start: the hard part of PyTorch isn't the concepts, it's the bugs. You'll write a training loop that's structurally perfect, hit run, and watch the loss sit there like a stone - no error, no clue, just a model that refuses to learn. Or you'll get a wall of red text about devices and memory that means nothing the first time you see it.

The good news, and the whole point of this phase: **almost all of those bugs come from a short, knowable list.** The difference between someone who loses a day to PyTorch and someone who shrugs and fixes it in two minutes isn't talent - it's having seen the bug before. So let's hand you that list. For each one: the symptom you'll actually see, the cause underneath, and the fix.

The mental model to carry through this phase: **PyTorch trusts you completely.** It won't stop you from logging tensors that leak memory, mixing devices, or forgetting `zero_grad()`. That freedom is why researchers love it - and why these pitfalls exist. Knowing the list is how you earn the freedom without paying the tax.

## 1. The pitfall cheat-card

Bookmark this table. When something breaks, scan it first - your bug is very likely sitting right here.

| Symptom | Cause | Fix |
|---------|-------|-----|
| Loss explodes to `nan` or never converges | Forgot `optimizer.zero_grad()` - gradients accumulate across steps ([Phase 6](06-the-training-loop.md)) | Add `optimizer.zero_grad()` before every `loss.backward()` |
| Predictions wrong/random at inference; eval slow | Forgot `model.eval()` and/or `torch.no_grad()` ([Phase 9](09-saving-loading-inference.md)) | Call `model.eval()` and wrap inference in `with torch.no_grad():` |
| `RuntimeError: ... found at least two devices` | Model and data on different devices ([Phase 2](02-tensor-operations-and-gpu.md)) | Move both to the same `device` with `.to(device)` |
| Loss stuck high; "model won't learn" | Wrong loss/label setup - e.g. softmax applied before `CrossEntropyLoss`, or labels the wrong shape/dtype ([Phase 5](05-loss-and-optimizers.md)) | Feed raw logits to `CrossEntropyLoss`; labels are `int64` class indices of shape `(N,)` |
| Loss decreases painfully slowly or oscillates wildly | Learning rate too low (slow) or too high (oscillates); data not shuffled; a logic bug | Tune `lr` (try 10× up/down); shuffle the `DataLoader`; overfit a tiny batch to isolate |
| `CUDA out of memory` | Batch too large, or tensors quietly keeping the graph alive across the loop | Smaller batch; use `.item()`/`.detach()` for logged values; `torch.cuda.empty_cache()` |

The first three are so common they each deserve a closer look. Let's expand them, because seeing the bug *and* the fix side by side is what makes it stick.

## 2. Device discipline - the same-device rule

⚠️ **The model and the data it processes must live on the same device.** This is the single most common GPU error, and everyone hits it exactly once. You moved your model to the GPU, felt good about it, and forgot that the batch coming out of your `DataLoader` is still sitting on the CPU. PyTorch can't do math across two devices - moving data between them is expensive, and it refuses to do that silently behind your back.

Here's the bug:

```python
import torch
import torch.nn as nn

device = "cuda" if torch.cuda.is_available() else "cpu"

model = nn.Linear(10, 2).to(device)   # model on the GPU
X = torch.randn(4, 10)                 # data left on the CPU -- oops

pred = model(X)                        # boom
```

```console
RuntimeError: Expected all tensors to be on the same device, but found at least
two devices, cuda:0 and cpu! (when checking argument for argument mat1 in method wrapper_addmm)
```

*What just happened:* The model's weights are on `cuda:0`, but `X` is still on the `cpu` where it was created. The forward pass tries to multiply them together, and PyTorch stops cold rather than guessing where the work should happen. The error even names both devices for you - `cuda:0` and `cpu` - which is the clue that points straight at the fix.

The fix is the device-agnostic pattern from [Phase 2](02-tensor-operations-and-gpu.md): pick `device` **once**, then `.to(device)` the model and **every batch** as it comes in.

```python
device = "cuda" if torch.cuda.is_available() else "cpu"
model = nn.Linear(10, 2).to(device)        # model -> device, once

for X_batch, y_batch in data_loader:
    X_batch = X_batch.to(device)           # every batch -> same device
    y_batch = y_batch.to(device)
    pred = model(X_batch)                   # now everything agrees
    # ... loss, zero_grad, backward, step ...
```

*What just happened:* Because every tensor is moved to the *same* `device` variable, nothing can drift onto the wrong one. The model went to `device` once at setup; each batch goes to the same `device` inside the loop. This one habit - one `device`, `.to(device)` everything - eliminates the entire category of cross-device errors. Notice the model moves once (its weights persist on the GPU), but data moves every iteration (each batch is freshly loaded on the CPU first).

💡 The tell-tale sign you forgot a `.to(device)` somewhere is that error naming two devices. When you see it, don't panic - just find the tensor that's on the wrong one. It's almost always a batch you forgot to move.

## 3. CUDA out of memory

📝 Your GPU has its own memory (VRAM), and it's smaller than you think - a consumer card might have 8–24 GB, and a model plus its activations plus its gradients all have to fit. When they don't, you get the dreaded:

```console
torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 2.00 GiB
(GPU 0; 8.00 GiB total capacity; 6.50 GiB already allocated; 1.20 GiB free)
```

There are two flavors of this bug, and they have very different causes.

**Flavor one: the batch is too big.** You asked the GPU to hold more than it can. The fix is direct - use a smaller batch size, or a smaller model. This one is straightforward: you're over budget, so spend less.

**Flavor two - the sneaky one: a memory leak across the loop.** This is the bug that confuses people because the model and batch *do* fit, yet memory climbs every iteration until it overflows. The cause is almost always logging. Remember from [Phase 6](06-the-training-loop.md) that `loss` isn't just a number - it's a tensor that holds onto the entire computation graph that produced it. If you accumulate the raw loss tensor for logging, you're quietly keeping *every* iteration's graph alive in memory:

```python
# BAD: total_loss keeps each iteration's whole graph alive
total_loss = 0
for X_batch, y_batch in data_loader:
    pred = model(X_batch)
    loss = loss_fn(pred, y_batch)
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    total_loss += loss          # <-- adds the TENSOR (graph and all)
```

*What just happened:* `total_loss += loss` adds the loss *tensor*, not its value. Each `loss` carries a reference back through the graph to the activations that made it, and `total_loss` now holds a growing chain of them. PyTorch can't free any of that memory because you're still pointing at it. Over a few hundred batches, VRAM fills up and you crash - even though each individual batch fit fine.

The fix is one method call. Use `.item()` (which extracts the plain Python float, severing the graph) for logging, or `.detach()` if you need a tensor without its history:

```python
# GOOD: .item() pulls out the plain number, graph gets freed
total_loss = 0.0
for X_batch, y_batch in data_loader:
    pred = model(X_batch)
    loss = loss_fn(pred, y_batch)
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    total_loss += loss.item()   # <-- adds a float; graph is released
```

*What just happened:* `loss.item()` returns a bare Python `float` with no tie to the computation graph. Once the loop moves on, nothing references this iteration's activations, so PyTorch frees them. Memory stays flat across the whole epoch. This is the same `.item()` habit from Phase 6, and now you can see *why* it matters beyond tidy printing - it's the difference between a stable run and a slow memory leak.

💡 If you ever want to clear cached GPU memory mid-program, `torch.cuda.empty_cache()` releases blocks PyTorch is holding for reuse - handy in notebooks where old tensors linger. But it's a band-aid: if memory grows every iteration, you have a leak (flavor two), and `empty_cache()` won't save you. Find the tensor you forgot to `.item()`.

## 4. Speed - keep the GPU fed

Once your code *runs*, the next question is whether it runs *fast*. The mental model here is counterintuitive: 💡 **a GPU is so fast at math that the bottleneck is usually getting data to it, not the math itself.** A GPU sitting idle waiting for the next batch is the most common performance problem, and it's invisible unless you look for it.

The first lever is the `DataLoader` from [Phase 7](07-datasets-and-dataloaders.md). Two arguments keep the pipeline flowing: `num_workers` spins up background processes that prepare the next batches while the GPU chews on the current one, and `pin_memory=True` makes the CPU→GPU transfer faster. Together they stop the GPU from starving:

```python
loader = DataLoader(dataset, batch_size=64, shuffle=True,
                    num_workers=4, pin_memory=True)
```

*What just happened:* `num_workers=4` lets four background processes load and transform data in parallel, so a batch is ready the instant the GPU asks for it. `pin_memory=True` parks those batches in a region of RAM that transfers to the GPU faster. The GPU stops idling between batches. (Start with `num_workers` around the number of CPU cores you have and tune from there; too many can backfire.)

The second lever is **mixed precision**. By default PyTorch does math in 32-bit floats, but modern GPUs run *much* faster in 16-bit - and for most of training, 16-bit precision is plenty. `torch.cuda.amp` (Automatic Mixed Precision) flips the heavy operations to 16-bit while keeping the parts that need full precision safe, often giving a large speedup and roughly halving memory use for a few extra lines. On a recent GPU it's close to free performance, and worth reaching for once your loop works.

The third, newest lever: **`torch.compile`** (PyTorch 2.x). Wrap your model in `model = torch.compile(model)` and PyTorch traces and optimizes the whole computation into faster fused operations - often a real speedup with a single line, no other changes.

⚠️ One discipline above all: **profile before you optimize.** Don't guess where the time goes - measure it (PyTorch ships `torch.profiler`). Nine times out of ten the answer is "the GPU is starved for data," and you'll fix the `DataLoader` instead of micro-optimizing math that was never the problem.

## 5. The debugging mindset

Let's end with the meta-skill, because it's worth more than any single fix. When a model "won't learn," beginners reach for the tuning knobs - more epochs, a fancier optimizer, a bigger network. 💡 **But the overwhelming majority of "it won't learn" bugs aren't tuning problems at all.** They're one of four things:

1. **Shapes are wrong** - a `(batch, features)` that should be `(features, batch)`, a label tensor with an extra dimension. (Print `.shape`, as drilled in [Phase 2](02-tensor-operations-and-gpu.md).)
2. **Loss and labels are mismatched** - softmax applied before `CrossEntropyLoss`, or labels as floats when they should be `int64` class indices ([Phase 5](05-loss-and-optimizers.md)).
3. **The learning rate is off** - too high and the loss oscillates or explodes; too low and it barely moves.
4. **A missing `zero_grad()`** - the silent killer from [Phase 6](06-the-training-loop.md).

Here's the single best diagnostic, the one trick that separates frustrating debugging from productive debugging: **overfit a tiny batch.** Take five examples and train the model on just those, over and over, with no shuffling. A working model should *memorize* five examples easily - the loss should drop to nearly zero.

```python
# Sanity check: can the model memorize 5 examples?
tiny_X, tiny_y = next(iter(data_loader))
tiny_X, tiny_y = tiny_X[:5].to(device), tiny_y[:5].to(device)

model.train()
for step in range(200):
    pred = model(tiny_X)
    loss = loss_fn(pred, tiny_y)
    optimizer.zero_grad()
    loss.backward()
    optimizer.step()
    if step % 50 == 0:
        print(f"step {step:3d} | loss {loss.item():.4f}")
```

```console
step   0 | loss 2.4127
step  50 | loss 0.0832
step 100 | loss 0.0041
step 150 | loss 0.0006
```

*What just happened:* The model was asked to do the easiest possible task - memorize five fixed examples - and the loss collapsed toward zero, exactly as a healthy model should. **That's a passing sanity check:** the wiring is sound. If instead the loss had stayed flat or refused to drop, you'd know the problem is a *bug* (shapes, labels, a missing `zero_grad`), not a tuning issue - and you'd go hunting in that short list above instead of wasting hours on hyperparameters. This test takes thirty seconds and saves entire afternoons.

💡 That's the whole secret, and it's worth saying plainly: **the difference between frustrating and productive PyTorch is knowing this short list.** Every error message points to one of these. You're not memorizing PyTorch - you're memorizing the half-dozen ways it goes wrong, and how to recognize each one fast.

## Recap

- **Most PyTorch bugs come from a short, knowable list** - the cheat-card in section 1 covers the classics: missing `zero_grad()`, missing `eval()`/`no_grad()`, device mismatches, loss/label setup, learning rate, and out-of-memory.
- **Same-device rule:** model and data must be on one device. Pick `device` once, `.to(device)` the model and *every batch* - that habit kills the `two devices` error.
- **CUDA out of memory** has two flavors: a batch that's genuinely too big (use a smaller batch), and a sneaky leak from logging the raw `loss` tensor (`total_loss += loss` BAD → `+= loss.item()` GOOD).
- **Keep the GPU fed:** tune `DataLoader` (`num_workers`, `pin_memory`), use `torch.cuda.amp` mixed precision and `torch.compile` for speed - and profile before optimizing, because the bottleneck is usually data, not math.
- **Debugging mindset:** "won't learn" is almost always shapes, loss/labels, learning rate, or a missing `zero_grad`. Overfit a tiny batch first - if the model can't memorize 5 examples, it's a bug, not a tuning problem.

## Quick check

```quiz
[
  {
    "q": "Your training loop crashes with 'CUDA out of memory' after a few hundred batches, even though batch size is small and the first batches ran fine. What's the most likely cause?",
    "choices": ["The GPU is too small for any training", "You're accumulating the raw loss tensor (e.g. total_loss += loss), keeping every iteration's graph alive", "torch.compile is using too much memory"],
    "answer": 1,
    "explain": "Growing memory across iterations is the classic logging leak. The loss tensor carries its whole computation graph; adding it to a running total keeps every graph alive. Use total_loss += loss.item() to add a plain float and let the graphs be freed."
  },
  {
    "q": "You get 'RuntimeError: Expected all tensors to be on the same device, but found cuda:0 and cpu'. What's the fix?",
    "choices": ["Restart the GPU", "Move both the model and every batch to the same device with .to(device)", "Wrap the forward pass in torch.no_grad()"],
    "answer": 1,
    "explain": "The model is on the GPU but a batch is still on the CPU. Pick one device variable and .to(device) the model (once) and every batch (each iteration) so all tensors agree."
  },
  {
    "q": "Your model 'won't learn' - the loss won't drop. What's the smartest first diagnostic?",
    "choices": ["Add more epochs and a bigger network", "Switch to a fancier optimizer", "Overfit a tiny batch of ~5 examples - if it can't memorize them, it's a bug, not a tuning problem"],
    "answer": 2,
    "explain": "A healthy model memorizes 5 fixed examples easily, driving loss near zero. If it can't, the problem is a bug (shapes, loss/labels, missing zero_grad), so you hunt there instead of wasting time on hyperparameters."
  }
]
```


---

# Where to Go Next

Stop and look at what you can actually do now. You understand a **tensor** - the GPU-ready, autograd-aware array that everything runs on. You understand **autograd** - how PyTorch tracks operations and computes the gradients that make learning possible. You can build a model as a `nn.Module`, define a `forward()`, pick a loss and an optimizer, and write the **training loop** that ties them together: forward, loss, backward, step. You can feed it data through a `Dataset` and `DataLoader`, move work onto a GPU, and you've trained, evaluated, **saved**, and run inference with a real classifier.

That is not a warm-up. That is the foundation *every* piece of modern deep learning is built on - from a three-line toy model to a model with hundreds of billions of parameters. The rest is the same three ideas at scale, plus tooling. This last phase isn't another layer type - it's the clear-eyed map of where to point what you know.

## Don't train from scratch - transfer learning

💡 Here's the single most practical next step, and the one that surprises people: most of the time, you **don't train a model from scratch.** You start from one someone else already trained on a giant dataset, and you fine-tune it on yours. It's called **transfer learning**, and it's how real projects get good results without a data center or a million labeled images.

The intuition is clean. A vision model trained on millions of images has already learned the universal stuff in its early layers - edges, textures, shapes, the general grammar of "what images look like." That knowledge transfers. What's specific to *your* problem (cats vs. dogs, healthy vs. diseased leaves) lives mostly in the final layers. So you **freeze the early layers**, keep their learned weights, and retrain only a fresh **head** on your data. Far less data, far less compute, far better results.

```python
import torch
import torchvision

# Load a pretrained network (weights trained on ImageNet)
model = torchvision.models.resnet18(weights="IMAGENET1K_V1")

# Freeze everything - these gradients won't update
for param in model.parameters():
    param.requires_grad = False

# Swap the final layer for one that fits YOUR classes (say, 3)
model.fc = torch.nn.Linear(model.fc.in_features, 3)

# Now only model.fc has requires_grad=True - train just the head
# with the exact same loop you already know.
```

📝 Notice what *isn't* new here: `requires_grad` is the autograd flag from Phase 3, the new `fc` layer is the `nn.Module` idea from Phase 4, and you'd train it with the same loop from Phase 6. Transfer learning isn't a new skill - it's your existing skills aimed at a pretrained starting point.

## The ecosystem

PyTorch sits at the center of a big, friendly ecosystem. You don't need all of it, but it helps to know which branch solves which problem.

```mermaid
flowchart TD
  P[PyTorch - tensors, autograd, the loop] --> HF[Hugging Face<br/>pretrained models, datasets, hub]
  P --> L[Lightning / fast.ai<br/>less loop boilerplate]
  P --> D[torchvision / torchaudio / torchtext<br/>domain toolkits]
  P --> S[Deployment<br/>TorchServe · ONNX · torch.compile]
```

*What this shows:* one core, four directions you'll actually reach for.

- **Hugging Face** is the center of modern practice. Its `transformers`, `datasets`, and `hub` libraries give you thousands of pretrained models for text, vision, and audio that you can download and fine-tune in a few lines. When people say "I used a pretrained model," they usually mean from the Hugging Face Hub.
- **PyTorch Lightning** and **fast.ai** remove the training-loop boilerplate. You learned to write the loop by hand on purpose - now that you understand it, these let you stop rewriting it for every project while keeping the same mental model underneath.
- **torchvision / torchaudio / torchtext** are domain toolkits: pretrained models, datasets, and the standard transforms for images, sound, and text.
- **Deployment** is how a trained model leaves your notebook. **TorchServe** serves it behind an endpoint, **ONNX** exports it to a portable format other runtimes can run, and `torch.compile` speeds it up with a single line.

## The LLM connection

💡 Now the big one. The large language models you've heard of - the ones that write code and hold conversations - **are PyTorch.** They're a specific architecture (the *transformer*) trained with the exact loop you learned in Phase 6: forward, loss, backward, step. The difference is scale - staggering amounts of data, compute, and parameters - not a different kind of magic. Knowing PyTorch is what turns an LLM from a mysterious oracle into "oh, it's a very large model trained the way I now understand."

And here's the freeing part: **you almost never need to train one to use one.** Training a frontier LLM costs millions; *using* one is an API call. When you want to put an LLM to work, [Using an LLM API](/guides/using-an-llm-api) shows you how. And when you're deciding whether you even need to customize a model's behavior, [Fine-Tuning vs Prompting](/guides/fine-tuning-vs-prompting) lays out the plain tradeoff - most of the time a good prompt beats fine-tuning, and fine-tuning (the transfer-learning idea from earlier, applied to language) is the heavier tool you reach for only when prompting genuinely isn't enough.

So PyTorch demystifies the whole stack. You learned the small loop; the giant models are that loop at scale; and you can use them without ever running it yourself.

## What to build - and the last word

The way this knowledge sticks is by building one real thing. Pick whichever pulls at you:

- **Fine-tune a pretrained image model on your own photos.** Grab `resnet18`, freeze it, retrain the head on a few hundred of your own pictures sorted into folders. This is transfer learning end to end, and it's genuinely useful.
- **Train a text classifier with Hugging Face.** Pull a pretrained model from the Hub, fine-tune it on a labeled dataset (sentiment, spam, topic), and watch how few lines it takes.
- **Push the MNIST classifier past 99%.** Take the one you built in Phase 8 and make it better - add convolutional layers, tune the optimizer, add augmentation. The loop stays the same; you're just refining the model.

When you want the canonical reference, the **official PyTorch tutorials** are excellent, and the **"Deep Learning with PyTorch: A 60 Minute Blitz"** is the best fast tour of everything you've learned, in one sitting. Bookmark both.

And remember the through-line of this whole guide. None of it was sorcery. A tensor is an array you do fast math on. Autograd is bookkeeping that hands you gradients. The training loop is four steps in a row, repeated. **Tensors, autograd, the loop** - you have the three ideas that sit under all of modern AI. Go fine-tune something, break it, fix it, and show someone. You're ready.

## Recap

1. **You own the foundation.** Tensors, autograd, models, the loop, data pipelines, saving and inference - that's what *all* deep learning is built on, from tiny models to giant LLMs.
2. **Don't train from scratch - fine-tune.** Transfer learning starts from a pretrained model (`torchvision.models`), freezes the early layers, and retrains only the head. Far less data and compute, far better results.
3. **Know the ecosystem.** Hugging Face for pretrained models, Lightning/fast.ai to shed loop boilerplate, the torch domain toolkits, and TorchServe/ONNX/`torch.compile` for deployment.
4. **LLMs are PyTorch.** They're transformers trained with the loop you learned, at massive scale - and you can *use* one via an API without training it. Prompt first; fine-tune only when you must.
5. **Build one real thing.** Fine-tune an image model on your photos, train a text classifier with Hugging Face, or push your MNIST model past 99% - then read the official tutorials and the 60-Minute Blitz.

## Quick check

Three questions on the decisions that matter most as you leave this guide:

```quiz
[
  {
    "q": "Why is transfer learning usually better than training a vision model from scratch?",
    "choices": [
      "It starts from a model that already learned general features (edges, textures), so you need far less data and compute",
      "It is the only way to use a GPU",
      "Training from scratch is impossible in PyTorch",
      "It skips the training loop entirely"
    ],
    "answer": 0,
    "explain": "A pretrained model's early layers already capture universal visual features. You freeze those, keep their weights, and retrain only the head on your data - so you get good results with a fraction of the data and compute."
  },
  {
    "q": "What is the relationship between large language models and PyTorch?",
    "choices": [
      "LLMs are a separate technology unrelated to PyTorch",
      "LLMs are transformer models trained with the same forward/loss/backward/step loop you learned, just at enormous scale",
      "PyTorch can only train image models, not language models",
      "You must train an LLM yourself before you can use one"
    ],
    "answer": 1,
    "explain": "LLMs are the transformer architecture trained with the exact loop from Phase 6, scaled up massively. And you can use one through an API without ever training it yourself."
  },
  {
    "q": "You want pretrained models for NLP and vision plus the datasets and hub to fine-tune them. Where do you look?",
    "choices": [
      "TorchServe",
      "ONNX",
      "Hugging Face (transformers, datasets, hub)",
      "torch.compile"
    ],
    "answer": 2,
    "explain": "Hugging Face is the center of modern practice - its transformers, datasets, and hub libraries give you thousands of pretrained models to download and fine-tune. TorchServe, ONNX, and torch.compile are deployment/optimization tools, not model hubs."
  }
]
```
