# C# From Zero

> Learn C# from nothing to genuinely advanced: install .NET and the basics - types, classes, objects, collections - then the deep half: generics, delegates and lambdas, LINQ, modern records and pattern matching, async/await, the .NET runtime and its garbage collector, testing, and performance. Mental-model-first, with clear explanations.


---

# C# From Zero

C# is Microsoft's flagship language, and over two decades it's grown into one of the most pleasant,
capable languages you can learn. It runs on **.NET**, a cross-platform runtime that's no longer
Windows-only - your C# runs on Linux, macOS, in the cloud, on phones, and in game engines. It powers a
huge share of enterprise backends (ASP.NET Core), desktop and mobile apps (MAUI), and - through Unity -
a large slice of the world's games. The language is famous for absorbing good ideas early: it had
LINQ and `async`/`await` years before most rivals, and it keeps evolving.

This guide takes you the whole way: from "I've never run `dotnet`" to understanding what C# and the .NET
runtime are *actually doing* underneath your code. We go mental-model-first the whole way: before any
command, you'll understand what the thing actually *is* and why C# made the choice it did.

> 📝 This guide teaches the **language**. If you've never programmed at all, start with
> [Programming From Zero](/guides/programming-from-zero) first - it covers the universal ideas (what a
> program is, variables, loops) every language shares. Then come back here.

It's one zero-to-hero journey in two halves. **Phases 1–9 are the basics** - enough to write real,
well-organized object-oriented programs. **Phases 10–17 are the deep half** - generics, delegates and
lambdas, LINQ, modern C# (records, pattern matching, nullable reference types), `async`/`await`, the
.NET runtime and garbage collector, testing, and performance, the stuff that separates "writes C#" from
"understands C#." Each phase carries a difficulty badge so you can see the climb.

## How to read this

- **Brand new to C#?** Read 1–9 in order, top to bottom - each builds on the last. Type the examples
  yourself; doing beats reading. Come back for 10+ when the basics feel comfortable.
- **Already know another language?** Skim phases 1–4 for C#'s spelling of ideas you have, then slow down
  at [Phase 5: Classes & Objects](05-classes-and-objects.md) - object-orientation is the spine of C#,
  and everything after assumes you think in objects.
- **Past the basics already?** Jump to the deep half - [Phase 10: Generics, Deep](10-generics-deep.md)
  onward is where C# stops being "a tidy OOP language" and becomes one you can reason about down to the
  garbage collector.

## The phases

**Part 1 - The basics (🟢 Basic → 🟡 Intermediate)**
1. **[Install & Your First Program](01-install-and-first-program.md)** 🟢 - the .NET SDK, the `dotnet` CLI, the CLR, a real program.
2. **[Syntax, Values & Types](02-syntax-values-and-types.md)** 🟢 - value vs reference types, `var`, static typing, strings, nullability.
3. **[Collections](03-collections.md)** 🟢 - arrays, `List<T>`, `Dictionary<K,V>`, `HashSet<T>`.
4. **[Control Flow & Methods](04-control-flow-and-methods.md)** 🟢 - `if`/`switch` (and switch expressions), loops, methods, overloading.
5. **[Classes & Objects](05-classes-and-objects.md)** 🟡 - **the spine of C#:** classes, properties, constructors, encapsulation.
6. **[Inheritance & Interfaces](06-inheritance-and-interfaces.md)** 🟡 - inheritance, interfaces, polymorphism, and `abstract`.
7. **[Errors & I/O](07-errors-and-io.md)** 🟡 - exceptions, `try`/`catch`/`finally`, `using`, and files.
8. **[Projects, NuGet & Tooling](08-projects-and-tooling.md)** 🟡 - projects/solutions, NuGet packages, the .NET ecosystem, the build.
9. **[Idioms & Gotchas](09-idioms-and-gotchas.md)** 🟡 - the C# way, and the traps (`null`, value vs reference equality, deferred `IEnumerable`) that bite everyone once.

**Part 2 - Beyond the basics (🔴 Advanced)**
10. **[Generics, Deep](10-generics-deep.md)** 🔴 - type parameters, constraints, and covariance/contravariance.
11. **[Delegates, Lambdas & Events](11-delegates-and-lambdas.md)** 🟡 - delegates, `Func`/`Action`, lambdas, and events.
12. **[LINQ](12-linq.md)** 🔴 - query and method syntax, deferred execution, and the LINQ pipeline.
13. **[Records, Pattern Matching & Modern C#](13-records-and-modern-csharp.md)** 🟡 - records, pattern matching, switch expressions, and nullable reference types.
14. **[async/await & Tasks](14-async-await-and-tasks.md)** 🔴 - the `Task` model, `async`/`await` in depth, and what really happens when you `await`.
15. **[The .NET Runtime: Memory, GC & JIT](15-the-dotnet-runtime-and-gc.md)** 🔴 - the managed heap, GC generations, value vs reference memory, JIT and IL.
16. **[Testing, Build & Profiling](16-testing-and-profiling.md)** 🟡 - xUnit, `dotnet test`, and profiling.
17. **[Performance & the Ecosystem](17-performance-and-ecosystem.md)** 🔴 - `Span<T>`, cutting allocations, measuring first, and the .NET ecosystem.

**Finale**
18. **[Where to Go Next](18-where-to-go-next.md)** 🟢 - ASP.NET Core, Blazor, MAUI, game dev with Unity, and what to build.

> Frameworks (ASP.NET Core, Blazor, MAUI, Unity) are their own world - this guide makes the *language and
> the .NET runtime* make sense, top to bottom.


---

# Install & Your First Program

Before C# is fun, you need a tool that turns your code into something runnable, and one tiny program that proves it works. With modern .NET both are quick - and the first program already teaches you how C# *thinks*: source isn't run directly, but compiled to a portable in-between form a runtime turns into real machine code as it runs.

## The mental model: your code becomes IL, the CLR runs it

When you "run" C#, three often-confused layers are involved:

- **The .NET SDK** - the developer toolkit: the C# compiler, the `dotnet` command-line tool, project templates, and everything needed to *build* software. This is what you install to write code.
- **The .NET runtime** - what your finished program needs to *run*. The SDK includes a runtime, so installing the SDK gives you both. (End users who only run .NET apps install just the runtime, not the whole SDK.)
- **The CLR (Common Language Runtime)** - the engine *inside* the runtime that actually executes your program: it loads code, manages memory (garbage collection), and turns the portable code into native instructions.

📝 **Terminology.** SDK = build tools. Runtime = what runs a finished app. CLR = the execution engine at its heart. Rule of thumb: install the **SDK** to develop, ship to people who have the **runtime**, and the **CLR** does the work at execution time.

Here's the part that surprises people coming from C or Go: the C# compiler does **not** produce a native `.exe` full of machine code for your CPU. It produces **IL** (Intermediate Language, also called CIL or MSIL): a compact, CPU-independent bytecode. At runtime, the CLR's **JIT** (Just-In-Time) compiler translates that IL into native machine code for the specific machine it's on, method by method, the first time each is called.

```mermaid
flowchart LR
  A["Program.cs"] --> B["C# compiler"]
  B --> C["IL (bytecode)"]
  C --> D["CLR + JIT"]
  D --> E["native code"]
```

*What just happened:* Your `.cs` source goes through the compiler once to become IL, which is what ships inside a .NET program. The CLR loads that IL and JITs it to native code on demand as the program runs - why the *same* compiled .NET program runs on Windows, macOS, and Linux: the IL is portable, and each platform's CLR JITs it to that platform's instructions.

💡 **Why this design.** Compiling to IL instead of straight to native buys portability (one build, many platforms) and a runtime that can optimize using information only known at execution time, like which CPU model you have. The cost: a runtime must be present, plus a brief JIT warm-up the first time code runs - invisible and worth it for most software.

One naming knot worth untangling: **Modern .NET** (.NET 5, 6, 7, 8, and onward) is cross-platform and open source; it used to be called ".NET Core." The **legacy .NET Framework** (4.x and earlier) is Windows-only and in maintenance mode. ⚠️ "Create a new project" almost always means modern .NET - that's what you're installing here. ".NET Framework" is the old Windows-only line, not what new code targets.

## Install the .NET SDK

Go to **[dotnet.microsoft.com/download](https://dotnet.microsoft.com/download)** and grab the **SDK** (not the runtime-only option) for the latest .NET version on your OS. Run the installer and accept the defaults. On macOS and Linux a package manager is often tidier:

```bash
# macOS (Homebrew)
brew install --cask dotnet-sdk

# Debian / Ubuntu (Microsoft package feed configured)
sudo apt-get install -y dotnet-sdk-10.0

# Windows (winget)
winget install Microsoft.DotNet.SDK.10
```

*What just happened:* Each fetches and installs the .NET SDK - compiler, `dotnet` CLI, and a runtime - and puts `dotnet` on your `PATH`. Pick the line matching your machine.

Now confirm it worked:

```bash
dotnet --version
```

```console
10.0.100
```

*What just happened:* `--version` reported the installed SDK version - anything 8.0 or newer is fine here. A version line means the toolchain is on your `PATH`.

⚠️ **Gotcha.** `dotnet: command not found` (or `'dotnet' is not recognized` on Windows) right after installing usually means a terminal window open *before* the install doesn't know about the updated `PATH` yet. Close it and open a fresh one; the install is fine.

## Your first program

Unlike single-file Java or Go, C# is **project-oriented** from the first program - you create a small project, not a lone source file. The CLI scaffolds one for you:

```bash
dotnet new console -o hello
cd hello
```

*What just happened:* `dotnet new console` created a **console** (text-based) app from a built-in template; `-o hello` put it in a new folder. Inside are two files that matter: `Program.cs` (your code) and `hello.csproj` (more in a moment). Open `Program.cs`:

```csharp
Console.WriteLine("Hello, World!");
```

*What just happened:* That single line is a complete, runnable C# program. `Console` is a type from .NET's standard library representing the text terminal; `WriteLine` prints its argument followed by a newline. No class, no `Main`, no `using` directive - yet it compiles and runs, thanks to **top-level statements**.

📝 **Top-level statements** let you write executable code directly at the top of a file, without wrapping it in a class and a method. The compiler treats the file's statements as the body of the entry point - the same machinery as the classic form below, with the boilerplate written for you.

To run it:

```bash
dotnet run
```

```console
Hello, World!
```

*What just happened:* `dotnet run` compiled `Program.cs` to IL, handed it to the CLR, and the JIT ran it - printing the text, with no intermediate files or manual compiler invocation. It's the "just try it" command you'll lean on constantly.

### What top-level statements desugar to

You'll meet the longer form in older code and most non-trivial programs. Under the hood, the compiler wraps your top-level statements in a class with a special `Main` method - historically *the* entry point of every C# program:

```csharp
using System;

class Program
{
    static void Main(string[] args)
    {
        Console.WriteLine("Hello, World!");
    }
}
```

*What just happened:* This is the classic, fully-spelled-out version your one-liner desugars into. `using System;` brings the `System` namespace into scope so you can write `Console` instead of `System.Console`. `class Program` is a container for your code (C# organizes everything into classes). `static void Main(string[] args)` is the **entry point** - the method the CLR calls when your program starts; `args` holds command-line arguments. The top-level form lets the compiler generate all this so a beginner's first program isn't a wall of ceremony.

💡 **Key point.** Top-level statements and the explicit `class Program { static void Main }` form are *the same thing* - one is compiler shorthand for the other. Use the short form for small programs and scripts; the long form appears in larger codebases and when you need fine control over the entry point.

## The project model: `.csproj` and why C# is project-first

That `hello.csproj` file the template created is the **project file**, central to how C# work is organized. Peek inside - it's surprisingly small:

```xml
<Project Sdk="Microsoft.NET.Sdk">

  <PropertyGroup>
    <OutputType>Exe</OutputType>
    <TargetFramework>net10.0</TargetFramework>
    <Nullable>enable</Nullable>
    <ImplicitUsings>enable</ImplicitUsings>
  </PropertyGroup>

</Project>
```

*What just happened:* The `.csproj` is an XML file that *describes* the project: that it builds an executable (`OutputType`), which .NET version it targets (`TargetFramework`), and a couple of language settings. `ImplicitUsings` is why your one-liner didn't need `using System;` - common namespaces import automatically. You rarely edit this by hand early on, but `dotnet run` reads it to know *how* to build your program.

📝 **Mental model.** A C# program is a **project**, not a loose file. `dotnet build` compiles the project; `dotnet run` builds *and* runs it. Unlike languages where you point a tool at a single source file, here the `.csproj` is the unit of work - extra structure that lets projects scale cleanly to many files and dependencies.

## Your everyday workflow

You can write C# in any editor, but a few setups are worth knowing:

- **Visual Studio** (Windows/Mac) - Microsoft's full, heavy IDE, deeply integrated with .NET.
- **VS Code + the C# Dev Kit extension** - lightweight, cross-platform, the most common modern choice.
- **JetBrains Rider** - a polished cross-platform IDE many professionals prefer.

All three give autocomplete, inline error highlighting, and a debugger. While learning, the `dotnet` CLI plus any editor is plenty.

One CLI command makes your loop smoother right away:

```bash
dotnet watch run
```

*What just happened:* `dotnet watch run` runs your program and *keeps watching* your source files - edit and save `Program.cs`, and it recompiles and re-runs automatically, no retyping `dotnet run`. It's the tightest feedback loop for experimenting.

We'll keep to single-project programs for now. Third-party libraries via **NuGet** and multi-project **solutions** come later, in Phase 8 - not needed to learn the language itself.

## Recap

1. **C# compiles to IL** (a portable bytecode), and the **CLR** runs it by JIT-compiling that IL to native machine code at execution time - which is why one build runs on Windows, macOS, and Linux.
2. **SDK vs runtime vs CLR**: install the SDK to develop, ship to machines that have the runtime, and the CLR is the engine that executes your code. **Modern .NET** is cross-platform (formerly ".NET Core"); legacy **.NET Framework** is the old Windows-only line.
3. Install the SDK from [dotnet.microsoft.com/download](https://dotnet.microsoft.com/download) (or a package manager) and confirm with **`dotnet --version`**.
4. **`dotnet new console`** scaffolds a project; **`dotnet run`** builds and runs it. **Top-level statements** let your first program be a single line, which desugars to the classic `class Program { static void Main }`.
5. A C# program is a **project** described by a **`.csproj`** file - C# is project-oriented from the start, unlike single-file languages.
6. Use an IDE (Visual Studio, VS Code + C# Dev Kit, Rider) and **`dotnet watch run`** for an auto-reloading feedback loop; NuGet and solutions come in Phase 8.

Next, we give the program something to work with: named values, the types C# stores them in, and the rules that govern how they behave.

## Quick check

Test yourself on the idea that underpins everything else - how C# actually runs:

```quiz
[
  {
    "q": "When the C# compiler builds your program, what does it produce?",
    "choices": [
      "IL (Intermediate Language) bytecode, which the CLR JIT-compiles to native code at runtime",
      "Native machine code for your exact CPU, ready to run with no runtime",
      "An interpreted script the runtime reads line by line every time",
      "A .csproj file describing how to build the program later"
    ],
    "answer": 0,
    "explain": "The compiler emits portable IL, not native code. At runtime the CLR's JIT compiler translates that IL into native machine code for the current machine - which is what makes one build run across platforms."
  },
  {
    "q": "What's the difference between the .NET SDK and the .NET runtime?",
    "choices": [
      "The SDK is the developer toolkit (compiler + CLI + a runtime); the runtime is just what's needed to run a finished app",
      "They're two names for the same download",
      "The SDK runs apps and the runtime builds them",
      "The SDK is for Windows and the runtime is for macOS and Linux"
    ],
    "answer": 0,
    "explain": "You install the SDK to develop - it bundles the compiler, the dotnet CLI, and a runtime. End users who only run .NET apps install just the runtime."
  },
  {
    "q": "The one-line `Console.WriteLine(\"Hello, World!\");` with no class or Main - what is it?",
    "choices": [
      "A top-level statement that the compiler desugars into a class with a static Main entry point",
      "A special scripting mode that skips compilation entirely",
      "An error that only works in the REPL, not in a real project",
      "A different language feature unrelated to the classic Main method"
    ],
    "answer": 0,
    "explain": "Top-level statements let you skip the boilerplate; the compiler wraps them in a generated class with a static Main method. The short form and the explicit class Program { static void Main } form are the same program."
  }
]
```


---

# Syntax, Values & Types

Phase 1 got a program to print and run - the warm-up. Real programs hold things: a name, a count, a price -
and the rules C# uses to *store* those things shape almost everything you'll write next. Two ideas do the
heavy lifting: the compiler knows the type of every value before your code runs, and C# splits all values
into two camps - ones that *copy* when passed around, and ones that *share*. Get that split into your head
now and a whole category of "why did this change?" bugs never happens to you.

## Static typing - the compiler checks before you run

**What it actually is.** C# is **statically typed**: every variable has a fixed type locked in when
declared, and the compiler verifies every use of it *before the program runs*. A variable holding a whole
number can never later hold text - try it, and the build fails with a red squiggle, no runtime crash.

📝 **Static** means "checked at compile time," as opposed to *dynamic* ("checked while running"). C# checks
up front, so by the time your program runs, "what type is this?" is already answered and can't go wrong.

```csharp
int count = 0;
count = count + 1;     // fine: an int plus an int is an int
count = "hello";       // compile error: cannot convert string to int
```
*What just happened:* `int count = 0;` told the compiler "`count` is an integer, forever." Adding to it is
fine; assigning the text `"hello"` is a contradiction the compiler catches instantly with
`error CS0029: Cannot implicitly convert type 'string' to 'int'`. A whole class of "I thought this was a
number but it was text" mistakes simply cannot ship.

## Value types vs reference types - copy or share?

This is *the* idea in this phase.

📝 C# splits every type into two families:

- **Value types** hold their data *directly*. When you assign or pass one, you get a **copy** - a separate,
  independent value. These include `int`, `double`, `bool`, `char`, every `struct`, and every `enum`.
- **Reference types** hold a *reference* (a pointer) to data that lives elsewhere in memory. When you assign
  or pass one, you copy the *reference*, not the data - so both names now point at the **same** object.
  These include every `class`, plus `string` and arrays.

Why the split? Small, simple values (a number, a flag) are cheapest to copy. Big or shared things (an
object with many fields, a list everyone needs the same version of) are cheapest to pass by reference. The
catch: copy-vs-share changes *behavior*, not just performance - what trips people up.

A `struct` is a value type; a `class` is a reference type:

```csharp
struct PointVal { public int X; }   // value type
class PointRef  { public int X; }   // reference type

var a = new PointVal { X = 1 };
var b = a;                          // COPY - b is independent
b.X = 99;
// a.X is still 1, b.X is 99

var c = new PointRef { X = 1 };
var d = c;                          // SHARE - d points at the same object as c
d.X = 99;
// c.X is now 99 too - same object
```
```console
PointVal: a.X = 1,  b.X = 99
PointRef: c.X = 99, d.X = 99
```
*What just happened:* `b = a` copied the *whole struct*, so `b` is independent - changing `b.X` left `a`
untouched. But `d = c` only copied the *reference*; `c` and `d` name one object, so writing through `d`
changed what `c` sees too. Same line of code (`x = y`), opposite result - the only difference is whether
the type is a `struct` (value) or a `class` (reference).

⚠️ **The classic surprise: passing to a method.** The same rule applies when you hand a value to a method.
Pass a struct and the method works on a *copy* - your original is safe. Pass a class and it works on *your
actual object* - changes leak back out.

```csharp
struct Counter { public int N; }
class  Box     { public int N; }

void BumpStruct(Counter c) { c.N++; }   // bumps a copy
void BumpClass(Box b)      { b.N++; }   // bumps the caller's object

var counter = new Counter { N = 0 };
BumpStruct(counter);
// counter.N is still 0 - the method changed its own copy

var box = new Box { N = 0 };
BumpClass(box);
// box.N is now 1 - the method reached the real object
```
```console
counter.N = 0
box.N = 1
```
*What just happened:* `BumpStruct` received a *copy* of `counter`, incremented it, and threw it away -
`counter` never changed. `BumpClass` received a copy of the *reference*, still pointing at the caller's
`box`, so the increment stuck. A method that "didn't change my data" is almost always a value type; one
that "changed my data behind my back" is a reference type.

💡 **Key point.** Ask two questions about any type: *Does assigning it copy or share? Does passing it to a
method protect my original or expose it?* Value types copy and protect; reference types share and expose.
`null` only enters the picture for reference types and nullable value types (next) - only a reference can
point at "nothing."

## `var` - let the compiler infer the type

Spelling out the type twice gets old fast: `Dictionary<string, List<int>> map = new Dictionary<string,
List<int>>();` is a mouthful. `var` lets the compiler **infer** the type from the right side:

```csharp
var name = "Ada";          // compiler sees "Ada" is text → name is a string
var age = 36;              // compiler sees 36 is a whole number → age is an int
var price = 9.99;         // → double
var ready = true;         // → bool
```
*What just happened:* Each `var` told the compiler to figure out the type from the right-hand side. `name`
is *still* a `string` - fully static, fully checked - you just didn't write the word. `var` isn't "any type"
or dynamic; it's shorthand for a type the compiler can already see. `name = 42;` afterward is still the
same compile error, because `name`'s type was fixed the moment it was inferred.

💡 **When to use it.** Reach for `var` when the type is obvious from the right side (`var user = new
User();`) or genuinely long to spell out. Prefer the explicit type when the value alone doesn't make it
clear (`var result = Process();` - what *is* `result`?). The goal is readability, not saving keystrokes.

## Nullability - the war on NullReferenceException

📝 `null` means "this reference points at nothing." Read a value off a `null` reference and you get a
`NullReferenceException` - historically C#'s single most common crash. Modern C# fights back on two fronts.

**Nullable value types.** Value types like `int` can't normally be `null` - they always hold a real number.
But sometimes you need "a number, or nothing yet" (an unanswered survey field, say). Add a `?` to make a
**nullable value type**:

```csharp
int score = 0;            // always a number; cannot be null
int? maybeScore = null;   // a number OR null - note the ?

if (maybeScore.HasValue)
    Console.WriteLine(maybeScore.Value);
else
    Console.WriteLine("no score yet");
```
```console
no score yet
```
*What just happened:* `int` refuses `null` outright. `int?` (shorthand for `Nullable<int>`) adds one extra
state - "nothing" - modeling "not answered yet" plainly instead of faking it with `-1` or `0`.

**Nullable reference types.** Reference types have *always* been able to be `null`, exactly why crashes
were so common. Modern C# (`<Nullable>enable</Nullable>` in your project file, default for new projects)
flips this: you must *opt in* to nullability. A plain `string` means "never null"; `string?` means "might
be null" - and the compiler *warns* you whenever you risk dereferencing a possible `null`.

```csharp
#nullable enable
string name = "Ada";       // promises: never null
string? note = null;       // allowed to be null

Console.WriteLine(name.Length);   // fine - name can't be null
Console.WriteLine(note.Length);   // warning CS8602: possible null dereference
```
*What just happened:* With nullability enabled, `name` is *non-nullable* - the compiler trusts it's never
`null` and lets you use it freely. `note` is a `string?`, so reading `note.Length` earns a warning: "this
might be `null`, you didn't check." The compiler moves the `NullReferenceException` from a 3-a.m.
production crash to a squiggle in your editor. ⚠️ These are *warnings*, not errors, so it's tempting to
ignore them. Don't: each is a real crash the compiler spotted for you. (Deep dive on nullability and the
`?.` / `??` operators in [Phase 13](13-records-and-modern-csharp.md).)

## Strings - text done the C# way

Strings get their own section for a couple of quietly important behaviors.

📝 A `string` is a **reference type**, but an **immutable** one: once created, its characters never change.
Every operation that looks like it modifies a string (uppercasing, replacing, concatenating) builds a
*brand-new* string and leaves the original alone.

The most useful everyday feature is **string interpolation** - prefix with `$` and drop expressions right
inside `{ }`:

```csharp
var name = "Ada";
var age = 36;
var greeting = $"Hello {name}, you are {age} years old.";
Console.WriteLine(greeting);
Console.WriteLine($"Next year: {age + 1}");   // expressions work too
```
```console
Hello Ada, you are 36 years old.
Next year: 37
```
*What just happened:* The `$` turns the string into a template; each `{ ... }` is evaluated and its result
dropped in. `{age + 1}` shows real expressions work inside, not just variable names - far cleaner than
gluing strings with `+`, and what you'll use almost every time.

⚠️ **The equality gotcha - `==` on strings compares *value*.** Coming from Java, this surprises people: in
Java, `==` compares *references* (same object?), so you need `.equals()`. In C#, `==` on strings is
overloaded to compare the actual *text*:

```csharp
var a = "hello";
var part = "hel";
var b = part + "lo";        // built at runtime - a genuinely different object, same text
Console.WriteLine(a == b);  // True - compares the characters
```
```console
True
```
*What just happened:* `a` and `b` are different objects in memory (`b` is built at runtime, so it isn't the
same interned literal as `a`), but `==` on `string` compares their *contents*, so you get `True`. This is special to `string` - your *own* classes compare references by
default, a different story told in [Phase 9](09-idioms-and-gotchas.md). For now: strings compare by value
with `==`, and that's the one you want.

Two more string forms: **verbatim strings** (`@"..."`) turn off escape sequences, so backslashes and line
breaks are literal - perfect for Windows paths and regex. **Raw string literals** (`"""..."""`) let you
paste multi-line text (JSON, SQL, HTML) without escaping anything:

```csharp
var path = @"C:\Users\Ada\file.txt";          // no doubled backslashes needed
var json = """
    { "name": "Ada", "age": 36 }
    """;                                       // quotes inside need no escaping
Console.WriteLine(path);
Console.WriteLine(json);
```
```console
C:\Users\Ada\file.txt
{ "name": "Ada", "age": 36 }
```
*What just happened:* In a normal string, `"C:\Users"` would treat `\U` as an escape and break. `@` says
"take this literally," so backslashes stand as-is. The `"""` raw string holds a block with embedded `"`
quotes and newlines, no escaping required - the compiler even strips leading indentation up to the closing
`"""`. Reach for `@"..."` for paths and `"""..."""` for chunks of structured text.

## Recap

1. **C# is statically typed** - every variable's type is fixed at declaration and checked at *compile time*,
   so type mismatches fail the build instead of crashing at runtime.
2. **Value vs reference is the big split.** Value types (`int`, `double`, `bool`, `struct`, `enum`) hold data
   directly and **copy** on assignment and method calls. Reference types (`class`, `string`, arrays) hold a
   reference and **share** - both names point at the same object.
3. ⚠️ Passing a **struct** to a method protects your original (it gets a copy); passing a **class** exposes
   it (the method reaches your real object).
4. **`var`** infers the type from the right-hand side - still fully static, just less typing. Use it when
   the type is obvious.
5. **Nullability:** `int?` adds "nothing" to a value type; modern **nullable reference types** make `string`
   non-null and `string?` maybe-null, with compiler warnings to kill `NullReferenceException` early.
6. **Strings** are immutable reference types. Use `$"..."` interpolation; remember `==` compares **value**
   (unlike Java); use `@"..."` for paths and `"""..."""` for multi-line text.

Next, we move from single values to *collections* - arrays, the `List<T>` you'll actually use, and
dictionaries for looking things up by key.

## Quick check

Test yourself on the idea that matters most - copy versus share:

```quiz
[
  {
    "q": "You have a `struct Point` and write `var b = a;` then `b.X = 99;`. What is `a.X`?",
    "choices": [
      "Unchanged - a struct is a value type, so `b = a` made an independent copy",
      "99 - `b` and `a` point at the same object",
      "0 - assigning a struct resets its fields",
      "A compile error - you can't copy a struct"
    ],
    "answer": 0,
    "explain": "A struct is a value type: assignment copies the whole value, so `b` is independent of `a`. Changing `b.X` leaves `a.X` exactly as it was. If `Point` were a `class` (reference type), `b = a` would share the object and `a.X` would become 99 too."
  },
  {
    "q": "With nullable reference types enabled, what's the difference between `string` and `string?`?",
    "choices": [
      "`string` is non-nullable (the compiler warns if it might be null); `string?` is allowed to be null",
      "`string` is a value type and `string?` is a reference type",
      "There is no difference - the `?` is just a style preference",
      "`string?` is faster because it skips null checks"
    ],
    "answer": 0,
    "explain": "Under `<Nullable>enable</Nullable>`, a plain `string` promises it's never null and the compiler warns when you risk dereferencing a possible null. `string?` explicitly permits null. The point is to surface NullReferenceException risks as compile-time warnings instead of runtime crashes."
  },
  {
    "q": "In C#, what does `==` do when both operands are strings?",
    "choices": [
      "Compares the actual text (their characters) for equality",
      "Compares whether they're the same object in memory, like Java",
      "Always returns false unless they're literally the same literal",
      "Throws an exception - you must use `.Equals()` for strings"
    ],
    "answer": 0,
    "explain": "`==` is overloaded for `string` to compare value (the characters), so two different string objects with the same text are equal. This differs from Java, where `==` compares references. Note that for your own classes, `==` compares references by default - string is the special case."
  }
]
```


---

# Collections - Arrays, Lists, Dictionaries & Sets

Up to now you've held one value in one variable. Real programs deal in *many*: items in a cart, scores in a
game, users by their IDs. C# gives you four shapes you pick between by asking one question: *how do I need
to get my data back out?* Get that question right and the choice makes itself.

The mental model: a collection is a box, and the boxes differ by what they're *good at*. An array is rigid,
with a fixed number of slots. A `List<T>` grows. A `Dictionary<K,V>` you reach into by label instead of
position. A `HashSet<T>` quietly refuses duplicates. Same idea - hold many things - different trade-offs.

## Arrays - the fixed-size box

**What it actually is.** An **array** is a fixed-length sequence of values, all the same type, laid out
back-to-back in memory. "Fixed-length" is the whole personality: the size is set at creation and never
changes. Want a fifth slot in a four-slot array? Make a new array.

```csharp
int[] nums = { 1, 2, 3 };
Console.WriteLine(nums[0]);      // first element - C# counts from 0
Console.WriteLine(nums[2]);      // third element
Console.WriteLine(nums.Length);  // how many slots
```
```console
1
3
3
```
*What just happened:* `int[] nums = { 1, 2, 3 }` declared an array of three integers and filled it in one
go. You read an element by its **index** in square brackets, and like most languages C# counts from zero -
`nums[0]` is the first item, `nums[2]` the third. `nums.Length` tells you the size. Reaching past the end
(`nums[3]`) throws an `IndexOutOfRangeException` at runtime - no slot to read.

⚠️ **Length is a property, not a method.** It's `nums.Length` with no parentheses. (Confusingly, *other*
collections below use `.Count` instead, and strings use `.Length` too - the naming isn't consistent.)

💡 **When do you actually use an array?** When the size is genuinely fixed and known up front - the seven
days of a week, the RGB channels of a pixel, a lookup table that never grows. Want to *add* items? You've
outgrown arrays.

## The generic collections mental model

Before the growable containers, one idea that unlocks all of them: **generics**.

📝 **Generics.** The `<T>` in `List<T>` is a *type parameter* - a blank you fill in. `List<int>` is "a list
of ints"; `List<string>` is "a list of strings." Collections in `System.Collections.Generic` are
**type-safe**: a `List<int>` only ever holds ints, enforced by the compiler. Stuff a string in and your
code won't compile.

This matters because of what came *before* generics. The old `System.Collections` namespace had types like
`ArrayList` that held `object` - "anything." That sounds flexible, but it was a trap: everything you pulled
out came back as a vague `object` you had to cast, value types got silently boxed (an extra allocation),
and a `string` accidentally added to your list-of-numbers blew up only at *runtime*.

⚠️ **Avoid the old non-generic collections.** `ArrayList`, `Hashtable`, and friends still exist for backward
compatibility, but in modern C# they're a code smell. Always reach for `List<T>`, `Dictionary<K,V>`,
`HashSet<T>`. If you see `ArrayList` in a tutorial, that tutorial is old.

💡 **Program to interfaces where it helps.** The generic collections all implement a layered set of
*interfaces* - names that describe *behavior* rather than a concrete type:

- **`IEnumerable<T>`** - "you can `foreach` over me." The most general; promises nothing but iteration.
- **`ICollection<T>`** - `IEnumerable<T>` plus a `Count` and the ability to `Add`/`Remove`.
- **`IList<T>`** - `ICollection<T>` plus indexed access (`[i]`) and ordering.

You don't need these yet to *use* a list. But accepting the *narrowest interface that does the job* in a
method makes it flexible: one that only loops should take `IEnumerable<T>`, so a caller can hand it a list,
array, or anything iterable. We'll lean on this heavily once we hit LINQ.

## `List<T>` - the growable box

**What it actually is.** A **`List<T>`** is an ordered, indexable, *growable* sequence - an array that
handles its own resizing: `Add` items and it makes room; read and write by index just like an array. It's
the collection you'll use more than all the others combined.

```csharp
var fruits = new List<string> { "apple", "banana" };

fruits.Add("cherry");          // grows by one
fruits.Add("date");

Console.WriteLine(fruits.Count);   // how many - note Count, not Length
Console.WriteLine(fruits[1]);      // index like an array

foreach (var fruit in fruits)
{
    Console.WriteLine(fruit);
}
```
```console
4
banana
apple
banana
cherry
date
```
*What just happened:* `new List<string> { "apple", "banana" }` used **collection initializer** syntax - the
`{ ... }` seeds the list with starting items, the same convenience the array got. `Add` appends to the end
and the list grows itself, no size declared anywhere. Ask how many it holds with `.Count` (lists use
`Count`; arrays used `Length` - the inconsistency is annoying). Read by index with `fruits[1]` like an
array, and `foreach` walks every element in order - the iteration `IEnumerable<T>` promises, for free.

💡 **`foreach` vs indexing.** Use `foreach` to just visit every item (cleaner, no off-by-one risk). Use the
indexer `list[i]` when you need the *position* - to modify a specific slot, or walk two lists in lockstep.

## `Dictionary<K,V>` - the box you reach into by label

A list is perfect when you care about *order* and *position*. But often you want "the item labeled `bob`,"
not "the third item." That's a **dictionary**.

**What it actually is.** A **`Dictionary<K,V>`** stores **key → value** pairs and looks up a value
*instantly* by its key, no scanning. `Dictionary<string, int>` reads as "keys are strings, values are
ints" - usernames to scores, say. (Other languages call this a hash map, hash, or associative array.)

```csharp
var scores = new Dictionary<string, int>
{
    ["alice"] = 50,
    ["bob"] = 30,
};

scores.Add("carol", 90);        // add a new pair
scores["alice"] = 75;           // indexer overwrites an existing key

Console.WriteLine(scores["bob"]);   // look up by key

// Safe lookup for a key that might not exist:
if (scores.TryGetValue("dave", out int daveScore))
    Console.WriteLine($"dave: {daveScore}");
else
    Console.WriteLine("dave not found");

// Iterate the pairs:
foreach (KeyValuePair<string, int> pair in scores)
{
    Console.WriteLine($"{pair.Key} = {pair.Value}");
}
```
```console
30
dave not found
alice = 75
bob = 30
carol = 90
```
*What just happened:* We built a dictionary with initializer syntax (`["alice"] = 50`), then grew it two
ways: `Add("carol", 90)` for a brand-new key, and `scores["alice"] = 75` where the indexer *overwrote* the
existing value. `scores["bob"]` returned its value instantly. `TryGetValue("dave", ...)` returned `false`
with no crash, so we printed the fallback. `foreach` over a dictionary hands you each entry as a
**`KeyValuePair<K,V>`**, with `.Key` and `.Value` on it.

⚠️ **The indexer throws on a missing key.** Reading `scores["dave"]` when `dave` isn't there does **not**
return zero or null - it throws a `KeyNotFoundException` and stops your program. When a key *might* not
exist, use **`TryGetValue`** (or `ContainsKey` first): the indexer is for keys you're certain are present,
`TryGetValue` for keys you're hoping are present.

```csharp
// ContainsKey is the other safe check - handy when you don't need the value yet:
if (scores.ContainsKey("alice"))
    Console.WriteLine("alice is on the board");
```
```console
alice is on the board
```
*What just happened:* `ContainsKey` answers a plain yes/no without fetching the value. Prefer `TryGetValue`
when you'll *use* the value right after (it does the lookup once); reach for `ContainsKey` when you only
need the boolean. Both spare you the exception.

## `HashSet<T>` - the box that refuses duplicates

The last container is the specialist. A **`HashSet<T>`** holds **unique** elements - add the same value
twice and the second add is silently ignored - and answers "do you contain this?" very fast.

```csharp
var seen = new HashSet<string>();

Console.WriteLine(seen.Add("apple"));   // true - newly added
Console.WriteLine(seen.Add("apple"));   // false - already present, ignored
seen.Add("banana");

Console.WriteLine(seen.Count);          // 2, not 3
Console.WriteLine(seen.Contains("banana"));  // fast membership test
```
```console
True
False
2
True
```
*What just happened:* `Add` returns a `bool` telling you whether the value was actually new - `true` the
first time `"apple"` went in, `false` the second time since the set already had it, so the count stayed at
`2`. `Contains` checks membership quickly, without scanning every element the way a `List`'s `Contains` would.

💡 **Pick the right box.** The takeaway that makes everything above click:

- **`List<T>`** - you need *order* and *access by position*; duplicates are fine. The default workhorse.
- **`Dictionary<K,V>`** - you need *fast lookup by a key* rather than by position.
- **`HashSet<T>`** - you need *uniqueness* and fast "is it in here?" checks; order and position don't matter.
- **`T[]` (array)** - the size is genuinely fixed and known up front.

Underneath all four sits **`IEnumerable<T>`** - every one is iterable, why a single `foreach` works on all
of them. That same interface is the foundation of **LINQ**, C#'s query toolkit, where you'll filter and
transform any collection with the same handful of operators. More in [Phase 12](12-linq.md).

## Recap

1. An **array** (`int[]`) is fixed-size and indexed from `0`; ask its size with `.Length`. Use it only when
   the count is genuinely fixed.
2. **Generics** (`List<T>`, the `<T>`) make collections **type-safe** - a `List<int>` holds only ints, checked
   at compile time. Avoid the old non-generic `ArrayList`/`Hashtable`.
3. **`List<T>`** is the growable, ordered, indexable workhorse: `Add`, `[i]`, `.Count`, and `foreach`. It's
   what you reach for most.
4. **`Dictionary<K,V>`** maps keys to values for instant lookup; ⚠️ the indexer throws on a missing key, so
   use **`TryGetValue`** or **`ContainsKey`** when a key might be absent. Iterate it as `KeyValuePair<K,V>`.
5. **`HashSet<T>`** keeps elements unique and tests membership fast; `Add` returns `false` when the value was
   already present.
6. Pick by *how you read data back out* - position (`List`), key (`Dictionary`), uniqueness (`HashSet`),
   fixed count (array) - and remember all of them are `IEnumerable<T>`, the foundation LINQ builds on.

Next, we put these collections to work: **control flow and methods** - the `if`, `switch`, and loops that
make decisions, and how to package logic into reusable methods.

## Quick check

Test yourself on the choices that matter - which box to pick, and the dictionary trap:

```quiz
[
  {
    "q": "You need to store users keyed by their unique ID and look one up instantly by that ID. Which collection fits best?",
    "choices": [
      "An array (T[])",
      "A List<T>",
      "A Dictionary<K,V>",
      "A HashSet<T>"
    ],
    "answer": 2,
    "explain": "A Dictionary<K,V> maps keys to values and looks up a value instantly by its key. A List would force you to scan every element to find the matching ID; a dictionary jumps straight to it."
  },
  {
    "q": "What happens when you read `scores[\"dave\"]` from a Dictionary<string,int> and the key \"dave\" doesn't exist?",
    "choices": [
      "It returns 0, the default for int",
      "It returns null",
      "It throws a KeyNotFoundException and stops the program",
      "It silently adds \"dave\" with value 0"
    ],
    "answer": 2,
    "explain": "The dictionary indexer throws KeyNotFoundException on a missing key - it does not return a default or add the key. Use TryGetValue or ContainsKey when a key might be absent."
  },
  {
    "q": "Why prefer the generic `List<T>` over the old non-generic `ArrayList`?",
    "choices": [
      "List<T> is type-safe - the compiler guarantees it holds only one type, catching mistakes before runtime",
      "ArrayList cannot grow, but List<T> can",
      "List<T> is the only one you can foreach over",
      "There is no real difference; they are interchangeable"
    ],
    "answer": 0,
    "explain": "ArrayList holds object (anything), so type errors surface only at runtime and value types get boxed. List<T> is type-safe: a List<int> holds only ints, enforced by the compiler. Both can grow and both are iterable."
  }
]
```


---

# Control Flow & Methods - Decisions, Loops & Reusable Logic

So far your programs run top to bottom. Real programs do three things straight-line code can't: **decide** (do this, but only if that's true), **repeat** (do this for every item), and **organize** logic into named, reusable pieces you call instead of copy-pasting. This phase is all three.

The mental model: control flow is about *choosing which lines run*, methods about *giving a chunk of lines a name so you can run it from anywhere*. C# has a slightly larger toolbox here than some languages - two flavors of `switch`, four loop keywords - each fitting a specific shape of problem. We'll focus on *when* to reach for each, not just *how*.

## `if` / `else` - the basic decision

The `if` statement runs a block only when a condition is true - any expression evaluating to a `bool` (a `true`/`false` value from [Phase 2](02-syntax-values-and-types.md)).

```csharp
int age = 20;

if (age >= 18)
{
    Console.WriteLine("adult");
}
else
{
    Console.WriteLine("minor");
}
```
```console
adult
```
*What just happened:* `age >= 18` evaluated to `true`, so the first block ran and printed `adult`; a `false` would have run `else` instead. C# *requires* parentheses around the condition. The braces `{ }`, optional for a single statement, are worth always keeping - they prevent "I added a second line and it silently ran every time" bugs.

For chains of conditions, stack `else if`:

```csharp
int score = 73;

if (score >= 90)
{
    Console.WriteLine("A");
}
else if (score >= 80)
{
    Console.WriteLine("B");
}
else if (score >= 70)
{
    Console.WriteLine("C");
}
else
{
    Console.WriteLine("needs work");
}
```
```console
C
```
*What just happened:* C# checked each condition top to bottom and ran the **first** true one (`score >= 70`), then skipped the rest. `90` and `80` failed, `70` matched, giving `C`. `else` is the catch-all when nothing matched. Boolean expressions combine with `&&` (and), `||` (or), `!` (not) - `if (age >= 18 && hasTicket)` runs only when *both* are true.

💡 **Key point.** A long `else if` ladder comparing *one variable* against several values is exactly what `switch` was built for.

## `switch` - comparing one value against many

Testing a single value against a list of possibilities makes a tower of `else if` noisy. `switch` flattens it.

### The classic `switch` statement

```csharp
string day = "Sat";

switch (day)
{
    case "Sat":
    case "Sun":
        Console.WriteLine("weekend");
        break;
    case "Fri":
        Console.WriteLine("almost there");
        break;
    default:
        Console.WriteLine("weekday");
        break;
}
```
```console
weekend
```
*What just happened:* `switch (day)` compared `day` against each `case` label, matched `"Sat"`, and printed `weekend`. `default` is the catch-all, like the final `else`. Stacking `case "Sat":` and `case "Sun":` with no code between them means "either matches the same block" - that's how you group values.

⚠️ **Gotcha (the good kind) - C# forbids implicit fall-through.** Notice every case ends in `break`. In C and older Java/JavaScript, forgetting that `break` lets execution silently "fall through" into the *next* case - a notorious bug source. **C# won't compile a non-empty case that doesn't explicitly end** (`break`, `return`, etc.), so that bug cannot happen. (Grouping empty cases like `case "Sat": case "Sun":` is still allowed - shared labels, not fall-through.)

### The modern `switch` *expression*

The statement above *does* something (prints). Often you want to *produce a value* from the input instead - the **switch expression** (C# 8+) does that, more compactly:

```csharp
string day = "Sat";

string kind = day switch
{
    "Sat" or "Sun" => "weekend",
    "Fri"          => "almost there",
    _              => "weekday",
};

Console.WriteLine(kind);
```
```console
weekend
```
*What just happened:* This is a `switch` written as an **expression** - it evaluates to a value, stored here in `kind`. The shape is `value switch { pattern => result, ... }`; each arm uses `=>` ("goes to") to map a pattern to a result. `"Sat" or "Sun"` matches either; `_` (discard) is the catch-all. No `break`, no `case`/`:` ceremony - the whole construct *is* one value.

📝 **Statement vs. expression.** A *statement* performs an action (no value); an *expression* evaluates to a value you can assign, return, or pass along. The classic `switch` is a statement; `x switch { ... }` is an expression - reach for it when every branch's job is "produce this value."

That `"Sat" or "Sun"` syntax is a taste of **pattern matching** - switch expressions can also match on types, ranges, and property values. Deep dive in [Phase 13](13-records-and-modern-csharp.md); for now, matching constant values covers most everyday use.

## Loops - doing something repeatedly

C# has four looping keywords. They overlap, but each has a sweet spot: know *how many times* up front, loop *until a condition changes*, or walk *every item in a collection*?

### `for` - when you're counting

Use `for` when you know the count or need the index - it bundles three parts into one line.

```csharp
for (int i = 0; i < 3; i++)
{
    Console.WriteLine(i);
}
```
```console
0
1
2
```
*What just happened:* The `for` header has three semicolon-separated parts: **init** (`int i = 0`, runs once), **condition** (`i < 3`, checked before each pass), and **post** (`i++`, "add one to `i`", runs after each pass). It printed `0`, `1`, `2` and stopped once `i` reached `3`. `i` exists only inside the loop.

### `while` - when you loop until something changes

Use `while` when repetitions depend on a condition, not a count. The condition is checked *before* each pass, so the body might run zero times.

```csharp
int n = 3;

while (n > 0)
{
    Console.WriteLine(n);
    n--;
}
```
```console
3
2
1
```
*What just happened:* `while (n > 0)` checked the condition first, ran the body while it held, and stopped when `n` hit `0`. (`n--` subtracts one from `n`.) Had `n` started at `0`, the body would never run - the check happens up front. ⚠️ Make sure something inside the loop changes the condition, or you've written an infinite loop.

### `do-while` - when you must run at least once

`do-while` is `while`'s twin, but checks the condition *after* the body, so it always runs at least once - the right tool for "prompt the user, then re-prompt if invalid."

```csharp
int countdown = 0;

do
{
    Console.WriteLine($"value is {countdown}");
    countdown--;
}
while (countdown > 0);
```
```console
value is 0
```
*What just happened:* Even though `countdown > 0` was already false, the body ran **once** before the check happened - the whole point of `do-while`. After printing, the condition tested false and the loop ended. A `while` loop here would have printed nothing.

### `foreach` - the workhorse for collections

Most real loops walk every item in a collection. `foreach` does that directly - no index bookkeeping, no off-by-one risk. This is how you iterate the collections from [Phase 3](03-collections.md).

```csharp
string[] names = { "Ada", "Linus", "Grace" };

foreach (string name in names)
{
    Console.WriteLine($"Hello, {name}!");
}
```
```console
Hello, Ada!
Hello, Linus!
Hello, Grace!
```
*What just happened:* `foreach (string name in names)` handed us each element in turn, binding it to `name`. No counter, no `names[i]`, no running past the end - `foreach` knows when the collection is exhausted and stops. This is the loop you'll write most often.

💡 **Key point - which loop when?** `for` for index or known count, `while` until a condition flips, `do-while` when the body must run at least once, `foreach` for "do this to every item" - most of the time. When in doubt over a collection, reach for `foreach` first.

## Methods - naming reusable logic

📝 **Method** - a named, reusable block of code that takes inputs (**parameters**) and optionally hands back an output (**return value**), so you can call it from anywhere instead of copying it. (In C#, all code lives inside methods, which live inside classes - see [Phase 5](05-classes-and-objects.md).)

Here's a method that adds two numbers:

```csharp
static int Add(int a, int b)
{
    return a + b;
}

Console.WriteLine(Add(3, 4));
```
```console
7
```
*What just happened:* The signature `static int Add(int a, int b)` reads piece by piece: `static` (more in a second), `int` is the **return type**, `Add` is the name, `(int a, int b)` are two `int` **parameters**. `return a + b` computes the sum and hands it back. The call `Add(3, 4)` passed `3` and `4` as **arguments**, got `7` back, and printed it. A method returning nothing uses `void`.

**Expression-bodied members.** When a method is just a single expression, `=> ` (same arrow as the switch expression) trims the braces and `return`:

```csharp
static int Add(int a, int b) => a + b;
static int Square(int x) => x * x;

Console.WriteLine(Square(5));
```
```console
25
```
*What just happened:* `=> a + b` means "this method returns `a + b`" - exactly equivalent to `{ return a + b; }`, just shorter. Use it for one-liners; keep braces for anything multi-step. This `=>`, the one in switch expressions, and lambdas (later) all share the "goes to / produces" meaning.

**`static` vs. instance - just enough for now.** A `static` method belongs to the class itself, called without creating an object (`Add(3, 4)`). An **instance** method belongs to a specific object, called through it (`myList.Add(x)`). Your entry point is `static void Main(...)` because the runtime calls it *before any object exists*. Full story in [Phase 5](05-classes-and-objects.md); for now: `static` = "call it on the type, no object needed."

## Parameters: optional, named, and `ref`/`out`

Plain parameters are just the start. C# offers several ways to make calls clearer and more flexible.

**Optional parameters** have a default value, so callers can skip them:

```csharp
static string Greet(string name, string greeting = "Hello")
{
    return $"{greeting}, {name}!";
}

Console.WriteLine(Greet("Ada"));                       // uses the default
Console.WriteLine(Greet("Linus", "Welcome"));          // overrides it
Console.WriteLine(Greet("Grace", greeting: "Hi"));     // named argument
```
```console
Hello, Ada!
Welcome, Linus!
Hi, Grace!
```
*What just happened:* `greeting = "Hello"` makes that parameter **optional** - call `Greet("Ada")` and it fills in `"Hello"`. The third call uses a **named argument** (`greeting: "Hi"`), labeling the argument by its parameter name - self-documenting, and lets you skip optional parameters you don't care about. Optional parameters must come *after* all required ones.

**`out` parameters - the "try" pattern you'll meet immediately.** Sometimes a method needs to hand back *more than one thing*: a result *and* whether it succeeded. `out` lets a parameter carry a value *out*, in addition to the return value - as in `int.TryParse`, which safely converts text to a number:

```csharp
string input = "42";

if (int.TryParse(input, out int number))
{
    Console.WriteLine($"Parsed: {number + 1}");
}
else
{
    Console.WriteLine("Not a valid number");
}
```
```console
Parsed: 43
```
*What just happened:* `int.TryParse` returns a `bool` (did it work?) *and* writes the parsed value into the `out` parameter. `out int number` declares `number` right inside the call; if parsing succeeds, `TryParse` fills it in and returns `true`. If `input` were `"banana"`, it would return `false` (no crash) and we'd hit `else`. This `bool` + `out` shape - also used by `Dictionary.TryGetValue` - is *the* idiomatic C# way to do "give me the value if it exists, but don't blow up if it doesn't."

📝 **`out` vs. `ref`.** `out` means "the method *will* assign this - its incoming value is ignored." `ref` means "the method can *read and modify* this existing variable in place." Both pass the variable itself, so changes are visible to the caller. `out` (the `Try` pattern) is common; `ref` is rarer, for methods that need to both see and update a caller's variable.

**Overloading - same name, different parameters.** Several methods can share a name as long as their parameter lists differ. C# picks the right one at compile time based on the arguments:

```csharp
static int Multiply(int a, int b) => a * b;
static double Multiply(double a, double b) => a * b;
static int Multiply(int a, int b, int c) => a * b * c;

Console.WriteLine(Multiply(3, 4));         // matches (int, int)
Console.WriteLine(Multiply(2.5, 2.0));     // matches (double, double)
Console.WriteLine(Multiply(2, 3, 4));      // matches (int, int, int)
```
```console
12
5
24
```
*What just happened:* Three methods named `Multiply`, distinguished by parameters - **overloading**. The compiler matched each call to the overload whose parameter types fit: `(3, 4)` to `(int, int)`, `(2.5, 2.0)` to `double`, the three-argument call to its own overload. This is **compile-time resolution**, based on argument types - why `Console.WriteLine` accepts a string, an int, a bool, and more: one name, many overloads.

## Recap

1. **`if` / `else`** runs a block based on a `bool` condition; `else if` chains check top to bottom and run the first match. Combine conditions with `&&`, `||`, `!`.
2. **`switch`** compares one value against many. The classic statement needs `break` on every case - ⚠️ C# *forbids* implicit fall-through, killing a classic bug. The modern **switch expression** (`x switch { v => result, _ => ... }`) returns a value, no `break` needed.
3. **Four loops:** `for` (counting / index), `while` (loop until a condition flips, checked first), `do-while` (runs at least once, checked after), and `foreach` (every item in a collection - the everyday workhorse).
4. **Methods** name reusable logic: `static returnType Name(params)`. Expression-bodied `=> ...` is shorthand for a one-line body; `static` means "call on the type, no object needed."
5. **Parameters** can be **optional** (`x = 0`), passed by **name**, or marked **`out`/`ref`** to pass values back. The `bool` + `out` `Try...` pattern (`int.TryParse`) is everywhere in C#.
6. **Overloading** lets several methods share a name with different parameters; the compiler resolves which to call at compile time from the argument types.

You can now make decisions, repeat work, and bundle logic into named, callable pieces. Next, we put methods and data together into **classes and objects** - the heart of how C# programs are structured.

## Quick check

Test yourself on the ideas most likely to trip you up - fall-through, the switch expression, and the `out` pattern:

```quiz
[
  {
    "q": "In a classic C# `switch` statement, what happens if you write a non-empty `case` block without a `break` (or other terminator)?",
    "choices": [
      "The code won't compile - C# forbids implicit fall-through",
      "Execution silently falls through into the next case, like in C",
      "Only the matching case runs, and the rest are skipped automatically",
      "It compiles but throws an exception at runtime"
    ],
    "answer": 0,
    "explain": "C# requires every non-empty case to end explicitly (with break, return, etc.). It will not compile a case that would fall through, which eliminates the classic 'forgot the break' bug found in C and older Java/JavaScript."
  },
  {
    "q": "What's the key difference between the classic `switch` statement and a switch *expression* (`x switch { ... }`)?",
    "choices": [
      "The switch expression evaluates to a value you can assign or return; the statement performs an action and has no value",
      "The switch expression is slower because it checks every arm",
      "The switch statement can match patterns but the expression cannot",
      "There is no difference - they are just two spellings of the same thing"
    ],
    "answer": 0,
    "explain": "A statement does something (it has no value); an expression produces a value. `x switch { v => result, _ => ... }` evaluates to a result you store, return, or pass along, while the classic `switch` runs side-effecting code."
  },
  {
    "q": "Why does `int.TryParse(\"42\", out int number)` use an `out` parameter instead of just returning the parsed number?",
    "choices": [
      "So it can return a bool for success/failure AND hand back the parsed value through the out parameter - without crashing on bad input",
      "Because out parameters are always faster than return values",
      "Because methods in C# can only return bool, never int",
      "To force the caller to create the variable before calling the method"
    ],
    "answer": 0,
    "explain": "TryParse needs to communicate two things: whether parsing succeeded (the bool return) and the value itself (the out parameter). This lets you safely attempt a conversion and check the result without an exception when the input isn't a valid number."
  }
]
```


---

# Classes & Objects - The Spine of C#

You've been writing methods inside a class since Phase 1, even if nobody made a fuss about it. Every C# program lives inside a class - there's no "outside." Classes and objects aren't a corner of C#; they *are* C#. Almost everything you'll touch - a list, a file, an HTTP client, a button - is an object built from a class.

The word "OOP" arrives wrapped in scary vocabulary: encapsulation, polymorphism, abstraction. Set that aside - there's one idea underneath, and once it lands the keywords stop being spells: **a class bundles data together with the behavior that acts on that data, and `new` stamps out individual copies you can use.**

## The mental model: blueprint vs. instance

📝 **Class** - a blueprint describing what *kind* of thing exists: what data it holds and what it can do. **Object** (also called an **instance**) - one concrete thing built from that blueprint with `new`. The class is the cookie cutter; the objects are the cookies.

A `class Account` doesn't *hold* any money - it's the *idea* of an account. Write `new Account(...)` and you get a real account, with its own balance, separate from every other account you make.

> 💡 **Key point.** The class is written once, at design time. Objects are created over and over, at run time. One blueprint, many independent instances - each remembering its own state.

## Fields, constructors, and `this`

A class needs somewhere to keep its data (**fields**), a way to set that data up when an object is born (a **constructor**), and a word for "the specific object I'm working on" (`this`). Let's build `Account` for real.

📝 **Field** - a variable that lives on the object, holding its data. **Constructor** - a special method that runs automatically when you write `new`, whose job is to set up the new object's starting data. It has no return type and shares the class's name. **`this`** - a reference to the current object, the one a method was called on.

```csharp
class Account
{
    private decimal balance;   // a field: data stored on each Account
    public string Owner;       // another field

    // Constructor: runs when you write new Account(...)
    public Account(string owner, decimal opening)
    {
        this.Owner = owner;        // this.Owner = the field; owner = the parameter
        this.balance = opening;
    }

    public void Deposit(decimal amount)
    {
        this.balance += amount;    // change THIS account's balance
    }

    public decimal Balance()
    {
        return this.balance;
    }
}

class Program
{
    static void Main()
    {
        Account ada = new Account("Ada", 100m);   // build one instance
        ada.Deposit(50m);
        Console.WriteLine($"{ada.Owner}: {ada.Balance()}");
    }
}
```
```console
Ada: 150
```
*What just happened:* `new Account("Ada", 100m)` ran the constructor, copying the two arguments into the new object's fields. `ada` is now one instance carrying its own `Owner` and `balance`. `ada.Deposit(50m)` changed *Ada's* balance specifically - `this.balance` means "the balance of the account `Deposit` was called on." A second account's balance would be untouched - that separateness is the whole point of instances.

Notice `this.Owner = owner`: the field is `Owner`, the parameter is `owner` - same word, different case. `this.` makes it unambiguous: left side is the object's field, right the parameter. You'll see this constantly in constructors.

💡 **Object initializer syntax.** C# gives a shortcut for setting public members right after construction - instead of (or alongside) constructor arguments, write the assignments in braces:

```csharp
Account ada = new Account("Ada", 0m)
{
    Owner = "Ada Lovelace"   // set a public member inline, after the constructor runs
};
```
*What just happened:* The constructor ran first (with `"Ada"` and `0m`), then the braces reassigned `Owner` to `"Ada Lovelace"`. Object initializers are pure convenience - they run *after* the constructor and can only touch members the caller is allowed to set. They read nicely for an object with several settable fields.

## Properties - C#'s signature feature

Here's where C# parts ways with most languages you've seen. You might make a field `public` and let callers read and write it directly. That works, but it's a trap: the day you need to *validate* a value, compute it, or log when it changes, a raw public field gives you no hook - you'd have to change every caller.

Java solved this with `getName()` / `setName(...)` methods everywhere - verbose, and the *caller* has to know whether it's reading a field or calling a getter. C# solved it differently with one of its defining features: the **property**.

📝 **Property** - a member that *looks* like a field from the outside (`account.Owner`) but *runs code* underneath. It has a `get` accessor (runs on read) and/or a `set` accessor (runs on assignment, with the value arriving as the keyword `value`). Callers can't tell a property from a field - that's the point.

The simplest form is an **auto-property**, where the compiler generates the hidden backing field:

```csharp
class Account
{
    public string Owner { get; set; }      // auto-property: read and write
    public string Bank  { get; init; }     // init-only: settable once, at creation
    public decimal Balance { get; private set; }  // public read, private write

    public Account(string owner)
    {
        Owner = owner;
        Balance = 0m;
    }
}
```
*What just happened:* `Owner { get; set; }` is a full read/write property in one line - the compiler wired both accessors to an invisible backing field. `Bank { get; init; }` is **init-only**: settable in the constructor or an object initializer, then frozen. `Balance { get; private set; }` lets *anyone* read the balance but only code *inside this class* change it - callers can't set it to a million.

When you need logic, spell the accessors out - here's a guarded setter that refuses bad data:

```csharp
class Account
{
    private decimal balance;   // the backing field, hidden

    public decimal Balance
    {
        get { return balance; }
        set
        {
            if (value < 0)
                throw new ArgumentException("Balance cannot be negative.");
            balance = value;    // 'value' is the incoming assigned amount
        }
    }
}

class Program
{
    static void Main()
    {
        var acc = new Account();
        acc.Balance = 200m;            // calls the set accessor; value = 200
        Console.WriteLine(acc.Balance); // calls the get accessor
        acc.Balance = -5m;             // set accessor throws
    }
}
```
```console
200
Unhandled exception. System.ArgumentException: Balance cannot be negative.
```
*What just happened:* `acc.Balance = 200m` looked like a plain assignment but ran the `set` block, `value` holding `200`. Since `200` passed the check, it landed in the hidden `balance` field. Reading `acc.Balance` ran the `get` block. The illegal `acc.Balance = -5m` failed the `value < 0` check and threw - the property *guarded its own data*, with no change to the caller's code.

You can also make a **computed, read-only property** - only a `get` calculating its value from other state:

```csharp
public string Summary => $"{Owner} has {Balance:C}";   // expression-bodied, get-only
```
*What just happened:* `Summary` has no backing field. Every read runs the `=>` expression and builds a fresh string from `Owner` and `Balance`. It's read-only because there's nothing to assign to - this is how you expose *derived* information without storing it.

💡 **Why this matters.** Properties are the idiomatic C# way to expose an object's state - not raw public fields, not Java-style `getX`/`setX`. They give a field's clean syntax *and* a method's power to validate, compute, and control access. Your reflex should be `public string Name { get; set; }`, not `public string Name;`. Reach for the full accessor form only when you need logic.

## Encapsulation & access modifiers

The guarded setter above hinted at the bigger idea: **encapsulation** - keeping an object's internal data private and exposing only a controlled surface. The reason isn't tidiness; *uncontrolled* state is where bugs breed. If any code anywhere can set `balance` to anything, any code anywhere can corrupt it. Make `balance` private with a guarded property, and there's exactly *one* place a bad balance can come from.

C# controls visibility with **access modifiers** - the four you'll use constantly:

📝 **`public`** - visible everywhere. **`private`** - visible only inside the same class (the default for class members if you write nothing). **`protected`** - visible inside this class and its subclasses (matters once you hit inheritance, Phase 6). **`internal`** - visible anywhere in the same project/assembly, but not to outside code that references your library.

The discipline: **make state `private`, expose it through `public` properties and methods.**

```csharp
class Thermostat
{
    private double celsius;   // private state - nobody touches this directly

    public double Celsius
    {
        get => celsius;
        set
        {
            if (value < -273.15)
                throw new ArgumentException("Below absolute zero.");
            celsius = value;
        }
    }

    // A public, read-only view derived from the private state
    public double Fahrenheit => celsius * 9 / 5 + 32;
}

class Program
{
    static void Main()
    {
        var t = new Thermostat();
        t.Celsius = 20;
        Console.WriteLine($"{t.Celsius}C = {t.Fahrenheit}F");
    }
}
```
```console
20C = 68F
```
*What just happened:* `celsius` is private, so no outside code can poke an impossible temperature into it. The only way in is `Celsius`'s setter, which rejects anything below absolute zero. `Fahrenheit` is a read-only computed view of that private value. The class decides exactly what the world can do to it - set a valid Celsius, read either scale, nothing else. That controlled surface is encapsulation, keeping an object trustworthy as your program grows.

⚠️ **Don't reflexively make everything `public`.** A class with all-public fields is a bag of variables with no defenses - it can't stop bad data or change its internals later without breaking callers. Start private, open up only what callers genuinely need.

## `static` members, and overriding `ToString`

So far every field and method has belonged to an *instance* - each `Account` has its own `balance`. But sometimes a member belongs to the *class itself*. That's what `static` means.

📝 **Instance member** - belongs to each object; you need one to use it (`ada.Deposit(...)`). **`static` member** - belongs to the class as a whole, used through the class name (`Account.Count`), existing even with zero objects made.

This is why your entry point is `static void Main()`: when the program starts, *no objects exist yet*, so `Main` can't belong to an instance - it has to belong to the class itself.

```csharp
class Account
{
    public static int Count = 0;   // shared across ALL accounts
    public string Owner { get; }

    public Account(string owner)
    {
        Owner = owner;
        Count++;                   // bump the shared counter on every new account
    }
}

class Program
{
    static void Main()
    {
        new Account("Ada");
        new Account("Grace");
        Console.WriteLine(Account.Count);   // through the class, not an instance
    }
}
```
```console
2
```
*What just happened:* `Count` is `static`, so there's exactly *one* shared by every `Account`. Each constructor incremented that shared counter, so after two accounts `Account.Count` is `2`. You read it through the class name, never through an object - static members don't belong to any instance.

The other thing nearly every class should do is teach itself how to print. By default, printing an object gives its type name - useless. **Overriding `ToString`** fixes that. Every C# object inherits a `ToString()` method (from the universal base type `object`, Phase 6), which you can replace with something meaningful:

```csharp
class Account
{
    public string Owner { get; }
    public decimal Balance { get; }

    public Account(string owner, decimal balance)
    {
        Owner = owner;
        Balance = balance;
    }

    public override string ToString()      // replace the default, useless version
    {
        return $"Account({Owner}: {Balance:C})";
    }
}

class Program
{
    static void Main()
    {
        var ada = new Account("Ada", 150m);
        Console.WriteLine(ada);            // Console.WriteLine calls ToString() for you
    }
}
```
```console
Account(Ada: $150.00)
```
*What just happened:* `Console.WriteLine(ada)` automatically called `ada.ToString()`. Because we **overrode** it (`override` tells the compiler we're deliberately replacing the inherited version), it returned our friendly string instead of the default `"Account"`. Overriding `ToString` makes objects readable in logs, debuggers, and quick `Console.WriteLine`s.

⚠️ **One gotcha to bank for later: `==` doesn't mean what you'd expect for classes.** For a class (a *reference type*), `==` and `.Equals()` compare *whether two variables point at the same object in memory* - not whether they hold the same data. Two separate `Account` objects with identical owner and balance are `==` only if literally the same object:

```csharp
var a = new Account("Ada", 150m);
var b = new Account("Ada", 150m);
Console.WriteLine(a == b);   // False - two different objects, despite identical data
Console.WriteLine(a == a);   // True  - same object
```
```console
False
True
```
*What just happened:* `a` and `b` describe the same account on paper, but they're two distinct objects at different spots in memory, so `==` says `False`. This is *reference equality*, the default for all classes (a few built-in types like `string` override it to compare contents, which is why `string ==` behaved the way you'd expect back in Phase 2, but your own classes get reference equality by default) - it trips people constantly. **Records** flip this to compare by *value* automatically (Phase 13); the full gotcha, including how to override equality yourself, is in Phase 9. For now: `==` on class objects asks "same object?", not "same contents?"

## Recap

1. A **class** is a blueprint written once; an **object/instance** is one concrete thing built from it with `new`, carrying its own data. One blueprint, many independent instances.
2. **Fields** hold an object's data; the **constructor** sets that data up when `new` runs; **`this`** means "the current object," and disambiguates a field from a same-named parameter.
3. **Properties** are C#'s signature feature: they look like fields but run code (`get`/`set`, with `value` as the incoming assignment). Auto-properties (`{ get; set; }`), `init`-only setters, `private set`, and computed get-only properties are the idiomatic way to expose state - not raw public fields.
4. **Encapsulation** means keeping state `private` and exposing a controlled surface via properties/methods; **access modifiers** (`public`, `private`, `protected`, `internal`) set visibility. Private-by-default prevents whole classes of bugs.
5. **`static`** members belong to the class, not any instance (which is why `Main` is static); override **`ToString`** so your objects print meaningfully.
6. ⚠️ For classes, **`==` compares object identity, not contents** - two objects with identical data are not equal unless they're the same object. (Records fix this; full story in Phase 9.)

You can now design your own types - the building blocks of every C# program. Next, we connect classes together: how one class can build on another, and how **interfaces** let unrelated classes promise the same behavior.

## Quick check

Test yourself on the ideas that define C# classes - properties and reference equality:

```quiz
[
  {
    "q": "What makes a C# property different from a plain public field?",
    "choices": [
      "A property looks like a field to callers but can run code (validation, computation) in its get/set accessors",
      "A property is always faster than a field because the compiler optimizes it",
      "A property can only be read, never written",
      "There is no real difference; 'property' is just another word for 'field'"
    ],
    "answer": 0,
    "explain": "A property exposes a field-like syntax (account.Owner) while running code underneath. Its get and set accessors let you validate, compute, log, or restrict access - without callers ever knowing it isn't a raw field. That's why C# uses properties instead of public fields or Java-style getX/setX."
  },
  {
    "q": "You write two separate accounts with identical data and compare them with ==. What do you get, and why?",
    "choices": [
      "False - for a class, == compares object identity (same object in memory), not contents",
      "True - == always compares the data inside two objects",
      "A compile error, because you can't use == on classes",
      "True, but only if both objects were created in the same method"
    ],
    "answer": 0,
    "explain": "Classes are reference types, so == (and the default .Equals) ask 'are these the same object?', not 'do they hold the same data?'. Two distinct objects with identical fields are not ==. Records change this to value comparison (Phase 13); the full gotcha is in Phase 9."
  },
  {
    "q": "Why is a program's entry point declared `static void Main()`?",
    "choices": [
      "Because when the program starts no objects exist yet, so Main must belong to the class itself rather than an instance",
      "Because static methods run faster than instance methods",
      "Because Main is not allowed to use the 'this' keyword for security reasons",
      "Because static is required on every method that returns void"
    ],
    "answer": 0,
    "explain": "A static member belongs to the class as a whole, not to any object. At startup there are no instances yet, so Main can't be an instance method - there'd be no object to call it on. It has to belong to the class itself, which is exactly what static means."
  }
]
```


---

# Inheritance & Interfaces - Sharing Behavior

Phase 5 built a single class - fields, constructors, properties, the whole spine of a type. But programs are rarely one type. You get families of related types: a `Dog` and a `Cat` both animals, a `Circle` and a `Rectangle` both shapes, a `FileLogger` and a `ConsoleLogger` that both, well, log. This phase answers: *how do related types share code and promise common behavior without copy-pasting?*

C# gives you two tools, and the phase is really about telling them apart. **Inheritance** is "this type *is a* more specific version of that type" - a `Dog` is an `Animal`, so it gets everything an `Animal` has, plus its own twist. **Interfaces** are "this type *can do* a certain thing" - anything drawable promises a `Draw()` method, no matter what it actually is. Inheritance shares *implementation*; interfaces share *a contract*.

## Inheritance - the "is-a" relationship

📝 **Inheritance** - defining a new class (the **derived**/**child** class) that automatically gets all the public and protected members of an existing class (the **base**/**parent** class), then adds or changes things. Written with a colon: `class Dog : Animal`. The derived class *is a* kind of the base class.

The litmus test is the **"is-a" test**: read it aloud as "a Dog is an Animal." If true and stays true, inheritance fits. "A Car *has an* Engine" is *has-a* - a field, not inheritance. Getting this wrong is the single most common OOP mistake.

When the base class has a constructor that needs arguments, the derived class passes them up with the `base(...)` keyword:

```csharp
class Animal
{
    public string Name { get; }

    public Animal(string name)
    {
        Name = name;
    }

    public void Eat()
    {
        Console.WriteLine($"{Name} is eating.");
    }
}

class Dog : Animal          // Dog IS AN Animal
{
    public Dog(string name)
        : base(name)        // pass 'name' up to Animal's constructor
    {
    }

    public void Fetch()     // Dog adds its own behavior
    {
        Console.WriteLine($"{Name} fetches the ball.");
    }
}

class Program
{
    static void Main()
    {
        Dog rex = new Dog("Rex");
        rex.Eat();      // inherited from Animal
        rex.Fetch();    // defined on Dog
    }
}
```
```console
Rex is eating.
Rex fetches the ball.
```
*What just happened:* `Dog : Animal` means `Dog` inherited `Animal`'s `Name` property and `Eat()` method for free - `rex.Eat()` works even though `Dog` never defines `Eat`. `: base(name)` handed `"Rex"` up to `Animal`'s constructor so the inherited `Name` got set; without it, the compiler wouldn't know how to initialize the `Animal` part of the `Dog`. Then `Dog` added `Fetch()`, which `Animal` knows nothing about - get everything the parent has, then extend it.

💡 **Every class already inherits something.** Even a plain `class Account` with no colon silently inherits from a universal base type called `object` - why every object has a `ToString()` you can override (Phase 5). Inheritance is already underneath everything.

## `virtual` and `override` - and the trap that bites Java refugees

Here's a rule that surprises people from other languages. Inheriting a method is one thing; *replacing* it with a more specific version is another - and C# makes you ask for that explicitly on **both** sides.

📝 **`virtual`** - a base-class method keyword saying "derived classes may replace this." **`override`** - a derived-class keyword saying "I am deliberately replacing the virtual base version." You need *both*: `virtual` opens the door, `override` walks through it.

⚠️ **This is the C#-specific rule that trips everyone.** In Java, *every* method is overridable by default - matching the signature is enough. In C#, methods are **sealed shut by default**: if the base method isn't `virtual`, the derived class cannot truly override it. A deliberate design choice - C# wants overriding intentional, not accidental.

```csharp
class Animal
{
    public string Name { get; }
    public Animal(string name) => Name = name;

    public virtual void Speak()      // 'virtual' opens this up for overriding
    {
        Console.WriteLine($"{Name} makes a sound.");
    }
}

class Dog : Animal
{
    public Dog(string name) : base(name) { }

    public override void Speak()     // 'override' replaces the base version
    {
        Console.WriteLine($"{Name} barks: Woof!");
    }
}

class Cat : Animal
{
    public Cat(string name) : base(name) { }

    public override void Speak()
    {
        Console.WriteLine($"{Name} meows: Meow!");
    }
}

class Program
{
    static void Main()
    {
        Animal a = new Dog("Rex");    // an Animal-typed variable holding a Dog
        a.Speak();                     // which Speak runs?

        Animal b = new Cat("Whiskers");
        b.Speak();
    }
}
```
```console
Rex barks: Woof!
Whiskers meows: Meow!
```
*What just happened:* This is **dynamic dispatch** - the payoff for the `virtual`/`override` ceremony. `a` is *declared* as `Animal`, but at run time it holds a `Dog`. Calling `a.Speak()`, C# looks at the *real* object, not the declared type, and runs `Dog`'s overridden `Speak`. Same for `b` and its `Cat`. The decision is made at run time based on the actual object - hence "dynamic."

⚠️ **The `new` keyword is a trap, not a fix.** Forget `virtual` on the base method and write `override` in the derived class, and the compiler errors - the good case. But write `new` instead of `override`, and the code compiles and looks like it works while doing something subtly wrong: `new` *hides* the base method rather than overriding it. The difference only shows up through a base-typed variable:

```csharp
class Animal
{
    public void Speak()              // NOT virtual
    {
        Console.WriteLine("Animal sound");
    }
}

class Dog : Animal
{
    public new void Speak()          // 'new' HIDES, it does not override
    {
        Console.WriteLine("Woof!");
    }
}

class Program
{
    static void Main()
    {
        Dog d = new Dog();
        Animal a = d;                // same object, two different declared types

        d.Speak();                   // uses Dog's version
        a.Speak();                   // uses Animal's version - surprise!
    }
}
```
```console
Woof!
Animal sound
```
*What just happened:* `d` and `a` point at the *exact same object*, yet print different things. With `new`, C# picks the method based on the **declared type of the variable**, not the real object - the opposite of dynamic dispatch, almost never what you want. When you mean to override, use `virtual` + `override` and verify through a base-typed variable. Seeing `new` on a method is a red flag.

## Polymorphism - one type, many behaviors

You just saw the mechanism - now the name.

📝 **Polymorphism** ("many forms") - treating different derived objects uniformly through a base-class (or interface) reference, each running *its own* overridden behavior at run time. A variable of type `Animal` can hold a `Dog`, `Cat`, or any other animal, and calling `.Speak()` does the right thing - *without your code knowing or caring which*.

The payoff shows up with a *collection* of mixed subtypes: one loop against the base type, each object bringing its own behavior:

```csharp
class Program
{
    static void Main()
    {
        // A list of Animals - but each element is really a Dog or a Cat.
        List<Animal> zoo = new List<Animal>
        {
            new Dog("Rex"),
            new Cat("Whiskers"),
            new Dog("Buddy")
        };

        foreach (Animal a in zoo)   // we only know they're Animals here
        {
            a.Speak();              // each runs ITS OWN Speak
        }
    }
}
```
```console
Rex barks: Woof!
Whiskers meows: Meow!
Buddy barks: Woof!
```
*What just happened:* The loop variable `a` is just an `Animal` as far as the code can tell - `foreach` has no idea it's juggling dogs and cats. But every `a.Speak()` dispatched to the real object's overridden method, so dogs barked and the cat meowed, all from one line of code. Add a `Hamster : Animal` next week, drop it in the list, and **this loop never changes** - code against the base type automatically handles types that didn't exist when you wrote it.

## Interfaces - a contract any class can sign

Inheritance has a hard limit in C#: a class can inherit **exactly one** base class. You can't be both a `Bird` and a `Swimmer` by inheritance. And often "is-a" is the wrong relationship - a `FileLogger` and a `Button` share nothing as *types*, yet both might need to be "savable" or "disposable." For sharing *capability* across unrelated types, you want an **interface**.

📝 **Interface** - a contract: a named list of members (methods, properties) a type promises to provide, with *no implementation*. A class signs it with the same colon syntax (`class Circle : IShape`) and must supply a body for every declared member. By convention interface names start with `I`: `IShape`, `IComparable`, `IDisposable`.

The crucial difference from inheritance: a class inherits **one** base class but implements **many** interfaces. An interface says nothing about *what* a type is - only what it can *do*.

```csharp
interface IShape
{
    double Area();        // just the signature - no body, no fields
    string Describe();
}

class Circle : IShape     // Circle promises to fulfil the IShape contract
{
    private double radius;
    public Circle(double radius) => this.radius = radius;

    public double Area() => Math.PI * radius * radius;
    public string Describe() => $"Circle with area {Area():F2}";
}

class Rectangle : IShape  // an UNRELATED class, same contract
{
    private double w, h;
    public Rectangle(double w, double h) { this.w = w; this.h = h; }

    public double Area() => w * h;
    public string Describe() => $"Rectangle with area {Area():F2}";
}

class Program
{
    static void Main()
    {
        List<IShape> shapes = new List<IShape>
        {
            new Circle(2),
            new Rectangle(3, 4)
        };

        foreach (IShape s in shapes)   // we only know they're IShapes
        {
            Console.WriteLine(s.Describe());
        }
    }
}
```
```console
Circle with area 12.57
Rectangle with area 12.00
```
*What just happened:* `IShape` declared *what* a shape must offer - `Area()` and `Describe()` - but not *how*. `Circle` and `Rectangle` supplied their own implementations, sharing no base class; their only connection is the contract. `List<IShape>` treated them uniformly, exactly like polymorphism with inheritance. Interfaces give the same "one loop, many behaviors" payoff *without* forcing types into a family tree - how unrelated things agree to be interchangeable.

💡 **Default interface methods (a modern wrinkle).** Since C# 8, an interface can include a default body for a member, so types that don't override it inherit that default - handy for extending an interface without breaking existing implementers. Treat it as a special-purpose escape hatch, not the norm; an interface's everyday job is still to declare a contract, not carry code.

## Abstract and sealed - the two ends of the dial

Two more keywords sit at opposite extremes: one forces inheritance, the other forbids it.

📝 **`abstract` class** - a base class you *cannot instantiate directly* (`new Animal(...)` is a compile error); it exists only to be inherited. It can mix concrete members (shared code) with **`abstract` members** - declared but unimplemented - which every concrete subclass is *forced* to override. Use it when there's no such thing as a generic instance ("only dogs, cats, ... never just an `Animal`") but subtypes share real code and state.

```csharp
abstract class Animal
{
    public string Name { get; }
    protected Animal(string name) => Name = name;   // shared setup

    public abstract void Speak();        // no body - subclasses MUST provide one

    public void Sleep()                  // shared concrete behavior
    {
        Console.WriteLine($"{Name} sleeps.");
    }
}

class Dog : Animal
{
    public Dog(string name) : base(name) { }
    public override void Speak() => Console.WriteLine($"{Name}: Woof!");
}

class Program
{
    static void Main()
    {
        // Animal a = new Animal("???");  // compile error: can't instantiate abstract
        Dog d = new Dog("Rex");
        d.Speak();
        d.Sleep();
    }
}
```
```console
Rex: Woof!
Rex sleeps.
```
*What just happened:* `abstract class Animal` can't be `new`-ed - there's no such thing as a generic animal, and the commented-out line proves the compiler enforces that. `Speak()` is `abstract`, so `Animal` declares it but refuses to implement it, *forcing* `Dog` to override it. Meanwhile `Sleep()` is fully written once and shared by every subclass. That blend - force some methods, share others, hold common state - is what separates an abstract class from an interface.

📝 **`sealed` class** - the opposite: `sealed` *cannot be inherited from* at all. `sealed class Receipt` slams the door so no one can subtype it. Reach for it when behavior must not be altered by subclassing - safety, guarantees, or signaling "this is final." You can also seal an individual `override` to stop *further* overriding down the chain.

💡 **So which do you actually pick?** The plain guidance:

- **Default to interfaces.** They model *capability*, a type can implement many, and they avoid a rigid family tree. When types just need to agree on a contract, an interface is almost always right.
- **Use an abstract class when subtypes genuinely share state and code.** If every subclass would copy-paste the same fields and helpers, an abstract base earns its keep by owning that shared implementation once. The cost is the one-base-class limit.

⚠️ **Favor composition and interfaces over deep inheritance hierarchies.** The classic beginner trap is building tall towers - `Animal → Mammal → Carnivore → Dog → Puppy` - then finding a change near the top ripples unpredictably to the bottom, or a new type doesn't fit the tree. Most reuse in real C# codebases comes from *small interfaces* plus *composition* (an object holding other objects, the "has-a" case). Use inheritance shallowly, only when "is-a" is unmistakably true.

## Recap

1. **Inheritance** (`class Dog : Animal`) gives a derived class everything its base has, then lets it add more. Use the **"is-a" test**, and pass base-constructor arguments with **`base(...)`**.
2. **`virtual` + `override`** are both required for true overriding - and ⚠️ this is C#-specific: methods are *not* virtual by default (unlike Java). The `new` keyword *hides* instead of overriding, which is a silent trap; verify behavior through a base-typed variable.
3. **Polymorphism** is the payoff: a base-typed variable (or a `List<Animal>`) runs each object's own overridden method at run time, so one loop handles every subtype - including ones you add later.
4. **Interfaces** declare a contract with no implementation; a class implements **many** interfaces but inherits **one** base class. Name them `IThing`. They share *capability* across unrelated types.
5. **`abstract`** classes can't be instantiated and can force subclasses to implement members (shared code + enforced contract); **`sealed`** forbids inheritance entirely.
6. 💡 Prefer **interfaces** for capability and **composition** over deep hierarchies; reach for an **abstract class** only when subtypes truly share state and code.

You can now connect your types together - let them build on each other, promise common behavior, and be used interchangeably. Next, we deal with the messy real world: what happens when things go wrong (exceptions) and how to read and write data (I/O).

## Quick check

Test yourself on the ideas that separate inheritance from interfaces - especially the override trap:

```quiz
[
  {
    "q": "In C#, what do you need for a derived class to truly override a base-class method (so a base-typed variable runs the derived version)?",
    "choices": [
      "The base method must be marked 'virtual' and the derived method marked 'override'",
      "Nothing special - every C# method is overridable by default, like in Java",
      "The derived method must use the 'new' keyword to replace the base version",
      "Both methods must be marked 'static'"
    ],
    "answer": 0,
    "explain": "C# methods are sealed by default. You need 'virtual' on the base method to allow overriding and 'override' on the derived method to do it. Unlike Java, matching the signature isn't enough - and 'new' only *hides* the method (picking by declared type), which is a common trap, not a real override."
  },
  {
    "q": "What is the key difference between inheriting a base class and implementing an interface in C#?",
    "choices": [
      "A class can implement many interfaces but inherit only one base class; interfaces share a contract while a base class shares implementation",
      "There is no difference - they are two words for the same feature",
      "Interfaces let you inherit from several base classes at once, replacing single inheritance",
      "A base class has no implementation, while an interface always provides full method bodies"
    ],
    "answer": 0,
    "explain": "Inheritance models 'is-a' and shares actual implementation, but C# allows only one base class. Interfaces model 'can-do' - a pure contract (traditionally with no bodies) that any number of unrelated classes can sign, so a class can implement many at once."
  },
  {
    "q": "When should you reach for an abstract class instead of an interface?",
    "choices": [
      "When the subtypes genuinely share state and concrete code, and a generic instance of the base makes no sense",
      "Always - abstract classes are strictly better than interfaces",
      "Whenever you want a type to be usable by code that didn't exist when you wrote it",
      "When you need a single type to inherit from several bases at once"
    ],
    "answer": 0,
    "explain": "An abstract class earns its keep when subtypes share real fields and helper methods (so the base owns that code once) and there's no sensible standalone instance of the base. Otherwise prefer interfaces: they model capability, allow multiple implementation, and avoid the one-base-class limit and brittle deep hierarchies."
  }
]
```


---

# Errors & I/O - Exceptions, Resources & Files

Up to now your programs have run the happy path: the file was there, the number parsed, the index was in range. Real programs spend half their lives off it - disks fill, users type nonsense, networks blink. This phase covers how C# tells you when something went wrong, and how you respond without leaving a mess behind.

If you've read the [Go guide](/guides/go-from-zero), you saw a language where *errors are values* - a function returns a failure alongside its result, checked on the spot. C# made the opposite bet: failure is a **thrown object** that interrupts normal flow and travels *up* the call stack looking for someone willing to handle it. That's the **exception**, and getting comfortable with it - when to catch, when to throw, how to always clean up - is the whole job.

## Exceptions - C#'s error model

**What it actually is.** An exception is an object (an instance of `Exception` or a subclass) *thrown* when something goes wrong. The instant it's thrown, the current method stops dead and the runtime starts **unwinding the stack** - abandoning the current method, then its caller, then *its* caller - until it finds a `catch` block willing to handle it. If nobody catches it, the program crashes and prints a **stack trace** showing the path the exception took.

📝 **Exception** - an object describing a failure, thrown at the point of trouble and caught (or not) up the call chain. **`try`** wraps risky code; **`catch`** handles a failure; **`finally`** runs cleanup either way. **Stack trace** - the printed list of method calls the exception unwound through, newest first.

⚠️ **C# exceptions are all *unchecked*.** Unlike Java's "checked exceptions," the compiler never makes you wrap a call in `try`/`catch`. Any method can throw anything, and you'll only find out at runtime or by reading its docs - freedom and a footgun in one.

The mental model: a thrown exception is a hot potato, rising until someone catches it or it falls out the top and crashes the program.

**A real example.**

```csharp
int[] scores = { 90, 80, 70 };

Console.WriteLine("about to read index 5");
Console.WriteLine(scores[5]);   // there is no index 5
Console.WriteLine("this line never runs");
```
```console
$ dotnet run
about to read index 5
Unhandled exception. System.IndexOutOfRangeException: Index was outside the bounds of the array.
   at Program.<Main>$(String[] args) in /app/Program.cs:line 4
```
*What just happened:* `scores[5]` reached past the end of a three-element array, so the runtime threw an `IndexOutOfRangeException`. The third `WriteLine` never ran - the throw aborted the method immediately. With no `catch` anywhere up the stack, the program printed "Unhandled exception," the type and message, and the stack trace pointing at `line 4` - your best clue for *where* a crash happened.

Now wrap it so the failure is handled instead of fatal:

```csharp
int[] scores = { 90, 80, 70 };

try
{
    Console.WriteLine(scores[5]);
}
catch (IndexOutOfRangeException ex)
{
    Console.WriteLine($"oops: {ex.Message}");
}
finally
{
    Console.WriteLine("cleanup always runs");
}

Console.WriteLine("program continues normally");
```
```console
$ dotnet run
oops: Index was outside the bounds of the array.
cleanup always runs
program continues normally
```
*What just happened:* the `try` block threw, and control jumped straight to the matching `catch`, which read `Message` and kept going. `finally` ran next - it runs whether the `try` succeeded, threw and was caught, or threw something *not* caught here. Execution flowed past the whole block and the program continued. The exception was contained.

## Catching well - be specific, don't swallow

**What it actually is.** A `catch` block can name the exception type it handles. Stack several, and the runtime picks the *first* matching type. The art is catching failures you can actually do something about, and letting everything else keep rising.

💡 **Catch specific types, not `Exception`.** A bare `catch (Exception ex)` grabs *everything*: the file-not-found you expected, but also out-of-memory, a null-reference bug, a typo you'd want to crash loudly. Swallowing it all hides real bugs behind a calm-looking but quietly broken program. Catch the narrowest type you're prepared to handle; if you can't recover, don't catch - let it propagate (or crash, which is the clean outcome).

**A real example.**

```csharp
string input = "not a number";

try
{
    int n = int.Parse(input);
    Console.WriteLine($"parsed {n}");
}
catch (FormatException ex)
{
    Console.WriteLine($"bad input, not a number: {ex.Message}");
}
catch (OverflowException)
{
    Console.WriteLine("that number is too big to fit in an int");
}
```
```console
$ dotnet run
bad input, not a number: The input string 'not a number' was not in a correct format.
```
*What just happened:* `int.Parse` throws a `FormatException` when the text isn't a number and an `OverflowException` when it's too large for `int`. We wrote a separate `catch` for each, and the runtime matched `FormatException` and skipped the other. We did *not* write `catch (Exception)` - an unrelated bug should crash and show us, not get mistaken for "bad input."

**Exception filters with `when`.** Sometimes you want to catch a type only *when* a condition holds - retry on some HTTP failures but not others. `when` adds that condition without forcing you to catch, inspect, and re-throw.

```csharp
try
{
    throw new InvalidOperationException("retryable: server busy");
}
catch (InvalidOperationException ex) when (ex.Message.Contains("retryable"))
{
    Console.WriteLine("caught a retryable error, will try again");
}
```
```console
$ dotnet run
caught a retryable error, will try again
```
*What just happened:* the `when (...)` filter ran *before* the catch body engaged. The message contained "retryable," so the filter returned `true` and this block handled it. Had it *not* matched, this `catch` would have been skipped entirely and the exception kept unwinding, as if the block weren't there. That's the win over catching-then-rethrowing: a declined exception never counts as "handled here," keeping a cleaner stack trace.

## Throwing & custom exceptions

**What it actually is.** You raise an exception yourself with `throw new SomeException("message")`, when your method is asked to do something it *can't* sensibly do. Built-in types cover most cases: `ArgumentException` (a parameter is wrong), `ArgumentNullException` (a required argument was null), `InvalidOperationException` (wrong state for this call).

**A real example - validate and throw.**

```csharp
decimal Withdraw(decimal balance, decimal amount)
{
    if (amount <= 0)
        throw new ArgumentException("amount must be positive", nameof(amount));
    if (amount > balance)
        throw new InvalidOperationException("insufficient funds");

    return balance - amount;
}

Console.WriteLine(Withdraw(100m, 30m));   // fine
Console.WriteLine(Withdraw(100m, 500m));  // throws
```
```console
$ dotnet run
70
Unhandled exception. System.InvalidOperationException: insufficient funds
   at Program.<Withdraw>g__Withdraw|0_0(Decimal balance, Decimal amount)
```
*What just happened:* the first call passed validation and returned `70`. The second asked to withdraw more than the balance - a state this method refuses to handle - so it threw `InvalidOperationException`. `throw` immediately stopped `Withdraw` and handed the failure to the caller. `nameof(amount)` passes the parameter's name to the exception so the message can say *which* argument was bad, without hard-coding a string that would rot if you renamed the parameter.

**A small custom exception.** When the built-ins don't capture your domain, derive your own from `Exception` - convention names it `...Exception` with a constructor that takes a message.

```csharp
public class InsufficientFundsException : Exception
{
    public decimal Shortfall { get; }

    public InsufficientFundsException(decimal shortfall)
        : base($"short by {shortfall:C}")
    {
        Shortfall = shortfall;
    }
}

decimal Withdraw(decimal balance, decimal amount)
{
    if (amount > balance)
        throw new InsufficientFundsException(amount - balance);
    return balance - amount;
}

try
{
    Withdraw(100m, 130m);
}
catch (InsufficientFundsException ex)
{
    Console.WriteLine($"declined - {ex.Message} (shortfall {ex.Shortfall})");
}
```
```console
$ dotnet run
declined - short by $30.00 (shortfall 30)
```
*What just happened:* `InsufficientFundsException` carries a typed `Shortfall` field, so callers can catch it specifically *and* read structured data off it - far better than parsing a string message. `: base(...)` hands a human-readable message up to `Exception`. A caller can write `catch (InsufficientFundsException ex)` to handle exactly this case.

💡 **Throw vs. return a result.** Throw for the *exceptional* - the genuinely unexpected, "I can't do my job" case. For normal, expected outcomes (a lookup that might miss, a parse that might fail on user input), prefer a non-throwing path: `int.TryParse` and `dictionary.TryGetValue` return a `bool` instead, because "the user typed letters" is a Tuesday, not a catastrophe. Exceptions are relatively expensive and interrupt flow.

## `using` / `IDisposable` - deterministic cleanup

**What it actually is.** Some objects hold resources the garbage collector can't tidy up promptly - open files, sockets, database connections. These must be *released* the moment you're done, not "eventually." C#'s answer is **`IDisposable`**: any type holding such a resource implements a `Dispose()` method that releases it. **`using`** guarantees `Dispose()` runs the instant the variable leaves scope - even if an exception is thrown partway through.

📝 **`IDisposable`** - an interface with one method, `Dispose()`, for releasing unmanaged resources. **`using`** - calls `Dispose()` automatically when its scope ends: "open this, and *no matter what happens*, close it on the way out."

**Why this exists.** You *could* do this by hand with `try`/`finally` - open the file, `Close()` in a `finally` block so it runs even on failure. `using` is that pattern compressed into one keyword you can't forget.

**A real example.**

```csharp
using (var writer = new StreamWriter("log.txt"))
{
    writer.WriteLine("first line");
    writer.WriteLine("second line");
}   // writer.Dispose() runs HERE - file flushed and closed automatically

Console.WriteLine("file is closed and saved");
```
*What just happened:* `StreamWriter` implements `IDisposable` because it holds an open file handle. The `using` block opened the file, wrote to it, and called `writer.Dispose()` automatically at the closing brace, flushing buffered text to disk and releasing the handle. Even if `WriteLine` had thrown, `Dispose()` would *still* run, so the file would never stay locked open.

Modern C# offers a tidier form, the **`using` declaration** - no braces, disposal at the end of the *enclosing* scope:

```csharp
void SaveReport()
{
    using var writer = new StreamWriter("report.txt");
    writer.WriteLine("totals: ...");
    // no closing brace block - writer.Dispose() runs when SaveReport() returns
}
```
*What just happened:* `using var writer = ...` is the same guarantee with less nesting - `Dispose()` fires when `writer` falls out of scope at the method's end. Use this form when the resource lives the whole method; use braced `using (...) { }` to release it partway through instead.

## File I/O - reading and writing

**What it actually is.** `System.IO` is C#'s toolbox for files and streams. For "read/write a whole file," the static `File` class has one-call helpers; for a large file piece by piece, `StreamReader` streams it line by line without loading it all into memory.

**A real example - the one-call helpers.**

```csharp
using System.IO;

File.WriteAllText("greeting.txt", "hello\nfrom C#");

string whole = File.ReadAllText("greeting.txt");
Console.WriteLine($"--- whole file ---\n{whole}");

string[] lines = File.ReadAllLines("greeting.txt");
Console.WriteLine($"--- line count: {lines.Length} ---");
foreach (string line in lines)
    Console.WriteLine($"> {line}");
```
```console
$ dotnet run
--- whole file ---
hello
from C#
--- line count: 2 ---
> hello
> from C#
```
*What just happened:* `File.WriteAllText` created (or overwrote) the file and wrote the string in one call - opens, writes, closes for you, no handle to dispose. `File.ReadAllText` slurped the entire contents as one string; `File.ReadAllLines` split on line breaks into a `string[]`. Perfect for small files. ⚠️ The catch is in the name: `ReadAllText` loads the *whole* file into memory, so a multi-gigabyte log needs streaming instead.

**Reading line by line with `StreamReader`.** For a big file, read it as a stream so only one line sits in memory at a time, wrapped in `using` so the handle always closes.

```csharp
using System.IO;

File.WriteAllLines("data.txt", new[] { "alpha", "beta", "gamma" });

using var reader = new StreamReader("data.txt");
string? line;
int count = 0;
while ((line = reader.ReadLine()) != null)
{
    count++;
    Console.WriteLine($"{count}: {line}");
}
```
```console
$ dotnet run
1: alpha
2: beta
3: gamma
```
*What just happened:* `StreamReader.ReadLine()` returns the next line each call, and `null` when nothing's left - the loop's exit condition. Only one line is held at a time, working on files far too large for RAM. `using var` releases the handle when the method ends, even if a read throws. `string? line` marks it possibly-null, since `ReadLine()` returns `null` at end of file (more in [Phase 13](13-records-and-modern-csharp.md)).

⚠️ **`NullReferenceException` - the error you'll hit most.** Calling a method or reading a property on a `null` reference throws `NullReferenceException` ("Object reference not set to an instance of an object") - C#'s single most common runtime crash, from a file read returning `null`, a missed dictionary lookup, an object you forgot to construct. The modern defense is **nullable reference types**, warning the compiler about possible nulls *before* you run - covered in [Phase 13](13-records-and-modern-csharp.md) and [Phase 9](09-idioms-and-gotchas.md). For now: check before you use a value that *could* be null.

## Recap

1. **Exceptions are C#'s error model** - a thrown object aborts the current method and unwinds the stack until a `catch` handles it; uncaught, it crashes with a stack trace. `try`/`catch`/`finally` is the structure, and `finally` always runs.
2. **All C# exceptions are unchecked** - the compiler never forces you to handle one, so knowing what can throw is on you. Catch *specific* types you can recover from; never swallow a bare `catch (Exception)` and hide real bugs. Use `when` filters to catch conditionally.
3. **Throw deliberately** - `throw new ArgumentException(...)` when your method can't do its job; define a custom `: Exception` to carry typed, domain-specific data. Reserve exceptions for the genuinely exceptional; prefer `TryParse`-style methods for expected failures.
4. **`using` / `IDisposable` guarantee cleanup** - `using var f = ...;` calls `Dispose()` at scope end no matter what, so files, connections, and sockets are always released. It's `try`/`finally` for resources, made unforgettable.
5. **File I/O lives in `System.IO`** - `File.ReadAllText`/`WriteAllText`/`ReadAllLines` for whole small files; `StreamReader.ReadLine()` in a `using` for large files read line by line.
6. ⚠️ **`NullReferenceException` is the crash you'll meet most** - calling into a `null` reference. Check for null, and lean on nullable reference types ([Phase 13](13-records-and-modern-csharp.md)) to catch it at compile time.

You can now fail gracefully and clean up after yourself - the difference between a toy and a program people trust. Next we step out of the language and into the toolbox around it: how real C# projects are structured, how to pull in libraries with NuGet, and the commands that build and run it all.

## Quick check

Test yourself on the ideas that matter most - how exceptions flow, and why `using` exists:

```quiz
[
  {
    "q": "What happens the moment an exception is thrown and nothing in the current method catches it?",
    "choices": [
      "The runtime unwinds the stack, abandoning the current method and rising to its caller, looking for a matching catch",
      "The method returns its default value and execution continues normally",
      "The compiler refuses to build the program until you add a try/catch",
      "The exception is silently ignored and the next line runs"
    ],
    "answer": 0,
    "explain": "A thrown exception aborts the current method and travels up the call stack, method by method, until a matching catch handles it - or it falls out the top and crashes the program. The compiler never forces you to handle it: C# exceptions are all unchecked."
  },
  {
    "q": "Why is catching `Exception` (the base type) usually a bad idea?",
    "choices": [
      "It grabs everything - including bugs you'd want to crash loudly - and hides them behind a program that looks fine but is quietly broken",
      "It is slower than catching a specific type",
      "The compiler emits an error when you catch the base Exception type",
      "It only works inside a finally block"
    ],
    "answer": 0,
    "explain": "A bare catch (Exception) swallows the failure you expected AND unrelated bugs like null-reference or out-of-memory. Catch the narrowest type you can actually recover from; let everything else keep propagating."
  },
  {
    "q": "What does a `using` statement guarantee about the object it wraps?",
    "choices": [
      "Its Dispose() method is called when the variable leaves scope, even if an exception is thrown",
      "The object can never be set to null",
      "The object is loaded entirely into memory before use",
      "The object's methods can only be called once"
    ],
    "answer": 0,
    "explain": "using is try/finally for resources: it calls Dispose() automatically at the end of scope no matter how you leave it - normal exit or exception. That is how files, sockets, and connections get released promptly and reliably."
  }
]
```


---

# Projects, NuGet & Tooling - From Files to a Real Solution

So far you've been writing C# the way you learn it: a few `.cs` files, one `Main`, run it, see output. But the moment you build something real - more than a handful of files, someone else's code pulled in, or something you need to *ship* - you step into the **toolchain**: how C# code is grouped, packaged, and turned into something you can hand to a server or teammate.

The mental model worth holding onto: **a C# project is not a folder of files - it's a `.csproj` file that *describes* a folder of files.** That XML file is the source of truth: which .NET version you target, which packages you depend on, what kind of thing you're building. Once you see the `.csproj` as the center of gravity, the `dotnet` commands, NuGet, and the IDEs click into place around it.

## Namespaces & `using` - giving your types an address

📝 **Namespace.** A *named container* for your types - grouping related classes, structs, and enums under one label so full names don't collide. `string` lives in `System`; `JsonSerializer` lives in `System.Text.Json`. Think of it as a postal address: `System.Text.Json.JsonSerializer` is the full address, the namespace everything before the last dot.

Declare one with `namespace`; *import* one with `using` to refer to its types by short names instead of typing the full address every time.

```csharp
namespace MyApp.Billing;   // everything in this file lives here

public class Invoice
{
    public decimal Total { get; set; }
}
```

```csharp
// In another file:
using MyApp.Billing;       // import the namespace...
using System;

Invoice inv = new() { Total = 42.50m };   // ...so we can say "Invoice", not "MyApp.Billing.Invoice"
Console.WriteLine(inv.Total);             // "Console" comes from the "using System;"
```

*What just happened:* The first file put `Invoice` at the address `MyApp.Billing.Invoice`. The second said `using MyApp.Billing;`, telling the compiler "look in there too" - so `Invoice` resolves without the full prefix; otherwise you'd write `MyApp.Billing.Invoice` every time. Same for `using System;`, why `Console` works unqualified. (The `;` after the namespace name is the modern *file-scoped* form; the older `namespace MyApp.Billing { ... }` brace-block style means the same thing.)

💡 **Global & implicit usings - why your files look so bare.** Modern C# projects (.NET 6+) use two conveniences to stop repeating imports. A **global using**, written once, applies to the entire project: `global using System.Text.Json;` in any file means every file can use it. **Implicit usings** - one line in your `.csproj` - auto-add namespaces nearly every program needs (`System`, `System.Collections.Generic`, `System.Linq`). That's why a fresh file can call `Console.WriteLine` with *no* `using`: the project already imported `System` behind the scenes.

## Projects & solutions - the `.csproj` and the `.sln`

A loose pile of `.cs` files isn't a project. A **project** is defined by a `.csproj` file, and that's what `dotnet` actually compiles.

📝 **Project (`.csproj`).** A small XML file that *is* your project's definition: target framework, package dependencies, output type (executable vs. library). The `.cs` files in the same folder tree are pulled in automatically - one `.csproj` = one buildable unit.

📝 **Solution (`.sln`).** A *grouping* of multiple projects so you can open and build them together - a web app, a shared library, a test project, three `.csproj` files tied into one `.sln` your IDE opens. Small programs need no solution at all.

Here's a complete, tiny `.csproj` - notice how little is in it:

```xml
<Project Sdk="Microsoft.NET.Sdk">

  <PropertyGroup>
    <OutputType>Exe</OutputType>
    <TargetFramework>net10.0</TargetFramework>
    <ImplicitUsings>enable</ImplicitUsings>
    <Nullable>enable</Nullable>
  </PropertyGroup>

</Project>
```

*What just happened:* This is the entire project definition for a runnable app. `Sdk="Microsoft.NET.Sdk"` brings in the default build machinery. `<OutputType>Exe</OutputType>` says "build an executable" (drop this and you get a library `.dll`). `<TargetFramework>net10.0</TargetFramework>` pins the .NET version. `<ImplicitUsings>enable</ImplicitUsings>` is the auto-import switch from above; `<Nullable>enable</Nullable>` turns on the nullable-reference warnings from Phase 9. No source file list - every `.cs` under this folder is included by convention.

You don't hand-write these. The `dotnet` CLI scaffolds, builds, and runs them:

```bash
dotnet new console -o HelloApp   # scaffold a new console project in ./HelloApp
cd HelloApp
dotnet run                       # compile and run in one step
dotnet build                     # compile only, produce the binaries
```

*What just happened:* `dotnet new console` generated a folder with a ready-to-go `.csproj` and a starter `Program.cs`. `dotnet run` is your fast inner loop - compiles and immediately runs, like "F5" from the command line. `dotnet build` compiles but stops there. (`dotnet new` has many templates - `classlib`, `web`, `xunit` - each scaffolding the right `.csproj`.)

## NuGet - the package manager you'll lean on constantly

You will not write everything yourself, and shouldn't try. **NuGet** gets you the .NET world's vast catalog of ready-made libraries.

📝 **NuGet.** The .NET package manager. A *package* is a bundle of compiled, reusable code (plus metadata) published to a registry - the public one is **nuget.org**. Add a package and your code can use it. The .NET equivalent of npm or pip.

Adding one is a single command:

```bash
dotnet add package Newtonsoft.Json
```

```console
info : Adding PackageReference for package 'Newtonsoft.Json' into project '...HelloApp.csproj'.
info : Restoring packages for ...HelloApp.csproj...
info : Installed Newtonsoft.Json 13.0.3 from https://api.nuget.org/v3/index.json
```

*What just happened:* `dotnet add package` looked up `Newtonsoft.Json` on nuget.org, picked the latest stable version, downloaded it, and - the key part - *recorded the dependency in your `.csproj`*. No files were manually copied. Next build, .NET **restores** it automatically (downloads if not cached), so a teammate who clones your repo just runs `dotnet build` and the package shows up too.

That recorded dependency looks like this inside the `.csproj`:

```xml
<ItemGroup>
  <PackageReference Include="Newtonsoft.Json" Version="13.0.3" />
</ItemGroup>
```

*What just happened:* `<PackageReference>` says "this project depends on this package at this version" - all `dotnet add package` did. Living in version-controlled XML, not copied binaries, keeps your repo small and versions reproducible from the `.csproj` alone.

💡 **Reach for a package before rolling your own.** JSON parsing, HTTP clients with retries, date/time handling, CSV, PDFs, database access - the odds someone's already built and battle-tested what you need are high. Search nuget.org first; a mature package has handled edge cases you haven't thought of, and "I added a well-maintained package" beats "I wrote my own and now maintain it forever."

## Build & publish - compile vs. ship

`dotnet build` and `dotnet publish` sound similar and are easy to mix up. The difference: *who the output is for*.

📝 **Build vs. publish.** `dotnet build` compiles for *you* to run and debug locally. `dotnet publish` produces a *deployable* bundle - everything needed on another machine, ready for a server or container.

```bash
dotnet build                                   # compile for local use
dotnet publish -c Release -o ./out             # produce a deployable bundle
```

*What just happened:* `dotnet build` produced binaries under `bin/`, ready to run on your dev machine where the runtime's already installed. `dotnet publish -c Release` did an optimized compile and gathered the result into `./out` - a shippable folder. Publish comes in two flavors: **framework-dependent** (smaller, but the target needs the matching runtime installed) and **self-contained** (larger, bundles the runtime, so the target needs nothing pre-installed) - what makes "copy this folder to a bare server and run it" possible.

⚠️ **The `bin/` and `obj/` folders - generated, never committed.** Every build creates `obj/` (intermediate compiler junk) and `bin/` (final binaries). Both are *generated output*, rebuilt from source any time, so they don't belong in version control. The default `.gitignore` already excludes them; committing them is a classic beginner mistake that bloats the repo and causes merge conflicts over files nobody edits.

## The wider toolchain - the batteries that come with C#

The `dotnet` CLI is the engine, but you'll spend your days inside richer tools built around it:

- **IDEs & editors.** **Visual Studio** (Windows-only, full-featured heavyweight), **VS Code** with the **C# Dev Kit** extension (lightweight, cross-platform, hugely popular), and **JetBrains Rider** (cross-platform, beloved for refactoring; the usual pick on Mac and Linux). All three give IntelliSense, a visual debugger, and one-click run/test - all driving the same `dotnet` build.
- **`dotnet format`.** The built-in formatter. Like `gofmt` or `black`, it enforces consistent whitespace so code review stops being about indentation. Run before committing, or wire it to run on save.
- **Analyzers & Roslyn.** C#'s compiler is **Roslyn**, exposing its understanding of your code to *analyzers* - plugins flagging bugs, style violations, and risky patterns *as you type*. Many ship with the SDK automatically; teams add more for their own rules.
- **Testing & profiling.** Unit testing (**xUnit** and friends) and profiling are first-class here, big enough for their own treatment in Phase 16. For now, know `dotnet test` exists.

💡 **Key point.** C# has one of the most *batteries-included*, mature toolchains in software. Build tool, package manager, formatter, analyzers, and test runner are all official, integrated, and driven by the same `dotnet` CLI - not third-party parts hoped to cooperate.

## Recap

1. **Namespaces** give types an address (`namespace MyApp;`) and **`using`** imports them so you can use short names; **global/implicit usings** mean modern files often need no imports at all.
2. **A `.csproj` *is* the project** - a small XML file naming the target framework, output type, and `<PackageReference>` dependencies; the `.cs` files come in by convention. A **`.sln`** groups multiple projects.
3. **`dotnet new` / `build` / `run`** scaffold, compile, and run; **NuGet** (`dotnet add package`) pulls in outside libraries from nuget.org, recording them in the `.csproj` so they restore automatically.
4. **`dotnet build`** compiles for local use; **`dotnet publish`** produces a deployable bundle (framework-dependent or self-contained). The generated **`bin/`** and **`obj/`** folders never get committed.
5. The wider toolchain - **Visual Studio / VS Code + C# Dev Kit / Rider**, **`dotnet format`**, **Roslyn analyzers**, and the testing/profiling tools coming in Phase 16 - is mature, official, and integrated around the same `dotnet` CLI.

You can now organize, package, and ship real C# - not just run loose files. Next we cover the idioms and gotchas that separate "compiles" from "looks like C# a pro would write."

## Quick check

Make sure the core mental model - the `.csproj` as the center of a project - stuck:

```quiz
[
  {
    "q": "What does a `.csproj` file actually define?",
    "choices": [
      "The project itself: its target framework, output type, and package dependencies - the .cs files are included by convention",
      "A list of every source file in the project, which you must keep up to date by hand",
      "Only the compiled output binaries, regenerated on each build",
      "The solution that groups several projects together"
    ],
    "answer": 0,
    "explain": "The `.csproj` is the project's definition - target framework, output type, and `<PackageReference>` dependencies. The `.cs` files under its folder are pulled in automatically, so you don't list them. Grouping multiple projects is the job of a `.sln`."
  },
  {
    "q": "After you run `dotnet add package Newtonsoft.Json`, how does the package end up available to a teammate who clones your repo?",
    "choices": [
      "The command records a `<PackageReference>` in the .csproj, and `dotnet build` restores the package automatically from it",
      "The package's binaries are copied into the repo and committed alongside your code",
      "Your teammate must manually download the package from nuget.org and place it in bin/",
      "The package is embedded directly into every .cs file that uses it"
    ],
    "answer": 0,
    "explain": "`dotnet add package` edits a `<PackageReference>` line into the `.csproj`. Because the dependency lives in version-controlled XML, a fresh clone just needs `dotnet build`, which restores (downloads) the recorded packages automatically - no binaries committed to the repo."
  },
  {
    "q": "What's the difference between `dotnet build` and `dotnet publish`?",
    "choices": [
      "`build` compiles binaries for you to run locally; `publish` produces a deployable bundle (optionally self-contained) for another machine",
      "`build` is for libraries and `publish` is for executables",
      "They are identical; `publish` is just the older name for `build`",
      "`build` downloads NuGet packages while `publish` removes them"
    ],
    "answer": 0,
    "explain": "`dotnet build` compiles binaries for local running and debugging. `dotnet publish` gathers everything needed to deploy elsewhere - framework-dependent (needs the runtime installed) or self-contained (bundles the runtime so the target needs nothing)."
  }
]
```


---

# Idioms & Common Gotchas - Write It Like a C# Dev, Dodge the Traps

You can write C# that compiles and runs. This phase covers the gap between *that* and code that looks like a seasoned C# developer wrote it - plus the traps that have caught every C# programmer who ever lived. (Genuinely - `NullReferenceException` alone has cost the industry more debugging hours than anyone wants to count.)

Two halves. First the **idioms** - conventions that make C# code feel coherent instead of arbitrary. The mental model: *C# rewards saying exactly what you mean, and letting the compiler catch your mistakes*. Expose state through properties, not raw fields; let the compiler infer obvious types; ask for "no value" out loud with nullable reference types instead of letting `null` sneak in unannounced.

Then a scannable **gotcha cheat-card** - surprises named *before* they bite, so you recognize them instead of staring at a stack trace. The mental model: *some types copy by value, some by reference, and queries don't run when you think they do*.

## Idioms - the way it's written today

### Use properties, not public fields

**What it actually is.** A *property* looks like a field from the outside (`user.Name`) but is really a pair of accessors you control. A public field hands out raw access to internals with no way to validate, compute, or change the implementation later.

```csharp
// Idiomatic: a property. Auto-implemented, but you can add logic anytime.
public class User
{
    public string Name { get; set; }
    public int Age { get; private set; }   // settable only from inside
}

// Not idiomatic: a public field. No validation, no encapsulation, ever.
public class UserBad
{
    public string Name;
}
```
*What just happened:* `Name { get; set; }` is an *auto-property* - the compiler generates a hidden backing field, no more typing than a public field, but a real property. The day you need to validate, log, or make it read-only, you change the property body and *no caller has to change*. A public field never grows those abilities without breaking everyone. Properties are also what data-binding and serializers expect - default to them.

💡 **Key point.** Use `{ get; private set; }` (or `{ get; init; }`) when a value should be set once and stay put. Expose *behavior and controlled state*, not raw memory.

### Prefer `var` when the type is obvious

**What it actually is.** `var` tells the compiler to infer the type from the right-hand side. It's still *statically typed* - `var count = 5;` is exactly an `int`, not dynamic - it just saves repeating a type the reader can already see.

```csharp
var names = new List<string>();        // clearly a List<string>
var count = names.Count;               // clearly an int
var user = new User();                 // clearly a User

// When the type ISN'T obvious from the right side, spell it out:
int total = Calculate();               // what does Calculate return? Be explicit.
```
*What just happened:* on the first three lines the type sits right there on the right of `=`, so `var` removes noise without removing clarity. On the last line `Calculate()`'s return type isn't visible, so naming it (`int`) helps. The idiom: *use `var` when the type is already obvious, name it when it isn't*. Clarity is the goal, not brevity.

### String interpolation with `$"..."`

**What it actually is.** Prefixing a string with `$` lets you drop expressions inside it with `{ }`, instead of gluing pieces with `+` or juggling positional `{0}` placeholders.

```csharp
string name = "Ada";
int age = 36;

string greeting = $"Hello, {name}! You are {age} years old.";
string math = $"Next year you'll be {age + 1}.";   // full expressions allowed
Console.WriteLine(greeting);
Console.WriteLine(math);
```
```console
Hello, Ada! You are 36 years old.
Next year you'll be 37.
```
*What just happened:* `$"...{name}..."` substituted `name`'s value where it appears, and `{age + 1}` evaluated a whole expression inline. Compare to `"Hello, " + name + "! You are " + age + ..."` - the interpolated version reads like the sentence it produces, no forgotten space or mismatched `{0}`. The default way to build strings in modern C#.

### Expression-bodied members and `nameof`

**What it actually is.** When a method or property is a single expression, `=>` writes it on one line instead of a full `{ return ...; }` block. `nameof(x)` turns a symbol into its *name as a string* - checked by the compiler, so a rename or typo becomes a build error instead of a stale string.

```csharp
public class Circle
{
    public double Radius { get; init; }

    // Expression-bodied property and method - concise, no braces needed.
    public double Area => Math.PI * Radius * Radius;
    public override string ToString() => $"Circle(r={Radius})";

    public void SetRadius(double r)
    {
        if (r < 0)
            throw new ArgumentOutOfRangeException(nameof(r), "must be >= 0");
    }
}
```
*What just happened:* `Area => ...` and `ToString() => ...` are *expression-bodied members* - `=>` replaces the `{ return ...; }` ceremony for one-liners. In `SetRadius`, `nameof(r)` produced the string `"r"`, but renaming the parameter updates it via your IDE and a typo won't compile. Use `nameof` anywhere you'd otherwise hard-code a name - exception messages, logging, property-change notifications.

### Null-handling operators: `?.`, `??`, `??=`

**What it actually is.** Three operators tame `null`. `?.` (null-conditional) reads a member only if not null, short-circuiting instead of throwing. `??` (null-coalescing) supplies a fallback when the left side is null. `??=` assigns *only if* the target is currently null.

```csharp
string? maybeName = GetName();   // might return null

int? length = maybeName?.Length;          // null if maybeName is null - no crash
string shown = maybeName ?? "(unknown)";  // fallback when null
maybeName ??= "default";                   // assign only if it was null

Console.WriteLine($"{length} / {shown} / {maybeName}");
```
*What just happened:* `maybeName?.Length` would normally throw a `NullReferenceException`, but `?.` short-circuits to `null` instead. `?? "(unknown)"` supplied a default, and `??=` set a value only because the variable was still null. Together these replace towers of `if (x != null)` checks: "read if present," "fall back if missing," "fill in if empty."

⚠️ **Gotcha - `?.` returns a *nullable*.** `maybeName?.Length` is an `int?`, not `int`. The compiler makes you handle that - don't be surprised you can't assign it straight to a plain `int`. Chain `?? 0` for a non-null result.

### Pattern matching

**What it actually is.** `switch` expressions and `is` patterns test an object's *shape* - type, values, properties - and pull data out in the same step, replacing long `if/else` ladders and clumsy casts with a decision table.

```csharp
object value = 42;

// Type pattern with `is` - test and capture in one move.
if (value is int n && n > 10)
    Console.WriteLine($"big int: {n}");

// switch expression - every input maps to a result, no break needed.
string describe(object x) => x switch
{
    null            => "nothing",
    int i when i < 0 => "negative",
    int              => "an int",
    string s        => $"text of length {s.Length}",
    _               => "something else"     // _ is the catch-all
};
Console.WriteLine(describe("hi"));
```
```console
big int: 42
text of length 2
```
*What just happened:* `value is int n` checked the type *and* gave a typed variable `n` in one step - no separate cast that might throw. The `switch` expression mapped each shape to a result, `when` adding extra conditions and `_` catching everything else. The compiler even warns if you forget a case. (Deep dive in [Phase 13](13-records-and-modern-csharp.md).)

### Enable nullable reference types

**What it actually is.** A project-level switch (`<Nullable>enable</Nullable>` in `.csproj`) making the compiler track which references *can* be null. `string` means "never null"; `string?` means "might be null" - it warns whenever you might dereference a null without checking.

```csharp
#nullable enable
string name = null;       // ⚠️ compiler warning: assigning null to non-nullable
string? maybe = null;     // fine - the ? says "this can be null"

void Print(string? text)
{
    Console.WriteLine(text.Length);  // ⚠️ warning: text might be null here
    Console.WriteLine(text?.Length); // fine - guarded with ?.
}
```
*What just happened:* with nullable reference types on, the compiler treats `null` as a *visible part of the type*. `string` promises non-null, so assigning `null` earns a warning; `string?` admits it might be null, so the compiler *insists* you guard every use. This turns `NullReferenceException` from a runtime surprise into a compile-time nag. Turn it on in every new project.

### Dispose with `using`, and program to interfaces

**What it actually is.** Anything holding an unmanaged resource - a file, socket, database connection - implements `IDisposable` and must be *closed* when done. `using` guarantees cleanup runs even if an exception is thrown. And as elsewhere, declare variables and parameters by their *interface* (`IEnumerable<T>`, `IList<T>`) rather than the concrete class, so callers depend on the contract, not the implementation.

```csharp
// `using` declaration: the file is closed automatically at the end of scope.
using var reader = new StreamReader("data.txt");
string firstLine = reader.ReadLine();
// reader.Dispose() runs here, even if ReadLine throws - no leak.

// Program to the interface: callers don't care it's really a List.
void PrintAll(IEnumerable<string> items)
{
    foreach (var item in items)
        Console.WriteLine(item);
}
```
*What just happened:* `using var reader = ...` ties the `StreamReader`'s lifetime to its scope - `Dispose()` runs automatically on exit, even on an exception. Forget `using` and you leak handles. Separately, `PrintAll` accepts `IEnumerable<string>` - the most general interface that does the job - so it works with a `List`, array, LINQ query, or anything iterable.

> 💡 The umbrella idiom: make intent explicit and let the compiler help. Properties expose controlled state, `var` removes noise only where the type is clear, nullable reference types make absence visible, `using` makes cleanup automatic, interfaces narrow the contract. Clarity over cleverness.

## The gotcha cheat-card

> **Hit something baffling? Find the symptom here.** These trap *everyone* - recognizing them on sight is most of the battle.

| The trap | What bites you | The fix |
|---|---|---|
| `==` vs `.Equals()` | For classes, `==` compares *references* - two objects with identical data are "not equal" | Override `Equals`/`==`, or use a `record`; note `string` already compares by value |
| `NullReferenceException` | Calling a member on a `null` reference throws at runtime | `?.`, `??`, nullable reference types, explicit null checks |
| Deferred LINQ execution | A query doesn't run when defined - it re-runs on each enumeration, capturing variables *late* | Materialize with `.ToList()` / `.ToArray()` when you need a stable snapshot |
| Integer division | `5 / 2` is `2`, not `2.5` - the fraction is silently discarded | Cast an operand: `5 / 2.0` or `(double)a / b` |
| `async void` | Exceptions vanish unobserved and callers can't await it | Use `async Task` everywhere except event handlers |
| Mutating a collection in `foreach` | Adding/removing during iteration throws `InvalidOperationException` | Loop a copy, use a `for` loop, or collect changes then apply |
| Struct copy semantics | Assigning or passing a `struct` copies it - your edit hits the copy, not the original | Know value vs reference (Phase 2); prefer classes for mutable shared state |

Now the *why* behind the sharpest ones.

### Deferred execution of LINQ

The one that makes people doubt their sanity. A LINQ query like `numbers.Where(...)` doesn't *run* when written - it builds a *recipe*. The work happens later, every time you enumerate it (`foreach`, `.ToList()`, `.Count()`), reading the source variables *at enumeration time*, not definition time.

```csharp
var numbers = new List<int> { 1, 2, 3 };

// This does NOT run yet - it just describes "evens of numbers".
var evens = numbers.Where(n => n % 2 == 0);

numbers.Add(4);   // we change the source AFTER defining the query

Console.WriteLine(string.Join(", ", evens));  // query runs NOW
Console.WriteLine(string.Join(", ", evens));  // ...and runs AGAIN
```
```console
2, 4
2, 4
```
*What just happened:* `evens` captured the *recipe* "give me the even numbers in `numbers`," not a snapshot. Adding `4` *after* defining the query meant the later enumeration saw it. Worse, each `Console.WriteLine` re-ran the whole query; against a database that's extra round-trips. For a stable result, *materialize* it: `var evens = numbers.Where(...).ToList();` runs the query once and freezes the answer.

⚠️ **Gotcha - deferred queries re-run and capture late.** Two classic bites: (1) enumerating the same query repeatedly redoes the work, and (2) a query inside a loop captures the *loop variable's final value*, not each iteration's. When in doubt, `.ToList()` for a snapshot.

### Integer division

Dividing two `int`s gives an `int` - C# discards the remainder rather than producing a fraction. Bites every beginner computing an average or percentage.

```csharp
int total = 5, count = 2;

Console.WriteLine(total / count);            // 2   - fraction discarded!
Console.WriteLine((double)total / count);    // 2.5 - cast first
Console.WriteLine(total / (double)count);    // 2.5 - either operand works
```
```console
2
2.5
2.5
```
*What just happened:* `5 / 2` is integer arithmetic - `2` with the `.5` silently dropped, no error or warning. Casting one operand to `double` forces *floating-point* division: `2.5`. If you want a fractional result, make at least one operand a `double` (or `decimal` for money) *before* the division runs. `(double)(total / count)` is too late - the integer division already happened.

### `==` vs `.Equals()` - value vs reference equality

For most *classes*, `==` asks "are these the *same object* in memory?" - reference equality. `.Equals()` *can* mean value equality, but a plain class defaults to reference equality until overridden. The big exception: `string` overrides both to compare *contents*, behaving like a value type - which lulls beginners into assuming `==` always compares values, until it breaks on their own class.

```csharp
class Point { public int X, Y; }
record PointR(int X, int Y);   // records compare by value, for free

var a = new Point { X = 1, Y = 2 };
var b = new Point { X = 1, Y = 2 };
Console.WriteLine(a == b);              // False - different objects

var p = new PointR(1, 2);
var q = new PointR(1, 2);
Console.WriteLine(p == q);              // True  - record value equality

Console.WriteLine("hi" == "hi");        // True  - string compares contents
```
```console
False
True
True
```
*What just happened:* the two `Point` objects hold identical data, but `==` on a plain class compares *references* - separate objects, `False`. The two `PointR` *records* compare by value automatically (C# generates `Equals`, `==`, `GetHashCode` from the fields), `True`. `"hi" == "hi"` is `True` because `string` overrides equality to compare characters. For value-like data compared by content, use a **`record`** (or override `Equals`/`GetHashCode`/`==`). Don't assume `==` means value equality just because it does for strings.

📝 **The other three, in one line each.** **`async void`** - can't be awaited, exceptions disappear instead of bubbling to the caller; use `async Task` everywhere except UI event handlers. **Mutating during `foreach`** - adding/removing while iterating throws `InvalidOperationException`; iterate a copy, use an index-based `for`, or collect changes and apply after. **Struct copy semantics** - a `struct` is a *value type* (Phase 2), so assigning or passing it copies the whole thing; mutating that copy leaves the original untouched. For shared mutable state, reach for a `class`.

## Recap

1. **Use properties** (`{ get; set; }`), not public fields - same syntax, but you keep control of validation and future change.
2. **Prefer `var` when the type is obvious**, name the type when it isn't; use **`$"..."`** interpolation, **expression-bodied members**, and **`nameof`** for compiler-checked names.
3. **Tame `null`** with `?.`, `??`, `??=`, and turn on **nullable reference types** so the compiler catches `NullReferenceException` before runtime. Dispose with **`using`** and **program to interfaces**.
4. ⚠️ **Deferred LINQ** doesn't run until enumerated, re-runs each time, and captures variables late - `.ToList()` to snapshot it.
5. ⚠️ **The cheat-card** - `==` on a class compares references (use a `record` or override `Equals` for value equality; `string` already compares contents); `5 / 2` is `2` (cast to `double`); `async void` swallows exceptions; mutating a collection in `foreach` throws; structs copy by value.

That's idiomatic C#. You can now read other people's C# and write code that looks like it belongs, having met the traps before they meet you. Next we go deep on **generics**: how `List<T>` really works, constraints, variance, and why the compiler sometimes argues with you about types.

## Quick check

Test yourself on the three traps that catch everyone:

```quiz
[
  {
    "q": "You have two plain-class objects with identical field values: `var a = new Point{X=1,Y=2};` and `var b = new Point{X=1,Y=2};`. What does `a == b` return, and how do you get value equality?",
    "choices": [
      "`False` - `==` compares references for a class; use a `record` or override `Equals`/`==` to compare by value",
      "`True` - C# always compares objects by their field values",
      "`True` - but only because Point has two fields",
      "It throws, because `==` isn't defined for custom classes"
    ],
    "answer": 0,
    "explain": "For a plain class, `==` compares references, so two distinct objects are not equal even with identical data. Use a `record` (which generates value equality) or override `Equals`/`GetHashCode`/`==`. Note `string` is the exception - it compares contents."
  },
  {
    "q": "You write `var evens = numbers.Where(n => n % 2 == 0);`, then `numbers.Add(4);`, then enumerate `evens`. Why does the new `4` show up?",
    "choices": [
      "LINQ is deferred - the query holds a recipe and runs at enumeration time, reading the (now updated) source",
      "`.Where` made a copy of the list, so it tracks changes automatically",
      "Adding to a list always re-runs every query that referenced it",
      "It's a bug; the query should have captured the original three values"
    ],
    "answer": 0,
    "explain": "A LINQ query is deferred: it doesn't run when defined, it runs each time you enumerate it, reading the source then. Since you added 4 before enumerating, the query sees it. Call `.ToList()` at definition time to freeze a snapshot."
  },
  {
    "q": "What does `5 / 2` evaluate to in C#, and how do you get `2.5`?",
    "choices": [
      "It's `2` (integer division discards the remainder); cast an operand to double, e.g. `5 / 2.0` or `(double)5 / 2`",
      "It's `2.5` already - C# promotes int division to double automatically",
      "It's `3` - C# rounds to the nearest integer",
      "It throws a DivideByZeroException for non-divisible numbers"
    ],
    "answer": 0,
    "explain": "Dividing two ints gives an int, dropping the fraction, so `5 / 2` is `2`. Force floating-point division by making at least one operand a double (or decimal) before the division runs. Casting after, like `(double)(5/2)`, is too late."
  }
]
```


---

# Generics, Deep - Type Safety Without Duplication

Back in [Phase 3](03-collections.md) you used `List<T>` the way everyone first meets generics: you wrote `List<string>`, it held strings, and the compiler stopped you shoving an `int` in - you took the `<T>` on faith. This phase covers what that angle-bracket actually *is*.

The mental model: a generic is **code with a hole in it where a type goes**. You write the logic once, leave the type as a blank labeled `T`, and the compiler fills that blank - `string`, `int`, `User`, whatever - when you use it. Write it *once*, keep full *compile-time type safety*, pay *zero runtime cost*.

## Why generics exist - the bad old days of `object`

Before generics, a list that could hold *anything* meant using `object` - the type every other type inherits from. It compiles fine. It's also a trap.

```csharp
// Pre-generics: a list of `object` holds anything... which is the problem.
var things = new System.Collections.ArrayList();
things.Add(42);
things.Add("not a number");   // compiler is fine with this - uh oh

int first = (int)things[0];   // cast back: works
int second = (int)things[1];  // cast back: BOOM at runtime
```
```console
Unhandled exception. System.InvalidCastException: Unable to cast object
of type 'System.String' to type 'System.Int32'.
```
*What just happened:* `ArrayList` stores everything as `object`, so it happily accepted both an `int` and a `string` - the compiler had no idea the second item wasn't a number. The mistake surfaced only at runtime, when the cast on `things[1]` blew up *in production*. There's a quieter cost too: stuffing `42` into an `object` slot **boxes** it - wraps it in a heap-allocated object - and casting it back **unboxes** it, an allocation and a copy for every value-type item.

Generics fix both problems at once: `List<int>` knows its contents are `int`s, so the mistake becomes a *compile error* and the boxing never happens - the `int`s are stored raw.

```csharp
var numbers = new List<int>();
numbers.Add(42);
numbers.Add("not a number");   // ⚠️ compile error - caught before you run
int first = numbers[0];        // no cast, no boxing
```
```console
error CS1503: Argument 1: cannot convert from 'string' to 'int'
```
*What just happened:* `List<int>` carries its element type in the type itself, so `Add("not a number")` is rejected at *compile time* - the bug can never reach a user. Because the list genuinely holds `int`s, not boxed `object`s, reading `numbers[0]` needs no cast and no heap allocation. **`object` gives you flexibility by throwing away type information; generics give you flexibility while keeping it.**

## Generic methods and classes - write the logic once

A generic puts a *type parameter* - conventionally `T` - into a method or class signature. Inside, `T` stands in for "whatever type the caller used"; the compiler checks the body against that placeholder and substitutes the real type at each call.

📝 **Type parameter** - a named placeholder for a type, written in angle brackets (`<T>`), filled in by the caller or the compiler's inference. By convention single letters: `T` for "type," `TKey`/`TValue` for paired roles, `TResult` for a return.

A generic *method* - "give me the first item" works for a list of anything:

```csharp
// <T> declares the placeholder; it then appears in the parameter and return types.
T First<T>(List<T> items)
{
    return items[0];
}

var words = new List<string> { "alpha", "beta" };
var sizes = new List<int> { 10, 20, 30 };

string w = First(words);   // T inferred as string - note: no First<string>(...) needed
int n = First(sizes);      // T inferred as int
Console.WriteLine($"{w}, {n}");
```
```console
alpha, 10
```
*What just happened:* `First<T>` is one method that works for any element type. You didn't write `First<string>(words)` - the compiler performed **type inference**, saw `words` was a `List<string>`, deduced `T` must be `string`, and filled it in, so `w` comes back a real `string` with no cast needed.

A generic *class* - a little box that holds one value of whatever type you choose:

```csharp
class Box<T>
{
    private T _value;

    public Box(T value) => _value = value;

    public T Get() => _value;

    // default(T): the "zero" value for whatever T is.
    public bool IsDefault() => EqualityComparer<T>.Default.Equals(_value, default(T));
}

var boxedInt = new Box<int>(0);
var boxedName = new Box<string>("Ada");
Console.WriteLine($"{boxedInt.IsDefault()} / {boxedName.Get()}");
```
```console
True / Ada
```
*What just happened:* `Box<T>` is a class with a type-shaped hole - `new Box<int>(0)` stamps out a box holding an `int`; `new Box<string>("Ada")` a box holding a `string`. The interesting bit is `default(T)`: every type has a *default value* - `0` for `int`, `false` for `bool`, `null` for reference types - and `default(T)` gives you that value generically, without knowing what `T` is. We used it to check whether the box holds its type's "zero."

💡 **Key point.** `default(T)` matters because inside a generic you often need a starting value but can't write a literal - you don't know if `T` is a number, a string, or a struct. Modern C# lets you shorten `default(T)` to just `default` wherever the target type is already known.

## Constraints - telling the compiler what `T` can do

There's a catch in `First<T>` and `Box<T>`: the compiler assumes `T` could be *literally any type*, so it only lets you do things every type supports - store it, pass it around, call `.ToString()`. Try `a > b`, `new T()`, or `a.SomeMethod()` and it refuses, since not every type has those.

The fix is a **constraint**: a `where` clause that narrows what `T` is allowed to be, which in turn *unlocks* the operations that narrower set of types supports.

📝 **Constraint** - a `where T : ...` clause restricting which types can be used for `T`. A two-way promise: you limit the callers' choices, and in exchange the compiler lets you use the capabilities all allowed types are guaranteed to have.

The common constraints:

| Constraint | Means "T must be…" | Unlocks |
|---|---|---|
| `where T : class` | a reference type | comparing to `null`, `null` defaults |
| `where T : struct` | a value type (non-nullable) | value semantics; `T` is never null |
| `where T : new()` | a type with a public parameterless constructor | calling `new T()` |
| `where T : IComparable<T>` | a type implementing that interface | calling `.CompareTo(...)` |
| `where T : SomeBaseClass` | that class or a subclass | that base's members |

Without a constraint you can't write a generic "max" - the compiler can't assume `T` is comparable. Add `where T : IComparable<T>` and it works:

```csharp
// IComparable<T> guarantees a.CompareTo(b) exists, which is what unlocks the comparison.
T Max<T>(T a, T b) where T : IComparable<T>
{
    return a.CompareTo(b) >= 0 ? a : b;
}

Console.WriteLine(Max(3, 9));          // works for int (int : IComparable<int>)
Console.WriteLine(Max("apple", "pear")); // works for string too
```
```console
9
pear
```
*What just happened:* `a.CompareTo(b)` returns a number - negative if `a` is smaller, positive if bigger, zero if equal. That method exists only because we *promised*, via `where T : IComparable<T>`, that every `T` implements it. `int` and `string` both do, so both calls compile. Drop the `where` clause and the compiler rejects `a.CompareTo(b)` outright - `T` might otherwise be a type that can't be compared at all.

The `new()` constraint lets a generic *manufacture* instances:

```csharp
// new() promises T has a parameterless constructor, so `new T()` is allowed.
T MakeOne<T>() where T : new()
{
    return new T();
}

var freshList = MakeOne<List<int>>();   // calls new List<int>()
Console.WriteLine(freshList.Count);
```
```console
0
```
*What just happened:* `new T()` is normally forbidden in a generic - the compiler can't know `T` even *has* a usable constructor. The `where T : new()` constraint guarantees it does, so `new T()` becomes legal.

When a caller violates a constraint, the error lands at compile time:

```csharp
class NoDefaultCtor
{
    public NoDefaultCtor(int required) { }   // only constructor needs an argument
}

var bad = MakeOne<NoDefaultCtor>();   // ⚠️ no parameterless constructor
```
```console
error CS0310: 'NoDefaultCtor' must be a non-abstract type with a public
parameterless constructor in order to use it as parameter 'T' in 'MakeOne<T>()'
```
*What just happened:* `NoDefaultCtor` only has a constructor that *requires* an argument, so it fails the `new()` constraint. The compiler refuses the call and tells you why - `T` doesn't meet the promise the method depends on.

## Covariance and contravariance - the genuinely tricky bit

A `Cat` is an `Animal`. So a `List<Cat>` is a `List<Animal>`… right? In C#, **no** - and the reason matters.

⚠️ The intuition is *half* true. Whether `Something<Cat>` can stand in for `Something<Animal>` depends on whether that "something" only ever *hands values out* or also *takes values in*.

Imagine the substitution were always allowed - you hand your `List<Cat>` to code that thinks it has a `List<Animal>`:

```csharp
List<Cat> cats = new List<Cat> { new Cat() };
List<Animal> animals = cats;   // PRETEND this were allowed...
animals.Add(new Dog());        // ...a Dog, into a list that's really all Cats. Corruption.
Cat c = cats[0];               // your "list of cats" now contains a Dog
```
*What just happened (in this hypothetical):* if `List<Cat>` could masquerade as `List<Animal>`, code holding the `List<Animal>` view could legally `Add` a `Dog` - because a `Dog` *is* an `Animal`. But the underlying list is genuinely all `Cat`s, so you've smuggled a `Dog` into it. That's why C# forbids the assignment: `List<T>` lets you *write* into it, and writing is where widening goes wrong.

Now flip it: a type that only ever *produces* values can widen safely - a thing that hands you `Cat`s can be treated as a thing that hands you `Animal`s, since every `Cat` it gives out *is* an `Animal`. That's **covariance**, and `IEnumerable<T>` (read-only iteration) is declared this way with the keyword `out`:

```csharp
IEnumerable<Cat> cats = new List<Cat> { new Cat(), new Cat() };
IEnumerable<Animal> animals = cats;   // ALLOWED - IEnumerable<out T> is covariant

foreach (Animal a in animals)         // every Cat handed out is, indeed, an Animal
    Console.WriteLine(a.GetType().Name);
```
```console
Cat
Cat
```
*What just happened:* `IEnumerable<T>` is declared `IEnumerable<out T>` - `out` marks `T` as **covariant**, "this type only ever *produces* `T`, never consumes one." Because iteration only *reads* items out, treating an `IEnumerable<Cat>` as an `IEnumerable<Animal>` is safe: every value pulled is a `Cat`, which is an `Animal`, and there's no `Add` to corrupt anything.

The mirror image is **contravariance**: a type that only ever *consumes* values can be treated as one that accepts a *narrower* type. `Action<in T>` (a function taking a `T`, returning nothing) is the classic case:

```csharp
// A consumer of Animals: it can handle ANY animal handed to it.
Action<Animal> describe = a => Console.WriteLine($"an {a.GetType().Name}");

// We need something that consumes Cats. An Animal-consumer qualifies - it eats cats too.
Action<Cat> describeCat = describe;   // ALLOWED - Action<in T> is contravariant
describeCat(new Cat());
```
```console
an Cat
```
*What just happened:* `Action<T>` is declared `Action<in T>` - `in` marks `T` as **contravariant**, "this type only ever *consumes* `T`." We needed something that consumes `Cat`s; `describe` consumes *any* `Animal`, so it handles a `Cat` fine. Consumers can safely *narrow*.

One sentence covers it: **`out` = produces-only = can widen (covariant); `in` = consumes-only = can narrow (contravariant); both = invariant, no substitution.** `List<T>` does both, which is why it's invariant and the first example was illegal.

⚠️ **Gotcha - arrays are unsafely covariant, a historical wart.** Arrays *do* allow `Animal[] a = new Cat[2];`, even though arrays are writable. C# inherited this from early .NET (before generics existed) for compatibility, and it's a known design mistake: every array write carries a hidden *runtime* type check, and writing the wrong type throws `ArrayTypeMismatchException` instead of failing at compile time. Generics learned from this - `List<T>` is invariant so the same bug can't happen. Treat array covariance as a trap, not a feature.

## C# generics are reified - a real edge over Java

Java implements generics by **type erasure**: `List<String>` and `List<Integer>` are *the same type at runtime* - the `<String>` is a compile-time fiction deleted before the program runs. That's why Java can't do `new T()`, can't ask `T.class`, and boxes every `int` into an `Integer`.

C# made the opposite choice: generics are **reified**, the type information is *real at runtime*, baked into the actual type.

📝 **Reified generics** - generic type information that survives to runtime, rather than being erased after compilation. In C#, `List<int>` and `List<string>` are genuinely *distinct types* at runtime, and `T` is a real, queryable type inside generic code.

💡 **Key point.** Three consequences fall out of reification, each impossible under Java's erasure:
- **`typeof(T)` works** - ask, at runtime, "what type is `T` right now?" and get a real answer.
- **`new T()` works** (with the `new()` constraint) - the runtime knows what `T` is, so it can construct one.
- **Value types avoid boxing** - `List<int>` stores raw `int`s, not boxed `Integer`-style objects. Java *must* box, since erased generics can't hold a primitive.

```csharp
void Inspect<T>() where T : new()
{
    T instance = new T();              // reification: runtime can build a T
    Console.WriteLine($"T is {typeof(T).Name}");   // and tell you what T is
    Console.WriteLine($"made: {instance}");
}

Inspect<List<int>>();

// And distinct runtime types - not the case in Java:
Console.WriteLine(typeof(List<int>) == typeof(List<string>));
```
```console
T is List`1
made: System.Collections.Generic.List`1[System.Int32]
False
```
*What just happened:* inside `Inspect<T>`, both `typeof(T)` and `new T()` work because the runtime genuinely *knows* what `T` is. The final comparison prints `False`: `List<int>` and `List<string>` are different runtime types. (``List`1`` is .NET's name for "a generic List with 1 type parameter.") Under Java's erasure, `T` would be unknowable, `new T()` impossible, `List<Integer>`/`List<String>` indistinguishable.

## Recap

1. **Why generics:** the old `object`-and-cast approach pushed type errors to *runtime* and boxed every value type. Generics keep type information, so mistakes become *compile errors* and value types stay unboxed.
2. **Generic methods and classes** (`T First<T>(...)`, `class Box<T>`) write logic once for many types; the compiler **infers** `T` from arguments, and `default(T)` / `default` gives a generic "zero" value when you can't write a literal.
3. **Constraints** (`where T : class / struct / new() / IComparable<T> / SomeBase`) narrow what `T` can be and *unlock* operations - like `a.CompareTo(b)` or `new T()` - those types are guaranteed to support. Violations fail at compile time.
4. **Covariance (`out`) and contravariance (`in`)**: a *producer-only* type (`IEnumerable<out T>`) can safely widen; a *consumer-only* type (`Action<in T>`) can safely narrow. `List<T>` reads *and* writes, so it's invariant. Arrays are unsafely covariant - a historical wart that throws at runtime.
5. **Reified generics:** C# keeps type info at runtime, so `typeof(T)`, `new T()`, and `List<int>` ≠ `List<string>` all work, and value types avoid boxing - an advantage over Java's erasure.

You can now write code with type-shaped holes the compiler fills safely and for free. Next: **delegates, lambdas, and events** - functions as values, the engine behind LINQ and most of the .NET event model.

## Quick check

Test yourself on the three ideas that matter most - constraints, variance, and reification:

```quiz
[
  {
    "q": "Why does `T Max<T>(T a, T b)` need the `where T : IComparable<T>` constraint?",
    "choices": [
      "Without it, the compiler can't guarantee `T` supports `.CompareTo(...)`, so the comparison in the body wouldn't be allowed",
      "It makes the method run faster by skipping a runtime type check",
      "It forces `T` to be a value type so it can be stored on the stack",
      "It's optional - the method compiles fine with no constraint at all"
    ],
    "answer": 0,
    "explain": "A constraint both restricts which types `T` may be and unlocks that capability inside the body. `IComparable<T>` provides `.CompareTo(...)`; without the constraint, `T` could be a non-comparable type, so the call would be rejected at compile time."
  },
  {
    "q": "Why does C# allow `IEnumerable<Animal> a = someIEnumerableOfCat;` but forbid the same assignment for `List<T>`?",
    "choices": [
      "`IEnumerable<out T>` only produces values (covariant), so widening is safe; `List<T>` also writes, so widening could smuggle a wrong type in",
      "`IEnumerable` is a class and `List` is an interface, and only classes support variance",
      "It's an arbitrary compiler rule with no real reason behind it",
      "`List<T>` is covariant too - the assignment actually is allowed"
    ],
    "answer": 0,
    "explain": "Covariance (`out`) is safe only for producer-only types: every value handed out is a Cat, which is an Animal. `List<T>` reads AND writes, so treating a `List<Cat>` as `List<Animal>` would let code Add a Dog into it - which is why `List<T>` is invariant."
  },
  {
    "q": "What does C#'s 'reified generics' (vs Java's type erasure) let you do that Java cannot?",
    "choices": [
      "Use `typeof(T)` and `new T()` at runtime, and store value types like `int` without boxing - because the type info survives to runtime",
      "Write generic methods with type parameters at all - Java has no generics",
      "Run generic code faster by deleting type information before execution",
      "Make `List<int>` and `List<string>` the same runtime type for compatibility"
    ],
    "answer": 0,
    "explain": "C# keeps generic type info at runtime (reification), so `typeof(T)` and `new T()` work and `List<int>` holds raw unboxed ints. Java erases generics, making T unknowable at runtime, forbidding `new T()`, and forcing boxing of primitives."
  }
]
```


---

# Delegates, Lambdas & Events - Functions as Values

Up to now, a method has been something you *call* by name. This phase flips that: a method can also be a **value** - handed to another method, stashed in a variable, called later. Once functions are first-class values, whole categories of code get shorter: sorting by a custom rule, retrying with a strategy, reacting when something happens.

The mental model: **a function can be passed around like data.** Delegates are the typed box you put a function in. Lambdas are the shortest way to write one on the spot. `Func`/`Action`/`Predicate` are the standard boxes everyone uses. Events are delegates with a safety rail bolted on. Get this clicking - [Phase 12: LINQ](12-linq.md) is built entirely on top of it.

## Delegates - a typed reference to a method

📝 **A delegate is a type-safe reference to a method** - a "function pointer, but with the parameter and return types checked by the compiler." Declaring a delegate type describes a *shape* of method: how many parameters, what types, what it returns. Any matching method can be stored in a variable of that delegate type and called through it.

A delegate type is declared like a method signature with the `delegate` keyword in front:

```csharp
// Declare a delegate TYPE: "any method taking two ints and returning an int".
delegate int Op(int a, int b);

class Calculator
{
    static int Add(int a, int b) => a + b;
    static int Multiply(int a, int b) => a * b;

    static void Main()
    {
        Op operation = Add;        // store a method in a delegate variable
        Console.WriteLine(operation(3, 4));   // call it - runs Add(3, 4)

        operation = Multiply;       // point the same variable at another method
        Console.WriteLine(operation(3, 4));   // now runs Multiply(3, 4)
    }
}
```
```console
7
12
```
*What just happened:* `delegate int Op(int a, int b);` defined a new *type* named `Op` whose values are "methods that take two ints and return an int." `Op operation = Add;` stored the `Add` method itself - not its result. Calling `operation(3, 4)` ran whatever method the variable held; reassigning `operation = Multiply` swapped behavior without touching the call site. The compiler enforces the shape throughout - a wrong signature won't compile, the "type-safe" part a raw C function pointer lacks.

💡 **Key point.** The payoff isn't reassigning a variable - it's *passing behavior into other code*. A method can take an `Op` parameter and let the caller decide what operation to run: one sort routine that sorts by any rule, one retry loop that runs any action.

## Lambdas & anonymous methods - functions written inline

Declaring a separate named method just to pass it somewhere is a lot of ceremony for one line of logic. A **lambda** is a function written *inline*, right where you need it, with no name.

📝 The syntax is `parameters => body`. The `=>` reads as "goes to." A single-expression body returns its value automatically (an **expression lambda**). Multiple statements need braces and an explicit `return` (a **statement lambda**).

```csharp
delegate int Op(int a, int b);

class Program
{
    static void Main()
    {
        // Expression lambda: one expression, value returned implicitly.
        Op add = (a, b) => a + b;

        // Statement lambda: braces, multiple lines, explicit return.
        Op maxPlusOne = (a, b) =>
        {
            int bigger = a > b ? a : b;
            return bigger + 1;
        };

        Console.WriteLine(add(3, 4));
        Console.WriteLine(maxPlusOne(3, 4));
    }
}
```
```console
7
5
```
*What just happened:* `(a, b) => a + b` created a function with no name, stored directly in an `Op` delegate - no separate `static int Add(...)` needed. The compiler inferred `a` and `b` as `int` from the `Op` type, so you didn't repeat them. The statement lambda used braces since it has more than one line, needing an explicit `return`. Both are the same idea as named methods, just written at the point of use.

(The older `delegate (int a, int b) { return a + b; }` form is an **anonymous method** - the pre-lambda way to write the same thing. Lambdas replaced it.)

## `Func`, `Action`, `Predicate` - the built-in delegate types

Declaring a custom `delegate` type for every shape gets tedious, and worse, your `Op` and my `Calc` would be incompatible even with identical signatures. .NET ships a standard set of generic delegate types instead - learn these three and you can read almost any modern C# API.

📝 **`Func<…, TResult>`** - a method that **returns a value**; the last type parameter is the return type, the rest are parameters. `Func<int, int, int>` takes two ints and returns an int.

📝 **`Action<…>`** - a method that returns **void**. `Action<string>` takes a string and returns nothing.

📝 **`Predicate<T>`** - a method taking one `T`, returning **`bool`** - a yes/no test. `Predicate<int>` asks a true/false question about an int.

```csharp
// Func: takes two ints, returns an int (return type is LAST).
Func<int, int, int> add = (a, b) => a + b;

// Action: takes a string, returns nothing.
Action<string> shout = msg => Console.WriteLine(msg.ToUpper());

// Predicate: takes an int, returns a bool - a test.
Predicate<int> isEven = n => n % 2 == 0;

Console.WriteLine(add(2, 5));
shout("hello");
Console.WriteLine(isEven(4));
```
```console
7
HELLO
True
```
*What just happened:* `Func<int, int, int>` held a value-returning function - read the type parameters left to right as "int, int → int," return type always last. `Action<string>` held a function with no return value, only a side effect (printing). `Predicate<int>` held a true/false test. Note `shout` and `isEven` skip parentheses around a single parameter - `msg => …` is shorthand for `(msg) => …`. Each is exactly a delegate, just pre-declared so you never write `delegate` yourself.

💡 **Key point.** These three are the *vocabulary* LINQ and modern APIs speak: `.Where(...)` wants a `Func<T, bool>` (a test), `.Select(...)` a `Func<T, TResult>` (a transform), `list.ForEach(...)` an `Action<T>`. A parameter typed `Func` or `Action` wants you to hand it behavior - usually a lambda.

## Closures - a lambda that captures its surroundings

📝 A **closure** is what you get when a lambda uses a variable from the scope where it was *defined*. The lambda doesn't copy the value - it captures the *variable itself*, reading its current value whenever it runs, even long after the surrounding method has returned.

```csharp
Func<int, int> MakeAdder(int amount)
{
    // The returned lambda "captures" the parameter `amount`.
    return x => x + amount;
}

var add10 = MakeAdder(10);
var add100 = MakeAdder(100);

Console.WriteLine(add10(5));    // 15
Console.WriteLine(add100(5));   // 105
```
```console
15
105
```
*What just happened:* `MakeAdder` returned a lambda referring to `amount`, a local parameter that would normally vanish when `MakeAdder` returns - but the lambda *captured* it, keeping it alive bundled with the function. `add10` and `add100` each closed over their own `amount`, which is why they behave differently.

⚠️ **Gotcha - capturing a loop variable.** Because closures capture the *variable*, not a snapshot, capturing a `for` loop's counter bites hard - every lambda shares the one loop variable and sees its *final* value:

```csharp
var funcs = new List<Func<int>>();

// for loop: ONE shared `i` captured by all three lambdas.
for (int i = 0; i < 3; i++)
    funcs.Add(() => i);

foreach (var f in funcs)
    Console.WriteLine(f());     // prints 3, 3, 3 - not 0, 1, 2!

// foreach: modern C# gives each iteration its OWN variable.
var fixedFuncs = new List<Func<int>>();
foreach (var n in new[] { 0, 1, 2 })
    fixedFuncs.Add(() => n);

foreach (var f in fixedFuncs)
    Console.WriteLine(f());     // prints 0, 1, 2 - fresh `n` each time
```
```console
3
3
3
0
1
2
```
*What just happened:* the `for` loop has a single variable `i` that lives for the whole loop; all three lambdas captured *that one variable*, and by the time you called them the loop had finished, leaving `i` at `3`. `foreach` is different: modern C# (5+) gives each iteration a *fresh* loop variable, so each lambda captured its own `n`. Fix the `for` case by copying into a local inside the loop: `int copy = i; funcs.Add(() => copy);` - now each lambda captures a distinct `copy`. ⚠️ `foreach` was fixed; `for` still bites.

## Events - publish/subscribe built on delegates

📝 An **event** is a publish/subscribe mechanism built on delegates. One object (the *publisher*) announces "something happened"; other objects (*subscribers*) register to be notified, attaching with `+=` and detaching with `-=`. The publisher *raises* the event to call everyone at once - the backbone of UI toolkits and reactive code, where a button doesn't know who's listening for its click, it just fires.

```csharp
class Button
{
    // An event: subscribers attach handlers; only Button can raise it.
    public event EventHandler? Clicked;

    public void SimulateClick()
    {
        Console.WriteLine("Button: raising Clicked");
        Clicked?.Invoke(this, EventArgs.Empty);   // notify every subscriber
    }
}

class Program
{
    static void Main()
    {
        var button = new Button();

        // Subscribe two handlers with +=.
        button.Clicked += (sender, e) => Console.WriteLine("Handler A: clicked!");
        button.Clicked += (sender, e) => Console.WriteLine("Handler B: also clicked!");

        button.SimulateClick();
    }
}
```
```console
Button: raising Clicked
Handler A: clicked!
Handler B: also clicked!
```
*What just happened:* `public event EventHandler? Clicked;` declared an event - `EventHandler` is the standard delegate type for "something happened" notifications (it carries a `sender` and an `EventArgs`). Two subscribers attached lambdas with `+=`. `SimulateClick` raised the event with `Clicked?.Invoke(...)` - the `?.` guards against nobody having subscribed (then `Clicked` is null and invoking it would throw). `Button` never named or knew about either handler - that decoupling is the point.

💡 **Key point.** An event is "a delegate with subscribe/unsubscribe semantics and safety." Under the hood it's a delegate holding a list of methods, but the `event` keyword restricts the outside world to only `+=` and `-=`: subscribers can add or remove *their own* handler, but can't overwrite the whole list, clear everyone else's, or raise the event themselves. Only the declaring class can do that - why events, not raw public delegates, model "X happened, react if you care."

## Recap

1. **A delegate is a type-safe reference to a method** - a typed box for a function, pass around, call later. `delegate int Op(int a, int b);` declares the shape; any matching method fits.
2. **Lambdas** write functions inline: `(a, b) => a + b` for a single expression (returned implicitly), or `(a, b) => { … return x; }` with braces for multiple statements.
3. **`Func`, `Action`, `Predicate`** are the built-in delegate types everyone uses: `Func<…,TResult>` returns a value, `Action<…>` returns void, `Predicate<T>` returns bool - the vocabulary LINQ speaks.
4. A **closure** captures the *variable* from its enclosing scope (not a snapshot), keeping it alive. ⚠️ Capturing a `for` loop variable makes every lambda share one variable and see its final value; `foreach` was fixed in modern C#, but `for` still bites - copy to a local inside the loop.
5. An **event** is publish/subscribe on top of delegates: subscribe with `+=`, unsubscribe with `-=`, raise with `?.Invoke(...)`. The `event` keyword adds safety so only the declaring class can raise or replace the handler list.

You can now treat functions as values - store, pass, react. That's the foundation [Phase 12: LINQ](12-linq.md) stands on: every query operator takes a function (a lambda) and applies it to a sequence.

## Quick check

Test yourself on the ideas that LINQ will lean on hardest:

```quiz
[
  {
    "q": "What is a delegate in C#?",
    "choices": [
      "A type-safe reference to a method - you can store a method in it, pass it around, and call it later",
      "A keyword that makes a method run on a background thread",
      "A way to inherit behavior from another class",
      "A read-only property that can't be reassigned"
    ],
    "answer": 0,
    "explain": "A delegate is a typed handle to a method - a 'function pointer with types.' You declare a shape (parameters and return type), store any matching method in a variable of that type, and call behavior through it. That's what lets functions be passed around as values."
  },
  {
    "q": "You want to pass a lambda that takes an `int` and returns a `bool` (a test). Which built-in delegate type fits?",
    "choices": [
      "`Func<int, bool>` (or `Predicate<int>`) - it takes an int and returns a bool",
      "`Action<int>` - it takes an int",
      "`Func<bool, int>` - bool first, then int",
      "`Op<int>` - the standard test delegate"
    ],
    "answer": 0,
    "explain": "`Func<int, bool>` reads 'int → bool': parameters first, return type last. `Predicate<int>` is the same shape. `Action<int>` is wrong because Action returns void. This is exactly the shape LINQ's `.Where(...)` expects."
  },
  {
    "q": "You add three lambdas `() => i` inside a `for (int i = 0; i < 3; i++)` loop, then call them all. What prints?",
    "choices": [
      "3, 3, 3 - all three lambdas captured the same `i`, which ended at 3",
      "0, 1, 2 - each lambda captured the value of `i` at that iteration",
      "0, 0, 0 - the lambdas captured `i` before the loop ran",
      "A compile error - you can't capture a loop variable"
    ],
    "answer": 0,
    "explain": "A closure captures the variable, not a snapshot. A `for` loop has one shared `i`, so all three lambdas point at it; by the time you call them the loop is done and `i` is 3. Copy to a local inside the loop (`int copy = i;`) to fix it. Note `foreach` was fixed in modern C# to give each iteration its own variable."
  }
]
```


---

# LINQ - Querying Anything, Declaratively

In [Phase 11](11-delegates-and-lambdas.md) you learned that a lambda like `n => n > 2` is just a value you can pass around - a packet of behavior wearing a `Func` delegate. That was the warm-up. This phase is the payoff, one of the things C# programmers genuinely brag about: **LINQ** (Language-Integrated Query). Once it clicks, you'll stop writing half the loops you used to write.

The shift in thinking: wanting "the even numbers, doubled" used to mean a loop - make a result list, walk the source, test each item, transform the survivors, append. You described *how* to get the answer, step by step. LINQ flips that: you describe *what* you want - "where it's even, select it doubled" - and let the machinery handle the stepping. That's **imperative** (a recipe of steps) versus **declarative** (a statement of intent).

## The mental model - describe the result, not the steps

📝 **LINQ** - a set of query operators, built into C#, that let you filter, transform, sort, and group *any* sequence using one consistent vocabulary. The same `Where`/`Select`/`OrderBy` words work on an in-memory `List<T>`, a database table, an XML document, or a CSV stream.

Put the two styles side by side on the same task: take a list of numbers, keep the ones greater than 2, and double each. The hand-written loop:

```csharp
var numbers = new List<int> { 1, 2, 3, 4 };

var result = new List<int>();
foreach (var n in numbers)
{
    if (n > 2)
    {
        result.Add(n * 10);
    }
}
// result is now { 30, 40 }
```

*What just happened:* You spelled out every mechanical step - allocate a list, loop, test, transform, append. It works, but the *intent* ("the ones over 2, times ten") is buried under bookkeeping. A reader has to mentally run the loop to see what you meant.

The LINQ version of the exact same thing:

```csharp
var numbers = new List<int> { 1, 2, 3, 4 };

var result = numbers
    .Where(n => n > 2)      // keep the ones over 2
    .Select(n => n * 10)    // transform each survivor
    .ToList();              // gather into a list
// result is now { 30, 40 }
```

*What just happened:* No loop, no temporary list, no manual `Add`. You named the *what* - filter with `Where`, transform with `Select` - and LINQ did the stepping. The two lambdas are exactly the `Func` values from Phase 11: `Where` takes a `Func<int, bool>` (keep-or-not), `Select` a `Func<int, int>` (the new value). LINQ is delegates all the way down.

💡 **The reading test.** The loop says "do this, then this, then this." The LINQ chain says "what I want is this." When the *what* matters more than the *how*, declarative wins.

## Method syntax - the workhorse

The form you just saw is **method syntax**: a chain of method calls, each taking a lambda. This is the one you'll write 90% of the time.

📝 **Method syntax** - calling LINQ operators as extension methods on a sequence, passing lambdas to say what each step does: `source.Where(...).Select(...).OrderBy(...)`. Each returns a new sequence, so they chain like a pipeline.

```csharp
using System.Linq;

var words = new List<string> { "apple", "fig", "banana", "kiwi", "cherry" };

var shortUpper = words
    .Where(w => w.Length <= 5)        // keep short words
    .Select(w => w.ToUpper())         // shout them
    .ToList();

foreach (var w in shortUpper)
    Console.WriteLine(w);
```

```console
APPLE
FIG
KIWI
```

*What just happened:* Read the chain top to bottom as a pipeline. `Where(w => w.Length <= 5)` let through `apple`, `fig`, and `kiwi` (the others are too long); `Select(w => w.ToUpper())` transformed each survivor to uppercase; `ToList()` collected the results. ⚠️ Note the `using System.Linq;` at the top - LINQ's operators are extension methods living in that namespace, invisible without it. Forget it and the compiler insists `List<T>` has no method called `Where`.

## Query syntax - the SQL-flavored sugar

C# offers a second way to write the same queries, and if you've touched SQL it'll look eerily familiar. It's called **query syntax**, and reads like a sentence:

📝 **Query syntax** - keyword-based query expressions (`from`, `where`, `select`, `orderby`, `group`) that the compiler translates into the exact same method-syntax calls under the hood. Pure syntactic sugar: nothing it does is impossible in method syntax.

The first example - keep numbers over 2, double them - written both ways so you can see they're the same query:

```csharp
var numbers = new List<int> { 1, 2, 3, 4 };

// Query syntax - reads like SQL
var viaQuery = from n in numbers
               where n > 2
               select n * 10;

// Method syntax - the same thing the compiler turns the above into
var viaMethod = numbers.Where(n => n > 2).Select(n => n * 10);

Console.WriteLine(string.Join(", ", viaQuery));
Console.WriteLine(string.Join(", ", viaMethod));
```

```console
30, 40
30, 40
```

*What just happened:* Both produce `30, 40` because they *are* the same query - the compiler rewrites `from n in numbers where n > 2 select n * 10` into precisely the `Where(...).Select(...)` chain below it. They're interchangeable; pick by readability. Query syntax shines with joins, grouping, or multiple `from` clauses; method syntax wins for simple chains and for operators - like `Count()`, `First()`, `ToList()` - that query syntax has no keyword for. Most C# code leans on method syntax.

## Deferred execution - the gotcha that bites everyone once

The single most important idea in this phase, and the one that catches every C# developer exactly once: a LINQ query is **lazy**. Writing it doesn't *run* it - it builds a description of work, a recipe, and does nothing until something asks it to produce values.

📝 **Deferred execution** - a LINQ query (the `IEnumerable<T>` you get from `Where`, `Select`, etc.) is not a result; it's a *plan* for computing one. The plan only executes when you **enumerate** it: a `foreach`, or a "consumer" like `ToList()`, `Count()`, `First()`, or `Sum()`. Until then, nothing has happened.

This is the same laziness you'd meet in Python's generators or Rust's iterators - adapters describe, consumers fire. In C# it produces a genuinely surprising result:

```csharp
var source = new List<int> { 1, 2, 3 };

// Build a query - note: this does NOT run yet.
var query = source.Where(n => n > 1);

source.Add(4);   // change the source AFTER defining the query

// NOW enumerate it.
foreach (var n in query)
    Console.WriteLine(n);
```

```console
2
3
4
```

*What just happened:* The `4` showed up even though you added it *after* writing the query, because the query didn't run when you wrote it - it ran when the `foreach` enumerated it, and by then the source list already had `4`. The query is a live recipe pointed at `source`, not a snapshot taken at definition time.

⚠️ **And it re-runs every single time you enumerate it.** A deferred query is not a cached result. Loop over it twice and the whole pipeline executes twice - every `Where` test, every `Select` transform, fresh. If those lambdas hit a database or do real work, you've silently paid for it twice (or N times).

```csharp
int calls = 0;
var numbers = new List<int> { 1, 2, 3 };

var query = numbers.Where(n =>
{
    calls++;                 // count how often the filter actually runs
    return n > 1;
});

var firstCount  = query.Count();   // enumeration #1
var secondCount = query.Count();   // enumeration #2

Console.WriteLine($"filter ran {calls} times for {firstCount} results");
```

```console
filter ran 6 times for 2 results
```

*What just happened:* Three items, two enumerations - so the filter lambda ran *six* times, not three. Each `Count()` walked the entire pipeline again from scratch: wasteful here, dangerous when the lambda is expensive. Fix it by **materializing** the query once into a real collection with `ToList()` (or `ToArray()`), which forces a single enumeration and stores the results.

```csharp
int calls = 0;
var numbers = new List<int> { 1, 2, 3 };

// .ToList() forces the query to run ONCE, right here, and stores the results.
var results = numbers.Where(n => { calls++; return n > 1; }).ToList();

var firstCount  = results.Count;   // just reads the stored list - no re-run
var secondCount = results.Count;

Console.WriteLine($"filter ran {calls} times for {firstCount} results");
```

```console
filter ran 3 times for 2 results
```

*What just happened:* `ToList()` executed the pipeline exactly once - one pass over all three source items, so the filter ran three times and let two through. Now `results` is concrete data, not a recipe - `results.Count` just reads the stored length, and the filter never runs again. Rule of thumb: enumerating more than once, or needing results to reflect the source *as it is now*, means ending the chain with `.ToList()`.

## The powerful operators - a realistic query

You've met `Where` and `Select`. Here's the rest of the toolkit you'll reach for, then a realistic example chaining several together.

- **`OrderBy` / `OrderByDescending`** - sort by a key (`ThenBy` for tie-breakers).
- **`GroupBy`** - bucket items by a key; you get groups you can loop over, each with its own `Key`.
- **`First` / `FirstOrDefault`** - grab the first item (optionally matching a condition).
- **`Any` / `All`** - does *any* item match? do *all* of them?
- **`Sum` / `Count` / `Average` / `Max`** - the aggregates, computed in one pass.
- **`SelectMany`** - flatten a sequence-of-sequences into one flat sequence.

Group a list of orders by customer and total each customer's spend - the kind of thing you'd otherwise write fifteen lines of loop for:

```csharp
var orders = new[]
{
    new { Customer = "Ada",  Amount = 30 },
    new { Customer = "Lin",  Amount = 50 },
    new { Customer = "Ada",  Amount = 20 },
    new { Customer = "Lin",  Amount = 10 },
    new { Customer = "Ada",  Amount = 15 },
};

var summary = orders
    .GroupBy(o => o.Customer)                       // bucket by customer
    .Select(g => new { Name = g.Key, Total = g.Sum(o => o.Amount) })
    .OrderByDescending(x => x.Total)               // biggest spender first
    .ToList();

foreach (var row in summary)
    Console.WriteLine($"{row.Name}: {row.Total}");
```

```console
Ada: 65
Lin: 60
```

*What just happened:* `GroupBy(o => o.Customer)` split the five orders into two buckets keyed by name - Ada's three and Lin's two. `Select` turned each bucket `g` into a summary object using `g.Key` (the customer name) and `g.Sum(o => o.Amount)` (that bucket's total). `OrderByDescending` sorted so the biggest spender leads, and `ToList()` froze the answer. Grouping, aggregation, and sorting in four readable lines.

⚠️ **`First` throws; `FirstOrDefault` doesn't.** A constant source of crashes: `First()` on a sequence with no match throws an `InvalidOperationException` ("Sequence contains no matching element"). When a miss is possible - most real queries - use `FirstOrDefault()`, which returns the type's default (`null` for reference types, `0` for `int`) instead, then check for it.

```csharp
var names = new List<string> { "Ada", "Lin" };

var found   = names.FirstOrDefault(n => n.StartsWith("L"));  // "Lin"
var missing = names.FirstOrDefault(n => n.StartsWith("Z"));  // null, not a crash

Console.WriteLine(found ?? "(none)");
Console.WriteLine(missing ?? "(none)");
```

```console
Lin
(none)
```

*What just happened:* The first call found "Lin." The second matched nothing, and because it's `FirstOrDefault`, it calmly returned `null` instead of throwing; `?? "(none)"` handled the null gracefully. Plain `First(n => n.StartsWith("Z"))` would have crashed on that line.

💡 **The same LINQ can run against a database.** A `List<T>` in memory means **LINQ to Objects** - `IEnumerable<T>`, lambdas running as C# code. A database table via an ORM like Entity Framework means an `IQueryable<T>` - the *identical* `Where`/`Select`/`GroupBy` syntax gets **translated into SQL** and run by the database engine. Same words, different engine: `users.Where(u => u.Age > 18)` filters a `List` in RAM or generates `WHERE Age > 18` against Postgres depending only on what `users` is. ⚠️ Not every C# expression translates to SQL (a call to your own helper method has no SQL equivalent), so `IQueryable` queries have rules `IEnumerable` ones don't - a story for the Entity Framework guide.

## Recap

1. **LINQ is declarative**: describe the result you want (`Where`/`Select`/`OrderBy`) instead of hand-writing the loop. One vocabulary queries lists, databases, XML, and more.
2. **Method syntax** (`.Where(...).Select(...)`) is the everyday workhorse, built on `Func` lambdas from Phase 11. **Query syntax** (`from ... where ... select ...`) is SQL-flavored sugar the compiler rewrites into method syntax - pick whichever reads better.
3. ⚠️ **Deferred execution** is the big trap: a query is a recipe, not a result. It does nothing until enumerated (`foreach`, `ToList`, `Count`, `First`), reads its source *late*, and **re-runs every enumeration**. Call `.ToList()` to materialize it once and stop the surprises.
4. The toolkit: `OrderBy`, `GroupBy`, `First`/`FirstOrDefault`, `Any`/`All`, `Sum`/`Count`/`Average` - chained together they replace whole loops, grouping and totaling in a few readable lines.
5. ⚠️ `First` throws on no match; prefer **`FirstOrDefault`** (returns `null`/`0`) whenever a miss is possible, then check the default.
6. 💡 The same LINQ runs on `IEnumerable` (in memory, as C#) or `IQueryable` (a database, translated to SQL) - same syntax, different engine.

You can now query collections the way you'd describe them to a colleague. Next: **records and modern pattern matching** - features that make the *data* you query just as concise as the queries themselves.

## Quick check

Test yourself on the idea that trips up everyone - laziness - plus the two operators most likely to bite you:

```quiz
[
  {
    "q": "You write `var q = list.Where(n => n > 1);` and then add an item to `list` before looping over `q`. The new item appears in the loop. Why?",
    "choices": [
      "Deferred execution - the query is a recipe that only runs when enumerated, so it reads `list` as it is at loop time, not at definition time",
      "LINQ automatically copies the list when you call Where, then keeps it in sync",
      "Where always re-reads the source from disk on every call",
      "It's a bug; the new item should not appear"
    ],
    "answer": 0,
    "explain": "A LINQ query is lazy. Defining it builds a plan but runs nothing; the plan executes when you enumerate (here, the foreach), reading the source at that moment. Since the item was added before enumeration, it's included. Call .ToList() at definition time if you want a snapshot."
  },
  {
    "q": "What's the practical difference between `Where(...).Count()` called twice and `Where(...).ToList()` called once?",
    "choices": [
      "The deferred query re-runs the entire pipeline on each Count(); ToList() executes it once and stores the results, so later reads don't re-run anything",
      "There is no difference; both run the filter exactly once",
      "ToList() is slower because lists use more memory than queries",
      "Count() caches its result, so the second call is free, unlike ToList()"
    ],
    "answer": 0,
    "explain": "A deferred query is not cached - each enumeration (each Count()) re-executes every Where and Select from scratch. ToList() forces a single execution and materializes a real collection, so subsequent reads are just reading stored data with no re-run."
  },
  {
    "q": "You query a list for the first name starting with 'Z', but none exist. Which call avoids crashing?",
    "choices": [
      "FirstOrDefault(n => n.StartsWith(\"Z\")) - it returns null instead of throwing",
      "First(n => n.StartsWith(\"Z\")) - it returns null when nothing matches",
      "Both throw, so you must wrap either one in try/catch",
      "Any(n => n.StartsWith(\"Z\")) - it returns the first match or null"
    ],
    "answer": 0,
    "explain": "First throws InvalidOperationException when no element matches. FirstOrDefault returns the type's default (null for a string) instead, so it's the safe choice whenever a miss is possible - just remember to check for that default afterward."
  }
]
```


---

# Records, Pattern Matching & Modern C# - Less Boilerplate, More Safety

Back in [Phase 5](05-classes-and-objects.md) you wrote a `class` by hand: fields, a constructor that copies parameters into them, maybe an overridden `ToString`, and - if you wanted two objects with the same data to count as equal - a pile of equality code you didn't even attempt. A lot of typing for a simple idea: "a little bundle of values."

Modern C# noticed, growing features whose whole purpose is to *delete boilerplate* while making code *safer*. The frame is the same throughout: **here's the old verbose way, and here's the new concise way the compiler now writes for you.**

The two headliners - **records** and **pattern matching** - pair up beautifully. Records let you *define* small immutable data types in a line; pattern matching lets you *ask questions* about that data - "what shape is this? what's inside it?" - without a ladder of `if`/`else`. Together they're C#'s answer to modeling "this value is one of a fixed set of possibilities" cleanly.

## Records - immutable data types in one line

📝 **Record** - a type (class or struct) declared with the `record` keyword, designed for holding data. From a single line, the compiler generates the constructor, the properties, value-based `Equals` and `GetHashCode`, a readable `ToString`, and deconstruction support. You write the *intent*; the compiler writes the ceremony.

Remember the `Account ==` gotcha from Phase 5? Two class objects with identical data were *not* equal, because classes compare by reference rather than by value. Records flip that default - built for exactly the "a point is its X and Y, nothing more" value that *should* compare by contents.

The old way - a hand-written class for a simple point:

```csharp
class PointClass
{
    public int X { get; }
    public int Y { get; }

    public PointClass(int x, int y)
    {
        X = x;
        Y = y;
    }

    public override string ToString() => $"PointClass {{ X = {X}, Y = {Y} }}";

    public override bool Equals(object? obj) =>
        obj is PointClass p && p.X == X && p.Y == Y;

    public override int GetHashCode() => HashCode.Combine(X, Y);
}
```
*What just happened:* Roughly twenty lines to say "a point is two ints, and two points with the same numbers are equal." The constructor copies parameters into properties, `ToString` builds a readable string, and `Equals`/`GetHashCode` do the value comparison by hand - easy to get subtly wrong (forget a field and equality silently breaks).

Now watch it collapse:

```csharp
record Point(int X, int Y);
```
*What just happened:* That one line - a **positional record** - generates *all* of the above. `(int X, int Y)` is the **primary constructor**: it declares two init-only properties `X` and `Y` *and* a constructor that fills them. Value equality, a tidy `ToString`, and deconstruction, for free. Twenty lines became one, and the generated version can't drift out of sync the way hand-written equality does.

```csharp
var a = new Point(1, 2);
var b = new Point(1, 2);

Console.WriteLine(a);          // ToString, generated
Console.WriteLine(a == b);     // value equality, generated
Console.WriteLine(a.X);        // property, generated

var (x, y) = a;                // deconstruction, generated
Console.WriteLine($"x={x}, y={y}");
```
```console
Point { X = 1, Y = 2 }
True
1
x=1, y=2
```
*What just happened:* `a == b` is `True` - the opposite of the class behavior in Phase 5 - because records compare by *contents*. `Console.WriteLine(a)` printed a readable summary instead of the type name, and `var (x, y) = a` pulled the two values straight out (deconstruction), no code written.

**Non-destructive mutation with `with`.** Records are immutable by default - those generated properties are init-only, so you can't change a `Point` after building it. For "the same thing, with one field different," use the `with` expression: it makes a *copy*, changing only what you name, leaving the original untouched.

```csharp
var p1 = new Point(1, 2);
var p2 = p1 with { Y = 99 };   // a copy of p1, but Y is 99

Console.WriteLine(p1);          // unchanged
Console.WriteLine(p2);          // the modified copy
```
```console
Point { X = 1, Y = 2 }
Point { X = 1, Y = 99 }
```
*What just happened:* `p1 with { Y = 99 }` produced a brand-new `Point` carrying `p1`'s `X` and an overridden `Y`. `p1` itself never changed - *non-destructive* mutation: the convenience of "tweak this value" without the bugs of objects mutating under you while other code holds a reference.

💡 **When to reach for a record.** Use records for **data defined by its values** - DTOs, value objects (`Money`, `Coordinate`, `DateRange`), events, query results. Use a regular `class` for things with *identity and behavior* that mutate over their lifetime - a `ShoppingCart`, a database connection, a game's `Player`. Quick test: if "two of these are the same when their contents match," it wants to be a record.

📝 **`record struct`.** `record` defaults to a reference type (a class). Write `record struct Point(int X, int Y);` for the same conveniences on a *value type* - copied on assignment, lives on the stack when local, no heap allocation. One difference to know: a positional `record struct`'s properties are mutable (`get; set;`) by default, so it is *not* immutable like a `record` class; write `readonly record struct Point(int X, int Y);` if you want the same init-only immutability. Reach for it for tiny, short-lived values to avoid heap pressure; plain `record` is the right default until you've measured a reason to switch.

## `init` and `required` members - immutability without a giant constructor

Records bake in init-only properties, but you can use the same tools on any class. You met `init` briefly in Phase 5 - here's why it matters.

📝 **`init` accessor** - like `set`, but it can *only* run during object construction (in a constructor or an object initializer); after that, the property is frozen. **`required` member** - a property the compiler *forces* the caller to set when creating the object, or the code won't compile.

The old tension: object initializers (`new Thing { A = 1, B = 2 }`) read beautifully, but a plain `set` property leaves the object mutable forever, and nothing stops a caller forgetting a field. `init` + `required` resolve both - initializer syntax, guaranteed-set fields, immutability afterward.

```csharp
class Config
{
    public required string Host { get; init; }   // must be set; frozen after
    public int Port { get; init; } = 8080;        // optional, with a default
}

var c = new Config { Host = "localhost", Port = 5432 };
Console.WriteLine($"{c.Host}:{c.Port}");
// c.Host = "other";   // compile error: init-only, can't assign after construction
// var bad = new Config { Port = 1 };  // compile error: required 'Host' not set
```
```console
localhost:5432
```
*What just happened:* `required string Host` made the compiler refuse any `new Config { ... }` that doesn't set `Host` - a compile error instead of a runtime `null`. Both properties are `init`, so once `c` is built its values are locked - readability of object initializers plus safety of immutability, no sprawling constructor.

## Pattern matching - asking "what shape is this data?"

**Pattern matching** is C#'s way of testing a value's *shape* - its type, contents, structure - and pulling pieces out of it in one concise expression. It replaces tall stacks of `if (x is SomeType) { var y = (SomeType)x; if (y.Prop > 5) ... }`.

**The `is` pattern.** The simplest form tests a type *and* binds a variable in one move:

```csharp
object o = "hello";

if (o is string s)                 // is it a string? if so, call it s
    Console.WriteLine(s.Length);   // s is already typed as string here
```
```console
5
```
*What just happened:* `o is string s` did two jobs at once: checked whether `o` is a `string`, and if so assigned it to a properly-typed variable `s` you can use immediately - no separate cast, no second declaration. The old way was a type check followed by an explicit `(string)o` cast; this fuses them.

**`switch` expressions with patterns.** The real power shows up in the `switch` *expression* (produces a value, distinct from the older `switch` statement): it matches a value against a series of patterns and returns the arm that fits.

```csharp
static string Classify(object value) => value switch
{
    int n when n < 0      => "negative int",
    int n                 => $"int: {n}",          // type pattern, binds n
    string { Length: 0 }  => "empty string",        // property pattern
    string s              => $"string of {s.Length}",
    null                  => "null",
    _                     => "something else"       // _ is the catch-all
};

Console.WriteLine(Classify(42));
Console.WriteLine(Classify(-3));
Console.WriteLine(Classify("hi"));
Console.WriteLine(Classify(""));
Console.WriteLine(Classify(null));
Console.WriteLine(Classify(3.14));
```
```console
int: 42
negative int
string of 2
empty string
null
something else
```
*What just happened:* The `switch` expression tested `value` top to bottom and returned the first matching arm. `int n` is a **type pattern** binding the matched value to `n`. `when n < 0` adds a **guard** - an extra condition. `string { Length: 0 }` is a **property pattern**: matches a string *and* checks its `Length` is 0. `_` is the discard, the catch-all. Compare this to a nest of `if`/`else if` with manual casts: same logic, a fraction of the noise, and the compiler warns if you've left a case unhandled.

**Relational and logical patterns.** Inside a pattern you can use `<`, `>`, `<=`, `>=` and combine patterns with `and`, `or`, `not` - reading like the math you'd say out loud:

```csharp
static string Grade(int score) => score switch
{
    < 0 or > 100 => "invalid",
    >= 90        => "A",
    >= 80        => "B",
    >= 70        => "C",
    _            => "F"
};

Console.WriteLine(Grade(95));
Console.WriteLine(Grade(72));
Console.WriteLine(Grade(150));
```
```console
A
C
invalid
```
*What just happened:* `< 0 or > 100` is a **logical pattern** combining two **relational patterns**. The arms below lean on order: once `>= 90` fails, `>= 80` only sees scores under 90, so you don't repeat the upper bound - clearer than `if (score < 0 || score > 100)` chains.

**Property and positional patterns - matching deep into data.** Property patterns shine on records; you can match nested fields, and because records support deconstruction you can also use **positional patterns** that match by position:

```csharp
record Person(string Name, int Age);

static string Describe(Person p) => p switch
{
    { Age: > 64 }            => $"{p.Name} is a senior",
    { Age: >= 18 and < 65 }  => $"{p.Name} is an adult",
    ("Sam", _)               => "Hi Sam, whatever your age",   // positional pattern
    { Age: < 0 }             => "invalid age",
    _                        => $"{p.Name} is a minor"
};

Console.WriteLine(Describe(new Person("Ada", 70)));
Console.WriteLine(Describe(new Person("Bo", 30)));
Console.WriteLine(Describe(new Person("Sam", 5)));
Console.WriteLine(Describe(new Person("Kit", 10)));
```
```console
Ada is a senior
Bo is an adult
Hi Sam, whatever your age
Kit is a minor
```
*What just happened:* `{ Age: > 64 }` is a **property pattern** reaching into the record's `Age`. `{ Age: >= 18 and < 65 }` combines a property pattern with a logical-relational pattern. `("Sam", _)` is a **positional pattern**: it deconstructs the `Person` by position (`Name`, `Age`), matching when the first slot is `"Sam"` and the second is anything (`_`).

💡 **Why this pairing matters.** Records + pattern matching is how C# models "this value is one of a fixed set of cases" - a shape that's an `Empty`, a `Circle`, or a `Rectangle`; a result that's a `Success` or a `Failure`. Define the cases as records, then `switch` over them with type and property patterns; the compiler can tell you when you've forgotten a case.

## Nullable reference types - the compiler hunts your null bugs

📝 **Nullable reference types (NRT)** - a compiler feature (on by default in new projects via `<Nullable>enable</Nullable>` in the `.csproj`) that splits reference types into two: `string` means "never null," `string?` means "might be null." The compiler then *warns* you wherever you might dereference something that could be null.

You met this in [Phase 2](02-syntax-values-and-types.md). The single most common crash in C# history is the `NullReferenceException` - calling `.Length` or `.Name` on something that turned out to be `null`. NRT's goal: move that discovery from *3 a.m. in production* to *right now, as a squiggle in your editor*.

```csharp
// With <Nullable>enable</Nullable>:
string name = "Ada";       // non-null: the compiler guarantees it
string? maybe = null;      // nullable: explicitly allowed to be null

Console.WriteLine(name.Length);    // fine - name can't be null

// Console.WriteLine(maybe.Length);   // WARNING: 'maybe' may be null here

if (maybe is not null)             // once you check...
    Console.WriteLine(maybe.Length);   // ...the warning is gone - compiler knows it's safe
```
```console
3
```
*What just happened:* `string name` is non-nullable, so the compiler lets you use `name.Length` freely. `string? maybe` is nullable, so dereferencing it without a check earns a *compile-time warning*. After `if (maybe is not null)`, the compiler's **flow analysis** understands `maybe` can't be null inside that block, so the warning disappears.

**The null-forgiving operator `!`.** Occasionally you know more than the compiler - certain a value isn't null even though it can't prove it. The `!` operator says "trust me, this isn't null" and silences the warning:

```csharp
string? maybe = GetItMaybe();
string definitely = maybe!;        // 'I promise this isn't null'
```
*What just happened:* `maybe!` tells the compiler to treat the value as non-null and drop the warning. ⚠️ Use this *rarely*, only when genuinely certain - it's a promise *you* make, not a check. If you're wrong, you get the very `NullReferenceException` the feature was meant to prevent, safety net deliberately switched off.

⚠️ **NRT is compile-time analysis, not a runtime guarantee.** `string` (non-nullable) is a *promise the compiler tries to verify*, not a wall enforced when the program runs. A `null` can still sneak in - from older code compiled without NRT, JSON deserialization, reflection, a library that ignores the annotations. NRT makes nulls *visible* and *much rarer*, not *impossible* - treat the warnings as an early-warning system, not a force field.

## Other modern niceties - small wins that add up

A grab-bag of recent features that quietly trim noise from everyday code, together making modern C# read cleaner than a few years ago.

**Target-typed `new()`.** When the type is already obvious from the left side, you don't repeat it on the right:

```csharp
// old: type written twice
List<string> names1 = new List<string>();

// new: the compiler infers it from the declared type
List<string> names2 = new();
Dictionary<string, int> counts = new();
Point origin = new(0, 0);
```
*What just happened:* `new()` fills in the type from the variable's declared type on the left - unambiguous, shorter, no `List<string>` echoed twice.

**`using` declarations.** Phase 11's `using` statement needed braces and an extra indent. A `using` *declaration* drops both - the resource is disposed automatically at the end of the enclosing scope:

```csharp
// old: using statement, extra braces and nesting
using (var reader = new StreamReader("data.txt"))
{
    Console.WriteLine(reader.ReadLine());
}

// new: using declaration - disposed at end of the method, no braces
using var reader2 = new StreamReader("data.txt");
Console.WriteLine(reader2.ReadLine());
```
*What just happened:* `using var reader2 = ...` schedules `reader2.Dispose()` for the end of the current scope, exactly like the block form, minus the braces and indentation.

**File-scoped namespaces.** Declare the namespace once with a semicolon and skip the wrapping braces that indented your whole file:

```csharp
// old: braces wrap (and indent) the entire file
namespace MyApp
{
    class Thing { }
}

// new: file-scoped - one line, no indentation tax
namespace MyApp;

class Thing { }
```
*What just happened:* `namespace MyApp;` applies to the entire file, so everything below sits at the left margin instead of indented inside braces - the modern default, since almost every file has exactly one namespace.

**Top-level statements (recap).** As seen in Phase 1, a simple program doesn't need `class Program { static void Main() { ... } }` scaffolding - write executable statements straight in `Program.cs`, and the compiler generates `Main` for you.

## Recap

1. **Records** (`record Point(int X, int Y);`) generate the constructor, properties, value equality, `ToString`, and deconstruction from one line. They compare by *contents* (fixing the Phase 5 `==` gotcha), are immutable by default, and support `with` for non-destructive copies. Use for DTOs and value objects; `record struct` for tiny value-type versions.
2. **`init` and `required`** give object-initializer syntax with immutability: `init` freezes a property after construction, `required` forces callers to set it - "forgot a field" becomes a compile error.
3. **Pattern matching** - `is` patterns, `switch` expressions, and type/property/relational/logical/positional patterns - tests a value's *shape* and extracts its parts declaratively, replacing tall `if`/cast ladders.
4. **Nullable reference types** split `string` (non-null) from `string?` (nullable) and warn at compile time about possible null dereferences; `!` silences a warning when certain (use rarely). ⚠️ Compile-time analysis, **not** a runtime guarantee - nulls can still slip in.
5. **Modern niceties** - target-typed `new()`, `using` declarations, file-scoped namespaces, top-level statements - each delete a bit of repetition.
6. 💡 **Records + pattern matching together** model "one of a fixed set of cases" - define the cases as records, `switch` over them by shape.

You can now write modern C# that's both shorter and safer than the verbose style - less boilerplate to maintain, more bugs caught before you run. Next: **async/await and Tasks**, the feature that lets your programs do many things at once without freezing.

## Quick check

Test yourself on the two ideas that define this phase - records' value equality and pattern matching:

```quiz
[
  {
    "q": "You define `record Point(int X, int Y);`, then create `var a = new Point(1, 2);` and `var b = new Point(1, 2);`. What does `a == b` return, and why?",
    "choices": [
      "True - records compare by value (their contents), so two points with the same X and Y are equal",
      "False - like all C# types, == compares whether they're the same object in memory",
      "A compile error, because you can't use == on a record",
      "True, but only because they were created in the same method"
    ],
    "answer": 0,
    "explain": "Records generate value-based equality, so `a == b` is True when their contents match - the opposite of a plain class, where == compares object identity. This is exactly why records suit DTOs and value objects."
  },
  {
    "q": "What does the pattern `{ Age: >= 18 and < 65 }` match in a switch expression over a record?",
    "choices": [
      "A value whose Age property is at least 18 and less than 65 (a property pattern combined with relational/logical patterns)",
      "A value that equals the literal 18 or 65",
      "Any record that has a property named Age, regardless of its value",
      "A list of ages between 18 and 65"
    ],
    "answer": 0,
    "explain": "It's a property pattern (`{ Age: ... }`) reaching into the record's Age, combined with relational patterns (`>= 18`, `< 65`) joined by the logical pattern `and`. It matches when Age is in the range 18 to 64 inclusive."
  },
  {
    "q": "Under `<Nullable>enable</Nullable>`, what does declaring a variable as `string` (no `?`) actually give you?",
    "choices": [
      "A compile-time promise the compiler tries to verify (warning you about possible null dereferences) - not a runtime guarantee that it can never be null",
      "A runtime guarantee that the value can never, under any circumstances, be null",
      "Exactly the same behavior as `string?` - the `?` is purely cosmetic",
      "A value that automatically converts null into an empty string"
    ],
    "answer": 0,
    "explain": "Nullable reference types are compile-time flow analysis: `string` means 'intended non-null' and the compiler warns where a null might slip through. But it's not enforced at runtime - nulls can still arrive from deserialization, reflection, or older code, so a NullReferenceException remains possible."
  }
]
```


---

# async/await & Tasks - Concurrency Without the Pain

You've written `var data = File.ReadAllText(path);` and it just worked. But while that line ran, your program sat on a thread, fully employed, doing absolutely nothing - staring at the disk, waiting for bytes to arrive. Multiply that across a web server handling a thousand requests, each parked on a thread waiting for a database, and you've got a thousand threads burning memory to wait. That's the problem `async`/`await` solves.

C# pioneered this syntax back in 2012, and nearly every language since (JavaScript, Python, Rust, Swift) borrowed it. Worth getting the *mental model* right, not just the keywords - the model underneath is where the real understanding lives.

The one idea to carry through this whole phase: **`await` lets the current thread go do other useful work while you wait, then resumes you later.** It is not a blocking wait, and not (by itself) a new thread.

## The problem - a blocked thread is a wasted thread

When you call a synchronous I/O method - reading a file, hitting a database, calling a web API - the thread that runs it can't do anything else until the bytes come back. Network and disk are *slow* compared to a CPU: a database query taking 50ms is an eternity in which a modern core could execute hundreds of millions of instructions. Instead, the thread just blocks.

A thread isn't free - each costs around a megabyte of stack memory plus scheduling overhead. On a server, "one blocked thread per in-flight request" is exactly how you run out of threads and grind to a halt.

📝 **Async ≠ parallel.** The distinction that trips everyone up. *Parallelism* is doing multiple things *at the same instant* on multiple CPU cores. *Asynchrony* is dealing with work that *completes later* - like I/O - without holding a thread hostage while you wait. Async is about *not wasting a thread on waiting*; parallelism is about *using more threads to go faster*. You can have one without the other - more on this at the end.

💡 **Key point.** The win from async I/O isn't that any single operation finishes faster - the database is just as slow either way. The win is that the waiting thread is *freed up* to serve other work meanwhile. Same hardware, far more throughput.

## `Task` and `Task<T>` - work that will finish later

Before `await` makes sense, you need the thing it waits *on*: a `Task`.

📝 **`Task`** - an object representing an operation that will complete in the future. C#'s name for what other languages call a *promise* or *future*. A bare `Task` produces no value (it just finishes); a `Task<T>` eventually hands back a value of type `T`. Think of it as a receipt: "your result isn't ready yet, but here's a handle to collect it when it is."

A method doing asynchronous work *returns a Task* instead of the value directly, so the caller gets the receipt immediately and decides when to wait for the real result.

```csharp
using System.Net.Http;

// Returns a Task<string> - the receipt - right away.
// The actual download finishes later.
Task<string> DownloadHomepageAsync(HttpClient client)
{
    return client.GetStringAsync("https://example.com");
}
```

*What just happened:* `GetStringAsync` kicks off a network download and hands back a `Task<string>` instantly - long before the HTML arrives. `DownloadHomepageAsync` passes that receipt straight up to its own caller; nothing has blocked. Somewhere up the chain, someone will turn that receipt into an actual string - that's where `await` comes in.

⚠️ **A `Task` is not a thread.** Creating a `Task` for I/O does *not* spin up a background thread to sit and wait - the OS signals completion via an I/O callback when the bytes land. (You *can* put CPU work on a thread with `Task.Run`, covered later, but that's a different use of `Task`.) Conflating "Task" with "thread" is the root of most async confusion.

## `async`/`await` - what `await` really does

You mark a method `async` and give it a return type of `Task` or `Task<T>`. Inside, you use `await` on any Task. The crucial part - what `await` *actually* does:

When execution hits `await someTask`:
1. If the task is already done, it grabs the result and keeps going.
2. If not, the method **suspends** - returning control to *its* caller right then. The thread is now free to do anything else.
3. When the task completes, the rest of your method (everything after the `await`) runs as a **continuation** - picking up exactly where it left off, locals intact.

`await` is *not* `Thread.Sleep`, not a busy-wait - it's "pause me, free the thread, wake me back up when the result is ready."

```mermaid
flowchart TD
  A[Caller calls GetPriceAsync] --> B[Run until 'await httpTask']
  B --> C{Task done?}
  C -->|no| D[Suspend method<br/>return to caller<br/>thread is FREE]
  D --> E[...time passes,<br/>download completes...]
  E --> F[Continuation scheduled:<br/>resume after the await]
  C -->|yes| F
  F --> G[Run rest of method<br/>return result via Task]
```

In action, with prints so you can see the suspend-and-resume happen:

```csharp
using System;
using System.Threading.Tasks;

async Task<int> GetNumberAsync()
{
    Console.WriteLine("2: inside, before await");
    await Task.Delay(100);                 // simulates slow I/O - suspends here
    Console.WriteLine("4: inside, after await (resumed)");
    return 42;
}

Console.WriteLine("1: before calling");
Task<int> task = GetNumberAsync();         // runs up to the await, then returns
Console.WriteLine("3: after calling, before awaiting result");
int result = await task;                   // wait for completion, unwrap the value
Console.WriteLine($"5: got {result}");
```
```console
1: before calling
2: inside, before await
3: after calling, before awaiting result
4: inside, after await (resumed)
5: got 42
```

*What just happened:* Calling `GetNumberAsync()` did *not* run the whole method - it ran synchronously up to `await Task.Delay(100)`, then suspended and handed control back, which is why `"3: after calling"` prints *before* `"4: after await"`. While `Task.Delay` ticked, the thread was free; when the delay completed, the continuation fired, `"4"` printed, and `42` flowed back through the Task that `await task` unwrapped into `result`. The interleaved order proves `await` pauses and resumes rather than blocking.

💡 **`await` unwraps the result and rethrows exceptions.** Two jobs in one keyword. `await task` on a `Task<int>` gives you the `int` directly, no `.Result` needed. If the awaited operation *threw*, `await` rethrows that exception right at the `await` line, so an ordinary `try`/`catch` handles it as if the code were synchronous - the magic that makes async code *read* like normal sequential code.

```csharp
try
{
    string html = await client.GetStringAsync("https://does-not-exist.invalid");
}
catch (HttpRequestException ex)
{
    Console.WriteLine($"download failed: {ex.Message}");
}
```

*What just happened:* The network call failed inside the awaited Task, so the exception was captured on the Task. When `await` saw a faulted Task, it rethrew the original `HttpRequestException` at the `await` line, letting an ordinary `catch` handle it. Without `await`, that exception would sit silently on the Task object, easy to miss - why you almost always *await* your tasks rather than letting them dangle.

## The pitfalls - async void, deadlocks, and fire-and-forget

Async is wonderful right up until you hit one of these. Each bites essentially every C# developer exactly once. Here they are, before they bite you.

### ⚠️ `async void` - almost never what you want

You *can* write `async void` instead of `async Task`. Don't, with one exception.

The problem: an `async void` method returns *nothing to await*. The caller can't wait for it, can't know when it finished, and - worst of all - **can't catch its exceptions**, which have nowhere to go; they're raised on whatever context is current and typically crash the process.

```csharp
async void DoWorkBad()          // ⚠️ exceptions here are unobservable
{
    await Task.Delay(10);
    throw new InvalidOperationException("boom");   // crashes - no one can catch this
}

async Task DoWorkGood()         // ✅ caller can await AND catch
{
    await Task.Delay(10);
    throw new InvalidOperationException("boom");   // surfaces normally via await
}
```

*What just happened:* Both methods throw, but the outcomes differ. `DoWorkBad` returns `void`, so its caller has no Task to await - the exception escapes and brings the program down. `DoWorkGood` returns a `Task`, so `await DoWorkGood()` receives the exception at the `await` and can `try`/`catch` it. The rule: **`async` methods return `Task` or `Task<T>`.** The *only* legitimate `async void` is an event handler, since the event signature demands a `void` return - and even there, `try`/`catch` inside it.

### ⚠️ Sync-over-async - the classic deadlock

The nastiest one: calling an async method, then *blocking* on its result with `.Result` or `.Wait()` instead of awaiting it. In UI apps (and older ASP.NET), this can deadlock your program solid.

The mechanism: some environments have a **SynchronizationContext** - a rule that says "continuations must resume on a *specific* thread" (the UI thread, so you can safely touch controls). Block that thread on `.Result` and it sits waiting for the task - but the task's continuation needs *that same thread* to resume, and it's busy blocking. Each waits on the other. Frozen forever.

```csharp
// In a UI app or legacy ASP.NET context - this DEADLOCKS:
async Task<string> GetDataAsync()
{
    await Task.Delay(100);          // continuation wants the UI thread back
    return "done";
}

void OnButtonClick()
{
    string data = GetDataAsync().Result;   // ⚠️ UI thread blocks waiting...
    // ...for a continuation that needs the UI thread. Deadlock.
    Console.WriteLine(data);               // never reached
}
```

*What just happened:* `OnButtonClick` ran on the UI thread and called `.Result`, which *blocks* that thread until the task finishes. But `GetDataAsync`'s continuation was scheduled to resume *on the UI thread* - now frozen inside `.Result`, unable to run anything. The task can never complete, so `.Result` never returns. The fix: **don't block - await all the way up.** Make `OnButtonClick` an `async void` event handler and write `string data = await GetDataAsync();`.

### 💡 `ConfigureAwait(false)` - for library code

The deadlock above exists because the continuation insists on returning to the original context. If your code doesn't need that context - library code usually doesn't - tell `await` to skip it:

```csharp
public async Task<string> FetchAsync(HttpClient client)
{
    // In a reusable library: we don't care which thread resumes us.
    string html = await client.GetStringAsync("https://example.com")
                              .ConfigureAwait(false);
    return html.Trim();   // resumes on a thread-pool thread, not the captured context
}
```

*What just happened:* `ConfigureAwait(false)` says "resume the continuation on any available thread-pool thread, don't bother returning to the original context." This makes library code faster (no context-hop) and immune to the sync-over-async deadlock, since the continuation no longer needs that one blocked thread. Rule of thumb: **use it in library code; skip it in application code** where you *do* want to land back on the right context. (Modern ASP.NET Core has no SynchronizationContext, so the deadlock doesn't occur there - but the habit still matters for libraries and desktop apps.)

### ⚠️ Don't forget to await - fire-and-forget

Call an async method and *don't* await it (or store its Task), and you've launched "fire-and-forget" work. The compiler usually warns you. The danger: no idea if it succeeded, and any exception it throws vanishes silently onto an unobserved Task.

```csharp
SaveToDatabaseAsync(record);          // ⚠️ no await - bug! warning CS4014
// execution continues immediately; if the save throws, you'll never know

await SaveToDatabaseAsync(record);    // ✅ wait for it, observe success or failure
```

*What just happened:* The first line starts the save and moves on without waiting - the returned Task is dropped on the floor. If the write fails, the exception lands on that abandoned Task and is never observed; your program continues as if it succeeded. The second line awaits it, so any failure surfaces at the `await`. Unless you have a deliberate, carefully-handled reason for fire-and-forget, **await every Task you create.**

## Parallelism vs async - and composing tasks

So far every `await` waited for *one* thing at a time. The real power shows up when you run several async operations *concurrently* and await them together.

### `Task.WhenAll` and `Task.WhenAny`

To fetch three URLs concurrently, *start* all three tasks first (don't await yet), then await them as a group:

```csharp
async Task<string[]> FetchAllAsync(HttpClient client)
{
    // Start all three NOW - they run concurrently. No await yet.
    Task<string> a = client.GetStringAsync("https://example.com/1");
    Task<string> b = client.GetStringAsync("https://example.com/2");
    Task<string> c = client.GetStringAsync("https://example.com/3");

    // Now wait for all of them together.
    return await Task.WhenAll(a, b, c);   // ~as slow as the slowest one, not the sum
}
```

*What just happened:* By starting `a`, `b`, and `c` before awaiting any of them, all three downloads were in flight at once. `Task.WhenAll` returns a single Task completing when *every* input task does, handing back an array of all results - total time roughly the duration of the *slowest* download, not the sum. Contrast with `await a; await b; await c;`, which waits for each in turn. ⚠️ Ordering matters: `await` each task where you create it, and you've serialized them, throwing away the concurrency.

`Task.WhenAny` is the sibling for "whichever finishes first" - useful for timeouts (race your real task against a `Task.Delay`) or querying several mirrors and taking the fastest reply.

### `Task.Run` - for CPU-bound work

Everything above was *I/O-bound* (waiting on the network). Sometimes you have genuinely heavy *computation* you don't want freezing your UI thread - deliberately push it onto a background thread with `Task.Run`:

```csharp
// CPU-bound: a heavy calculation. Move it off the UI thread.
long sum = await Task.Run(() =>
{
    long total = 0;
    for (int i = 0; i < 1_000_000_000; i++) total += i % 7;
    return total;
});
```

*What just happened:* `Task.Run` hands the lambda to a thread-pool thread and returns a Task representing it. Awaiting that Task frees the calling thread (e.g. the UI) while the CPU work churns in the background - the *one* time you genuinely *do* want a new thread, since the work is real computation, not waiting. ⚠️ Don't wrap I/O calls in `Task.Run`; that wastes a thread babysitting work that was already non-blocking.

### Data parallelism - `Parallel.ForEach` and PLINQ

For a CPU-heavy operation applied across a *collection*, using *all* your cores, reach for the data-parallel tools:

```csharp
using System.Threading.Tasks;
using System.Linq;

// Run a CPU-bound transform across all cores:
Parallel.ForEach(images, img => Resize(img));

// Or with PLINQ - parallel LINQ:
var results = numbers.AsParallel().Select(n => ExpensiveCompute(n)).ToArray();
```

*What just happened:* `Parallel.ForEach` splits the collection across multiple threads, processing chunks simultaneously on different cores - true parallelism. `AsParallel()` does the same for a LINQ query. Both are about *going faster by using more cores*, a different goal from async I/O.

💡 **The dividing line.** Use **async/await** for **I/O-bound** work (network, disk, database) - the goal is *not wasting a thread while waiting*. Use **parallelism** (`Task.Run`, `Parallel.ForEach`, PLINQ) for **CPU-bound** work - the goal is *using more cores to finish faster*. They look similar but solve opposite problems: async on CPU work just adds overhead, parallelism on I/O just wastes threads.

## Recap

1. **A blocked thread is a wasted thread.** Synchronous I/O parks a thread doing nothing while it waits; async I/O frees that thread to serve other work.
2. **A `Task`/`Task<T>` is a receipt for work that finishes later** - C#'s promise/future. A Task is *not* a thread; I/O Tasks don't burn a thread to wait.
3. **`await` suspends your method and returns to the caller**, freeing the thread; when the Task completes, a continuation resumes right after the `await`, locals intact. It also unwraps the result and rethrows exceptions so async code reads like sync code.
4. ⚠️ **Avoid `async void`** (exceptions are unobservable - only for event handlers), and ⚠️ **never block on async** with `.Result`/`.Wait()` on a captured context (deadlocks). Use `ConfigureAwait(false)` in libraries, and await every Task you start.
5. **Compose with `Task.WhenAll`/`WhenAny`** by starting tasks *before* awaiting, so they run concurrently - total time tracks the slowest, not the sum.
6. 💡 **Async for I/O-bound, parallelism for CPU-bound.** `async`/`await` stops you wasting threads on waiting; `Task.Run`/`Parallel.ForEach`/PLINQ use more cores to compute faster. Different problems, different tools.

You can now write code that handles thousands of concurrent operations without a thread per wait - the foundation of every responsive UI and scalable server in .NET. Next: one level deeper into the runtime itself - how memory, the garbage collector, and the JIT actually make your C# run.

## Quick check

Test yourself on the one idea that makes all of this work - what `await` really does:

```quiz
[
  {
    "q": "When execution hits `await someTask` and the task is NOT yet complete, what happens?",
    "choices": [
      "The method suspends and returns control to its caller, freeing the thread; the rest runs later as a continuation when the task completes",
      "The current thread blocks and sleeps until the task finishes, doing nothing else",
      "A new thread is always spawned to run the awaited work in parallel",
      "The task is cancelled and the method returns its default value immediately"
    ],
    "answer": 0,
    "explain": "await is not a blocking wait. It suspends the method and hands control back to the caller, leaving the thread free to do other work. When the task completes, the code after the await resumes as a continuation, with locals intact. That's why it scales - no thread is held hostage while waiting."
  },
  {
    "q": "Why does blocking on an async method with `.Result` on a UI thread (with a SynchronizationContext) deadlock?",
    "choices": [
      "The UI thread blocks waiting for the task, but the task's continuation needs that same UI thread to resume - so each waits on the other forever",
      "`.Result` is not a valid way to get a value from a Task and throws an exception",
      "The task runs on a thread that the garbage collector pauses indefinitely",
      "Async methods can only ever run on background threads, never the UI thread"
    ],
    "answer": 0,
    "explain": "The captured SynchronizationContext requires the continuation to resume on the UI thread. But .Result blocks that very thread waiting for the task to finish. The continuation can't run (thread is busy blocking), so the task never completes, so .Result never returns. The fix is to await all the way up instead of blocking."
  },
  {
    "q": "You have a CPU-heavy calculation (a billion-iteration loop) that's freezing your UI. Which tool fits?",
    "choices": [
      "`Task.Run` to move the computation onto a background thread, then await it",
      "Plain `async`/`await` on the loop, which will run it off-thread automatically",
      "`Task.WhenAll`, since it always runs work in parallel across cores",
      "`ConfigureAwait(false)`, which moves CPU work to the thread pool by itself"
    ],
    "answer": 0,
    "explain": "This is CPU-bound work, so it needs a real thread to run on - that's exactly what Task.Run provides, handing the work to a thread-pool thread and freeing the UI thread (which you await). Plain async/await is for I/O (it doesn't move CPU work off-thread by itself), and ConfigureAwait only controls where continuations resume, not where the work runs."
  }
]
```


---

# The .NET Runtime: Memory, GC & JIT - What Runs Your IL

For fourteen phases you've written C# and it ran. You never called `malloc`, never called `free`, never thought about where an `int` or a `Customer` lives in memory - the whole point of a *managed* language. Underneath your code sits the **CLR**, the .NET runtime, quietly doing three big jobs: deciding where your data lives, cleaning up data you stop using, and turning your code into native machine instructions on the fly.

You can ship C# for years without opening this hood. But the day a service's memory climbs and never comes back down, or a request stalls for 50ms with no code running, or you wonder why your value-type benchmark is 10x faster than the reference-type one - you're asking runtime questions. The big idea: **.NET trades a little control you don't want for a lot of bookkeeping you don't have to do** - and it pays off everywhere except the few spots where you need to know the deal you signed.

## The CLR, recapped and deepened

Back in [Phase 1](01-install-and-first-program.md) you learned that C# compiles to **IL** (Intermediate Language), not machine code, and that something called the CLR runs it. Time to deepen that.

📝 **CLR (Common Language Runtime)** - the .NET runtime: it loads your assemblies, manages their memory, JIT-compiles their IL to native instructions, and supervises them while they run. "Managed code" means code whose memory and execution the CLR oversees.

The shape to hold in your head: your `.cs` files compile to IL stored in a `.dll` - portable, not directly runnable by any CPU, instructions for an imaginary stack machine. When you run the program, the CLR loads that IL and does the real work: allocates memory for your objects, tracks which are still in use, runs the **garbage collector** to reclaim the rest, and **JIT-compiles** each method to native code the first time it's called.

```mermaid
flowchart LR
  A[C# source] --> B[csc compiler]
  B --> C[IL in .dll]
  C --> D[CLR loads it]
  D --> E[JIT to native]
  D --> F[Manages heap + GC]
  E --> G[CPU runs it]
```

*One idea:* the compiler's job ends at IL. Everything between "IL on disk" and "instructions executing on your CPU" is the CLR's job - memory management and JIT are the two biggest parts of it.

## Stack vs heap, and value vs reference types

Back in [Phase 2](02-syntax-values-and-types.md) you met the split between **value types** (`int`, `bool`, `double`, `struct`, `enum`) and **reference types** (`class`, `string`, arrays, `object`) - framed then as *copying behavior*. The deeper reason: they live in different *places*, and those places behave very differently.

📝 **Stack** - a small, fast, per-thread region of memory that grows and shrinks with method calls. When a method is called, its locals get a slice of the stack ("stack frame"); when it returns, that slice vanishes instantly and for free - allocation is a pointer bump, cleanup is automatic and costs nothing. **Managed heap** - a larger shared pool where reference-type objects live, not freed when a method returns; reclaiming it is the garbage collector's job.

The rough rule: **value types live where they're declared; reference types live on the heap with a reference pointing to them.** A local `int` sits right in the method's stack frame. A local `Customer c = new Customer()` puts the `Customer` *object* on the heap and keeps a *reference* (a pointer) to it on the stack. An `int` field inside a class rides along on the heap inside that object; a local `int` is on the stack. ("Often on the stack" is the accurate phrasing - the CLR can put value types on the heap when captured by a lambda or boxed, the next gotcha.)

```csharp
struct Point { public int X, Y; }          // value type
class Box    { public int Value; }          // reference type

void Demo()
{
    int n = 42;                 // value: lives in this stack frame
    Point p = new Point();      // value: also lives in this stack frame, inline
    Box b = new Box();          // reference: the Box object is on the heap;
                                //   `b` (the reference to it) is on the stack
}                               // frame vanishes: n and p are gone instantly.
                                //   the Box on the heap waits for the GC.
```

*What just happened:* `n` and `p` are value types, so they sit directly in `Demo`'s stack frame - when `Demo` returns, they're swept away with the frame at zero cost. `b` is a *reference*: the `Box` object itself was allocated on the managed heap, and `b` just holds its address. When `Demo` returns, `b` disappears with the frame, but the `Box` object lingers until the GC notices nothing points to it. This is why the distinction matters for performance: stack allocation is nearly free and self-cleaning, while heap allocation costs more *and* creates future work for the GC.

One trap quietly drags a value type onto the heap: **boxing**.

📝 **Boxing** - wrapping a value type in a heap object so it can be treated as a reference type (typically `object` or an interface). **Unboxing** reverses it: copying the value back out. Each box is a fresh heap allocation plus a copy.

```csharp
int n = 7;
object boxed = n;        // BOXING: a heap object is allocated to hold a copy of 7
int back = (int)boxed;   // UNBOXING: copies the value back out of the heap object
```

*What just happened:* assigning an `int` to an `object` can't just store the number - `object` is a reference type, so the CLR allocates a little box on the heap, copies `7` into it, and points `boxed` at it: a heap allocation you didn't ask for. The cast back copies the value out again. One box is cheap; a million in a loop (easy to do accidentally with old non-generic collections like `ArrayList`, or formatting value types into `object[]`) is a measurable GC stampede.

⚠️ **Gotcha - boxing hides in plain sight.** It's why generics (`List<int>`) exist and `ArrayList` is discouraged: `List<int>` stores `int`s without boxing, while `ArrayList` boxes every one. If a hot path is allocating and you can't see `new` anywhere, suspect a hidden box - a value type slipping into an `object`, an interface, or a non-generic API.

## Garbage collection

You allocate heap objects constantly - every `new` on a class, every `string` concatenation, every list that grows - and never free any of them. What stops the heap from filling up forever? The garbage collector.

📝 **Garbage collector (GC)** - the CLR component that automatically finds heap objects your program can no longer reach and reclaims their memory, so you never call `free` yourself. .NET's GC is a **generational, tracing, compacting** collector: it marks what's reachable, reclaims the rest, and slides the survivors together so free space stays in one contiguous block (which is what makes a new allocation a cheap pointer bump).

The core idea is **reachability**. The GC starts from a set of **roots** - static fields, local variables and method arguments live on every thread's stack right now - and traces every reference it can follow. Every object it reaches is *live* and kept; every object it *can't* reach is unreachable - garbage - and its memory goes back into the pool. Reachability *is* the lifetime: the moment nothing points to an object, it's eligible for collection (not necessarily immediately, but eventually).

The "generational" part is the clever optimization, resting on one observation that holds for almost every program: **most objects die young.** The temporary string you built to log a line, the small list inside a method, become garbage almost as soon as they're born. A few objects (your cache, your config, your long-lived service) live for the whole program. So the GC sorts the heap by age:

```mermaid
flowchart LR
  N[new object] --> G0[Gen 0<br/>youngest, collected often]
  G0 -->|survives a GC| G1[Gen 1<br/>survivors]
  G1 -->|survives again| G2[Gen 2<br/>long-lived]
  L[object >= 85 KB] --> LOH[Large Object Heap<br/>collected with Gen 2]
```

📝 **Generations.** New objects start in **Gen 0**, small and collected very frequently and fast. Anything surviving a Gen 0 collection is promoted to **Gen 1**; survivors of Gen 1 move to **Gen 2**, the long-lived region, collected rarely - cheap, frequent work on a tiny slice of the heap, occasionally a full **Gen 2** sweep. Separately, objects 85 KB or larger go on the **Large Object Heap (LOH)**, collected together with Gen 2, since moving big objects around is expensive.

Two more dials worth knowing by name. **Workstation vs server GC**: workstation GC (default for desktop/client apps) is tuned for low latency on one or two cores; **server GC** (typical for ASP.NET and high-throughput services) runs parallel collection threads across many cores for higher throughput. Every collection involves a brief **stop-the-world (STW)** pause - threads freeze for at least part of the work so the object graph doesn't change mid-trace. Modern .NET keeps these small (often sub-millisecond for Gen 0), but never *zero*.

Step through a collection yourself - roots, what's reachable, and what gets reclaimed:

```playground-gc
```

💡 **The insight.** GC is automatic, but it isn't free. Every object you allocate is future work for the collector, and heavy **allocation pressure** (churning through short-lived objects in a hot path) means more frequent Gen 0 collections and more STW pauses. You rarely tune the GC directly - instead you *allocate less*: reuse buffers, prefer value types and `Span<T>`, avoid hidden boxing. That's the bridge to [Phase 17](17-performance-and-ecosystem.md).

## `IDisposable`, finalizers, and the limit of GC

Here's the trap that catches people who learn "the GC cleans up for me" and stop there: **the GC manages memory, and only memory.** It knows nothing about the file handle, network socket, database connection, or OS lock your object is holding. Those are *unmanaged resources*, and the GC won't release them for you - at least not when you need it to.

⚠️ **The GC frees memory; it does not free everything else.** An object holding an open file can become unreachable and the GC will happily reclaim its *memory* - but the file might stay locked until some indeterminate later moment (or until the process exits). For anything scarce or externally visible, release it *deterministically*, which is what `IDisposable` is for.

You met `using` back in [Phase 7](07-errors-and-io.md). Now you can see *why* it exists: the deterministic counterpart to the non-deterministic GC.

```csharp
// `using` calls Dispose() the instant the block ends - deterministic cleanup,
// not "whenever the GC gets around to it."
using (var file = new StreamReader("data.txt"))
{
    string first = file.ReadLine();
    Console.WriteLine(first);
}   // file.Dispose() runs HERE, releasing the OS file handle right now

// modern "using declaration" form - disposes at the end of the enclosing scope
using var conn = new SqlConnection(connectionString);
conn.Open();
// ... conn.Dispose() runs when the method's scope ends
```

*What just happened:* `using` guarantees `Dispose()` is called the moment control leaves the block - even if an exception is thrown - so the file handle is released *right then*, deterministically. Without it, the `StreamReader` would eventually become unreachable and the GC would reclaim its memory, but the OS handle could stay open far longer than you'd want, and a handful of leaked handles can exhaust an OS limit while you still have gigabytes of free memory.

What about **finalizers** (the `~ClassName()` method)? They're the GC's last-resort backstop: if an object holding an unmanaged resource is collected *without* anyone calling `Dispose()`, its finalizer runs during GC and can release the resource - but they're bad news as a primary strategy.

⚠️ **Avoid relying on finalizers.** A finalizable object survives an *extra* GC cycle (promoted so the finalizer can run, then collected later), runs your cleanup on a separate finalizer thread at an unpredictable time, and delays reclaiming that memory. They exist as a safety net for the rare case of wrapping a raw unmanaged handle where someone forgets to `Dispose`. The right pattern is the standard `Dispose` pattern with `GC.SuppressFinalize(this)` - "I cleaned up properly, skip the finalizer."

💡 **The one-line rule.** The GC frees *memory*; you free *everything else* - with `using`. Memory is the runtime's job; file handles, sockets, and connections are yours.

## JIT compilation and memory errors

The last piece of CLR magic is how your IL becomes machine code - not all at once at startup, but **just in time**.

📝 **JIT (Just-In-Time) compiler** - the part of the CLR that translates a method's IL into native machine code the *first time that method is called*, then caches the result so later calls run the already-native version. Methods you never call are never JIT-compiled.

This is why .NET apps "warm up": the first request to a freshly started service is often noticeably slower, since every method on that code path is JIT-compiled as it's hit, for the only time. Subsequent requests run the cached native code at full speed. Modern .NET sharpens this with **tiered compilation**: the JIT first produces code *quickly* with few optimizations (Tier 0) so the app starts fast, then recompiles hot methods (called many times) in the background with full optimizations (Tier 1) - fast startup *and* fast steady-state.

If you can't afford warm-up at all - a CLI tool that must be instant, or a serverless function billed per millisecond - there are ahead-of-time options: **ReadyToRun** (R2R) bakes pre-JITted native code into the assembly so startup skips most JIT work, and **Native AOT** compiles the whole app to a standalone native binary with no JIT machinery at all. The trade-off is larger, less flexible binaries and some feature restrictions; most apps stick with the JIT.

Two memory errors you'll eventually meet. The CLR throws `OutOfMemoryException` when it genuinely can't satisfy an allocation. Far more common in practice is the **managed memory leak**:

⚠️ **GC'd does not mean leak-proof.** The GC only frees what's *unreachable*. If your program keeps a reference alive by accident, the object stays in memory forever - and it looks exactly like a leak, because it is one. The classic culprits:

```csharp
// 1. A static collection that only ever grows.
static readonly List<byte[]> _cache = new();
void Remember(byte[] data) => _cache.Add(data);   // never removed → grows without bound

// 2. An event subscription you never unsubscribe.
publisher.DataChanged += subscriber.OnDataChanged;
// The publisher now holds a reference to `subscriber`. Even when you're "done"
// with `subscriber`, the publisher keeps it reachable - and alive - forever,
// until you do:  publisher.DataChanged -= subscriber.OnDataChanged;
```

*What just happened:* in both cases the GC is working *perfectly* - the memory genuinely is still reachable, so it's genuinely still alive. The static `_cache` keeps every array reachable for the life of the process. The event subscription is sneakier: a publisher's event holds a reference to every subscriber, silently pinning short-lived subscribers in memory until you unsubscribe. The fix is never a GC setting - it's bounding your collections (eviction policies, `WeakReference`) and matching every `+=` with a `-=`. Managed languages move the leak from "forgot to `free`" to "forgot to drop the reference," but it's just as real.

## Recap

1. The **CLR** runs your IL: manages the heap, runs the GC, and JIT-compiles methods to native code. "Managed code" means the CLR does this bookkeeping for you.
2. **Value types** typically live inline on the **stack** (cheap, auto-freed on return); **reference types** live on the **managed heap** with a reference on the stack. **Boxing** drags a value type onto the heap - a hidden cost worth hunting for.
3. The **GC** reclaims **unreachable** heap objects automatically. It's **generational** - Gen 0 is small and collected often because most objects die young; Gen 2 and the LOH are collected rarely. Every collection has a brief **stop-the-world** pause, and allocation pressure causes more of them.
4. ⚠️ The GC frees **memory, not other resources**. Files, sockets, and connections need deterministic cleanup via **`IDisposable`/`using`**; **finalizers** are a last-resort backstop, not a strategy.
5. The **JIT** compiles IL to native on first call (with **tiered compilation** for fast startup *and* fast hot paths), which is why .NET "warms up"; **ReadyToRun**/**Native AOT** skip JIT when warm-up is unacceptable.
6. ⚠️ GC'd ≠ leak-proof: an accidental lingering reference (static collections, un-unsubscribed events) keeps memory alive forever. Fix by dropping the reference, not tuning the GC.

You now know what happens beneath `new`, `using`, and every method call: where your data lives, who cleans it up, and how your IL becomes instructions. Next: making all of this measurable - testing, building, and profiling, where you'll watch allocations and timings with real tools instead of reasoning in the abstract.

## Quick check

Test yourself on the three ideas that matter most - where values live, what the GC does and doesn't free, and what JIT means:

```quiz
[
  {
    "q": "What is boxing in C#, and why does it cost something?",
    "choices": [
      "Wrapping a value type in a heap-allocated object so it can be used as a reference type - it costs a heap allocation plus a copy",
      "Putting a class instance on the stack to make it faster, which costs nothing",
      "Compiling IL to native code on first use, which costs startup time",
      "Marking an object unreachable so the GC can collect it"
    ],
    "answer": 0,
    "explain": "Boxing wraps a value type (like an int) in a fresh heap object so it can be treated as `object` or an interface. That's a heap allocation and a copy you didn't ask for - harmless once, but a GC stampede in a hot loop. It's the reason `List<int>` is preferred over the boxing `ArrayList`."
  },
  {
    "q": "Your object holds an open file handle and becomes unreachable. What does the garbage collector guarantee?",
    "choices": [
      "It will reclaim the object's memory eventually, but it does NOT promptly release the file handle - that needs IDisposable/using",
      "It immediately closes the file and frees the memory at the same instant",
      "Nothing - the GC never touches objects that hold unmanaged resources",
      "It throws OutOfMemoryException because file handles can't be collected"
    ],
    "answer": 0,
    "explain": "The GC manages memory and only memory. It'll reclaim the object's bytes once nothing references it, but it knows nothing about the OS file handle and won't release it deterministically. That's exactly what `IDisposable` and `using` are for - and why finalizers are only a last-resort backstop."
  },
  {
    "q": "Why is the very first request to a freshly started .NET service often slower than later ones?",
    "choices": [
      "The JIT compiles each method's IL to native code the first time it runs; later calls reuse the cached native code",
      "The garbage collector always runs a full Gen 2 collection at startup",
      "The CLR re-downloads the assemblies on the first request",
      "Value types are boxed on the first call and unboxed afterward"
    ],
    "answer": 0,
    "explain": "Methods are JIT-compiled to native code on first call, then cached. The first request pays that one-time compilation cost across its whole code path, so it 'warms up.' Tiered compilation softens this (fast Tier 0 first, optimized Tier 1 later), and ReadyToRun/Native AOT can skip it entirely."
  }
]
```


---

# Testing, Build & Profiling - Proving It Works, Finding the Slow Part

Back in [Phase 8](08-projects-and-tooling.md) I told you `dotnet test` exists and runs your tests, then promised the real treatment later. This is later. A project you can't *prove* works is a project you're shipping on faith, and a project you *think* is slow but never measured is one you're about to optimize in the wrong place.

The mental model to carry through this phase: **stop operating on vibes.** "It works" becomes a test you can run. "It's slow" becomes a number you can measure. "This function is the bottleneck" becomes a profile that points at the actual hotspot - which, more often than your ego would like, is somewhere you never suspected. Once wired in, correctness and performance stop being arguments and start being commands you run.

## xUnit basics - your first real test

📝 **Unit test.** A small, automated check that exercises one piece of your code - usually one method - with known inputs and asserts the output is what you expect. "Unit" because it tests a unit in isolation, not the whole app wired together. Write it once; it runs forever, catching the day someone (often future-you) breaks that behavior.

.NET has three mainstream test frameworks - **xUnit**, **NUnit**, **MSTest** - more alike than different. NUnit is older and capable; MSTest ships from Microsoft; **xUnit is the de-facto default for new projects** and what most open-source .NET code uses. Scaffold a test project with `dotnet new xunit`; it shows up as its own `.csproj` referencing the project under test.

A test is a plain method tagged `[Fact]` - "a fact that should always be true." Inside, follow the **AAA shape**: **Arrange** the inputs, **Act** by calling the thing, **Assert** the result. Testing this method:

```csharp
namespace Mathx;

public static class Calc
{
    public static int Clamp(int n, int min, int max)
    {
        if (n < min) return min;
        if (n > max) return max;
        return n;
    }

    public static int Divide(int a, int b) => a / b;
}
```

```csharp
using Xunit;
using Mathx;

public class CalcTests
{
    [Fact]
    public void Clamp_BelowMin_ReturnsMin()
    {
        // Arrange
        int n = -5, min = 0, max = 10;

        // Act
        int result = Calc.Clamp(n, min, max);

        // Assert
        Assert.Equal(0, result);
    }

    [Fact]
    public void Divide_ByZero_Throws()
    {
        Assert.Throws<DivideByZeroException>(() => Calc.Divide(10, 0));
    }
}
```

```console
$ dotnet test
Passed!  - Failed: 0, Passed: 2, Skipped: 0, Total: 2, Duration: 12 ms
  CalcTests.Clamp_BelowMin_ReturnsMin [PASS]
  CalcTests.Divide_ByZero_Throws [PASS]
```

*What just happened:* Each `[Fact]` method is one test the runner discovers automatically - you never register them. `Clamp_BelowMin_ReturnsMin` arranged its inputs, called `Clamp`, and used `Assert.Equal(expected, actual)` to demand the answer be `0`. The second test used `Assert.Throws<T>`, which *expects* an exception: it runs the lambda, passes if the right exception type is thrown, and **fails if no exception comes** - testing error paths without a `try/catch` of your own. Method names read like sentences (`Method_Scenario_Expectation`) on purpose, so a failure's name alone tells you what broke.

💡 **Key point.** A good assertion failure tells you *what* was wrong without opening a debugger. `Assert.Equal(0, result)` prints `Expected: 0, Actual: 5` on failure - the whole reason tests beat manually eyeballing output.

## Parameterized tests - one method, many cases

`Clamp` needs more than one case: below min, above max, inside the range, exactly on a boundary. Copy-pasting `[Fact]` four times with different numbers is the obvious move and the wrong one - four near-identical methods that drift apart over time. xUnit's answer is the **`[Theory]`**: C#'s flavor of table-driven testing.

📝 **Theory.** A test method that runs *once per data row*. Tag it `[Theory]` instead of `[Fact]`, give the method parameters, and feed rows of arguments with `[InlineData(...)]` - each row an independent test case with its own pass/fail.

```csharp
public class ClampTheoryTests
{
    [Theory]
    [InlineData(5, 0, 10, 5)]    // inside range
    [InlineData(-3, 0, 10, 0)]   // below min
    [InlineData(99, 0, 10, 10)]  // above max
    [InlineData(0, 0, 10, 0)]    // exactly on min
    public void Clamp_PinsIntoRange(int n, int min, int max, int expected)
    {
        int result = Calc.Clamp(n, min, max);
        Assert.Equal(expected, result);
    }
}
```

```console
$ dotnet test
Passed!  - Failed: 0, Passed: 4, Skipped: 0, Total: 4, Duration: 9 ms
  ClampTheoryTests.Clamp_PinsIntoRange(n: 5, min: 0, max: 10, expected: 5) [PASS]
  ClampTheoryTests.Clamp_PinsIntoRange(n: -3, min: 0, max: 10, expected: 0) [PASS]
  ClampTheoryTests.Clamp_PinsIntoRange(n: 99, min: 0, max: 10, expected: 10) [PASS]
  ClampTheoryTests.Clamp_PinsIntoRange(n: 0, min: 0, max: 10, expected: 0) [PASS]
```

*What just happened:* One method body, four runs. Each `[InlineData]` supplied a row of arguments matched positionally to the parameters, and xUnit treated every row as its own test - the runner prints the actual values, so a failure tells you *exactly which row* broke. Adding a fifth scenario is one new line, not a new method - a *visible list* a reviewer can scan and ask "where's the case for `min > max`?"

💡 **`[MemberData]` for non-constant cases.** `[InlineData]` only takes compile-time constants - numbers, strings, `true`. When cases need real objects (a `DateTime`, a custom type, a computed value), switch to `[MemberData(nameof(Source))]`, which pulls rows from a static property or method returning `IEnumerable<object[]>`. Same idea, sourced from code instead of attributes.

## Mocking - isolating the unit under test

Real code has dependencies: a method you want to test calls a database, an HTTP API, a clock, a payment gateway. You don't want your *unit* test hitting a live database - slow, flaky, testing the database instead of your logic. The fix is a **mock**: a fake stand-in you control completely, so the only real code in the test is the thing you're actually testing.

In .NET the common tools are **Moq** and **NSubstitute** (both NuGet packages; pick one per project). The pattern with Moq: create a `Mock<T>` of the dependency's *interface*, use `.Setup(...)` to script what its methods return, inject it into your class, then use `.Verify(...)` to assert it was called as expected.

```csharp
using Moq;
using Xunit;

public interface IPriceFeed { decimal GetPrice(string symbol); }

public class Portfolio
{
    private readonly IPriceFeed _feed;
    public Portfolio(IPriceFeed feed) => _feed = feed;

    public decimal ValueOf(string symbol, int shares)
        => _feed.GetPrice(symbol) * shares;
}

public class PortfolioTests
{
    [Fact]
    public void ValueOf_MultipliesPriceByShares()
    {
        // Arrange: a fake feed that always returns 10.00 for "ACME"
        var feed = new Mock<IPriceFeed>();
        feed.Setup(f => f.GetPrice("ACME")).Returns(10.00m);
        var portfolio = new Portfolio(feed.Object);

        // Act
        decimal value = portfolio.ValueOf("ACME", 3);

        // Assert
        Assert.Equal(30.00m, value);
        feed.Verify(f => f.GetPrice("ACME"), Times.Once);
    }
}
```

*What just happened:* `new Mock<IPriceFeed>()` created a controllable fake of the `IPriceFeed` interface. `.Setup(f => f.GetPrice("ACME")).Returns(10.00m)` scripted its behavior: "when someone asks for ACME's price, hand back 10.00 - no real feed involved." We passed `feed.Object` into `Portfolio`, so the only genuine logic running is `Portfolio.ValueOf`'s multiplication. `Verify(..., Times.Once)` asserted the feed was queried exactly once. Note `Portfolio` depends on the *interface*, not a concrete class - what makes it mockable.

⚠️ **Don't over-mock.** Mocking is for slow, external, or non-deterministic dependencies - networks, databases, clocks, the file system. Mock your *own* simple classes and your test stops verifying real behavior, instead verifying that your code calls methods in the order you said it would - breaking the instant you refactor, even when nothing's actually wrong. Mock at the boundaries; use the real thing inside them.

## Benchmarking with BenchmarkDotNet - measuring accurately

Now the speed question. Your instinct will be to wrap the code in a `Stopwatch`, loop it a million times, and print the elapsed milliseconds. ⚠️ **That number will lie to you**, and understanding *why* ties straight back to the runtime internals from [Phase 15](15-the-dotnet-runtime-and-gc.md).

A naive `Stopwatch` loop is wrong for three runtime reasons at once. First, **JIT warmup**: the first call to a method is *interpreted or freshly compiled*, far slower than steady state - early iterations measure compilation, not execution. Second, **tiered compilation**: the JIT initially produces quick-but-unoptimized code, then *recompiles hot methods* partway through your loop, so the method literally changes speed mid-measurement. Third, **the GC**: a collection can fire mid-window and charge its pause to whatever code happened to be running. Add dead-code elimination (the JIT may delete a result you never use) and your hand-rolled benchmark measures noise.

**BenchmarkDotNet** is the NuGet library that handles all of this for you. Tag methods with `[Benchmark]`; it runs warmup iterations until the JIT has settled, runs enough measured iterations for statistical confidence, isolates runs, prevents dead-code elimination, and - with one attribute - reports memory allocations too.

```csharp
using BenchmarkDotNet.Attributes;
using BenchmarkDotNet.Running;
using System.Text;

[MemoryDiagnoser] // adds the allocation columns
public class StringBuildBench
{
    private readonly string[] _parts = { "a", "b", "c", "d", "e", "f", "g", "h" };

    [Benchmark(Baseline = true)]
    public string Concat()
    {
        string s = "";
        foreach (var p in _parts) s += p;   // re-allocates the whole string each time
        return s;
    }

    [Benchmark]
    public string Builder()
    {
        var sb = new StringBuilder();
        foreach (var p in _parts) sb.Append(p);
        return sb.ToString();
    }
}

// In Program.cs:
// BenchmarkRunner.Run<StringBuildBench>();
```

```console
$ dotnet run -c Release
| Method  | Mean      | Ratio | Allocated | Alloc Ratio |
|-------- |----------:|------:|----------:|------------:|
| Concat  | 142.6 ns  |  1.00 |     360 B |        1.00 |
| Builder |  61.3 ns  |  0.43 |     200 B |        0.56 |
```

*What just happened:* BenchmarkDotNet ran each method through warmup-then-measure, reporting **Mean** (average time per call) and, thanks to `[MemoryDiagnoser]`, **Allocated** (bytes of heap garbage per call). `[Benchmark(Baseline = true)]` made `Concat` the reference, so **Ratio** reads directly: `Builder` runs at `0.43` - under half the time - and allocates `0.56` of the garbage. The story: `s += p` in a loop re-allocates the entire string every iteration (strings are immutable), generating garbage the GC must later collect, while `StringBuilder` grows one buffer. ⚠️ Note `dotnet run -c Release` - benchmarking a Debug build measures unoptimized code and is its own kind of lie.

💡 **Watch allocations, not just time.** The `Allocated` column is often the one to fix first: every byte allocated is future work for the garbage collector ([Phase 15](15-the-dotnet-runtime-and-gc.md)), and GC pauses wreck a server's tail latency under load. A change that shaves nanoseconds but doubles allocations can be a *net loss* in production even though the microbenchmark looks faster.

## Profiling & coverage - measure, don't guess

A benchmark tells you *that* a method is slow. It doesn't tell you *where the time goes* across a whole running program - which method, called from where, is eating the CPU or churning the heap. For that, you profile.

📝 **Profiler.** A tool that watches your program *as it runs* and records where it actually spends time (CPU) or memory (allocations), then ranks the results so you see true hotspots instead of guessing. The .NET toolbox:

- **`dotnet-trace`** - captures CPU/event traces from a running process; the cross-platform CLI workhorse.
- **`dotnet-counters`** - live dashboard of runtime metrics (CPU, GC frequency, allocation rate, thread-pool queue) - your first "what's it doing right now?" look.
- **`dotnet-gcdump`** - snapshots the managed heap so you can see *what's holding memory* (chasing leaks and bloat).
- **Visual Studio Profiler** and **PerfView** - rich GUI analyzers (PerfView is the deep, free, Windows-focused one the .NET team uses).

Install the CLI tools as global .NET tools and point them at a process:

```bash
dotnet tool install -g dotnet-trace
dotnet-trace collect --process-id 12345 --duration 00:00:30
# produces trace.nettrace - open it in Visual Studio, PerfView, or Speedscope
```

*What just happened:* `dotnet-trace collect` attached to a live process by PID and sampled 30 seconds of execution into a `.nettrace` file. Opened in an analyzer, that file ranks methods by time spent - the ranking is the entire point. The function you were *sure* was the bottleneck routinely isn't; the profile shows the one that actually is, often something boring like JSON serialization or a chatty database call in a loop.

💡 **Measure, don't guess - this is the whole discipline.** Every engineer has a confident hunch about the slow part, and that hunch is wrong often enough to waste real days. You optimize the function you suspected, ship it, and the app is exactly as slow as before, because the cost was somewhere you never looked. Profile *first*, then optimize the thing the profile points at. ([Phase 17](17-performance-and-ecosystem.md) covers what to *do* once the profile has spoken.)

**Coverage** answers a different question: how much of your code did the tests actually run? In .NET the standard tool is **coverlet**, which plugs into `dotnet test`:

```bash
dotnet test --collect:"XPlat Code Coverage"
# writes a coverage.cobertura.xml report; turn it into HTML with ReportGenerator
```

*What just happened:* coverlet instrumented the build, tracked which lines executed while the tests ran, and wrote a coverage report. Fed to a viewer, it color-codes your source: green lines ran, red lines never did - the *red* is the useful part, a map of the branches your tests forgot.

⚠️ **Coverage is not correctness.** This trap catches everyone. Coverage tells you a line *executed* - nothing about whether you *checked the result*. A test that calls `Clamp` and asserts nothing lights the whole method green. Treat coverage as a map of the untested (chase the red), never a score to maximize - high coverage with weak assertions is *more* dangerous than plain medium coverage, since it feels safe while proving almost nothing.

## Recap

1. **xUnit** is the common default (NUnit/MSTest also exist). A `[Fact]` is one test in **Arrange-Act-Assert** shape; `Assert.Equal` checks values, `Assert.Throws<T>` checks error paths. Run everything with `dotnet test`.
2. **`[Theory]` + `[InlineData]`** is C#'s table-driven testing: one method, many cases, each row an independent pass/fail. Use **`[MemberData]`** for non-constant values.
3. **Mocking** (Moq/NSubstitute) replaces a real dependency with a controllable fake - `Mock<T>`, `.Setup`, `.Verify` - so you test *your* unit in isolation. ⚠️ Mock at the boundaries (network, DB, clock); over-mocking tests your wiring, not your behavior.
4. **BenchmarkDotNet** (`[Benchmark]`) measures accurately because ⚠️ a `Stopwatch` loop can't account for JIT warmup, tiered recompilation, and GC pauses ([Phase 15](15-the-dotnet-runtime-and-gc.md)); `[MemoryDiagnoser]` adds allocation columns - often the number to fix first.
5. **Profilers** (`dotnet-trace`, `dotnet-counters`, `dotnet-gcdump`, Visual Studio, PerfView) find the *real* hotspot. 💡 Measure, don't guess - profile first, optimize second.
6. **Coverage** via coverlet maps which lines ran. ⚠️ Coverage ≠ correctness: it proves a line executed, not that you checked the result. Chase the red, never the score.

You can now prove your C# is correct, measure how fast it really is, and find the slow part with evidence instead of instinct. Next: turning that evidence into action - performance techniques and a tour of the ecosystem you'll lean on for the rest of your C# life.

## Quick check

Test yourself on the ideas that separate "I ran it once and it looked fine" from "I measured it":

```quiz
[
  {
    "q": "Why use `[Theory]` with `[InlineData]` instead of writing four separate `[Fact]` methods?",
    "choices": [
      "One method runs once per data row - each row is an independent case, and adding a scenario is one new line, not a new method",
      "`[Theory]` runs faster than `[Fact]` because it skips the assertion step",
      "`[Theory]` is required whenever a test method has more than one parameter",
      "It automatically generates random inputs to fuzz the method"
    ],
    "answer": 0,
    "explain": "A `[Theory]` runs the same method body once per `[InlineData]` row, each as its own pass/fail with the actual values printed. The behavior under test becomes a visible list, and adding a case is a single line rather than a whole new copy-pasted method."
  },
  {
    "q": "Why can't you trust a `Stopwatch` wrapped around a million-iteration loop to benchmark .NET code?",
    "choices": [
      "JIT warmup, tiered recompilation, and GC pauses distort the timing - BenchmarkDotNet handles warmup and isolation so the number is meaningful",
      "Stopwatch only has millisecond resolution, which is always too coarse",
      "Loops are optimized away entirely, so the code never actually runs",
      "Stopwatch measures wall-clock time instead of CPU time, which is never useful"
    ],
    "answer": 0,
    "explain": "The first calls are JIT-compiled (slow), hot methods get recompiled to faster code mid-loop (tiered compilation), and a GC can fire inside your timing window. BenchmarkDotNet runs warmup until the JIT settles, then enough measured iterations for confidence, and prevents dead-code elimination."
  },
  {
    "q": "Your test suite reports 95% code coverage. What does that actually guarantee?",
    "choices": [
      "That 95% of lines executed during the tests - nothing about whether the results were checked",
      "That 95% of possible bugs have been found",
      "That every method has at least one assertion",
      "That the code is 95% likely to be correct"
    ],
    "answer": 0,
    "explain": "Coverage only measures which lines ran. A test that calls a method and asserts nothing still marks those lines green. Treat coverage as a map of the untested code (chase the red); high coverage with weak assertions feels safe but proves almost nothing about correctness."
  }
]
```


---

# Performance & the Ecosystem

Most of your code is already fast enough, and most of your guesses about *where* it's slow will be wrong. Performance work isn't a bag of clever tricks - it's a discipline: measure, find the one place that actually matters, fix that, and stop. The tricks are the easy part; the discipline is what separates a real speedup from an afternoon of busywork that moved nothing.

This phase caps the deep half of the guide, leaning on the memory model from [Phase 15](15-the-dotnet-runtime-and-gc.md) and the benchmarking/profiling tools from [Phase 16](16-testing-and-profiling.md). We'll put those to work in the order that pays off: measure, fix the algorithm, then cut allocations. Then we'll point you at the .NET ecosystem you're finally ready for.

## Measure first, always

**The mental model.** Your intuition about performance is a liar - not because you're bad at this, but because modern CPUs, caches, the JIT, and the garbage collector interact in ways no human predicts reliably. The method you're *sure* is the bottleneck is often a rounding error, while the real cost hides in a string concatenation you never thought twice about. The only way to know is to look.

⚠️ **The number-one rule of optimization: never optimize on a hunch.** Every time you "speed something up" without a measurement proving it was slow *and* a measurement proving your change helped, you're gambling - and the usual prize is uglier code that runs the same speed (or slower).

The workflow is the one from Phase 16, used in anger:

1. Write a BenchmarkDotNet benchmark exercising the real, representative work.
2. Run it (and, for sharper detail, a profiler) to find where time and allocations go.
3. Fix the single biggest cost.
4. Re-run to *prove* the fix helped. Repeat from step 2.

```console
$ dotnet run -c Release

| Method          | Mean        | Allocated |
|---------------- |------------:|----------:|
| FindDuplicates  | 31,847.2 us |  1.2 MB   |
| ParseRecords    |    412.0 us |  88.4 KB  |
| FormatOutput    |     19.3 us |   2.1 KB  |
```
*What just happened:* BenchmarkDotNet ranked each method by `Mean` time and `Allocated` bytes. One method, `FindDuplicates`, burns roughly 77x more time than the next and allocates the most - that's your hot spot, and nothing else is worth touching until it's handled. (Numbers vary by machine; the *shape* - one method dominating - is typical.)

💡 **Key insight.** In almost every program, a tiny fraction of the code accounts for the overwhelming majority of the runtime. Your job isn't to make *everything* fast - it's to find that 3% and leave the other 97% alone. Optimizing the rest is wasted effort that only adds risk.

## Algorithmic cost dominates

**The mental model.** Before fiddling with a single allocation, ask: *is the approach itself right?* The largest performance wins almost never come from micro-tweaks - they come from replacing a fundamentally expensive strategy with a cheaper one, turning an O(n²) nested scan into an O(n) pass with a `HashSet` or `Dictionary`. No low-level cleverness rescues a quadratic algorithm; it just makes the cliff arrive slightly later.

If "O(n²)" and "O(n)" feel fuzzy, the dedicated primer [Big-O Without the Math Panic](/guides/big-o-without-the-math-panic) walks through exactly what they mean and why they decide who wins as your data grows.

The classic example: find which items in a list appear more than once. The naive version compares every element against every other:

```csharp
// O(n²): for each item, scan all the others looking for a match.
static List<string> FindDuplicatesSlow(List<string> items)
{
    var dups = new List<string>();
    for (int i = 0; i < items.Count; i++)
    {
        for (int j = i + 1; j < items.Count; j++)
        {
            if (items[i] == items[j])
            {
                dups.Add(items[i]);
                break;
            }
        }
    }
    return dups;
}
```

The `HashSet` version makes one pass, remembering what it's already seen:

```csharp
// O(n): one pass, a HashSet remembers what we've already seen.
static List<string> FindDuplicatesFast(List<string> items)
{
    var seen = new HashSet<string>(items.Count);
    var dups = new List<string>();
    foreach (var item in items)
    {
        if (!seen.Add(item))   // Add returns false if it was already there
        {
            dups.Add(item);
        }
    }
    return dups;
}
```

*What just happened:* both methods answer the same question, but the cost curves are nothing alike. The slow version's inner loop means work grows with the *square* of the input - double the items, quadruple the comparisons. The fast version trades a little memory (the `seen` set) for a single linear pass: a `HashSet` lookup is roughly constant-time, so doubling the input only doubles the work. `seen.Add(item)` does double duty, inserting and telling you whether the item was new in one hashed probe. On 10 items the difference is invisible; on 100,000 it's instant versus a coffee break.

Benchmark them side by side and the gap is brutal:

```console
$ dotnet run -c Release

| Method               | Mean         | Allocated |
|--------------------- |-------------:|----------:|
| FindDuplicatesSlow   | 31,847.2 us  |   1.05 MB |
| FindDuplicatesFast   |    233.0 us  |   1.31 MB |
```
*What just happened:* on 10,000 items the `HashSet` version is over a hundred times faster (`us` is microseconds - lower is better), and that multiplier *grows* with the input: at 100,000 items the quadratic version is thousands of times slower. The fast version even allocates a touch more (it builds the `seen` set) and still wins by two orders of magnitude - proof algorithm choice dwarfs allocation tuning. (Exact numbers vary by machine and runtime; the order-of-magnitude gap does not.)

Play with how each growth curve behaves as `n` climbs - it makes the gap concrete in a way numbers on a page can't:

```playground-bigo
```

## Allocations & GC pressure - the usual .NET bottleneck

**The mental model.** Once your algorithm is sound, the most common remaining drag in .NET is *heap allocation*. Recall from [Phase 15](15-the-dotnet-runtime-and-gc.md): every heap object is something the garbage collector must later track and reclaim. Short-lived objects are cheap to create but not free to collect - churn out millions and the GC keeps interrupting your program to clean up. More allocations means more GC work, which steals CPU from your actual code. In C#, "make it faster" very often means "make it allocate less."

📝 **`Allocated` (and Gen 0 collections)** - BenchmarkDotNet's `[MemoryDiagnoser]` reports bytes allocated per run plus how many garbage collections it triggered. Frequently a *better* optimization target than raw time, since cutting allocations cuts GC pressure, lowering time *and* making performance steadier under load.

The single most common waste: building a string by concatenating in a loop. Strings in C# are *immutable* - every `+=` throws away the old string and allocates a brand-new one with the combined contents. Build a string from 10,000 pieces that way and you allocate 10,000 ever-larger strings, nearly all instant garbage.

```csharp
// Wasteful: each += allocates a whole new string and copies everything so far.
static string JoinSlow(string[] parts)
{
    string result = "";
    foreach (var p in parts)
    {
        result += p;   // new string allocated every iteration
    }
    return result;
}

// Lean: StringBuilder writes into one growing buffer, no per-step garbage.
static string JoinFast(string[] parts)
{
    var sb = new StringBuilder();
    foreach (var p in parts)
    {
        sb.Append(p);  // mutates the buffer in place
    }
    return sb.ToString();   // one final string
}
```

*What just happened:* `JoinSlow` allocates a fresh string on every iteration, and since each copies all characters gathered so far, the total work is quadratic *and* the heap fills with discarded strings. `JoinFast` uses a single `StringBuilder` owning one resizable buffer, appending without throwing anything away, producing exactly one string at the end. Same output, a tiny fraction of the garbage.

```console
$ dotnet run -c Release

| Method     | Mean       | Allocated  | Gen0     |
|----------- |-----------:|-----------:|---------:|
| JoinSlow   | 9,214.0 us | 95.4 MB    | 11,000.0 |
| JoinFast   |    38.1 us |  256.2 KB  |     40.0 |
```
*What just happened:* `[MemoryDiagnoser]` added the `Allocated` and `Gen0` columns. The slow version allocated ~95 MB and triggered thousands of Gen 0 collections; the `StringBuilder` version allocated a fraction of that and barely troubled the GC - roughly 240x faster here. (Numbers vary by machine; the direction is reliable.)

The same principle shows up in three other everyday spots:

- **Presize collections you'll fill.** `new List<int>(10_000)` or `new Dictionary<string,int>(capacity)` reserves the backing array once, instead of starting tiny and reallocating a bigger array (copying everything) each time it outgrows itself.
- **Avoid boxing value types.** Storing an `int` (or any `struct`) in an `object`, a non-generic collection, or `params object[]` *boxes* it - a fresh heap allocation per value in a hot loop. Generics (`List<int>`, not `ArrayList`) keep value types unboxed - no per-item heap allocation.
- **Prefer a `struct` for small, hot, short-lived values.** A small value type used heavily in a tight loop lives on the stack and never touches the GC. (Mind the copy cost - keep structs small and ideally immutable, or the copying outweighs the saving.)

## A taste of `Span<T>` and low-allocation APIs

📝 **`Span<T>`** is a window onto a chunk of existing memory - a slice of an array, a string, or a stack buffer - that you can read and write *without copying or allocating*. `text.AsSpan(0, 5)` gives you the first five characters as a view, not a new string. `Memory<T>` is its heap-storable cousin for holding a slice across an `await`. These, with `ArrayPool<T>` (which lends out reusable arrays so you stop allocating fresh buffers), are the modern high-performance toolkit powering fast .NET libraries and ASP.NET Core itself.

```csharp
ReadOnlySpan<char> line = "2026-06-22T09:30".AsSpan();
ReadOnlySpan<char> date = line.Slice(0, 10);   // "2026-06-22" - no new string
ReadOnlySpan<char> time = line.Slice(11);      // "09:30"      - no new string
Console.WriteLine($"{date} at {time}");
```
*What just happened:* `AsSpan()` viewed the existing string's memory, and `Slice` carved out the date and time as *windows* into that memory - zero new allocations, where `Substring` would have allocated two. Parsing this way in a hot path can erase the allocations entirely.

⚠️ **Reach for these only in measured hot paths.** `Span<T>` carries real constraints (it's a `ref struct` - can't be a class field, can't cross an `await`, can't be boxed), and bending code around them costs readability. In the 97% of your code that isn't a bottleneck, a plain `Substring` or `List<T>` is clearer and plenty fast. Spans are a precision tool for the spot the profiler flagged, not a default style.

## Knowing when to stop - and where to go

**The mental model.** Optimization has a point of diminishing - then *negative* - returns. Every clever rewrite makes code harder to read, harder to change, and easier to break, a cost paid by every future reader, including you in six months. The goal is never "as fast as physically possible" - it's "fast enough for the actual requirement, and no more twisted than it has to be."

The discipline that makes this work:

- **Optimize the measured hot path. Leave the rest clear.** The 3% that profiling flagged earns the right to be clever - `Span<T>`, a pooled buffer, a hand-tuned loop. The other 97% should stay as straightforward as possible; that's where you'll spend your reading and debugging life.
- **Define "fast enough" before you start, then stop when you hit it.** A target ("p99 under 50ms," "10k records in under a second") tells you when you're done. Without one, optimization never ends and you keep paying readability for speed nobody needs.
- **Re-measure after every change.** A change that doesn't move the benchmark isn't an optimization - it's a complication. Revert it.

💡 **The closing rule of the deep half.** Readable code that's fast enough beats clever code that's unmaintainable, every time. Performance is a *measured* requirement, met with the *smallest* change that meets it, in the *one* place that needed it. Measure, fix the algorithm, cut the allocations that matter, and then - the hardest part - stop.

**Where C# goes from here - the ecosystem.** You now understand C# from `dotnet run` to the garbage collector. The language was always the foundation; the reason people reach for .NET is the platform built on it. Phase 18 maps it in full, but here's the lay of the land so the names are familiar:

- **ASP.NET Core** - the dominant framework for web APIs and server apps in .NET, and almost certainly how you'd build a backend in C#. It's also where the `Span<T>`/pooling lessons above pay off most directly - tuned to the bone.
- **Entity Framework Core** - the standard object-relational mapper. Write LINQ queries (Phase 13) against C# objects; EF Core turns them into SQL and maps the rows back.
- **Blazor** - build interactive web UIs in C# instead of JavaScript, running on WebAssembly or the server.
- **.NET MAUI** - one C# codebase for native desktop and mobile apps (Windows, macOS, iOS, Android).
- **Unity** - the leading game engine, scripted in C#. A huge share of the world's games are written in the language you just learned.

## Recap

1. **Measure first, always.** Never optimize on a hunch. Use BenchmarkDotNet and a profiler to find the real hot spot; most code is already fast enough, so hunt the 3% that isn't.
2. **Algorithmic cost dominates.** The biggest wins come from a better approach (an O(n) `HashSet`/`Dictionary` lookup over an O(n²) nested scan), not micro-tweaks - you can't optimize out of the wrong complexity class.
3. **Allocations drive GC pressure.** Fewer short-lived heap objects means less GC work, faster and steadier code. Use `StringBuilder` over `+=` in loops, presize collections, avoid boxing value types, prefer small `struct`s for hot values; watch `Allocated` with `[MemoryDiagnoser]`.
4. **`Span<T>` and friends** (`Memory<T>`, `ArrayPool<T>`) slice and reuse memory without allocating - the modern high-perf toolkit. Reach for them only in measured hot paths; they cost readability everywhere else.
5. **Know when to stop, then step out.** Optimize the measured hot path, define "fast enough" up front, revert anything that didn't move the number. Beyond the language lies the ecosystem - ASP.NET Core, EF Core, Blazor, MAUI, Unity - which Phase 18 maps.

That's the deep half done. You can now reason about how C# runs your code *and* make it faster on purpose, with evidence instead of guesses. The final phase steps back: where C# genuinely shines, and where to go next.

## Quick check

Test yourself on the discipline that makes performance work actually pay off:

```quiz
[
  {
    "q": "Before changing any code to make a C# program faster, what should you do first?",
    "choices": [
      "Profile with BenchmarkDotNet and a profiler to find where time and allocations actually go",
      "Replace every class with a struct",
      "Rewrite the slowest-looking method from memory",
      "Wrap every loop in Span<T>"
    ],
    "answer": 0,
    "explain": "Intuition about bottlenecks is unreliable. Measure first so you optimize the real hot spot - the small fraction of code that actually dominates runtime - instead of guessing and adding risk for no gain."
  },
  {
    "q": "You need to find duplicates in a 100,000-item list. Which change usually delivers the biggest speedup?",
    "choices": [
      "Choosing a better algorithm - an O(n) HashSet pass instead of an O(n²) nested scan",
      "Marking the method as static",
      "Renaming variables so the JIT optimizes better",
      "Removing comments from the inner loop"
    ],
    "answer": 0,
    "explain": "Algorithmic complexity dominates. Turning a quadratic nested scan into a linear HashSet pass wins by a margin that grows with the input - no micro-optimization can rescue the wrong complexity class."
  },
  {
    "q": "Why does building a long string with `+=` in a loop hurt performance in C#?",
    "choices": [
      "Strings are immutable, so each += allocates a brand-new string and copies everything so far - creating heaps of garbage for the GC",
      "The += operator is not supported on strings and throws at runtime",
      "It silently converts the string to a struct, which is slower",
      "Each += blocks the thread until the garbage collector runs"
    ],
    "answer": 0,
    "explain": "C# strings are immutable. Every += discards the old string and allocates a new one with the combined contents, so a loop produces many ever-larger throwaway strings - heavy GC pressure. A single StringBuilder mutates one buffer and allocates once at the end."
  }
]
```


---

# Where to Go Next - Putting C# to Work

Notice what you've got: C# the language - types, classes, generics, LINQ, async/await - *and* .NET, the runtime and library it lives on. Most people conflate the two; you now know the difference and can read real code without flinching. Everything from here is *application*.

This last phase is a map, not more syntax. The C# world spans web servers, mobile apps, desktop apps, and games - genuinely rare for one language - and that breadth can feel like pressure to learn it all at once. You don't. Here are the branches, what each is *for*, and the one I'd point most people at first.

## The branches from here

```mermaid
flowchart TD
  You[You: solid C# + .NET] --> Web[ASP.NET Core<br/>web APIs & sites]
  You --> Blazor[Blazor<br/>web UIs in C#]
  You --> MAUI[.NET MAUI<br/>mobile & desktop]
  You --> Unity[Unity<br/>game dev]
```

*What this shows:* four directions lead out from where you stand, all using the same C# you already write. You don't have to pick one forever, but pick *one to go deep on next* - depth beats breadth when learning. For most people, that one is ASP.NET Core.

## Web with ASP.NET Core - the highest-leverage next step

This is the C# job. Companies hiring C# developers are overwhelmingly hiring for **ASP.NET Core** - the framework for building web APIs and server-rendered sites on .NET. Fast, mature, cross-platform, and the center of gravity for the whole ecosystem.

You'll meet a few styles under one roof: **minimal APIs** (a handful of lines for a JSON endpoint - the gentlest on-ramp), **MVC** (classic controllers-and-views for larger apps), and **Razor Pages** (page-focused server rendering). Alongside them lives **Entity Framework Core** - EF Core - mapping your C# classes to database tables so you query with LINQ instead of raw SQL. The `async/await` from [Phase 14](14-async-await-and-tasks.md) is everywhere here, since web servers mostly wait on I/O.

💡 If you only go deep on one branch, make it this one - the most direct path to employability, exercising the most of what you already know.

## Blazor - building web UIs in C#

**Blazor** lets you build interactive browser UIs in *C#* instead of JavaScript - buttons, forms, live-updating components, sharing the same models and validation logic as your backend.

It comes in flavors (server-side, where the UI runs on the server over a live connection, and WebAssembly, where your C# runs *in the browser*), but the headline is the same: if you'd rather not context-switch into a separate JavaScript framework, Blazor keeps you in one language across the whole stack, and pairs naturally with ASP.NET Core.

## Desktop & mobile with .NET MAUI

Want to ship an app running on Android, iOS, Windows, and macOS from one codebase? That's **.NET MAUI** (Multi-platform App UI): write your UI and logic once in C#, and MAUI builds native apps for each platform.

It's the path if you're drawn to building *products people install* rather than websites they visit - a real, supported framework, though web tends to teach more transferable fundamentals for a first deep dive. Consider MAUI once you've got a backend under your belt.

## Game dev with Unity

For a lot of people, *this* is why they're here. **Unity** is one of the most widely used game engines in the world, scripted in C#. The C# you learned here is the same C# you'll write to move a character, spawn enemies, or wire up a menu - Unity layers its own engine APIs on top, but the language is yours already.

If games are the dream, you're not starting over - you're starting at "now I learn the engine."

## Why C# is a great bet

💡 Step back and look at that list: web, desktop, mobile, cloud, *and* games - all reachable from one language. Few languages span that range. C# is also a pleasure day to day: first-class tooling (Visual Studio and the VS Code extension are excellent), genuinely cross-platform (runs happily on Linux and macOS, not Windows-only), Microsoft-backed with serious long-term investment, and evolving fast with new features every year. A sound bet for your time.

## What to actually build

Reading got you here; *building* turns knowledge into skill. The trick is something small enough to finish but real enough to teach you the messy parts:

- **A REST API with ASP.NET Core + EF Core.** Three or four endpoints backed by a real database - exercises minimal APIs, async, LINQ, and EF Core at once, and it's the single most job-relevant thing on this list.
- **A CLI tool.** Turn a chore you do by hand into a console app - small, finishable, and it cements the fundamentals without framework noise.
- **A Blazor to-do app, or a tiny Unity game.** Pick by what excites you: a to-do app teaches components and state, a small game teaches the engine loop. Either is a great first "I made a thing that runs."

Whatever you pick, **finish it**. A finished rough project teaches more than three polished half-projects abandoned at 80%.

## A last word

The official **Microsoft Learn** path for C# and the **C# language docs** are genuinely excellent - well-written, thorough, completely free. Bookmark them. And if you ever want to think about *why* languages make the trade-offs they do - why C# reaches for a runtime and garbage collector while another reaches for manual memory - that's the subject of [Languages, Explained Like a Human](/guides/languages-explained-like-a-human), a good companion now that you've lived inside one language end to end.

You started not knowing what `Console.WriteLine` did. You're leaving able to read real C#, reason about async, model a problem with types, and choose your next step on purpose instead of by panic. Go build the small thing. You're ready.

## Recap

1. **You learned two things** - the C# language *and* the .NET runtime. Everything next is applying them.
2. **ASP.NET Core is the highest-leverage branch** - minimal APIs, MVC, Razor Pages, and EF Core; the dominant C# job, using the most of what you know.
3. **Blazor** builds web UIs in C# instead of JavaScript; **.NET MAUI** builds cross-platform mobile and desktop apps; **Unity** uses C# as its scripting language for games.
4. **C# is a great bet** - one language across web, desktop, mobile, cloud, and games, with first-class tooling, true cross-platform support, and fast Microsoft-backed evolution.
5. **Build one real thing and finish it** - a REST API (most job-relevant), a CLI tool, or a Blazor/Unity app - leaning on Microsoft Learn and the official docs.

## Quick check

Test yourself on the map you just drew:

```quiz
[
  {
    "q": "For most people learning C#, which branch is the highest-leverage next step to go deep on?",
    "choices": [
      "ASP.NET Core - it's the dominant C# job and exercises the most of what you already know",
      "Unity - every C# developer is expected to ship a game first",
      "It doesn't matter; you must learn all four branches at once before building anything",
      ".NET MAUI - because mobile is the only platform worth targeting"
    ],
    "answer": 0,
    "explain": "ASP.NET Core (minimal APIs, MVC, Razor Pages, EF Core) is the most direct path to employability and uses the most of the language and runtime you've already learned. Pick one branch to go deep on - depth beats breadth when learning."
  },
  {
    "q": "What does Blazor let you do?",
    "choices": [
      "Build interactive web UIs in C# instead of JavaScript",
      "Compile C# down to a single static binary with no runtime",
      "Write iOS apps in Swift from inside Visual Studio",
      "Replace the .NET runtime with a faster custom one"
    ],
    "answer": 0,
    "explain": "Blazor lets you build interactive browser UIs in C# - sharing models and logic with your backend - rather than switching into a separate JavaScript framework. It comes in server-side and WebAssembly flavors."
  },
  {
    "q": "Why is the C# you learned in this guide a head start for Unity game development?",
    "choices": [
      "Unity's scripting language is C#, so you already know the language - you just add the engine APIs",
      "Unity automatically converts your console apps into 3D games",
      "Unity uses a different language, but the syntax happens to look similar",
      "Unity only runs on Windows, which is where you learned C#"
    ],
    "answer": 0,
    "explain": "Unity scripts are written in C#. The language is the same one you already know; Unity layers its own engine APIs on top. You start at 'now I learn the engine,' not from scratch."
  }
]
```
