# Testing, Benchmarks & Profiling - Proving It Works and Finding the Slow Part

Back in [Phase 8](08-ecosystem-and-tooling.md) you saw the *shape* of a Go test: a function named `TestXxx(t *testing.T)`, a plain `if`, a `t.Errorf` when reality disagreed. Enough to write your first test - not enough to test a real codebase without drowning.

This phase is the rest of the iceberg. The same `testing` package - no plugins, no DSL - also gives a clean way to run dozens of cases through one function, a way to *measure* how fast and memory-hungry your code is, and a way to *find* the slow part instead of guessing. The mental model: **the Go toolchain refuses to let you operate on vibes.** It pushes you to write the cases down, measure the numbers, look at where the time went. Once internalized, "is it correct?" and "is it fast?" stop being arguments and become commands you run.

## Table-driven tests - the idiomatic Go pattern

**What it actually is.** A table-driven test holds a *slice of test cases* and loops over them, running the same assertion logic against each. Each case is a little struct: inputs, expected output, a name. Instead of copy-pasting the same three lines per scenario, you add a row.

📝 **Table-driven test** - a test where the cases live in data (a slice of structs), and one loop runs every case through the same check. Adding a scenario means a new row, not a new function.

**Why Go prefers this over assertion libraries.** Coming from other languages, you might expect `assertThat(x).isEqualTo(6)`. Go deliberately doesn't ship that, and the community mostly doesn't want it: an assertion library is a second little language to learn, with failure messages written by someone else. A table-driven test is *just Go* - a slice, a `for` loop, an `if`. Data and logic stay separate, so the table reads like a spec you could hand to a stranger, and you control exactly what a failure says.

**A real example.** Testing a `Clamp` function that pins a number into a `[min, max]` range:

```go
// clamp.go
package mathx

func Clamp(n, min, max int) int {
	if n < min {
		return min
	}
	if n > max {
		return max
	}
	return n
}
```

```go
// clamp_test.go
package mathx

import "testing"

func TestClamp(t *testing.T) {
	cases := []struct {
		name           string
		n, min, max    int
		want           int
	}{
		{"inside range", 5, 0, 10, 5},
		{"below min", -3, 0, 10, 0},
		{"above max", 99, 0, 10, 10},
		{"equal to min", 0, 0, 10, 0},
	}

	for _, tc := range cases {
		t.Run(tc.name, func(t *testing.T) {
			got := Clamp(tc.n, tc.min, tc.max)
			if got != tc.want {
				t.Errorf("Clamp(%d, %d, %d) = %d; want %d",
					tc.n, tc.min, tc.max, got, tc.want)
			}
		})
	}
}
```

```console
$ go test -v ./...
=== RUN   TestClamp
=== RUN   TestClamp/inside_range
=== RUN   TestClamp/below_min
=== RUN   TestClamp/above_max
=== RUN   TestClamp/equal_to_min
--- PASS: TestClamp (0.00s)
    --- PASS: TestClamp/inside_range (0.00s)
    --- PASS: TestClamp/below_min (0.00s)
    --- PASS: TestClamp/above_max (0.00s)
    --- PASS: TestClamp/equal_to_min (0.00s)
PASS
ok  	example/mathx	0.004s
```

*What just happened:* We declared an anonymous struct slice - the table - with one row per scenario. The `for` loop walked every row and ran the *same* check. Adding a fifth case is one new line, not a whole new function. `-v` printed each case as it ran, and because each row got its own `t.Run`, the output is a neat tree, not a wall of undifferentiated PASSes.

💡 **Key point.** The win isn't just less typing - the *behavior under test is now a visible list*. A reviewer can scan the table and ask "where's the case for `min > max`?" - much harder when each scenario is buried in its own function.

## Subtests, helpers, and parallelism

The `t.Run(name, func)` you just used unlocks three more things.

**Subtests (`t.Run`).** Each `t.Run` is an independent sub-test with its own name and pass/fail. Crucially, you can run *just one* without running the rest - names are addressable:

```console
$ go test -run TestClamp/below_min -v ./...
=== RUN   TestClamp
=== RUN   TestClamp/below_min
--- PASS: TestClamp (0.00s)
    --- PASS: TestClamp/below_min (0.00s)
PASS
```

*What just happened:* `-run` takes a regex matched against test names, and `/` reaches inside a parent into one subtest - how you re-run *only* a failing row in a tight loop instead of the whole suite.

**Helpers (`t.Helper`).** When you extract repeated assertion logic into a helper function, call `t.Helper()` at the top - it tells the testing package "blame the *caller's* line when I fail, not mine."

```go
func assertEqual(t *testing.T, got, want int) {
	t.Helper() // failures report the caller's line number, not this one
	if got != want {
		t.Errorf("got %d; want %d", got, want)
	}
}
```

*What just happened:* Without `t.Helper()`, every failure points inside `assertEqual` - useless, since they all look identical. With it, the failure points at the line in your test that called the helper.

**Parallelism (`t.Parallel`).** Calling `t.Parallel()` inside a subtest signals it's safe to run alongside other parallel tests, speeding up an I/O-heavy suite. ⚠️ But parallel tests share the process: two touching the same global or file creates a race. Reach for it only when each case is genuinely independent, and run `go test -race` to catch the cases where it isn't.

## Benchmarks - measuring speed and allocations

Correctness is one question. *Speed* is different, and Go answers it with the same `testing` package.

📝 **Benchmark** - a function named `BenchmarkXxx(b *testing.B)` that the toolchain runs many times to measure how long one operation takes. The key piece: `for i := 0; i < b.N; i++` - put the code-under-test inside it, and Go chooses `b.N`, dialing it up until timing is statistically stable.

**Why the `b.N` loop is shaped that way.** You can't time a single function call reliably - it's over in nanoseconds, swamped by clock noise. Go runs your operation `b.N` times (maybe millions), measures the total, and divides. You never set `b.N` yourself; the framework starts small, sees how long that took, and scales up until it trusts the average.

💡 **Newer idiom (Go 1.24+).** `for b.Loop() { ... }` does the same job and is now the recommended form: it manages the count for you, keeps any setup before the loop out of the timed region, and stops the compiler from optimizing the measured code away. `b.N` still works everywhere and fills most existing code, so it's worth reading both - the examples below use `b.N`.

**A real example.** Benchmarking two ways of building a string from a slice - naive `+=` concatenation versus `strings.Builder`:

```go
// build_test.go
package strbuild

import (
	"strings"
	"testing"
)

var parts = []string{"a", "b", "c", "d", "e", "f", "g", "h"}

func BenchmarkConcat(b *testing.B) {
	for i := 0; i < b.N; i++ {
		s := ""
		for _, p := range parts {
			s += p // re-allocates the whole string every time
		}
		_ = s
	}
}

func BenchmarkBuilder(b *testing.B) {
	for i := 0; i < b.N; i++ {
		var sb strings.Builder
		for _, p := range parts {
			sb.WriteString(p)
		}
		_ = sb.String()
	}
}
```

```console
$ go test -bench=. -benchmem ./...
goos: linux
goarch: amd64
pkg: example/strbuild
cpu: Intel(R) Core(TM) i7-9750H CPU @ 2.60GHz
BenchmarkConcat-12     8123140    142.6 ns/op    104 B/op    7 allocs/op
BenchmarkBuilder-12   19543210     61.3 ns/op     56 B/op    3 allocs/op
PASS
ok  	example/strbuild	2.913s
```

*What just happened:* `go test -bench=.` ran every benchmark, `-benchmem` added the memory columns:

- **`ns/op`** - nanoseconds per operation. `Builder` is ~61 ns vs `Concat`'s ~143 ns: more than twice as fast.
- **`B/op`** - bytes allocated per operation. `Builder` allocates 56 bytes; `Concat` allocates 104.
- **`allocs/op`** - *number* of heap allocations per operation. The killer column: `Concat` does 7 (one per `+=`, since strings are immutable so each builds a brand-new string), `Builder` does 3.

The `-12` suffix is the `GOMAXPROCS` the benchmark ran with. The lesson: `+=` in a loop quietly re-allocates the whole string every iteration, and the benchmark makes that invisible cost visible.

💡 **Key point.** `allocs/op` is usually the number to watch first. Allocations create GC work ([Phase 14](14-runtime-scheduler-and-memory.md)), so cutting them often improves the whole program's tail latency. A change that halves `ns/op` but doubles `allocs/op` may be a bad trade under real load.

## Profiling with pprof - measure, don't guess

A benchmark tells you *that* a function is slow, not *where inside it* the time goes. For that, Go has **pprof**.

📝 **pprof** - a profiler (and its analysis tool, `go tool pprof`) that samples your running program to record where it spends CPU time or allocates memory, then lets you inspect the result as a ranked list, call graph, or annotated source. "Profile" is the recorded data; "pprof" is the tool that reads it.

**The mental model: stop guessing.** Every engineer has a confident hunch about the bottleneck, wrong often enough to be dangerous. You optimize the function you *suspected*, ship it, and the program is exactly as slow as before, because the real cost was somewhere you never looked. A profiler replaces the hunch with a measurement, reporting in ranked order where time and memory truly went. The discipline: **profile first, then optimize the thing the profile points at.**

**Capturing a profile from a benchmark.** `go test` can dump a CPU profile while it runs:

```console
$ go test -bench=BenchmarkConcat -cpuprofile=cpu.out ./...
...
$ go tool pprof cpu.out
File: strbuild.test
Type: cpu
Entering interactive mode (type "help" for commands)
(pprof) top5
Showing nodes accounting for 2.31s, 91.3% of 2.53s total
      flat  flat%   sum%        cum   cum%
     1.04s 41.1%  41.1%      1.04s 41.1%  runtime.concatstrings
     0.62s 24.5%  65.6%      0.71s 28.1%  runtime.mallocgc
     0.31s 12.3%  77.9%      0.31s 12.3%  runtime.memmove
     0.21s  8.3%  86.2%      0.21s  8.3%  runtime.nextFreeFast
     0.13s  5.1%  91.3%      2.18s 86.2%  strbuild.BenchmarkConcat
(pprof) 
```

*What just happened:* `-cpuprofile=cpu.out` wrote a profile during the benchmark; `go tool pprof cpu.out` opened it interactively. `top5` ranked functions by **`flat`** time - time spent *in that function itself*, not its callees: 41% into `runtime.concatstrings`, 25% into `runtime.mallocgc` (the allocator). The profile *confirms* the benchmark's hint: this code's cost is string concatenation and the allocations it triggers. `cum` (cumulative) counts a function plus everything it calls, why `BenchmarkConcat` shows a low flat but high cum.

**Profiling a live server.** For long-running services, import `net/http/pprof`, registering profiling endpoints on your HTTP server:

```go
import (
	"net/http"
	_ "net/http/pprof" // blank import: registers /debug/pprof/ handlers as a side effect
)

// ... with your server running, profiles are now served under /debug/pprof/
```

```console
$ go tool pprof http://localhost:8080/debug/pprof/profile?seconds=30
```

*What just happened:* The blank import (`_`) pulls in the package for its `init()` side effect, wiring up handlers under `/debug/pprof/`. `go tool pprof` then collected 30 seconds of live CPU samples and dropped you into the same interactive analysis. (Inside pprof, `web` renders a visual call graph and `list <func>` shows source annotated with per-line cost.)

⚠️ Don't expose `/debug/pprof/` on a public interface - it leaks internals and lets anyone trigger expensive profiles. Bind it to localhost, an admin port, or behind authentication.

This sets up the next phase: profiling is the *evidence-gathering* step. [Phase 17](17-performance-and-optimization.md) covers what to *do* once the profile tells you where the time is.

## Coverage and fuzzing - what tests miss, and finding inputs you didn't

**Coverage** measures how much of your code your tests actually exercise. The toolchain tracks which lines ran:

```console
$ go test -cover ./...
ok  	example/mathx	0.004s	coverage: 87.5% of statements

$ go test -coverprofile=cover.out ./...
$ go tool cover -html=cover.out   # opens a browser: green = covered, red = not
```

*What just happened:* `-cover` printed the share of statements that ran. `-coverprofile` wrote per-line detail to a file, and `go tool cover -html` turned it into a color-coded view: green lines exercised, red never ran. The red is the useful part - it shows the branches your tests forgot.

⚠️ **100% coverage is not "bug-free."** The trap everyone walks into. Coverage tells you a line *executed*, nothing about whether you *checked the result* or tried the inputs that break it. A test that calls `Clamp` once and asserts nothing lights up the whole function green. Treat coverage as a *map of the untested* - chase the red - not a score to max out. High coverage with weak assertions is worse than plain medium coverage, because it *feels* safe.

**Fuzzing** attacks the other blind spot: inputs you never thought to write down. A **fuzz test** generates *random, mutating* inputs, hunting for one that crashes your function or breaks an asserted property.

📝 **Fuzzing** - automated testing where the framework feeds your function a flood of generated inputs (evolved from "seed" examples) to discover ones that cause a panic or violate an invariant. In Go it's built in: a `FuzzXxx(f *testing.F)` function.

```go
// reverse_test.go
package strbuild

import (
	"testing"
	"unicode/utf8"
)

func Reverse(s string) string {
	r := []rune(s)
	for i, j := 0, len(r)-1; i < j; i, j = i+1, j-1 {
		r[i], r[j] = r[j], r[i]
	}
	return string(r)
}

func FuzzReverse(f *testing.F) {
	f.Add("hello")        // seed corpus: starting examples
	f.Add("Göteborg")     // includes multi-byte runes on purpose
	f.Fuzz(func(t *testing.T, s string) {
		if !utf8.ValidString(s) {
			return // []rune can't round-trip invalid UTF-8, so skip it
		}
		// property: reversing twice returns the original
		if Reverse(Reverse(s)) != s {
			t.Errorf("double reverse changed %q", s)
		}
	})
}
```

```console
$ go test -fuzz=FuzzReverse
fuzz: elapsed: 0s, gathering baseline coverage: 0/192 completed
fuzz: elapsed: 3s, execs: 412903 (137621/sec), new interesting: 18
...
fuzz: elapsed: 12s, execs: 1843201 (no new interesting), 0 failures
^C
```

*What just happened:* `f.Add` seeded a few starting strings, and `f.Fuzz` ran a *property check* - "reversing twice gets the original back" - against a torrent of generated inputs. We asserted a property rather than specific outputs, since we don't know in advance what Go will invent. We skip invalid UTF-8 because `[]rune` can't round-trip it - drop that one guard and the fuzzer finds a failing input within seconds, which is exactly the kind of edge case it exists to surface. If a generated string ever broke the property, Go would stop, *shrink* it to the smallest failing input, and save it as a permanent regression test. (A normal `go test` still runs the seed corpus as ordinary cases - fuzzing only does open-ended generation under `-fuzz`.)

💡 **Key point.** Coverage asks "what code did my chosen inputs reach?" Fuzzing asks "what inputs did I never think to choose?" Coverage finds untested *code*; fuzzing finds untested *cases* - the empty string, the lone multi-byte rune, the overflow. Complementary; neither alone makes your code safe.

## Recap

1. **Table-driven tests** put your cases in a slice of structs and loop one assertion over all of them - adding a scenario is a new row, not a new function. Go prefers this to assertion libraries because it's *just Go*: readable as a spec, with failure messages you control.
2. **`t.Run`** creates addressable subtests (re-run one with `-run TestX/case`), **`t.Helper()`** makes a helper's failures point at the caller, and **`t.Parallel()`** opts a test into concurrent execution (pair it with `-race`).
3. **Benchmarks** (`BenchmarkXxx(b *testing.B)`) run your code `b.N` times - a count Go chooses - and `go test -bench=. -benchmem` reports **ns/op**, **B/op**, and the all-important **allocs/op**.
4. **pprof** replaces guesswork with measurement: capture a profile (`-cpuprofile` from a benchmark, or `net/http/pprof` on a server), open it with `go tool pprof`, and let `top`/`list` point you at the real hot spot before you optimize anything.
5. **Coverage** (`go test -cover`) maps which lines ran - use it to find the *red*, never as a score, because ⚠️ 100% coverage with weak assertions proves nothing.
6. **Fuzzing** (`FuzzXxx(f *testing.F)`) generates inputs you'd never write by hand and checks a *property*; it finds the edge case, shrinks the failure, and saves it as a regression.

You can now prove your Go is correct, measure how fast it is, and find the slow part with evidence instead of instinct. Next: that lens turned on the standard library itself, a masterclass in Go's design philosophy.

## Quick check

Test yourself on the ideas that separate "I ran a test" from "I measured my program":

```quiz
[
  {
    "q": "In a benchmark, why do you wrap the code under test in `for i := 0; i < b.N; i++`?",
    "choices": [
      "Go runs the operation b.N times - a count it picks automatically - and divides, so a single fast call isn't swamped by clock noise",
      "b.N is a constant you must set to the number of CPU cores",
      "The loop makes the benchmark allocate less memory",
      "It's required syntax with no effect on the measurement"
    ],
    "answer": 0,
    "explain": "A single nanosecond-scale call can't be timed reliably. Go dials b.N up until the total run time is statistically stable, then reports the per-operation average. You never set b.N yourself - you just make the loop body do exactly the work you want measured."
  },
  {
    "q": "Your test suite reports 100% coverage. What does that actually guarantee?",
    "choices": [
      "Every line of code executed at least once during the tests - nothing about whether results were checked",
      "The code has no bugs",
      "Every possible input was tested",
      "All assertions in the tests passed correctly"
    ],
    "answer": 0,
    "explain": "Coverage only measures which lines ran. A test that calls a function and asserts nothing still marks those lines green. Treat coverage as a map of the untested (chase the red); high coverage with weak assertions feels safe but proves almost nothing."
  },
  {
    "q": "You suspect a function is your bottleneck. What does pprof give you that a benchmark doesn't?",
    "choices": [
      "It samples the real run and ranks where time and memory actually went, so you optimize the proven hot spot instead of your hunch",
      "It automatically rewrites the slow function to be faster",
      "It guarantees the function has no allocations",
      "It only works on functions named BenchmarkXxx"
    ],
    "answer": 0,
    "explain": "A benchmark tells you that something is slow; pprof tells you where inside it the time goes. The discipline is to profile first and then optimize the thing the profile points at - because the function you suspected is wrong often enough to waste real effort."
  }
]
```
