# Probability & Statistics

> Probability measures how likely something is; statistics makes sense of data once it's collected. Together they're the everyday superpower for reasoning under uncertainty - and the single best defense against being fooled by numbers.


---

# Probability & Statistics

Almost every important decision is made without certainty - will this launch work, is this result real,
is that headline's scary number actually scary? **Probability** is the math of *how likely*, and
**statistics** is the math of *making sense of data* once you have it. They are the two halves of
thinking clearly about an uncertain world, and they're the last foundation in the Mathematics track for
a reason: nearly everything downstream - A/B tests, machine learning, risk, polls - runs on them.

There's a second reason this guide matters more than most. Numbers are the favorite costume of
misleading arguments: a cherry-picked average, a biased sample, a chart with a doctored axis, a
"correlation" sold as a cause. The same statistics that help you understand the world are the ones used
to bamboozle you about it. So this guide does both jobs - it teaches the tools straight, and it teaches
you how those tools get *abused*, so the next misleading number doesn't get past you.

## How to read this
- **Want the core idea?** [Phase 1](01-probability-measuring-uncertainty.md) is probability from zero.
- **Want the self-defense?** [Phase 3](03-how-statistics-mislead-you.md) is the catalog of how numbers
  lie - but Phases 1–2 are what let you spot it.

## The phases
1. **[Probability: Measuring Uncertainty](01-probability-measuring-uncertainty.md)** - likelihood from 0
   to 1, combining events, and expected value.
2. **[Reading Data: Statistics That Don't Lie](02-reading-data-statistics.md)** - mean, median, spread,
   and the shape of a distribution (and when the mean betrays you).
3. **[How Statistics Mislead You](03-how-statistics-mislead-you.md)** - correlation vs causation,
   sampling bias, base rates, and the charts built to deceive.

> This builds on [Counting & Combinatorics](/guides/counting-and-combinatorics) (probability is counting
> favorable outcomes) and pairs with
> [Critical Thinking & Fallacies](/guides/critical-thinking-and-fallacies). It's the final guide of the
> Mathematics foundations.


---

# Probability: Measuring Uncertainty

You flip a coin. You don't know how it lands until it lands - yet you'd happily bet even
money on it, and you'd refuse that same bet on a die showing a 4. Some unknowns feel more
likely than others, and you sense it in your gut. Probability turns that gut feeling into a
number you can reason with, compare, and combine.

This isn't about predicting the future. It's about measuring uncertainty in plain terms: "I don't
know what will happen, but here's exactly how much I don't know."

## A probability is a number from 0 to 1

Every probability is a single number between **0 and 1**:

- **0** means the event is impossible. It will not happen.
- **1** means the event is certain. It will happen.
- **0.5** means it's a coin flip - as likely to happen as not.

Anything in between is a shade of "maybe." A probability of 0.9 means very likely; 0.1 means
unlikely but not ruled out. People often write these as percentages - 0.9 is 90%, 0.25 is
25% - the same number wearing a different outfit. Multiply by 100 for a percentage; divide by
100 to get back.

📝 If someone hands you a "probability" of 1.4 or −0.2, something is wrong. Probabilities live
in the range 0 to 1, full stop. Less a rule to memorize than a free sanity check.

## Equally likely outcomes: favorable over total

The cleanest case is when every outcome is **equally likely** - a fair coin, a fair die, a
shuffled deck. Then probability is a counting problem:

```
P(event) = number of favorable outcomes / total number of outcomes
```

Roll one fair die. There are 6 possible outcomes (1 through 6), all equally likely. Exactly
one is a 4. So:

```
P(rolling a 4) = 1 / 6 ≈ 0.167
```

This is why [counting came first](/guides/counting-and-combinatorics). Equally-likely
probability is careful counting: count the outcomes you want, count all the outcomes, divide.
When the counting gets hard - "what's the probability of a full house?" - the probability is
hard for the same reason. Same skill underneath.

💡 The phrase "equally likely" does real work. This formula applies only when no outcome is
favored over another. A weighted die or loaded coin breaks it: you can't merely count, you'd
need each outcome's weight.

## The complement: the easy way in

Sometimes the thing you want is awkward to count, but its **opposite** is easy. The complement
of event A is "A does not happen," and the two always add up to certainty:

```
P(not A) = 1 − P(A)
```

A 1/6 chance of rolling a 4 means a 5/6 chance of *not* rolling a 4. That seems too simple to
matter - until you hit "what's the chance of at least one 6 in several rolls?" Counting every
way to get "at least one" is a headache. Counting the *one* way to get "none" is trivial. So
compute the easy opposite and subtract from 1. You'll see this pay off in the runnable example
below.

## Combining events: multiply for "and," add for "or"

Most real questions involve more than one event. There are two core moves, mirroring the
and-vs-or split from counting.

**Independent events multiply (the "and" case).** Two events are independent when one happening
doesn't change the odds of the other. Two separate coin flips are independent - the first coin
has no memory. For the probability that *both* happen, multiply:

```
P(A and B) = P(A) · P(B)        (when A and B are independent)
```

Two coin flips both landing heads:

```
P(heads and heads) = 1/2 · 1/2 = 1/4
```

That matches intuition: of the four equally likely pairs (HH, HT, TH, TT), exactly one is HH.

**Mutually exclusive events add (the "or" case).** Two events are mutually exclusive when they
can't both happen at once. Rolling a 4 and rolling a 5 on one die can't both be true. For the
probability that *one or the other* happens, add:

```
P(A or B) = P(A) + P(B)         (when A and B can't both happen)
```

```
P(rolling a 4 or a 5) = 1/6 + 1/6 = 2/6 = 1/3
```

A way to keep them straight: **"and" narrows down** (you need both, so the chance shrinks -
two numbers below 1 multiply to something smaller), while **"or" opens up** (more ways to win,
so the chance grows).

## Expected value: the long-run average

Probability tells you how likely each outcome is. **Expected value** tells you what you'd
average over many repeats. Multiply each possible value by its probability, then add:

```
expected value = sum of (each value × its probability)
```

Say a game pays $10 if a die shows a 6, and $0 otherwise:

```
expected value = $10 · (1/6) + $0 · (5/6) = $10/6 ≈ $1.67
```

Any single play gives you $10 or nothing, but over thousands of plays you'd average about
$1.67 each. That number tells you whether a $2 entry fee is a bad deal (it is - you'd lose
about $0.33 per play on average).

## Try it: at least one 6 in two rolls

Here's the complement trick in action. Computing "at least one 6 in two rolls" directly means
juggling several cases; computing "no 6 at all" is one clean multiplication.

```python runnable
# P(at least one 6 in two dice rolls) via the complement
p_no_six_one_roll = 5/6
p_no_six_twice = p_no_six_one_roll ** 2
print(round(1 - p_no_six_twice, 4))   # ~0.3056
```

*What just happened:* Instead of counting every way to get at least one 6, we counted the
single easy opposite - getting *no* 6. Each roll misses the 6 with probability 5/6, and the
two rolls are independent, so both miss with probability (5/6)² ≈ 0.6944. Subtract that from
1 and you get ≈ 0.3056: there's about a 31% chance of seeing at least one 6 across two rolls.
Notice we used all three ideas at once - the complement, independence (multiply), and the
0-to-1 range.

## For builders

This shows up in code more than you might expect:

- **Randomness and sampling.** When you pick a random element, shuffle a list, or A/B-test a
  feature, you're generating outcomes with known probabilities. The math lets you check
  whether your "random" distribution is actually fair.
- **Retries and failure probabilities.** If one network call fails 1% of the time
  (probability 0.01) and failures are independent, the chance that three independent retries
  *all* fail is 0.01³ = 0.000001 - one in a million. That's the entire argument for retry
  logic, in one multiplication.
- **Rare events at scale.** A "1-in-a-million" event sounds like it'll never happen. But if
  your service handles a million requests an hour, *expect* it roughly once an hour. Low
  probability times huge volume equals a regular occurrence. This is why large-scale systems
  hit bugs that "can't happen" - at scale, rare is routine.

## ⚠️ Gotcha: independence is a precondition, not a default

Multiplying probabilities is valid only when the events are genuinely independent - when one
truly doesn't affect the other. Drawing two cards *without* replacing the first changes the
deck, so those draws aren't independent, and naive multiplication gives the wrong answer.

The most famous version is the **gambler's fallacy**: a coin lands heads five times in a row,
and someone "feels" tails is overdue. It isn't. The coin has no memory; the next flip is still
exactly 1/2. The past flips and the next flip are independent, so the past tells you nothing
about the next outcome. Before you multiply, always ask: does the first event change the odds
of the second? If yes, you can't blindly multiply.

## Recap

- A probability is a single number from **0 (impossible) to 1 (certain)**, often shown as a
  percentage.
- For **equally likely** outcomes, `P(event) = favorable / total` - it's counting in disguise.
- The **complement** rule, `P(not A) = 1 − P(A)`, is often the shortcut, especially for "at
  least one" questions.
- **Independent** events **multiply** ("and"); **mutually exclusive** events **add** ("or").
- **Expected value** is the long-run average: sum of (value × its probability).
- Independence is a condition you must check - forgetting it is the gambler's fallacy.

A quick check before you move on:

```quiz
[
  {
    "q": "A fair eight-sided die has faces 1 through 8. What is the probability of rolling a 3?",
    "choices": ["1/8", "3/8", "1/3", "8/3"],
    "answer": 0,
    "explain": "Equally likely outcomes: favorable over total. There's one face showing 3 out of 8 total faces, so P = 1/8. (Notice it sits between 0 and 1, as every probability must.)"
  },
  {
    "q": "You flip a fair coin twice. What is the probability of getting heads both times?",
    "choices": ["1/2", "1/4", "1", "2/3"],
    "answer": 1,
    "explain": "The flips are independent, so you multiply: P(heads and heads) = 1/2 · 1/2 = 1/4."
  },
  {
    "q": "If the probability of rain tomorrow is 0.3, what is the probability of no rain?",
    "choices": ["0.3", "0.5", "0.7", "1.3"],
    "answer": 2,
    "explain": "The complement rule: P(not A) = 1 − P(A) = 1 − 0.3 = 0.7."
  }
]
```


---

# Reading Data: Statistics That Don't Lie

You have a pile of numbers - salaries, page-load times, exam scores - and someone asks the
most natural question in the world: *so, what's it like?* You can't read out all 10,000
values. You need to compress them into a few clear numbers that capture the shape of the pile
without lying about it.

That's this phase. Three questions, three kinds of answer: **where is the center?**, **how
spread out is it?**, and **what shape does it make?** Get those three and you can describe
almost any dataset in a sentence. Get the first one wrong and every downstream conclusion
inherits the lie.

## Where the data sits: measures of center

The "center" of a dataset is your one-number summary of a typical value. There are three
common ways to find it, and they don't always agree.

Take this tiny dataset of quiz scores: `[3, 5, 5, 7, 10]`.

- **Mean** (the "average"): add everything up, divide by how many there are.
  `(3 + 5 + 5 + 7 + 10) / 5 = 30 / 5 = 6`. The mean is the balance point.
- **Median** (the middle value): sort the numbers, take the one in the middle. Sorted, the
  middle of `[3, 5, 5, 7, 10]` is `5`. Half the data sits below it, half above. With an even
  count, average the two middle values.
- **Mode** (the most common value): the value that shows up most. Here `5` appears twice and
  everything else once, so the mode is `5`.

For this gentle dataset, mean (6), median (5), and mode (5) sit close together. That's the
comfortable case. The interesting - and dangerous - case is when they pull apart.

## Mean vs median: the one idea that matters most

Here's the single most useful thing in this phase, so slow down for it.

Imagine a small room with ten people. Nine earn around \$50,000 a year. The tenth is a
billionaire who earns \$1,000,000,000.

What's the **average** (mean) income? Add the nine \$50,000 salaries (\$450,000) to the
billion, divide by ten - roughly \$100,045,000. So the headline reads: *"Average income in
this room: over \$100 million."*

Every word of that is arithmetically true and completely useless. Nobody in that room lives
like they earn \$100 million. Nine earn \$50k; one earns a fortune. The mean got **dragged**
toward the one extreme value, and now it describes nobody.

The **median** doesn't flinch. Line up all ten incomes and look at the middle: still about
\$50,000. The median gives a straight answer to "what does a typical person here earn?" because a single
huge value can't move the middle of the line - it sits at the far end.

This is the core lesson: **the mean is sensitive to outliers and skew; the median resists
them.** When data is lumpy - a few values far from the rest - the mean tells you about the
lump, and the median tells you about the people.

That's why you hear "median household income" and "median home price" in the news, almost never
"average." For money, the median is the trustworthy number.

## How spread out: range and standard deviation

Knowing the center isn't enough. Two datasets can share a mean and feel completely different.
`[50, 50, 50]` and `[0, 50, 100]` both average to 50, but one is identical everywhere and the
other is all over the place. You need a number for **spread**.

The simplest is the **range**: largest value minus smallest. For `[0, 50, 100]` the range is
`100 - 0 = 100`. Quick, but fragile - one freak value blows it up, and it says nothing about
what happens in between.

The workhorse is **standard deviation**. The intuition is all you need right now:

> Standard deviation is the *typical distance* a value sits from the mean.

A small standard deviation means values cluster tightly around the average. A large one means
they're scattered. If exam scores average 70 with a standard deviation of 4, almost everyone
scored near 70. If it's 25, scores are flung from near-zero to near-perfect, and "the average
is 70" hides a lot.

(Under the hood: measure how far each value is from the mean, square those distances so
positives and negatives don't cancel, average them - that average is the **variance** - then
take the square root to get back to the original units. You don't grind the formula by hand; a
calculator or one line of code does it. What you need is the mental model: *typical distance
from the average*.)

## What shape: distributions and percentiles

A **distribution** is the shape the data makes when you sort it into buckets and ask "how many
values land here, how many there?" Picture a histogram: tall bars where values pile up, short
bars where they're rare.

The most famous shape is the **normal distribution** - the "bell curve." Most values cluster
around the center and thin out symmetrically as you move away in either direction. Adult
heights, measurement errors, and many natural quantities land roughly here. When data is
bell-shaped, the mean and median sit together in the middle, and standard deviation describes
the width of the bell.

Plenty of real data is **not** symmetric. When a long tail stretches off to one side, the
distribution is **skewed**. Incomes, latencies, and file sizes are classic right-skewed shapes:
most values are modest, with a few very large ones trailing off to the right. That long tail is
exactly what drags the mean away from the median - which is why those are the cases where the
median earns its keep.

To talk about position inside a distribution, use **percentiles**. The *p*th percentile is the
value below which *p* percent of the data falls. The **median is the 50th percentile** - half
the data sits below it. The 95th percentile (**p95**) is the value 95% of your data comes in
under; only the slowest or largest 5% exceed it. Percentiles let you describe the tail
precisely instead of hand-waving about "the big ones."

## See it move

Watch the mean and median react differently to a single outlier. This uses Python's
standard-library `statistics` module - nothing to install.

```python runnable
import statistics
data = [2, 3, 3, 4, 4, 4, 100]   # one big outlier
print(statistics.mean(data))      # pulled up by 100
print(statistics.median(data))    # resists the outlier
print(round(statistics.pstdev(data), 2))
```

*What just happened:* the mean comes out to about `17.14`, even though six of the seven values
are between 2 and 4. That single `100` dragged the average into a range where no actual data
point lives. The median is `4` - the genuine middle of the pile - and it doesn't budge no
matter how extreme that outlier gets. The standard deviation is large precisely because one
value sits so far from the mean. Same data, two very different stories about "the typical
value," and only one tells it straight.

## For builders: why teams track p95, not the average

If you run a web service, you care how fast it responds. The tempting metric is **average
response time** - but request latencies are right-skewed, and the average hides the pain.

Say 99 requests return in 50 ms and one stalls at 5,000 ms. The average is about 99.5 ms,
which sounds fine. But one user waited five full seconds, and the average quietly buried them.
Latency data almost always has this long tail, so the mean flatters you while real users
suffer.

That's why serious teams watch **percentiles** instead:

- **p50** (the median): the typical experience - half of requests are faster.
- **p95**: 95% of requests come in under this. A reasonable "most people are okay."
- **p99**: the slowest 1%. This is where your worst real experiences live - the spinning
  loaders, the rage-quits, the support tickets.

Optimizing the average can mean shaving milliseconds off requests that were already fast.
Optimizing p99 means fixing the requests that are actually hurting people. The percentile tells
you where the suffering is; the average tells you a comforting story. This same instinct shows
up anywhere you summarize counts - and counting itself has its own toolkit in
[/guides/counting-and-combinatorics](/guides/counting-and-combinatorics).

> ⚠️ **"Average" almost always means the mean.** And for skewed data - incomes, latencies, file
> sizes, response times - the mean is the number most likely to mislead you. When you see an
> average reported on lumpy data, ask for the median (and a percentile or two) before you trust
> it. The mean isn't wrong; it's answering a different question than the one you probably care
> about.

## Recap

- **Center**: the **mean** is the balance point, the **median** is the middle value, the
  **mode** is the most common value. On gentle data they agree.
- **Mean vs median** is the key skill: the mean gets pulled by outliers and skew; the median
  resists them. For money and other lumpy data, the median is the reliable summary.
- **Spread**: the **range** is max minus min; **standard deviation** is the typical distance of
  values from the mean. Same center, different spread = different data.
- **Shape**: a **distribution** describes how values pile up. The **normal** curve is the
  symmetric bell; **skewed** data has a long tail. **Percentiles** pin down position - the
  median is the 50th percentile, and p95/p99 expose the tail.

A quick check before you move on:

```quiz
[
  {
    "q": "Nine people earn about $50k and one earns $1 billion. Which number better describes a typical income in the room?",
    "choices": ["The mean, because it uses every value", "The median, because the outlier barely moves it", "The range, because it shows the gap", "They're equally good summaries here"],
    "answer": 1,
    "explain": "The single huge income drags the mean to over $100 million - a value nobody actually earns. The median stays near $50k because one extreme value can't shift the middle of the sorted list."
  },
  {
    "q": "What does the standard deviation of a dataset describe?",
    "choices": ["The most common value", "The middle value when sorted", "How far values typically sit from the mean (the spread)", "The difference between the largest and smallest value"],
    "answer": 2,
    "explain": "Standard deviation is the typical distance of values from the mean. A small one means tight clustering; a large one means values are scattered widely. (The largest-minus-smallest gap is the range.)"
  },
  {
    "q": "The median is equivalent to which percentile?",
    "choices": ["The 50th percentile", "The 95th percentile", "The 100th percentile", "The 0th percentile"],
    "answer": 0,
    "explain": "The pth percentile is the value below which p percent of the data falls. The median splits the data in half, so exactly 50% sits below it - making it the 50th percentile."
  }
]
```


---

# How Statistics Mislead You

In Phase 2 you learned to read data clearly: averages, spread, distributions. This phase is
the other half. The same tools - averages, percentages, correlations, charts - are the favorite
weapons of anyone who wants to sell you something, win an argument, or move a dashboard number.

Here's the uncomfortable part: most misleading statistics aren't lies. The numbers are real.
The arithmetic checks out. What's broken is the *story wrapped around the numbers* - what got
left out, who got counted, which slice got shown. That's why this fools smart people. You can't
catch it by checking the math. You catch it by knowing the moves.

So this phase is a field guide. Each section is one trick: what it looks like, why it works on
you, and the question that defuses it. Once you've seen them, you don't un-see them.

## Correlation is not causation

Two things move together. Ice cream sales rise; so do drownings. Plot them month by month and
the lines track beautifully. A tempting conclusion writes itself: ice cream causes drowning.

It doesn't, and you know why - summer. Hot weather drives *both* ice cream sales and swimming
(and therefore drownings). The two numbers are linked, but neither causes the other. They share
a hidden third cause. Statisticians call it a **lurking variable** or **confounder**, and it's
behind a huge share of "shocking finding" headlines.

Three things can be true when two numbers correlate:

- A causes B (smoking → cancer).
- B causes A (you assumed the arrow points the wrong way).
- C causes both A and B (the ice cream / summer trap above).

And a fourth, sneakier possibility: nothing causes anything. With enough data series in the
world, some line up by pure chance. The number of films a certain actor appears in per year
might track national cheese consumption for a decade - perfectly, and meaninglessly. That's a
**spurious correlation**, and the lesson is brutal: a tight line on a chart proves two numbers
*moved together*, and nothing about *why*.

*What just happened:* you learned the single most abused word in statistics. "Linked to,"
"associated with," and "tied to" are how careful writers say *correlation* - and how careless
ones smuggle in *causation*. When you read "X linked to Y," ask: what's the third thing that
could cause both? For the broader pattern of mistaking sequence and coincidence for cause, see
[/guides/critical-thinking-and-fallacies](/guides/critical-thinking-and-fallacies).

## Sampling bias: who got counted

A statistic describes a *sample* - the people or things you actually measured - then claims to
speak for a *population*, everyone you care about. That leap is valid only if the sample looks
like the population. When it doesn't, the conclusion is broken before the math starts.

The classic example is the WWII planes. Engineers studied bombers returning from missions to
decide where to add armor. The returning planes were riddled with bullet holes on the wings and
tail, almost none on the engines. The obvious move: armor the wings and tail. The statistician
Abraham Wald said the opposite - armor the engines. Why? The sample was only the planes that
*came back*. Planes hit in the engines didn't return to be measured. The holes they could see
marked where a plane could survive being hit; the clean spots marked where it couldn't.

That's **survivorship bias**, and it's everywhere: "every successful founder dropped out of
college" (you're not counting the dropouts who failed and vanished), "this old building was
built to last" (the flimsy ones from that era are long gone). You're studying the survivors and
mistaking them for the whole story.

The everyday version is the **self-selected survey**. An online poll asking "are you angry about
this policy?" only hears from people angry enough to click. "We surveyed our own users and 90%
love the feature" - the users who hated it already left. The sample selected itself, and it
selected *for the answer you got*.

The defusing question is always the same: **who is missing from this data, and would they answer
differently?**

## Base rate neglect: forgetting how rare things are

This is the most counterintuitive trick in the catalog, so we'll walk it slowly with round,
hypothetical numbers.

Imagine a disease that affects 1 in 1,000 people. There's a test that is "99% accurate" - if
you have the disease it says yes 99% of the time, and if you don't it correctly says no 99% of
the time. You take the test. It comes back positive. What's the chance you actually have the
disease?

The instinct screams "99%." The real answer is about **9%**. Here's why, with a hypothetical
population of 100,000 people:

```text
Population:                    100,000
Actually have the disease:          100   (1 in 1,000)
Don't have it:                   99,900

Of the 100 sick people:
  test catches 99% →                 99 true positives

Of the 99,900 healthy people:
  test wrongly flags 1% →           999 false positives

Total positive results:        99 + 999 = 1,098
Chance a positive is real:     99 / 1,098 ≈ 9%
```

The 1% error rate sounds tiny - until it's applied to the *enormous* group of healthy people. A
small slice of a huge number (999) dwarfs the true cases (99). What the "99% accurate" headline
forgot to mention is the **base rate**: how common the condition is to begin with. When the base
rate is low, even an excellent test produces mostly false alarms.

This isn't a math curiosity. It's how to think about rare events generally: airport-screening
hits, fraud-detection flags, a rare-bug alert that fires constantly. Before you trust a
positive, ask how rare the real thing is. The rarer it is, the more a "positive" is probably
noise.

> Rule of thumb: accuracy is meaningless without the base rate. "99% accurate" and "mostly
> wrong" can both be true about the same test at the same time.

## Cherry-picking and p-hacking: torturing the data

Run one coin-flipping experiment and getting 8 heads out of 10 is mildly surprising. Run a
*thousand* and some will hit 8, 9, even 10 heads - guaranteed, by pure chance. Now report only
those. You've "proven" the coin is loaded, using nothing but luck and selective reporting.

That's the engine behind two related tricks:

- **Cherry-picking** - showing only the data that flatters your point. "Sales are up 40%!"
  (since the worst month last year, conveniently chosen as the starting line). The same dataset
  with a fairly chosen start date might show flat or falling sales.
- **p-hacking** - testing so many things that something crosses the "statistically significant"
  line by accident, then presenting that one result as if it were the question all along. Does
  this food cause cancer? Test it against 20 diseases and one will look significant at the usual
  threshold roughly 1 time in 20 *even if the food does nothing*. Report that one. Bury the
  nineteen.

The tell is **the missing denominator**: how many things did you test, how many time windows did
you try, before you found the one you're showing me? A result chosen *after* looking at the data
is a much weaker claim than one predicted *before*. (If you've read
[/guides/counting-and-combinatorics](/guides/counting-and-combinatorics), you'll feel the trap
in your gut: enough chances, and even a rare outcome becomes near-certain to appear *somewhere*.)

## Misleading charts: true data, lying pictures

A chart can be completely accurate and still deliberately deceive, because your eye reads the
*picture*, not the numbers. The most common moves:

- **Truncated y-axis (no zero baseline).** A bar chart of values 98, 99, 100 looks flat if the
  axis starts at 0 - and looks like a dramatic 3x cliff if it starts at 97. Same numbers,
  opposite story. Bar charts almost always need a zero baseline, because the *length* of the bar
  is the message.
- **Cropped time range.** Show only the slice where the line does what you want. A stock that's
  flat over five years can look like a rocket if you zoom into its best three months.
- **Mismatched or dual scales.** Two lines on two different y-axes, scaled so they appear to
  "track" each other - manufactured correlation, on purpose.
- **Inverted or unlabeled axes.** Rare, but devastating: an axis flipped upside down so a *rise*
  in deaths reads visually as a *decline*.

The defense is mechanical. Before you trust a chart, read the axes out loud. Where does the
y-axis start? What's the time range, and why that range? Is each axis labeled with real units? A
chart that won't answer those questions is hiding something.

```text
Same data, two charts:

  y starts at 0          y starts at 97
  100 ┤ ▇ ▇ ▇            100 ┤        ▇
      │ ▇ ▇ ▇             99 ┤    ▇   ▇
   50 ┤ ▇ ▇ ▇             98 ┤▇   ▇   ▇
      │ ▇ ▇ ▇             97 ┤▇   ▇   ▇
    0 └────────              └────────
   "basically flat"      "exploding growth!"
```

## Small samples: numbers that swing wildly

"100% of customers recommend us!" - three customers. "This treatment doubled survival!" - from
two patients to four.

Small samples are unstable by nature. Flip a fair coin four times and getting all heads happens
about 1 time in 16 - often enough that it'll happen to *somebody*, who will then swear the coin
is magic. The fewer the data points, the more any single result is dominated by luck rather than
truth. This is the flip side of the base-rate and p-hacking traps: rare flukes are common when
you have lots of small samples lying around.

So a headline rate means little without the **sample size** behind it. A 70% success rate from
1,000 trials is a real signal. A 100% success rate from 3 trials is a coin that happened to land
heads three times. Always ask: *out of how many?* A percentage with no denominator is a mood, not
a measurement.

## For builders: A/B tests and vanity metrics

Everything above shows up in your own work the moment you start measuring a product.

**A/B test pitfalls** - running variant A against variant B and reading the result:

- **Peeking early.** You check on day two, variant B is "winning," you ship it. But early
  numbers swing wildly (small samples again), and significance thresholds assume you decide
  *when* to stop *before* you start. Stopping the moment you like the result manufactures false
  wins. Pick your sample size up front and wait for it.
- **Too-small samples.** A "12% lift" from 40 users is noise wearing a suit. Underpowered tests
  produce confident-looking numbers that evaporate when you run them again.
- **Multiple comparisons.** Testing ten button colors at once is p-hacking by another name - one
  will "win" by chance. The more variants and metrics you check, the more accidental "winners"
  you'll find.

**Vanity metrics** - numbers that look great and mean nothing:

- Total signups (a number that only goes up, even as everyone churns).
- Page views without engagement (a spike from one viral link that converts no one).
- Followers, downloads, raw totals - any cumulative count that can't go down and so can never
  tell you something got worse.

The real replacements measure a *rate* or a *retained* behavior: active users this week,
conversion percentage, whether people came back. If a metric can't get worse, it can't teach you
anything.

## The catalog, and where you go from here

Step back and look at what you can now do. Someone shows you a statistic. You run the checklist:

- Are they sliding from *correlation* to *causation*? What's the lurking third cause?
- Who's in the **sample** - and more importantly, who got left out?
- What's the **base rate**? Is a "positive" actually rare enough to trust?
- How many things did they test before showing me this one? (the missing denominator)
- What do the **chart's axes** actually say - zero baseline, full time range, real labels?
- Out of how many? What's the **sample size** behind that percentage?

None of these require advanced math. They're questions. That's the whole point of this guide,
and of the Mathematics foundations. We started in
[/guides/why-math-isnt-your-enemy](/guides/why-math-isnt-your-enemy) with one promise: math
isn't a wall keeping you out, it's a set of tools that make the world legible. Probability taught
you to reason about uncertainty instead of fearing it. Statistics taught you to read data
clearly. And this phase taught you the defensive half - because the same numbers that
illuminate are used, constantly, to manipulate.

That's the mission in one line: **numeracy is self-defense.** Not so you can win arguments with
spreadsheets, but so nobody can move you with a number you didn't understand. A truncated axis, a
self-selected survey, a "99% accurate" test on a rare condition - these are designed to work on
people who trust numbers without questioning them. You're no longer one of those people.

And these foundations feed everything ahead. When you reach the data and analytics material,
you'll already know why a clean average can lie and why the sample matters more than the size.
When you study AI and machine learning, base rates and false positives stop being a puzzle and
become the daily reality of every classifier you build or trust. When you tune performance, the
same instincts - read the distribution, distrust the small sample, ask out of how many - are how
you tell a real speedup from a lucky benchmark. You built the lens. The rest of the library is
what you point it at.

Quick gut-check before you go:

```quiz
[
  {
    "q": "A city finds that neighborhoods with more firefighters at a blaze tend to have more fire damage. What's the most reasonable conclusion?",
    "choices": [
      "Firefighters cause fire damage and should be sent in smaller numbers",
      "A third factor - the size of the fire - drives both the number of firefighters and the amount of damage",
      "The data must be wrong, since firefighters reduce damage",
      "More firefighters always means a more dangerous neighborhood"
    ],
    "answer": 1,
    "explain": "Classic correlation-without-causation. Bigger fires summon more firefighters AND cause more damage - the fire size is the lurking variable. The firefighters aren't the cause; they're a symptom of the same thing causing the damage."
  },
  {
    "q": "An online poll on a company's homepage shows 92% of respondents love the new redesign. Why should you be cautious?",
    "choices": [
      "92% isn't a high enough number to be meaningful",
      "Polls are always rigged by the company",
      "The sample is self-selected - only people still visiting and motivated to respond are counted, and people who hated it may have already left",
      "Online polls can't measure opinions accurately at all"
    ],
    "answer": 2,
    "explain": "This is sampling bias by self-selection. The poll only hears from people who stayed on the site and chose to answer. The users most upset by the redesign may have already churned and aren't in the sample at all - so it can't speak for everyone."
  },
  {
    "q": "A bar chart shows your competitor's revenue towering over yours - but you notice the y-axis starts at $9.8M, not $0, and the real values are $9.9M vs $10.1M. What's going on?",
    "choices": [
      "The chart is fine; your competitor really is dominating",
      "A truncated (non-zero) baseline exaggerates a tiny 2% difference into a visual landslide",
      "The numbers are fabricated and can't be trusted",
      "Bar charts should never start at zero"
    ],
    "answer": 1,
    "explain": "The data is true but the picture lies. A bar chart's message is bar length, so starting the axis at $9.8M instead of $0 turns a ~2% gap into what looks like a 3x gap. Read the axis before you trust the impression."
  }
]
```

You've finished the Probability & Statistics guide - and the Mathematics foundations. The numbers
work for you now, not the other way around.
