# Metrics That Lie

> How plausible-looking numbers mislead: averages hiding skew, survivorship bias, Simpson's paradox, and vanity metrics. Read a dashboard without being fooled.


---

# Metrics That Lie

Someone drops a number in a slide - "average response time is 200ms," "90% of users love it," "revenue is up and to the right" - and the room nods. The number is real, nobody made it up, and it's still steering you off a cliff. The lie isn't in the arithmetic; it's in what the number quietly leaves out.

This guide hands you the small set of tricks that fool almost everyone: an average that hides a long tail, data that only contains the survivors, a trend that flips when you split it by group, and pretty numbers that move without meaning. Once you've seen each one, you can't unsee it - and you stop getting played.

## How to read this
- **Got a number in front of you right now that smells off?** Skim [Phase 3: Reading a Dashboard Without Getting Fooled](03-reading-without-getting-fooled.md) - it's a field checklist for the moment of suspicion.
- **Want it to actually stick?** Read in order. We start with the one mistake under all the others (averages), then the biases hiding in the data, then how to defend yourself live.

## The phases
1. **[The Average Is Lying to You](01-the-average-is-lying.md)** - why the mean breaks the moment data is skewed or has outliers, what the median does instead, and why "average" is the most over-trusted word in analytics.
2. **[The Data You Never See](02-the-data-you-never-see.md)** - survivorship bias, base rates, and Simpson's paradox: three ways a dataset can be accurate and still point you the wrong way because of who's missing or how it's split.
3. **[Reading a Dashboard Without Getting Fooled](03-reading-without-getting-fooled.md)** - vanity vs actionable metrics, cherry-picked date ranges, truncated axes, and a quick interrogation you can run on any number in under a minute.

> This guide is the applied, paranoid sibling of [Probability and Statistics](/guides/probability-and-statistics) - that one builds the machinery, this one shows you where people abuse it. For turning clean numbers into decisions, see [Building a BI Dashboard That's Actually Useful](/guides/bi-dashboards-that-work).


---

# The Average Is Lying to You

You ask "how's the checkout flow doing?" and someone says "average load time is 1.2 seconds." Feels fine. Ship it. Except a chunk of your users are sitting at 9 seconds, rage-quitting, and that pain got blended into a number that sounds healthy. The average didn't lie about the math. It lied by smoothing a cliff into a gentle slope.

This is the first trick to learn because it's underneath all the others: **the mean assumes the world is symmetric and well-behaved, and the world almost never is.**

## What "average" actually does

The arithmetic mean adds everything up and divides by the count - a move that assumes every value pulls with equal, fair weight. Fine when the data is roughly symmetric (heights, say). But the moment a few extreme values show up, they yank the mean toward themselves with no resistance.

```text
Salaries on a small team (in thousands):
  40, 42, 45, 47, 50, 52, 900   ← the founder

  mean   = (40+42+45+47+50+52+900) / 7 = 168
  median = 47   (the middle value when sorted)
```

*What just happened:* the mean says the "average" person earns 168k - nobody on the team earns near that. The single 900 dragged it 120k away from where everyone actually lives. The median, unmoved by extremes, says 47k: a real person on the team.

## Skew is the tell

A distribution is **skewed** when it has a long tail on one side. Income, response times, file sizes, time-on-page, revenue per customer - nearly everything interesting in tech and business is skewed, usually with a long tail to the right (a few huge values).

```text
Right-skewed (long tail to the right):

  count
   |####
   |######
   |####
   |##         tail ─────────────────►
   |#    .   .      .          .
   +----------------------------------------
      median   mean
        ▲        ▲
        |        └─ dragged right by the tail
        └─ sits where most values actually are
```

*What just happened:* in any right-skewed distribution the mean sits to the *right* of the median, pulled by the tail. Reliable gut check: **if the mean is noticeably higher than the median, your data is skewed and the mean is over-reporting the typical case.** Close together means roughly symmetric - the mean can be trusted.

## Why this matters more than it looks

"Average" leaks into decisions everywhere - each leak is a chance to be wrong:

- **Performance:** "average API latency is 120ms" hides the user stuck at 4 seconds. This is exactly why engineers track **p95 / p99** (the value 95% or 99% of requests come in under) instead of the mean - the tail is where the suffering lives.
- **Money:** "average order value is $80" can mean everyone spends ~$80, or it can mean most spend $20 and a few whales spend $5,000. Those are two completely different businesses with the same average.
- **People:** "average time to close a ticket is 2 days" sounds great until you learn half close in an hour and a long tail festers for three weeks.

> The fix isn't "never use the mean." It's: **always ask for the median alongside it, and look at the spread.** Mean and median together tell you the shape. Either one alone is half a sentence.

Let's make the gut check concrete:

```python runnable
data = [40, 42, 45, 47, 50, 52, 900]

mean = sum(data) / len(data)
ordered = sorted(data)
median = ordered[len(ordered) // 2]   # odd count: the middle value

print(f"mean:   {mean:.0f}")
print(f"median: {median}")
print("skewed!" if mean > median * 1.2 else "roughly symmetric")
```

*What just happened:* the code flags skew when the mean runs more than 20% above the median, printing `mean: 168`, `median: 47`, `skewed!` - a one-line alarm that the average isn't describing the typical case here.

### For builders

When you expose a metric in a dashboard or an API, default to showing **median plus a high percentile** (p90 or p95), not the mean. It's the same query cost and it stops your own team from making the average mistake. If you must show one number, the median is the safer default for anything skewed - which is most of what you'll measure.

```quiz
[
  {
    "q": "Seven response times (ms): 80, 85, 90, 95, 100, 110, 5000. Which statement is true?",
    "choices": [
      "The mean describes the typical request well",
      "The mean is dragged far above the typical value by the 5000ms outlier",
      "The median is 5000",
      "Mean and median will be nearly equal"
    ],
    "answer": 1,
    "explain": "One huge outlier yanks the mean upward; the median (~95ms) describes the typical request, the mean does not."
  },
  {
    "q": "For a right-skewed distribution (long tail to the right), where does the mean sit relative to the median?",
    "choices": [
      "To the left of the median",
      "Exactly on the median",
      "To the right of the median",
      "It is impossible to say"
    ],
    "answer": 2,
    "explain": "The long right tail pulls the mean toward the high values, so the mean lands to the right of the median."
  },
  {
    "q": "Why do engineers track p95/p99 latency instead of the mean?",
    "choices": [
      "Percentiles are easier to compute",
      "The mean hides the slow tail where users actually suffer",
      "p99 is always smaller than the mean",
      "The mean cannot be calculated for latency"
    ],
    "answer": 1,
    "explain": "A healthy-looking mean can hide a painful tail; high percentiles expose the worst experiences real users hit."
  }
]
```


---

# The Data You Never See

Phase 1 was about a number describing the data badly. This phase is scarier: the data itself can be accurate and complete-looking, and still point you the wrong way - because of who's *missing* from it, what the background rate is, or how it's sliced. No bad math required. These three biases fool smart people daily.

## Survivorship bias: the dataset only contains the winners

The classic story: in World War II, analysts looked at returning bombers, mapped where they'd been shot, and proposed armoring those spots. Statistician Abraham Wald pointed out the obvious-once-said: those are the planes that *came back*. The bullet holes mark where a plane can take a hit and survive. The armor belongs on the spots with *no* holes - planes hit there never returned to be measured.

```text
Planes that returned:    holes on wings, tail, fuselage
Planes that didn't:      (not in your dataset at all)

  Naive read:  "wings get hit most → armor the wings"
  Wald's read: "we only see survivors → armor the UNmarked spots"
```

*What just happened:* the data was real and carefully collected - and still misleading, because the most important cases (the planes that went down) were silently absent. **Survivorship bias is the error of analyzing only the things that made it through some filter, while the filtered-out cases hold the answer.**

You hit this constantly:
- "Successful founders all dropped out of college" - you're not counting the far larger pile of dropouts who failed and never got interviewed.
- "Our long-time customers love feature X" - the people who hated it churned and aren't in your survey.
- "This trading strategy returned 30%/year backtested" - funds that blew up got dropped from the database you tested against.

The defense is one reflex question: **who or what got filtered out before this data was collected?**

## Base rates: a number means nothing without its denominator

A test for a rare disease is "99% accurate." You test positive. Are you 99% likely to be sick? Almost everyone says yes. The real answer is often *no* - because of the **base rate**, how common the thing is to begin with.

```text
Disease affects 1 in 1,000 people.
Test: 99% accurate (1% false-positive rate).
Test 100,000 people:

  Actually sick:     100  →  ~99 test positive   (true positives)
  Actually healthy:  99,900 → ~999 test positive  (false positives)

  Of everyone who tests positive:  99 / (99 + 999) ≈ 9%
```

*What just happened:* even with a "99% accurate" test, a positive result means only ~9% chance you're actually sick - the disease is so rare that false positives from the huge healthy group swamp the true positives. **A rate (accuracy, conversion, click-through) is meaningless until you anchor it to how common the underlying event is.** Ignore the base rate and you'll wildly over-react to any positive signal.

## Simpson's paradox: the trend flips when you split it

This is the one that breaks people's brains. A trend that's clearly true in the combined data can *reverse* in every subgroup. Same numbers, opposite conclusions.

```text
Treatment success rates, combined:
  Treatment A:  78%   ← looks better overall
  Treatment B:  83%

Now split by case severity:

              Mild cases        Severe cases
  Treatment A   93% (81/87)       73% (192/263)
  Treatment B   87% (234/270)     69% (55/80)

  → A wins in mild AND in severe, but loses overall.
```

*What just happened:* Treatment A beats B in mild cases *and* severe cases - yet loses on the combined number. The trick is **which group each treatment was given to.** A got mostly the hard, severe cases (lower success rates for anyone); B got mostly the easy mild ones. The combined average reflects the *case mix*, not the treatment. Aggregate first and you'd pick the worse treatment.

```text
The lurking variable here is CASE SEVERITY.
It correlates with both:
  - which treatment you got (assignment)
  - how likely you were to succeed (outcome)
That hidden third variable is what flips the sign.
```

*What just happened:* Simpson's paradox always comes from a **lurking variable** - a hidden third thing tied to both the group and the outcome. When the groups aren't comparable (different case mixes, traffic sources, time periods), the combined number can say the opposite of the truth. The defense: **when comparing two things, ask whether the groups are actually alike - and split by the obvious confounder to see if the story holds.**

> Survivorship bias, base rates, and Simpson's paradox share one root: **the number is fine, but the data behind it isn't representative of the question you're asking.** Who's missing, what's the background rate, and are the groups comparable. Three questions, most of your defense.

### For builders

When you build a comparison view - A/B test results, cohort tables, regional performance - give people a **"split by" control** for the obvious confounders (device, plan tier, signup channel, time period). A single combined KPI is precisely where Simpson's paradox hides; letting a user break the number down by one dimension turns an invisible reversal into a visible one. The deeper machinery - confidence intervals, controlling for variables - lives in [Probability and Statistics](/guides/probability-and-statistics).

```quiz
[
  {
    "q": "A VC says 'most unicorn founders are under 30, so we only fund young founders.' What bias is most likely at work?",
    "choices": [
      "Simpson's paradox",
      "Survivorship bias - failed young founders aren't in the 'unicorn' dataset",
      "Truncated axis",
      "Base-rate neglect only"
    ],
    "answer": 1,
    "explain": "Looking only at unicorns ignores the far larger pool of young founders who failed; the dataset contains only survivors."
  },
  {
    "q": "A disease affects 1 in 1,000. A test is 99% accurate. You test positive. Roughly how likely are you to actually have it?",
    "choices": [
      "About 99%",
      "About 90%",
      "About 9%",
      "Exactly 50%"
    ],
    "answer": 2,
    "explain": "Because the disease is rare, false positives from the large healthy group dominate, dropping the real probability to roughly 9%."
  },
  {
    "q": "Treatment A beats B in both mild and severe cases, yet loses on the combined rate. The cause is:",
    "choices": [
      "A calculation error in one subgroup",
      "A lurking variable (case mix) that differs between the groups",
      "The mean being pulled by an outlier",
      "A truncated chart axis"
    ],
    "answer": 1,
    "explain": "This is Simpson's paradox: a hidden variable tied to both group assignment and outcome flips the combined result."
  }
]
```


---

# Reading a Dashboard Without Getting Fooled

You now know the deep traps: skewed averages, missing data, flipped trends. This phase is the street-level version - the things people do to a chart to make a number land harder than it should, plus a fast interrogation to run the moment a metric makes you feel something. Because that feeling - "wow, that went up a lot" - is exactly when you're easiest to fool.

## Vanity metrics vs actionable metrics

A **vanity metric** goes up, makes you feel good, and changes nothing you do. An **actionable metric** ties to a decision: when it moves, you do something different.

```text
Vanity                         Actionable
─────────────────────────      ─────────────────────────
Total registered users         Weekly active users
Total downloads                7-day retention
Page views                     Conversion rate (signups / visits)
Followers                      Engagement per post
"$1M raised"                   Monthly recurring revenue, churn
```

*What just happened:* the left column only ever grows and never tells you to change course - total signups can't go down, so it always looks like progress even while the product dies. The right column can fall, has a denominator, and points at an action. **The test for a vanity metric: can it go down, and if it moved, would you do anything differently?** If the answer to both is no, it's decoration.

## Cherry-picked date ranges

The same data tells opposite stories depending on where you start and stop the window - the most common plausible-looking lie in business reporting.

```text
Full year:                        Cherry-picked window:

  $ |    .                          $ |        ___/
    |  ./ \.    .__.                   |    ___/
    |./     \__/                       | __/
    +------------------ time           +------------ time
    Jan            Dec                 Mar      May
    "flat, then declining"            "explosive growth!"
```

*What just happened:* both charts are drawn from the identical dataset. By starting at a local low (March) and ending at a local high (May), the second chart manufactures a growth story the full year contradicts. **Defense: ask "why this date range?" and demand to see a longer window.** If someone resists showing you more history, that resistance is your answer.

## Truncated axes

A bar or line chart whose y-axis doesn't start at zero exaggerates differences - a 2% change can be drawn to look like a doubling.

```text
Truncated (y starts at 95):       Accurate (y starts at 0):

 100 |        ███                  100 |   ███   ███
  98 |  ███   ███                   75 |   ███   ███
  96 |  ███   ███                   50 |   ███   ███
  95 +--------------                25 |   ███   ███
       A      B                      0 +--------------
   "B towers over A!"                    A      B
                                     "A=97, B=99, nearly identical"
```

*What just happened:* the left chart starts its axis at 95, so a difference of 2 (97 vs 99) fills most of the frame and screams "huge gap." The right chart starts at zero and shows the truth: the two bars are almost the same height. **For bar charts, the y-axis should start at zero - full stop.** (Line charts tracking change over time are a reasonable exception, but a truncated bar chart is almost always a manipulation.) When you see a dramatic-looking bar chart, check the axis before you react.

## A few more quick tells

- **Percentages with no denominator.** "Engagement up 200%!" From 1 user to 3. A percentage change on a tiny base is noise dressed as a headline.
- **No comparison or target.** A number alone ("4,200 signups") means nothing. Up or down from last month? Above or below goal? Context is the metric; the raw number is trivia.
- **Combined metric, no breakdown.** Remember Simpson's paradox from Phase 2 - a single combined KPI is where a reversal hides. Ask to split by the obvious dimension.
- **Mean with no median or spread.** Phase 1's trap. If they show you only the average, ask for the median.

## The one-minute interrogation

When a number makes you feel something, run this before you believe it:

```text
1. DENOMINATOR - a rate or a raw count? Out of how many? Can it go down?
2. DISTRIBUTION - is this a mean? Show me the median and the spread.
3. WHO'S MISSING - survivorship: what got filtered out before collection?
4. BASE RATE - how common is the thing this number is about?
5. THE WINDOW - why this date range? Show me a longer one.
6. THE AXIS - does the y-axis start at zero? (bar charts especially)
7. SO WHAT - if this moved, would any decision change? (vanity test)
```

*What just happened:* that's the whole guide compressed into seven questions. You don't need to remember which bias has which name - just the reflex to ask these before you nod. Most misleading metrics fail at least one of them immediately.

> The goal isn't cynicism, where you trust no number. It's calibration: trust numbers that survive the interrogation, and ask one sharp question about the ones that don't. A good analyst welcomes these questions - they're how solid work proves itself.

### For builders

If you build the dashboards, you set the defaults that decide whether your org reasons clearly. Bake the defenses in: bar charts that start at zero, comparison-to-target built into every tile, medians shown beside means, a "split by" control on combined KPIs, and date pickers that default to a sensible long window instead of a flattering short one. The full craft of doing this well is its own guide - see [Building a BI Dashboard That's Actually Useful](/guides/bi-dashboards-that-work). The easiest way to stop a team from being fooled is to never build the misleading view in the first place.

```quiz
[
  {
    "q": "Which of these is a vanity metric?",
    "choices": [
      "7-day retention rate",
      "Total cumulative downloads",
      "Conversion rate from visit to signup",
      "Monthly churn"
    ],
    "answer": 1,
    "explain": "Cumulative downloads can only go up and rarely changes any decision - the hallmarks of a vanity metric."
  },
  {
    "q": "A bar chart shows B towering over A. What should you check first?",
    "choices": [
      "Whether the bars are the right color",
      "Whether the y-axis starts at zero",
      "The font size of the labels",
      "Whether there are gridlines"
    ],
    "answer": 1,
    "explain": "A truncated y-axis (not starting at zero) exaggerates small differences into dramatic-looking gaps."
  },
  {
    "q": "Someone shows 'explosive growth' over a two-month window. The best response is:",
    "choices": [
      "Accept it - two months is plenty of data",
      "Ask why that specific date range, and request a longer window",
      "Assume the underlying data is fake",
      "Ask for the chart in a different color scheme"
    ],
    "answer": 1,
    "explain": "Cherry-picked windows manufacture trends; seeing a longer history reveals whether the growth is real."
  }
]
```
