# A/B Testing, Explained

> How randomized splits let you measure which variant actually performs better, and the three ways teams accidentally fool themselves with the results.


---

# A/B Testing, Explained

Someone on your team changes the signup button from blue to green. Signups go up 8% the next week. Was it the button? Or was it a Tuesday, a marketing email that went out the same day, or nothing at all - just the normal week-to-week wobble your numbers always have? Without a way to isolate the change, you cannot tell. A/B testing is how you isolate it: split your users into groups at random, show each group a different version, and measure which one actually moves the number you care about.

This guide walks through why randomizing matters, how a real test is put together, and the ways teams talk themselves into believing a result that isn't there.

## How to read this

Read it in order. Phase 1 is the core idea - why you need a control group at all. Phase 2 is how a real test is structured, including just enough of "statistical significance" to make the phrase mean something. Phase 3 is the part that actually matters day to day: the mistakes that turn a fair test into a misleading one.

## The phases

1. [Why you'd randomize at all](01-why-randomize.md) - the core idea, and why "before vs. after" is a trap.
2. [How a real test is structured](02-structuring-a-test.md) - control vs. variant, one metric, sample size, and what "significant" means.
3. [How teams fool themselves](03-how-teams-fool-themselves.md) - peeking early, metric fishing, and novelty effects.


---

# Why you'd randomize at all

Here's the trap almost everyone falls into before they learn better: you ship a change, watch the metric for a week, see it go up, and conclude the change worked. That's "before vs. after," and it feels like evidence. It is not.

## The problem with before vs. after

The week after you shipped isn't a clean copy of the week before it with one variable changed. It's a different week, full of things that have nothing to do with your button color:

- **Seasonality.** Traffic on a Monday doesn't behave like traffic on a Saturday. A holiday week doesn't behave like a normal one. If your "before" period and "after" period land on different points in that cycle, the difference you're seeing might just be the calendar.
- **Other changes happening at the same time.** Marketing sent an email. A competitor had an outage. Your app store rating ticked up. Another team shipped a different feature the same week. Any of these can move your metric, and "before vs. after" has no way to separate their effect from yours.
- **Confounds.** A confound is anything that changes alongside your treatment and could explain the result instead of it. If you rolled the new button out only to mobile users first, and mobile users already convert differently than desktop users, you're not measuring the button - you're measuring mobile vs. desktop.

The real problem is that "after" always differs from "before" in dozens of ways you didn't control, and your change is just one of them. You have no way to isolate its effect from all the rest.

> A metric moving after you shipped something is not evidence the thing you shipped caused it. Time moves things on its own.

## What randomizing actually buys you

The fix is to stop comparing across time and start comparing across groups that exist in the *same* time. Take your incoming users and randomly assign each one to see version A (the **control** - what already exists) or version B (the **variant** - the new thing). Both groups experience the same day, the same marketing email, the same outage, the same everything - except the one thing you're testing.

Randomization is what makes this work. If you let users pick which version they see, or assign based on any trait (signup date, browser, region), you've reintroduced a confound - now the groups differ by more than just the variant. Random assignment means, on average, the two groups are identical in every way *except* which version they saw. Any difference in outcomes has one remaining explanation: the variant.

```mermaid
flowchart LR
  U[Incoming users] -->|random 50/50 split| A[Control: existing version]
  U -->|random 50/50 split| B[Variant: new version]
  A --> M1[Measure signup rate]
  B --> M2[Measure signup rate]
  M1 --> C{Compare}
  M2 --> C
```

*What just happened:* both groups are drawn from the exact same population, at the exact same time, experiencing the exact same external world. The random split is what lets you attribute any difference in the two "measure" boxes to the one thing that differs between them - which version they saw.

## Why "at the same time" is the whole point

A proper A/B test never compares control-in-March to variant-in-April. Both groups run **concurrently** - same week, same days, same external conditions. That's the piece "before vs. after" is missing entirely, and it's what makes A/B testing worth the extra setup. You're not asking "did the metric change over time." You're asking "of two groups living through the identical moment, which one did better" - and that question has a clean answer.

This doesn't mean before/after comparisons are worthless everywhere. They're a reasonable first signal when you have no way to run a real test. But if you're deciding something that matters - a pricing change, a checkout redesign, anything you'd want to defend later - a random split is what turns "I think it worked" into "I can show it worked."

Phase 2 covers how to actually structure that split: what to hold fixed, what to measure, and how big the groups need to be before the comparison means anything.


---

# How a real test is structured

Randomizing the split (Phase 1) is the foundation. On top of it, a real test needs three more things nailed down before you start: what you're calling "control," what single number counts as winning, and how many people you need before the answer is trustworthy.

## Control vs. variant

The **control** group sees the thing that already exists - your current button, your current pricing page, your current onboarding flow. It's the baseline. The **variant** (sometimes called the "treatment") sees the new thing you're proposing. You can run more than one variant at once (A vs. B vs. C), but the core comparison is always "new thing vs. what we already know works."

The control isn't a formality - it's the only thing that tells you whether the variant actually *beat* something, rather than just landed on a number that sounds fine in isolation. "12% of the variant group converted" means nothing on its own. "12% of the variant group converted, vs. 9% of the control group, in the same week" means something.

## Pick ONE metric before you start

This is the step people skip and regret. Before you launch the test, decide the single metric that determines whether the variant wins: signup rate, checkout completion, seven-day retention - one number, chosen in advance.

Why one, and why in advance? If you watch ten metrics at once, at least one will move by chance even if your change does nothing - and it's tempting to declare *that* one "the real result" after the fact. Phase 3 covers this trap in detail; the prevention starts here: write down the metric before you see any data, and judge the test by that number alone.

That doesn't mean you can't *look* at other metrics - you should, for context and to catch a variant that wins on signups but tanks retention. It means only one of them was the pre-registered judge of "did this work."

## Why small tests are noisy

Say you run a test for one day, 40 people per group: control converts at 10%, the variant at 15%. That's roughly 4 conversions vs. 6 - a single person switching groups would change the whole story. Small samples swing wildly for the same reason a coin flipped 4 times can land heads 3 times without being unfair: there isn't enough data yet for randomness to average itself out.

This is **noise**: the natural bounce in a metric even when nothing real is happening underneath. The larger your sample, the smaller that bounce gets relative to a real effect - and the more you can trust that a gap you're seeing is real, not a lucky coin flip.

```mermaid
flowchart TD
  S1[40 users per group] --> N1[Large noise relative to any real effect]
  S2[4,000 users per group] --> N2[Small noise relative to any real effect]
  N1 --> R1[Hard to tell signal from luck]
  N2 --> R2[Easier to trust the gap is real]
```

*What just happened:* the sample size doesn't change how big the real effect is - it changes how confidently you can tell that effect apart from ordinary randomness. That's why serious tests are planned to run until they hit a target sample size, not just "run for a while and see."

## What "statistically significant" actually means

You'll hear this phrase constantly, and most people who say it can't explain it. Here's the plain version, no formulas: a result is called **statistically significant** when the gap between control and variant is large enough, given how much data you collected, that it's unlikely to have happened from ordinary random noise alone.

It is *not* a claim that the result is important, large, or permanent - it's a narrower one: "if there were truly no difference between these two versions, we probably wouldn't have seen a gap this big by chance." That's it.

A tiny, meaningless improvement can be "statistically significant" with enough users. A genuinely large improvement can fail to reach significance if tested on too few. Significance is about confidence that a gap is real - not whether it's worth caring about; ask both questions separately.

> "Statistically significant" means "probably not noise." It does not mean "big," "important," or "definitely permanent." Ask those separately.

The practical takeaway: before running a test, decide roughly how many users you need - most experimentation tools or a sample-size calculator do this math for you, given your baseline conversion rate and the smallest improvement worth detecting - then run until you hit that number and only then check significance. Checking before that point is where Phase 3 begins.

Watch it animated: [A/B testing](/explainers/ABTesting.dc.html)

```quiz
[
  {
    "q": "Why does a real A/B test run the control and variant groups at the same time, rather than control last month and variant this month?",
    "choices": [
      "It's faster to set up that way",
      "Running them concurrently controls for everything else that changes over time, like seasonality or other launches",
      "It uses less server capacity",
      "Users prefer seeing both versions in the same month"
    ],
    "answer": 1,
    "explain": "Concurrent groups experience identical external conditions, so any outcome difference can be attributed to the variant instead of the calendar."
  },
  {
    "q": "Why should you pick one metric before the test starts, instead of after?",
    "choices": [
      "Analytics tools only support tracking one metric at a time",
      "Choosing after the fact lets you cherry-pick whichever metric happened to move, even by chance",
      "One metric is required by law for regulated industries",
      "It makes the dashboard load faster"
    ],
    "answer": 1,
    "explain": "Pre-registering the judging metric prevents picking a winner after seeing which number happened to move."
  },
  {
    "q": "A test with 40 users per group shows the variant winning. What should you be most cautious about?",
    "choices": [
      "The variant is definitely worse in reality",
      "The sample is too small for the gap to reliably reflect a real difference rather than noise",
      "40 users is always enough for any test",
      "The control group probably had a bug"
    ],
    "answer": 1,
    "explain": "Small samples bounce around a lot by chance alone, so an early-looking win may just be noise rather than a real effect."
  }
]
```


---

# How teams fool themselves

A properly randomized, well-structured test can still produce a false conclusion - not because the setup was wrong, but because of what people do while it's running and after it ends. These three mistakes account for most of the "the test said it worked but it didn't" stories you'll hear.

## Peeking early and stopping the moment it looks good

You launch a test planned to run for two weeks and hit a target sample size. On day three, someone checks the dashboard and the variant is ahead. It "looks like a winner," so the test gets shut down early and shipped.

The problem: a metric that's genuinely tied (no real difference between control and variant) will still drift ahead and behind randomly as data comes in, the same way a fair coin can run heads-heavy for a stretch before evening out. If you check constantly and stop the instant the variant is ahead, you're not measuring "did the variant win" - you're measuring "did the variant get ahead at some point during a noisy walk," which happens far more often than a genuine effect does. This is sometimes called **peeking**, and it's the single most common way a "significant" result turns out to be nothing once the full sample is in.

The fix isn't "never look" - checking for a broken test (errors, a metric collecting zero events) is fine and normal. The fix is: decide the sample size or the end date in advance, and don't treat an early lead as the answer. Let the test finish the run it was planned for.

## Testing so many metrics that something looks significant by chance

Say you track twenty metrics on a test instead of the one you pre-registered: signup rate, time on page, scroll depth, button color engagement, session length, and so on. Pure chance says some of those twenty will show an apparent "win" for the variant even if the variant does absolutely nothing differently from the control - that's how randomness spreads across many measurements.

The trap is picking whichever metric happened to move and presenting it as "the winner," as if it had been the point of the test all along. It wasn't - it was one of twenty lottery tickets, and one of twenty tickets winning isn't a story about the variant, it's a story about how many tickets you bought.

```text
Track 1 metric  -> a "significant" result is meaningful
Track 20 metrics -> expect a few to look "significant" by pure chance
```

*What just happened:* the more things you measure, the more chances randomness has to hand you a coincidence that resembles a real effect. This is exactly why Phase 2 insisted on choosing one metric before the test starts - it's not a formality, it's the thing that prevents this exact mistake.

## Novelty effects fading once the new thing stops being new

Sometimes a redesigned feature genuinely performs better for the first week, then three weeks later the gap has closed or reversed - nothing about the code changed. Users were curious about something new and clicked around more, or paid closer attention because it looked different, not because it was better. Once the novelty wears off and the new version becomes the new normal, behavior settles back to whatever it would have been anyway.

This is a **novelty effect**, and it's a real risk for anything visually or structurally different, especially with users who visit repeatedly. A test that only runs for a few days can mistake temporary curiosity for a permanent improvement. The practical guard: for changes likely to trigger novelty (redesigns, new UI patterns, anything visually loud), run long enough to see whether the effect holds after curiosity fades, and watch whether the gap shrinks over the test rather than holding steady.

> A win in week one that's gone by week three isn't a failed test - it's the test doing its job and revealing that the effect wasn't durable.

## The thread connecting all three

Every mistake in this phase comes from the same root: treating a noisy process as if a single glance at it tells the whole truth.

- **Peeking** trusts a glance mid-run.
- **Metric fishing** trusts whichever glance happened to look good.
- **Novelty effects** trust a glance taken before the effect had time to settle.

The discipline that fixes all three is the same one from Phase 2: decide the metric and the sample size before you start, and let the test run its planned course before drawing a conclusion.
