# Flaky Tests

> Why a test passes and fails with no code change, the usual culprits, and how to kill flakiness for good instead of hiding it behind retries.


---

# Flaky Tests

You wrote a test. It passed. You changed nothing - and on the next run it failed. Re-run it, green again. That little flip of the stomach, the "is it me or is it the test?" - that's a flaky test, and it's quietly the most demoralizing thing in a test suite. This guide gives you the mental model for *why* a test does this, a field guide to the usual culprits, and a calm playbook for killing flakiness instead of papering over it.

## How to read this

- **Need to understand what's happening?** Read [Phase 1](01-what-a-flaky-test-is.md) - it's the whole mental model in one phase: a flaky test is nondeterministic, and here's where the nondeterminism sneaks in.
- **Staring at one right now?** Phase 2 is the field guide to the specific culprits - time, async, order, shared state, real externals - with the shape each one has.
- **Want to fix it for good?** Phase 3 is the diagnose-and-fix-and-quarantine playbook.

## The phases

1. **[What a Flaky Test Actually Is](01-what-a-flaky-test-is.md)** - the mental model: a passing test is a deterministic function, and flakiness is hidden nondeterminism leaking into it. Once you see tests this way, every culprit in Phase 2 has the same shape.
2. **[The Usual Culprits](02-the-usual-culprits.md)** - a field guide to the five everyday sources of flakiness: timing and sleeps, async not awaited, test order and shared state, real network/clock/randomness, and leaked resources. How to recognize each by its fingerprint.
3. **[Diagnose, Fix, Quarantine](03-diagnose-fix-quarantine.md)** - the playbook: rerun and isolate and seed to find the cause, control time/randomness/state/externals to fix it, and quarantine (never silently ignore) when you can't fix it today.


---

# What a Flaky Test Actually Is

Here's the moment. The build goes red on a test you didn't touch. You frown, re-run it, and it goes green. No code changed. Nothing you did fixed it - it *decided* to pass this time. Your first instinct is relief, your second is unease, and your third, if you're being straight with yourself, is to never trust that test again.

That test isn't broken in the normal sense. It's *flaky*. And flaky isn't a vibe or a mystery - it has a precise, almost boring definition. Once you have that definition in your head, every flaky test you'll ever meet stops being spooky and starts being a thing you can reason about.

## A passing test is supposed to be a pure function

Think about what a good test promises. You give it the same code, and it gives you the same answer - pass or fail - every single time. A healthy test behaves like a **pure function**: its result depends *only* on the code under test.

```text
   healthy test:    code  ──►  [ test ]  ──►  PASS   (always)
                    code  ──►  [ test ]  ──►  PASS   (still)
                    code  ──►  [ test ]  ──►  PASS   (again)
```

*What just happened:* The same input went in three times and the same result came out three times. That repeatability is the entire point - it's what lets a green check *mean* something. When a test is deterministic, "green" is a fact about your code.

## Flaky means a hidden input snuck in

A flaky test isn't actually getting different inputs from *your code* - the code is identical run to run. It's getting different inputs from somewhere you didn't notice: the clock, the order other tests ran in, a random number, a network response, a leftover row in a database. Those are real inputs to the test. You never wrote them down.

```text
   flaky test:    code + (clock? order? random? network?)  ──►  [ test ]  ──►  PASS
                  code + (clock? order? random? network?)  ──►  [ test ]  ──►  FAIL
                       └── same code, DIFFERENT hidden inputs ──┘
```

*What just happened:* The visible input (your code) stayed the same, but invisible inputs changed between runs, so the result flipped. The test was never really a function of your code alone - it was secretly a function of your code *plus the universe*, and the universe doesn't sit still. **That's the whole disease: nondeterminism leaking into something that was supposed to be deterministic.**

📝 **Terminology.** *Deterministic* means same input, same output, every time. *Nondeterministic* means the output can vary even when the input you care about doesn't. "Flaky" is the standard industry word for a nondeterministic test. They're the same idea wearing different clothes.

## Why this is worse than a failure that tells the truth

It's tempting to file a flaky test under "minor annoyance." It isn't. A normal failing test is straight with you - it says "something's broken," you fix it, it goes green, the system told you the truth. A flaky test lies. Sometimes it cries wolf when nothing's wrong; sometimes it stays silent when something *is*. Either way, it teaches everyone the same poisonous lesson:

```text
   week 1:   red build → "probably flaky" → re-run → green → merge
   week 6:   red build → "ugh, flaky again" → re-run → merge (nobody read it)
   week 12:  red build that is a REAL bug → "just re-run it" → ships the bug
```

*What just happened:* Each flaky failure trained the team to discount red. By week 12 the muscle memory was "red means re-run," so when red finally *meant something*, nobody listened. **A flaky test doesn't cost you one test - it slowly costs you trust in every test you have.** That's why we treat flakiness as a real bug, a bug in the test suite, not as background noise.

💡 **Key point.** The fix for flakiness is never "make the test pass more often." It's "find the hidden input and pin it down." You're not chasing luck - you're hunting a specific source of nondeterminism, and there's a short list of them. That list is Phase 2.

## The reframe that makes everything easier

Hold onto this single sentence and the rest of the guide unlocks:

> A flaky test is a test that depends on something it doesn't control.

The clock it didn't freeze. The random seed it didn't set. The database it didn't clean up. The async operation it didn't wait for. The API it called for real. Every cause of flakiness is one specific uncontrolled dependency - and the cure is always the same shape: **take control of it.** Freeze the clock. Seed the randomness. Isolate the state. Await the operation. Fake the network.

🪖 **War story.** A test that "only fails on CI, never locally" panicked a team I worked with for a week - they suspected the CI machine was cursed. It wasn't cursed. It was *slower* than their laptops, which exposed a race the fast laptops always won. The CI box wasn't broken; it was hardware telling the truth, revealing an uncontrolled dependency on timing. The flaky test had been lying on their laptops the whole time - the slow machine stopped covering for it.

For builders: when you see "passes locally, fails in CI," don't reach for "CI is flaky." Reach for "CI runs in a different, often slower or more parallel environment, and that difference is *surfacing* a real nondeterminism my fast quiet laptop was hiding." CI is usually the messenger, not the problem.

## Recap

- A healthy test is **deterministic**: same code in, same result out, every time. That repeatability is what makes "green" mean something.
- A **flaky** test is **nondeterministic**: the result flips with no code change because a hidden input - time, order, randomness, the network, leftover state - leaked in.
- Flakiness is worse than a failure that tells the truth because it trains people to **ignore red**, so real failures get ignored too.
- The universal reframe: **a flaky test depends on something it doesn't control.** Every fix is "take control of that thing." Phase 2 names the usual suspects.

```quiz
[
  {
    "q": "What is the precise definition of a flaky test?",
    "choices": ["A test that is slow to run", "A test that fails the first time you write it", "A test whose result varies even when the code under test doesn't change", "A test with no assertions"],
    "answer": 2,
    "explain": "Flaky = nondeterministic: the same code can produce pass or fail because a hidden input changed between runs."
  },
  {
    "q": "Why is a flaky failure considered worse than a failure that tells the truth?",
    "choices": ["It takes longer to run", "It trains the team to ignore red, so real failures get ignored too", "It uses more CI minutes", "It can't be skipped"],
    "answer": 1,
    "explain": "A flaky test erodes trust in every test, because people learn that red doesn't necessarily mean broken - and then ignore a real red."
  },
  {
    "q": "A test passes on your laptop but fails on CI with no code change. What's the most likely explanation?",
    "choices": ["The CI server is broken", "CI's different (often slower, more parallel) environment is surfacing a real nondeterminism your laptop was hiding", "The test should be deleted", "CI runs an older version of the code"],
    "answer": 1,
    "explain": "CI is usually the messenger: a slower or more parallel environment exposes an uncontrolled dependency (often timing or order) that a fast, quiet laptop happened to mask."
  }
]
```


---

# The Usual Culprits

You've got the mental model from Phase 1: a flaky test depends on something it doesn't control. Now you're staring at an actual flaky test and need to know *which* something. Flakiness comes from a short list of recurring sources, and each one leaves a recognizable fingerprint. Learn the five and you'll start diagnosing flaky tests on sight, before you've even read the stack trace.

## Culprit 1: Timing and sleeps

The most common source by a mile. The test assumes some work has finished "by now" - usually by sleeping a fixed number of milliseconds and then checking. On your laptop, the work finishes in time. On a loaded CI box, sometimes it doesn't, and the check runs against half-done state.

```javascript
// Flaky: hope the data loads within 100ms.
await sleep(100);
expect(screen.getText()).toBe('Loaded');

// Reliable: wait for the actual condition, however long it takes.
await waitFor(() => expect(screen.getText()).toBe('Loaded'));
```

*What just happened:* The first version bets that 100ms is enough, and a busy machine slows everything down, so the bet sometimes loses. The second waits for the *thing you actually care about* instead of a fixed delay. A `sleep` in a test is almost always a guess wearing a number.

**Fingerprint:** the failure clusters on slow or busy machines, gets *more* frequent under load or parallelism, and you can "fix" it by bumping the sleep number (which is the tell - if a bigger sleep helps, it's a timing race).

## Culprit 2: Async not awaited

Subtler and nastier. The test kicks off asynchronous work but doesn't actually *wait* for it before asserting - or worse, the test function returns before the async work runs at all. The assertion races the operation, and sometimes the assertion wins.

```javascript
// Flaky: the test function returns before saveUser() resolves.
test('saves the user', () => {
  saveUser({ name: 'alice' });          // returns a Promise, but it's dropped
  expect(db.count()).toBe(1);           // runs before the save finishes
});

// Reliable: await the work, so the assertion runs after it completes.
test('saves the user', async () => {
  await saveUser({ name: 'alice' });
  expect(db.count()).toBe(1);
});
```

*What just happened:* In the first version, the Promise from `saveUser` is created and immediately abandoned - the test moves to the assertion while the save is still in flight, so it passes or fails depending on which finishes first. `async`/`await` forces the assertion to wait for the save, removing the race.

**Fingerprint:** failures are erratic and don't correlate cleanly with machine speed; you often see "expected 1, got 0" or a value that's *one step behind* what you expected. A missing `await` (or a forgotten `return` of a Promise) is frequently the cause.

⚠️ **Gotcha.** Many test runners won't warn you when you forget to `await`. The test returns early and passes *most* of the time, so it looks healthy until the day the timing shifts. Treat any unawaited Promise in a test as a latent flake.

## Culprit 3: Test order and shared state

A test passes when run alone and fails when run after some other test - because the two share something they shouldn't: a database row, a global variable, a file on disk, a cache, an env var. One test leaves a mess; the next test trips over it.

```text
   test A:  creates user "alice"  (and does NOT clean up)
   test B:  asserts "no users exist"   ← passes alone, FAILS after A

   order [A, B] → B fails.    order [B, A] → B passes.
   CI shuffles or parallelizes → order changes → result flips → "flaky."
```

*What just happened:* Neither test is wrong on its own, but A leaks state into B, so the result flips the moment your runner shuffles order, parallelizes, or someone adds a test in between. A test that secretly depends on what ran before it isn't really one test - it's a hidden dependency.

**Fingerprint:** the test passes in isolation (`run it alone` → green) but fails in the full suite, or the failure moves around when you change the order or the parallelism. If "run it alone and it passes" is true, this is almost always your culprit.

## Culprit 4: Real network, clock, or randomness

The test reaches outside its own little world: it calls a live API, hits a real database over the network, reads the actual system clock, or uses unseeded randomness. Every one of those is an input you don't control, and any of them can hiccup or shift, failing your test for a reason that has nothing to do with your code.

```javascript
// Flaky: depends on the real wall clock - fails at the boundary, e.g. near midnight.
expect(formatTimestamp(Date.now())).toBe('2026-06-30');

// Flaky: depends on a live third-party API that can be slow, down, or rate-limited.
const res = await fetch('https://api.example.com/price');
expect(res.status).toBe(200);

// Flaky: depends on unseeded randomness - passes ~most of the time.
expect(pickRandom([1, 2, 3])).toBe(1);
```

*What just happened:* Each line lets the outside world decide the result - the clock rolls over a day boundary, the API rate-limits you, the random pick lands elsewhere. None of these failures mean your code is wrong; the test asked an uncontrolled source a question and didn't like the answer.

**Fingerprint:** failures correlate with *external events* - time of day, a flaky third party, network blips - and often can't be reproduced on demand because the external condition has moved on.

## Culprit 5: Leaked resources

The quiet one. Tests open things - file handles, database connections, ports, timers, background tasks - and don't close them. For a while nothing breaks. Then you hit the OS file-handle limit, or the connection pool is exhausted, or a leftover timer fires during the *next* test and corrupts it. The failure shows up far from the test that actually caused it.

```text
   test 1..40:  each opens a DB connection, none closes it
   test 41:     "could not get connection: pool exhausted"  ← blamed unfairly

   The leak is in tests 1–40. Test 41 is just the one that ran out of room.
```

*What just happened:* The pool filled up gradually, and the test unlucky enough to need connection 41 fails - even though it did nothing wrong. Leaks make flakiness depend on how many tests ran before, which is maddening because the failing test is innocent.

**Fingerprint:** failures appear deep into a run, not at the start; they move to a *different* test when you add or remove tests; and error messages mention exhausted pools, "too many open files," or "address already in use." Suspect a leak when the victim is well downstream and changes with suite size.

```mermaid
flowchart TD
  A[Test flakes] --> B{Passes alone?}
  B -- No --> C{Bigger sleep helps?}
  C -- Yes --> D[Timing / sleep]
  C -- No --> E{Touches network/clock/random?}
  E -- Yes --> F[Real external]
  E -- No --> G[Async not awaited]
  B -- Yes --> H{Fails deep in run / by suite size?}
  H -- Yes --> I[Leaked resource]
  H -- No --> J[Order / shared state]
```

*What just happened:* Same five culprits as a triage tree. "Passes alone?" splits the world: yes points to order/state or a leak, no points to timing, async, or an external. It won't be right every time, but it points you at the likely suspect fast.

For builders: when you write a test, ask "what does this depend on besides the code?" A clock, a real service, randomness, test order, or an opened resource - any of those is a future flake you can prevent now. The cheapest flaky test to fix is the one you never wrote.

## Recap

The five usual culprits, each with its tell:

1. **Timing / sleeps** - fixed `sleep` then assert; fails under load; bigger sleep "helps."
2. **Async not awaited** - assertion races the operation; "expected 1, got 0," value one step behind.
3. **Order / shared state** - passes alone, fails in the suite; result moves with order or parallelism.
4. **Real externals** - live network, real clock, unseeded randomness; failures track time of day or a third party.
5. **Leaked resources** - unclosed handles/connections/timers; fails deep in the run, moves with suite size.

```quiz
[
  {
    "q": "A test passes when you run it by itself but fails when you run the whole suite. Which culprit is this almost always?",
    "choices": ["A fixed sleep that's too short", "Test order / shared state", "Unseeded randomness", "A leaked file handle"],
    "answer": 1,
    "explain": "Passes alone, fails together is the signature of order/shared-state: another test leaves state behind that this one trips over."
  },
  {
    "q": "Why is replacing `await sleep(100)` with `await waitFor(() => condition)` more reliable?",
    "choices": ["It runs faster on every machine", "It waits for the actual condition instead of betting a fixed delay is enough", "It disables parallelism", "It retries the whole test on failure"],
    "answer": 1,
    "explain": "A fixed sleep is a guess that a busy machine can lose; waiting for the real condition removes the race regardless of machine speed."
  },
  {
    "q": "Test #41 fails with 'connection pool exhausted' while tests #1–40 pass. Where is the bug most likely?",
    "choices": ["In test #41's assertions", "In tests #1–40 leaking connections they never close", "In the test runner itself", "In an unawaited Promise in #41"],
    "answer": 1,
    "explain": "Leaked resources make a downstream, innocent test the victim. The leak is in the earlier tests that opened connections without closing them."
  }
]
```


---

# Diagnose, Fix, Quarantine

You know the disease (Phase 1) and the suspects (Phase 2). This phase is the playbook for actually killing a flaky test: **diagnose** (find the hidden input), **fix** (take control of it), **quarantine** (a holding cell for when you can't fix it this minute). Order matters - most people skip straight to fixing and guess wrong. Diagnosis is what makes the fix land.

## Step 1: Diagnose - make it fail on demand

You can't fix what you can't reproduce. Diagnosis turns an intermittent failure into a predictable one, so you can confirm the cause and confirm the fix. Three tools do most of the work.

**Rerun in a loop.** A test that fails 1-in-50 times is invisible if you run it once. Run it many times in a row and the failure rate becomes a fact you can measure - and later, a fact you can watch go to zero.

```bash
# Run one test 50 times; stop on the first failure so you can inspect it.
for i in $(seq 1 50); do
  npm test -- flaky.spec.js || { echo "FAILED on run $i"; break; }
done
```

*What just happened:* Instead of hoping to catch the flake by luck, you forced 50 attempts. A failure on run 23 gives you a reproduction and a baseline rate; 50-for-50 green after your fix is your evidence it worked.

**Isolate.** Run the suspect test completely alone. If it passes alone but fails in the suite, you've confirmed it's order or shared state (Culprit 3) - you don't even need to read the code yet. If it fails alone too, the cause lives inside the test itself (timing, async, externals).

```bash
# Run ONLY this test, nothing before it.
npm test -- --runTestsByPath flaky.spec.js
```

*What just happened:* This single comparison splits your search in half. Passes alone, fails together → look outward at other tests and shared state. Fails alone too → look inward at the test's own dependencies.

**Seed and pin the inputs.** If you suspect randomness or order, stop letting them vary. Fix the random seed and fix the run order, then re-run. If pinning them makes the flake disappear or become 100% reproducible, you've found your hidden input.

```bash
# Force a fixed test order and a fixed random seed, then loop.
npm test -- --seed=12345 --no-shuffle
```

*What just happened:* Freezing the things you suspected converts nondeterminism into determinism. A flake that vanishes under a fixed seed *was* a randomness flake; one that becomes 100% reproducible under a fixed order *was* an order flake. Pinning turns "sometimes" into "always" or "never," and both answers are useful.

💡 **Key point.** Diagnosis isn't about reading code harder. It's about *changing one variable at a time* - order, seed, isolation - until the failure becomes predictable. A flake you can reproduce on demand is already half-fixed.

## Step 2: Fix - take control of the dependency

Every fix is the same shape: take control of the thing the test didn't control. Here's the cure for each culprit from Phase 2.

**Control time.** Don't read the real clock - inject a fake one your test owns. Then "now" is whatever you say it is, on every machine, forever.

```javascript
// Before: depends on the real clock, fails near a day boundary.
expect(formatTimestamp(Date.now())).toBe('2026-06-30');

// After: freeze time so the test is deterministic.
jest.useFakeTimers().setSystemTime(new Date('2026-06-30T12:00:00Z'));
expect(formatTimestamp(Date.now())).toBe('2026-06-30');
```

*What just happened:* `setSystemTime` makes `Date.now()` return a value *you* chose, so the test no longer cares what the wall clock says or what day CI runs. A frozen clock cures every "fails near midnight" flake.

**Control randomness.** Seed the generator (or inject the value) so "random" is reproducible inside the test.

```python
import random

def test_pick_is_deterministic():
    random.seed(42)            # same seed -> same sequence, every run
    assert random.choice([1, 2, 3]) == 1
```

*What just happened:* With a fixed seed, the "random" choice is identical on every run, so the assertion is stable. You're not removing randomness from production - you're pinning it *in the test*.

**Await properly.** For async flakes, the fix is usually one keyword: actually wait for the work, or wait for the condition, before asserting.

```javascript
// Before: assertion races the save.
test('saves the user', () => {
  saveUser({ name: 'alice' });
  expect(db.count()).toBe(1);
});

// After: await the work; the assertion runs after it completes.
test('saves the user', async () => {
  await saveUser({ name: 'alice' });
  expect(db.count()).toBe(1);
});
```

*What just happened:* The `await` forces the assertion to run only after `saveUser` finishes, eliminating the race. For UI/eventual conditions, the same idea is `await waitFor(() => ...)` - never a fixed `sleep`.

**Isolate state.** Each test must set up its own world and tear it down, assuming nothing about what ran before. Reset shared state in a teardown hook so no test can leak into the next.

```javascript
afterEach(async () => {
  await db.clear();        // every test starts from a clean slate
});
```

*What just happened:* Clearing the database after each test removes the shared state that made order matter, so the suite produces the same result in order, shuffled, or parallel.

**Fake the externals.** Replace the live network/service with a controlled fake so the test only fails when *your* code is wrong, not when a third party hiccups.

```javascript
// Replace the real HTTP call with a stub that returns a fixed response.
jest.spyOn(api, 'getPrice').mockResolvedValue({ status: 200, price: 42 });
```

*What just happened:* The test no longer touches the real network, so a slow or down API can't make it flake.

> ⏭️ Faking the network, clock, and services properly - stubs, mocks, fakes, and when *not* to use them - is its own skill. See [Mocking & Test Doubles](/guides/mocking-and-test-doubles) for the full picture. Here, the takeaway: **control the edges so only your code decides the result.**

**Close what you open.** For leaks, the fix is disciplined teardown - close connections, handles, timers, and servers in an `afterEach`/`afterAll`, ideally automatically.

```javascript
afterAll(async () => {
  await pool.end();        // release connections so later runs don't exhaust the pool
});
```

*What just happened:* Releasing the connection pool at the end means a long suite can't run itself out of resources, so an innocent downstream test stops being blamed for an earlier test's leak.

⚠️ **Gotcha.** The tempting non-fix is to wrap a flaky test in a retry - "run it 3 times, pass if any pass." That doesn't fix flakiness; it *hides* it, slows the suite, and lets the underlying race survive to corrupt something subtler later. Retries belong around genuinely unreliable *externals* in end-to-end tests, not around races in your own code.

## Step 3: Quarantine - when you can't fix it today

Sometimes you find a flaky test and genuinely can't fix it this minute - it's deep, you're mid-incident, the owner is out. Two of your three options are wrong: leaving it failing randomly **poisons trust** (Phase 1's warning), and deleting it **loses coverage** silently. The right move is the third: **quarantine** it.

```javascript
// Skip until fixed - keeps the build reliably green without losing the test.
test.skip('reconnects after network drop - FLAKY, see issue #482', () => {
  // ...
});
```

*What just happened:* The test is pulled out of the gate *with a breadcrumb* - a tracking issue so it's remembered, not forgotten. The build's green means something again, with a paper trail back to the work. Quarantine is a holding cell, not a graveyard.

🪖 **War story.** The difference between a healthy team and a doomed one isn't whether they have flaky tests - everyone does. It's what they do with them. The doomed team adds a retry and moves on; a year later "CI is flaky" is a fact of life nobody questions. The healthy team quarantines with a ticket, fixes them off the critical path, and keeps the green check meaning something. Same flakes, opposite outcomes.

💡 **Key point.** Quarantine is explicit and tracked; silent ignoring is implicit and forgotten. Never `// eslint-disable` a flake into oblivion or comment it out with no trail. A skipped test with an issue number is a promise; a deleted or silently-disabled test is a lie of omission.

## The whole playbook in one breath

```text
   DIAGNOSE   rerun in a loop → isolate (alone vs suite) → pin seed + order
   FIX        control time · seed randomness · await · isolate state · fake externals · close resources
   QUARANTINE can't fix now? skip + tracking issue (never silently ignore)
```

*What just happened:* That's the entire response to any flaky test. Diagnose until reproducible, fix by taking control of the uncontrolled dependency, and if you truly can't fix it today, quarantine with a trail instead of letting it rot trust.

For builders: in CI, a flaky test is a slow leak in the thing protecting your `main` branch. Treat each flake as a bug ticket the moment it appears, not after it's worn everyone down. [Testing in CI](/guides/testing-in-ci) covers required checks and branch protection - a gate only as trustworthy as the suite behind it.

## Recap

- **Diagnose first.** Rerun in a loop, isolate (alone vs. in-suite), and pin seed + order to turn "sometimes" into "always/never." A reproducible flake is half-fixed.
- **Fix by taking control:** freeze the clock, seed randomness, `await` async work, isolate and tear down state, fake externals, close every resource you open.
- **Don't retry away a race.** Retries hide flakiness and slow the suite; they belong around unreliable externals in E2E, not your own code's races.
- **Quarantine, never silently ignore.** Can't fix today? Skip with a tracking issue so the build stays reliably green and the work is remembered.

```quiz
[
  {
    "q": "What's the FIRST thing to do with a flaky test, before trying to fix it?",
    "choices": ["Wrap it in a retry", "Make it fail on demand - rerun in a loop, isolate, and pin seed/order", "Delete it", "Add a longer sleep"],
    "answer": 1,
    "explain": "You can't confirm a fix you can't reproduce. Diagnose first: loop it, isolate it, pin the suspected inputs until the failure is predictable."
  },
  {
    "q": "Why is wrapping a flaky test in an automatic retry a bad fix for a race in your own code?",
    "choices": ["Retries are not supported by most runners", "It hides the flakiness, slows the suite, and lets the underlying race survive", "It makes the test deterministic", "It deletes the test's coverage"],
    "answer": 1,
    "explain": "A retry sweeps the coin-flip under the rug instead of removing it. Retries belong around genuinely unreliable externals in E2E, not around your own races."
  },
  {
    "q": "You found a flaky test but can't fix it right now. What's the correct move?",
    "choices": ["Delete it to keep the build green", "Leave it failing randomly", "Quarantine it: skip it with a tracking issue", "Wrap the whole suite in a retry"],
    "answer": 2,
    "explain": "Quarantine keeps the build reliably green without losing track: skip plus a tracking issue is a tracked promise to fix, unlike deleting or silently ignoring."
  }
]
```
