# Playwright, From Zero

> Reliable browser end-to-end tests: auto-waiting locators that kill flakiness, cross-browser runs, tracing, and codegen to record a test by clicking.


---

# Playwright, From Zero

You wrote an end-to-end test. It passed five times, then failed on CI for no reason you can name. You added a `sleep(2)`, it passed, you moved on, and it failed again next week. That's not your fault - it's the tool fighting you. Playwright was built by people who lived that exact pain, and it removes most of it by design.

This guide gets you from nothing to a test suite you actually trust: tests that wait for the right thing automatically, run across Chromium, Firefox, and WebKit, and hand you a time-travel recording when something breaks.

## How to read this

Read the phases in order the first time. Phase 1 is the mental model - what Playwright is and why auto-waiting locators are the whole point. Phase 2 is the daily loop: writing, running, and debugging tests. Phase 3 is the stuff that bites you in real projects. Type the commands as you go; that's how this sticks.

## The phases

1. [The mental model: a browser you can boss around](01-the-mental-model.md) - what Playwright is, and why auto-waiting kills flakiness.
2. [The everyday loop: write, run, debug](02-the-everyday-loop.md) - locators, web-first assertions, codegen, and the trace viewer.
3. [Production reality: the things that bite](03-production-reality.md) - auth state, parallelism, network mocking, CI, and the classic traps.


---

# The mental model: a browser you can boss around

Here's the situation you're probably in. Your app works when you click through it by hand. But you want a machine to click through it too - log in, add an item to the cart, check the total - so you find out the moment a deploy breaks the checkout flow. That's an **end-to-end (E2E) test**: it drives a real browser the way a real user would, and asserts that the right things happened.

Playwright is a library and test runner for exactly that. You write code that says "go to this URL, fill this box, click that button, and the page should now show 'Order confirmed'." Playwright opens an actual browser, performs those actions, and checks the result.

## What Playwright actually is

Strip away the marketing and Playwright is three things bolted together:

- A **driver** that controls real browser engines - Chromium (Chrome/Edge), Firefox, and WebKit (Safari) - over a fast, low-level protocol.
- A **test runner** (`@playwright/test`) that finds your test files, runs them in parallel, retries failures if you ask it to, and produces reports.
- A pile of **debugging tooling** - codegen, the inspector, and the trace viewer - that makes a broken test cheap to diagnose instead of a mystery.

You install it with one command, and it downloads the browser binaries it controls so you're not depending on whatever Chrome happens to be on the machine.

```bash
# In an existing project
npm init playwright@latest

# It scaffolds a config, an example test, and downloads browsers
# Then run the example suite:
npx playwright test
```

*What just happened:* the init command created `playwright.config.ts`, a `tests/` folder with a sample spec, and pulled down pinned Chromium, Firefox, and WebKit builds. `npx playwright test` ran every spec it found against those browsers.

## Why auto-waiting is the entire point

This is the part that matters, so slow down here.

The classic E2E nightmare is **flakiness**: a test that passes and fails without the code changing. The root cause is almost always timing. The page hasn't finished loading, a button isn't clickable yet because a spinner is on top of it, an element exists in the DOM but is still `display: none`. Older tools made you guess how long to wait. You'd write `sleep(2000)` and pray. Too short and it's flaky; too long and your suite crawls.

Playwright's answer is the **locator**, and locators **auto-wait**. A locator isn't the element - it's a *recipe* for finding the element, evaluated fresh every time you act on it. Before Playwright clicks, it automatically waits for the element to be present, visible, stable (not animating), enabled, and actually able to receive the click. Only then does it click. If those conditions aren't met within the timeout, you get a clear error instead of a silent misclick.

```js
// A locator: a description, not a snapshot
const submit = page.getByRole('button', { name: 'Submit' });

// click() auto-waits: present + visible + stable + enabled + not obscured
await submit.click();
```

*What just happened:* `getByRole(...)` built a locator but touched nothing yet. `click()` ran the actionability checks, waited until the button was genuinely clickable, then clicked - no `sleep`, no manual wait, no flake from "the spinner was still up."

> The mental shift: you stop telling the browser *when* to act and start telling it *what must be true* before it acts. Playwright fills in the waiting. Most flakiness you've ever fought disappears at this layer.

For the bigger picture of where E2E sits next to unit and integration tests, see /guides/unit-integration-e2e. For the deeper anatomy of why tests flake and how to fight it, /guides/flaky-tests goes further than we will here.

## Why it largely displaced Selenium for new projects

Selenium pioneered browser automation and still runs huge suites in production. But for *new* projects, teams keep reaching for Playwright, and the reasons are concrete, not fashion:

- **Auto-waiting is built in.** With Selenium you assemble explicit waits yourself; getting them right everywhere is the hard part. Playwright bakes the right waits into every action.
- **One install, pinned browsers.** No separate driver binaries to version-match against your browser.
- **First-class tooling.** Codegen records a test by watching you click; the trace viewer lets you scrub through a failed run frame by frame. We'll use both in phase 2.
- **Cross-browser from day one.** The same test runs on Chromium, Firefox, and WebKit by flipping a config option.

None of this means Selenium is wrong. It means Playwright removed the most painful day-to-day friction, and that's why greenfield suites tend to start here.

## A whole test, start to finish

Here's a complete, realistic test so the shape is concrete before phase 2 zooms in.

```js
import { test, expect } from '@playwright/test';

test('user can log in', async ({ page }) => {
  await page.goto('https://example.com/login');
  await page.getByLabel('Email').fill('ada@example.com');
  await page.getByLabel('Password').fill('correct-horse');
  await page.getByRole('button', { name: 'Log in' }).click();

  // web-first assertion: retries until true or times out
  await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();
});
```

*What just happened:* the test got a fresh `page` (an isolated browser tab), navigated, filled two fields by their labels, clicked Log in, then asserted the Dashboard heading appears. Every step auto-waited, so there's not a single `sleep` and nothing to make it flaky.

**For builders:** notice the test reads like a sentence describing user behavior. That's deliberate - locators like `getByRole` and `getByLabel` target what the user sees and what assistive tech announces, not brittle CSS classes that change every time someone touches the styling.

```quiz
[
  {
    "q": "What is a Playwright locator?",
    "choices": [
      "A snapshot of an element captured the moment you create it",
      "A reusable recipe for finding an element, re-evaluated each time you act on it",
      "A CSS selector that must be unique on the page",
      "A screenshot used for visual comparison"
    ],
    "answer": 1,
    "explain": "A locator describes how to find an element and is resolved fresh on each action, which is what makes auto-waiting possible."
  },
  {
    "q": "Why does auto-waiting reduce flakiness?",
    "choices": [
      "It runs tests more slowly so the page has time to settle",
      "It disables animations globally",
      "Before acting, it waits for the element to be present, visible, stable, and actionable instead of you guessing a fixed delay",
      "It retries the entire test file three times"
    ],
    "answer": 2,
    "explain": "Auto-waiting replaces fixed sleeps with condition checks, removing the timing guesswork that causes most flaky failures."
  },
  {
    "q": "What does `npm init playwright@latest` set up?",
    "choices": [
      "Only a config file, no browsers",
      "A config, a sample test, and pinned browser binaries it downloads",
      "A connection to your system's installed Chrome only",
      "A cloud account for running tests remotely"
    ],
    "answer": 1,
    "explain": "The init command scaffolds config and a sample spec and downloads its own pinned Chromium, Firefox, and WebKit builds."
  }
]
```


---

# The everyday loop: write, run, debug

Phase 1 gave you the mental model. Now we live in the tool. The daily rhythm with Playwright is short: pick good locators, assert with web-first assertions, run a focused subset while you work, and when something breaks, open the trace instead of staring at a stack trace. Let's walk each piece.

## Picking locators that don't rot

A test is only as durable as the way it finds elements. Reach for locators in roughly this order, most resilient first:

```js
// Best: by accessible role + name (what the user perceives)
page.getByRole('button', { name: 'Save' });
page.getByRole('link', { name: 'Settings' });

// Great for forms: by associated label or placeholder
page.getByLabel('Email');
page.getByPlaceholder('Search products');

// User-visible text
page.getByText('Welcome back');

// Explicit test hook when nothing semantic fits
page.getByTestId('cart-total'); // matches data-testid="cart-total"
```

*What just happened:* each call returned a locator keyed to something stable - a role, a label, visible text, or an explicit test id. None of them depend on CSS classes or DOM nesting, so a restyle or a wrapper `<div>` won't break them.

> Avoid `page.locator('.btn-primary')` and deep CSS/XPath chains when you can. They're glued to your markup's current shape and snap the moment a developer refactors the HTML. `getByTestId` is the escape hatch when nothing semantic fits - add a `data-testid` attribute rather than reaching for a class.

When a locator could match several elements, narrow it instead of guessing an index:

```js
// Scope inside a region, then find within it
const row = page.getByRole('row', { name: 'Widget Pro' });
await row.getByRole('button', { name: 'Delete' }).click();
```

*What just happened:* chaining `row.getByRole(...)` searched only inside that one table row, so "Delete" resolved unambiguously even though the page has many Delete buttons.

## Web-first assertions: the other half of auto-waiting

You met auto-waiting on *actions*. Assertions get it too. The `expect()` calls that take a locator are **web-first** - they retry until the condition holds or the timeout hits. This is the difference between checking a value once (and racing the UI) and checking it patiently.

```js
// Retries automatically until visible - or fails with a clear timeout
await expect(page.getByText('Order confirmed')).toBeVisible();

// Other common web-first assertions
await expect(page.getByRole('button', { name: 'Save' })).toBeEnabled();
await expect(page.getByLabel('Email')).toHaveValue('ada@example.com');
await expect(page).toHaveURL(/\/dashboard/);
await expect(page.getByTestId('cart-total')).toHaveText('$42.00');
```

*What just happened:* each `await expect(locator)...` polled the page until the assertion passed. No `sleep` before checking the confirmation message - the assertion itself does the waiting.

The trap to avoid: don't pull a value out and assert on the plain value, because that snapshots a single moment.

```js
// Fragile: reads once, races the UI
const text = await page.getByTestId('cart-total').textContent();
expect(text).toBe('$42.00'); // may run before the total updates

// Solid: web-first, retries until the total settles
await expect(page.getByTestId('cart-total')).toHaveText('$42.00');
```

*What just happened:* the first version grabbed the text immediately and compared once - flaky if the total updates a beat later. The second keeps re-checking until the text matches or it times out.

## Codegen: record a test by clicking

You don't have to write the first draft by hand. `codegen` opens a browser, watches what you do, and emits the matching Playwright code with sensible locators already chosen.

```bash
npx playwright codegen https://example.com
```

*What just happened:* a browser window and an inspector opened side by side. As you clicked and typed, the inspector filled with real test code - `getByRole`, `getByLabel`, and the actions you performed - which you copy into a spec and clean up.

Treat codegen output as a **starting point**, not the final test. It captures the actions; you still add the assertions that say what *should* be true, and you trim any noise.

## Running tests while you work

You rarely run the whole suite during development. Run a slice:

```bash
# Everything
npx playwright test

# One file
npx playwright test tests/login.spec.ts

# By title substring
npx playwright test -g "log in"

# Watch it happen in a real browser window
npx playwright test --headed

# One browser only (faster feedback loop)
npx playwright test --project=chromium

# The interactive UI mode - the nicest way to develop
npx playwright test --ui
```

*What just happened:* each flag narrowed or changed how the run executes. `--ui` is the standout: it opens a panel where you pick tests, watch them step through, and inspect each action - the tightest write-run-debug loop Playwright offers.

By default tests run **headless** (no visible window) and in **parallel**, which is why a suite finishes fast. `--headed` and a single `--project` slow things down on purpose so you can see what's going on.

## The trace viewer: time-travel debugging

This is the feature that pays for itself the first time a test fails on CI and you can't reproduce it locally. A **trace** is a recorded bundle of everything that happened during a run - a DOM snapshot before and after every action, console logs, network requests, and screenshots. Open it and you scrub through the run like a video, clicking any step to see the page exactly as it was.

Turn it on in `playwright.config.ts`:

```js
// playwright.config.ts
export default defineConfig({
  use: {
    // Record a trace only when a test retried and still needs diagnosing
    trace: 'on-first-retry',
  },
});
```

*What just happened:* Playwright now captures a trace whenever a test fails and gets retried, so you get a recording of real failures without bloating every green run with trace files.

Then open whatever it captured:

```bash
# After a failing run, open the report (traces are linked from it)
npx playwright show-report

# Or open a specific trace file directly
npx playwright show-trace trace.zip
```

*What just happened:* `show-trace` launched the viewer with a timeline of every action. Clicking a step shows the before/after DOM snapshot, the locator that was used, network activity, and console output at that exact moment - so "why did this fail on CI?" becomes a thing you can watch instead of guess.

**In the wild:** the common workflow is `trace: 'on-first-retry'` in CI plus the HTML report uploaded as a build artifact. A test goes red, you download the report, open the trace, and within a minute you see the spinner that was still covering the button - no re-running CI ten times.

```quiz
[
  {
    "q": "Which locator strategy is the most resilient to a CSS restyle?",
    "choices": [
      "page.locator('.btn-primary')",
      "page.getByRole('button', { name: 'Save' })",
      "An XPath like //div[3]/button",
      "page.locator('#app > div > button:nth-child(2)')"
    ],
    "answer": 1,
    "explain": "Role + accessible name targets what the user perceives, not the markup's current classes or nesting, so styling changes don't break it."
  },
  {
    "q": "Why prefer `await expect(locator).toHaveText('$42.00')` over reading textContent and comparing?",
    "choices": [
      "It is shorter to type",
      "It only works in headed mode",
      "It is web-first and retries until the text matches or times out, avoiding a race with the UI",
      "It captures a screenshot automatically"
    ],
    "answer": 2,
    "explain": "Web-first assertions poll until the condition holds, so they don't snapshot a single moment before the UI has updated."
  },
  {
    "q": "What does the trace viewer give you that a stack trace does not?",
    "choices": [
      "A faster test run",
      "A scrubbable recording with before/after DOM snapshots, network, and console for each action",
      "Automatic fixing of the failing test",
      "A way to run tests in the cloud"
    ],
    "answer": 1,
    "explain": "A trace records the full run so you can step through every action and see the page exactly as it was at the moment of failure."
  }
]
```


---

# Production reality: the things that bite

A handful of passing tests on your laptop is one thing. A suite that runs on every pull request, stays fast, and doesn't cry wolf is another. This phase is the gap between the two: how to skip logging in on every test, how parallelism really works, how to stop depending on a flaky backend, and the traps that catch almost everyone once.

## Don't log in a hundred times - reuse auth state

If every test starts by filling the login form, your suite is slow and the login flow is tested a hundred redundant times. The fix: log in once, save the browser's storage (cookies and localStorage) to a file, and load it into every test. Playwright calls this **storage state**.

```js
// auth.setup.ts - runs once before the tests that need a logged-in user
import { test as setup, expect } from '@playwright/test';

const authFile = 'playwright/.auth/user.json';

setup('authenticate', async ({ page }) => {
  await page.goto('https://example.com/login');
  await page.getByLabel('Email').fill('ada@example.com');
  await page.getByLabel('Password').fill('correct-horse');
  await page.getByRole('button', { name: 'Log in' }).click();
  await expect(page.getByRole('heading', { name: 'Dashboard' })).toBeVisible();

  // Persist cookies + localStorage to disk
  await page.context().storageState({ path: authFile });
});
```

*What just happened:* a setup test logged in like a normal user, then wrote the resulting cookies and localStorage to `user.json`. Every test that loads this file starts already authenticated - the login form runs once, not once per test.

Wire it up in the config so real tests depend on it and inherit the saved state:

```js
// playwright.config.ts
export default defineConfig({
  projects: [
    { name: 'setup', testMatch: /auth\.setup\.ts/ },
    {
      name: 'chromium',
      use: { ...devices['Desktop Chrome'], storageState: 'playwright/.auth/user.json' },
      dependencies: ['setup'],
    },
  ],
});
```

*What just happened:* the `chromium` project depends on `setup`, so Playwright runs the login once first, then runs every chromium test with the saved storage state already loaded. (Add `playwright/.auth/` to `.gitignore` - those files hold session credentials.)

## Fixtures: the `page` you've been getting for free

Every test so far destructured `{ page }`. That `page` is a **fixture** - a piece of test environment Playwright builds for you, fresh per test, and tears down afterward. Built-in fixtures include `page`, `context` (an isolated browser session), and `browser`. The point of per-test fixtures is **isolation**: each test gets a clean context with no leftover cookies or state from the last one, which is a huge source of cross-test flakiness eliminated by default.

You can define your own to remove repeated setup:

```js
// Provide an already-on-the-dashboard page to any test that asks for it
import { test as base } from '@playwright/test';

export const test = base.extend({
  dashboardPage: async ({ page }, use) => {
    await page.goto('/dashboard');     // setup
    await use(page);                   // hand it to the test
    // any teardown would go here, after use()
  },
});
```

*What just happened:* `dashboardPage` is a custom fixture. A test that asks for `{ dashboardPage }` receives a page already navigated to the dashboard, and the setup lives in one place instead of being copy-pasted into every test.

## Parallelism and isolation

By default Playwright runs test **files** in parallel across multiple worker processes, and each test gets its own isolated browser context. Fast, but it has consequences you need to respect:

- **Tests must not depend on each other or on order.** A worker may run any file at any time. Shared state between tests is a bug waiting to surface.
- **Watch for shared backend data.** If two parallel tests both create a user named `test@example.com`, they collide. Generate unique data per test (a timestamp or random suffix) or scope to per-worker data.

You control the degree of parallelism when you need to:

```bash
# Limit workers (e.g. a resource-constrained CI box)
npx playwright test --workers=2

# Force fully serial for one stubborn file (last resort)
# test.describe.configure({ mode: 'serial' }) inside the file
```

*What just happened:* `--workers=2` capped parallelism for the whole run. `mode: 'serial'` is the in-file escape hatch when tests genuinely must share order - use it sparingly, because serial mode also means one failure skips the rest of that group.

## Stop depending on a flaky backend: mock the network

E2E tests that hit a real API inherit that API's flakiness and slowness. When you're testing the *front end's* behavior - does it render this data, does it handle this error - intercept the request and return a fixed response. This makes the test fast, deterministic, and able to exercise error states you can't easily trigger for real.

```js
test('shows empty state when no orders', async ({ page }) => {
  // Intercept the API call and return controlled data
  await page.route('**/api/orders', async (route) => {
    await route.fulfill({ json: [] });
  });

  await page.goto('/orders');
  await expect(page.getByText('No orders yet')).toBeVisible();
});
```

*What just happened:* `page.route(...)` caught the request to `/api/orders` and answered with an empty array - no real backend involved. The test then asserted the empty-state message, deterministically, every run.

> Mock for front-end behavior tests; keep a few unmocked, full-stack "smoke" tests that hit the real system end to end. The mocked tests give you speed and coverage of edge cases; the smoke tests prove the pieces actually connect. You want both, not one or the other.

## Cross-browser, for real

Phase 1 promised cross-browser. Here's the cost-benefit. Add the engines as projects:

```js
// playwright.config.ts
projects: [
  { name: 'chromium', use: { ...devices['Desktop Chrome'] } },
  { name: 'firefox',  use: { ...devices['Desktop Firefox'] } },
  { name: 'webkit',   use: { ...devices['Desktop Safari'] } },
],
```

*What just happened:* every test now runs three times, once per engine, catching the Safari-only and Firefox-only bugs that Chromium-only suites miss. The tradeoff is roughly triple the runtime - many teams run all three on the main branch and a single engine on routine pull requests.

## CI and the classic traps

Running in CI is a config detail and a few hard-won lessons:

```bash
# In CI: install browsers with their OS dependencies
npx playwright install --with-deps

# Run; CI mode enables retries/reporters per your config
npx playwright test
```

*What just happened:* `--with-deps` installed the system libraries the browsers need on a bare CI image - the single most common reason a suite that works locally explodes on first CI run.

The traps that catch nearly everyone at least once:

- **Real waits sneaking back in.** `await page.waitForTimeout(2000)` is the new `sleep` - it's in the API for rare cases, but if it's in normal tests you've reintroduced flakiness. Use web-first assertions instead. See /guides/flaky-tests for the full pattern.
- **Asserting on a value instead of a locator.** Re-read phase 2's `toHaveText` vs `textContent` example; this is the number-one source of "passes locally, fails in CI."
- **Order-dependent tests.** They pass serially on your machine and fail under parallel workers. Make each test self-contained.
- **Committing auth/storage-state files.** They contain live session tokens. Gitignore them and regenerate in CI.
- **Strict-mode violations.** If a locator matches more than one element, Playwright errors on purpose rather than silently picking the first. That error is a feature - narrow the locator (scope it, or use an accessible name) instead of suppressing it.

**For builders:** a healthy suite is mostly mocked behavior tests for speed, a thin layer of real-backend smoke tests for confidence, `trace: 'on-first-retry'` so failures are diagnosable, and storage-state auth so it stays fast. Get those four right and your E2E tests become something the team trusts instead of mutes.

```quiz
[
  {
    "q": "Why save and reuse storage state across tests?",
    "choices": [
      "To run tests in parallel",
      "To log in once and start every test already authenticated, instead of running the login flow in every test",
      "To record a trace of the login",
      "To mock the network"
    ],
    "answer": 1,
    "explain": "Storage state persists cookies and localStorage so tests load an authenticated session rather than re-running the slow login each time."
  },
  {
    "q": "What is the main risk introduced by running tests in parallel?",
    "choices": [
      "Tests run too slowly",
      "Traces stop being recorded",
      "Tests that depend on order or share backend data collide because workers run files in any order",
      "Locators stop auto-waiting"
    ],
    "answer": 2,
    "explain": "Parallel workers run files in arbitrary order, so order-dependent or shared-state tests break. Each test must be self-contained."
  },
  {
    "q": "When does `page.route(...)` to mock an API make the most sense?",
    "choices": [
      "For every test, to avoid ever touching the backend",
      "When testing front-end behavior deterministically, including error states, while keeping a few real-backend smoke tests",
      "Only in CI, never locally",
      "Only to speed up the login flow"
    ],
    "answer": 1,
    "explain": "Mock to make front-end behavior tests fast and deterministic and to exercise hard-to-trigger states, but keep some unmocked smoke tests that prove the real system connects."
  }
]
```
