# Build a Web Scraper (Python)

> Build a real web scraper in Python with requests and BeautifulSoup - fetch, parse, extract structured data, paginate politely, and save it - run on your own machine.


---

# Build a Web Scraper (Python)

This weekend you and I are going to build a web scraper that actually pulls
data off real pages and saves it to a file you can open in a spreadsheet. Not
a toy that prints "hello" - a small program you'd be comfortable pointing at a
catalog of books, a list of quotes, or a directory of listings, and walking
away while it does the boring work.

You build this **on your own machine**. The code here is meant to be saved to
files and run from your terminal with `python script.py`. There's no browser
sandbox doing the work for you, because real scraping needs the network, the
filesystem, and a couple of libraries that don't live in your browser. That's
the real version of the skill, and it's the version that's useful on Monday.

## What you'll build

A command-line scraper, built up one piece at a time, that:

- fetches a web page over HTTP and checks it actually came back OK,
- parses the HTML into something you can search,
- pulls named fields (title, price, rating) into clean Python dictionaries,
- walks from one page to the next, with delays so you're a good guest,
- and writes everything to CSV and JSON.

By the last phase it's one working program. Each phase before that leaves you
with a real, runnable piece of it.

## The stack

| Piece | What it does | Why this one |
|-------|--------------|--------------|
| Python 3.10+ | the language | batteries-included, great for this |
| `requests` | fetches pages | the friendliest HTTP client around |
| `beautifulsoup4` | parses HTML | forgiving with messy real-world markup |
| `csv` / `json` | save the output | both ship with Python, nothing to install |

We use a practice site built for exactly this - `https://books.toscrape.com`
and `https://quotes.toscrape.com`. They exist so people can learn scraping
without bothering anyone's production servers. We'll talk about why that
matters in Phase 4.

## Rough time

A focused weekend. Call it three to four hours if you type along. Phases 1–3
are the core and move quickly; Phase 4 is where the judgment lives; Phase 5 is
the satisfying part where data lands in a file.

## What you'll learn

- How an HTTP request and response actually work, in code you can read.
- How to find the elements you want using both BeautifulSoup's `find` family
  and CSS selectors - and when to reach for which.
- How to write extraction that doesn't crash the moment a field is missing.
- The etiquette and the law of scraping: robots.txt, rate limits, terms of
  service, and what "be polite" means in practice.
- How to persist structured data and where you'd take this next - a database,
  a schedule, or a headless browser for sites that build themselves with
  JavaScript.

## The shape of it

```mermaid
graph LR
  A[Fetch page] --> B[Parse HTML]
  B --> C[Extract fields]
  C --> D[Next page?]
  D -->|yes| A
  D -->|no| E[Save CSV / JSON]
```

That loop is the whole game. Everything we build hangs off those five boxes.
Let's get a machine ready and pull down our first page.


---

# Setup and Fetch a Page

Before we parse anything, we need two things: a tidy place for this project to
live, and proof that we can pull a page off the internet at all. This phase
gets you both. By the end you'll run a script and watch a real HTML document
land in your terminal.

A quick note since you're building this on your machine: everything here
assumes you can open a terminal (Terminal on macOS, PowerShell or Git Bash on
Windows, your shell of choice on Linux) and that `python --version` prints
something that starts with `3.10` or higher. If it doesn't, install a current
Python from python.org first, then come back.

## Make a home for the project

Create a folder and step into it. A scraper is a project, not a one-off snippet,
and it'll grow over the weekend.

```bash
mkdir book-scraper
cd book-scraper
```

## The virtual environment, and why you want one

A virtual environment is a private copy of Python's package area that belongs to
this project alone. Install `requests` into it and you haven't touched the
Python your operating system relies on, and you haven't left a mess for the next
project. It's a habit worth keeping for every Python project you start.

Create one and turn it on:

```bash
python -m venv .venv
```

Now activate it. The command differs by platform:

```bash
# macOS / Linux
source .venv/bin/activate

# Windows PowerShell
.venv\Scripts\Activate.ps1

# Windows Git Bash
source .venv/Scripts/activate
```

Once it's active your prompt picks up a `(.venv)` prefix. That prefix is your
sign that `pip` and `python` now point inside the project. Any time you come
back to work on this, activate it again first.

## Install the two libraries

With the environment active, install what we need:

```bash
pip install requests beautifulsoup4
```

`requests` handles the talking-to-servers part. `beautifulsoup4` (you import it
as `bs4`) handles the reading-the-HTML part - we'll meet it properly in Phase 2,
but installing it now keeps us from a second trip to pip.

It's worth recording exactly what you installed so you - or anyone you share
this with - can recreate the environment:

```bash
pip freeze > requirements.txt
```

That writes a file listing every package and its version. Later, on another
machine, `pip install -r requirements.txt` rebuilds the same setup.

## Fetch your first page

Now the part you came for. Create a file called `fetch.py` and put this in it:

```python
import requests

URL = "https://books.toscrape.com/"

response = requests.get(URL)

print("Status code:", response.status_code)
print("Content type:", response.headers.get("Content-Type"))
print("Body length:", len(response.text), "characters")
print("First 300 characters:")
print(response.text[:300])
```

Run it:

```bash
python fetch.py
```

You should see something like a `200` status code, a content type mentioning
`text/html`, a body that's many thousands of characters long, and the opening of
an HTML document. That `200` is the whole point of this phase - the server
heard you and sent back the page.

## Read what came back

Let's slow down on `response`, because it's the object every later phase builds
on. A few of its parts you'll use constantly:

| Attribute | What it gives you |
|-----------|-------------------|
| `response.status_code` | the HTTP result number (200 = OK) |
| `response.text` | the page body as a string (decoded for you) |
| `response.content` | the raw bytes, if you need them |
| `response.headers` | a dict-like of response headers |
| `response.url` | the final URL, after any redirects |

Here's the flow of what happens when you call `requests.get`:

```mermaid
graph LR
  A[requests.get URL] --> B[Server]
  B --> C[Response object]
  C --> D[status_code]
  C --> E[text - the HTML]
```

## Don't trust a 200 you didn't check

Right now our script prints the status and moves on. In a real run you want to
*stop* if the page didn't come back. A 404 means the page is gone; a 500 means
the server broke; a 403 might mean you're being blocked. `requests` gives you a
one-line way to turn any of those into an error you can't ignore.

Update `fetch.py` to guard the request:

```python
import requests

URL = "https://books.toscrape.com/"

response = requests.get(URL, timeout=10)
response.raise_for_status()   # raises an exception on 4xx / 5xx

print("OK:", response.status_code)
print("Got", len(response.text), "characters of HTML")
```

Two changes earn their keep here. `timeout=10` means the request gives up after
ten seconds instead of hanging your program forever when a server goes quiet -
never make a real request without a timeout. And `raise_for_status()` turns a
bad status into a loud crash, so you find out something's wrong immediately
rather than parsing an error page as if it were data.

Run it again. A clean `OK: 200` and a character count means your environment is
solid and the network path works.

## Where we are

You have a project folder, an isolated environment with both libraries, a
recorded `requirements.txt`, and a script that fetches a live page and refuses
to continue on a bad response. That's the foundation. Next we take that big
string of HTML and turn it into something we can actually search through.


---

# Parsing the HTML

Last phase we ended with a giant string of HTML. A string is hard to work with -
you can't ask a string "give me every book title on this page." This phase turns
that string into a tree you *can* ask questions of, using BeautifulSoup, and
shows you the two ways to find things in it.

We're working with `https://books.toscrape.com/` - a fake bookstore made for
practice. Open it in your browser and right-click a book, then choose "Inspect,"
so you can see the HTML we're about to navigate. Scraping is half code, half
reading someone else's markup.

## Load the HTML into BeautifulSoup

Create `parse.py`. We fetch the page (same as before) and hand the body to
BeautifulSoup:

```python
import requests
from bs4 import BeautifulSoup

URL = "https://books.toscrape.com/"

response = requests.get(URL, timeout=10)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")

print(soup.title)        # the <title> tag
print(soup.title.text)   # just the text inside it
```

Run `python parse.py`. You should see the `<title>` tag and then its text. That
`soup` object is the whole page as a navigable tree. `"html.parser"` is Python's
built-in parser - nothing extra to install. (There are faster parsers like
`lxml`, but the built-in one is right for learning and fine for most jobs.)

## The find family

BeautifulSoup gives you two close cousins: `find` returns the **first** matching
element, and `find_all` returns a **list** of every match. You'll lean on these
constantly.

```python
# The first <h3> on the page
first_h3 = soup.find("h3")
print("First h3:", first_h3.text)

# Every <article> with class "product_pod" - each one is a book
books = soup.find_all("article", class_="product_pod")
print("Books found on this page:", len(books))
```

Two things to notice. You match by tag name (`"h3"`, `"article"`), and you can
narrow by attribute. Class is special: because `class` is a reserved word in
Python, BeautifulSoup spells the keyword `class_` with a trailing underscore.
You'll hit that one a lot.

On this page you should see 20 books - that's how many fit on a page before
pagination kicks in (Phase 4's problem).

## Reach inside a matched element

`find` and `find_all` work on any element, not only the whole soup. So once you
have a single book, you search *within* it for the title and price. Look at the
inspected HTML: the title sits in an `<a>` inside the `<h3>`, and the actual
title is in that link's `title` attribute. The price sits in a `<p>` with class
`price_color`.

```python
first_book = books[0]

# The link inside this book's <h3>
link = first_book.find("h3").find("a")
print("Title:", link["title"])     # read an attribute with [ ]

# The price paragraph
price = first_book.find("p", class_="price_color")
print("Price:", price.text)
```

Reading an attribute uses square brackets, like a dict: `link["title"]`,
`link["href"]`. Reading the visible text uses `.text`. Mixing those two up is
the most common early stumble, so it's worth saying out loud: brackets for
attributes, `.text` for what's between the tags.

## The other way: CSS selectors

There's a second style, and once it clicks many people never go back. If you
know CSS - the selectors you'd write in a stylesheet - you can use the exact same
syntax to find elements, with `select` (returns a list) and `select_one`
(returns the first).

```python
# Every book, via CSS selector
books = soup.select("article.product_pod")
print("Books:", len(books))

# Title link inside the first book
link = soup.select_one("article.product_pod h3 a")
print("Title:", link["title"])

# Price inside the first book
price = soup.select_one("article.product_pod p.price_color")
print("Price:", price.text)
```

Same results, different spelling. `article.product_pod` means "an `<article>`
with class `product_pod`." A space means "descendant of" - so
`article.product_pod h3 a` reads as "an `<a>` somewhere inside an `<h3>` somewhere
inside that article." If you can read CSS, you can read these.

## Which one should you use?

Neither is "correct." Here's how I choose:

| Situation | Reach for |
|-----------|-----------|
| One condition, by tag or class | either; `find` reads plainly |
| Deeply nested path | `select` - one selector beats nested `find` calls |
| Matching by class only | `select(".price_color")` is shorter than `find_all` |
| Logic between steps (loop, branch) | `find` - you stay in Python |
| You already think in CSS | `select` will feel like home |

A handy trick from your browser: inspect an element, right-click it in the
elements panel, and many browsers offer "Copy → Copy selector." That hands you a
CSS selector you can paste straight into `select_one`. Trim it down - the copied
version is often longer than it needs to be - but it's a fast start.

## See the whole page's structure

To get a feel for the tree, print every book's title in one pass:

```python
import requests
from bs4 import BeautifulSoup

response = requests.get("https://books.toscrape.com/", timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")

for book in soup.select("article.product_pod"):
    title = book.select_one("h3 a")["title"]
    print("-", title)
```

Run it. Twenty titles scroll past. You read a real page's worth of data out
of raw HTML - that's the parsing skill, and it's the heart of every scraper.

## Where we are

You can load HTML into a searchable tree and pull out exactly the elements you
want, two different ways, both within the whole page and within a single item.
Right now we're printing loose pieces. Next phase we gather those pieces into
clean, structured records - one tidy dictionary per book - and make the code
survive a page where a field is missing.


---

# Extracting Structured Data

So far we've printed pieces - a title here, a price there. A real scraper
produces *records*: one structured object per thing, with the same fields every
time, clean enough to drop into a spreadsheet without hand-fixing. This phase
builds that. By the end you'll have a function that turns one book's HTML into
one tidy dictionary, and a loop that gives you a list of them.

The dictionary is our record. Each book becomes
`{"title": ..., "price": ..., "rating": ..., "in_stock": ..., "url": ...}`.
Same keys, every book. That sameness is what makes the next phase - saving -
trivial.

## Extract one book into a dict

Create `extract.py`. We'll write a function that takes a single book element and
returns a dictionary. Look at the page's HTML again: the rating lives in a
class like `star-rating Three`, the stock status is text in a `p.instock`, and
the link is a relative `href` we'll need to fix up.

```python
import requests
from bs4 import BeautifulSoup

BASE = "https://books.toscrape.com/"


def parse_book(book):
    link = book.select_one("h3 a")
    title = link["title"]

    price = book.select_one("p.price_color").text

    # Rating is encoded in the class, e.g. "star-rating Three"
    rating_classes = book.select_one("p.star-rating")["class"]
    rating = rating_classes[1]   # ["star-rating", "Three"] -> "Three"

    in_stock = book.select_one("p.instock.availability").text

    url = BASE + link["href"]

    return {
        "title": title,
        "price": price,
        "rating": rating,
        "in_stock": in_stock,
        "url": url,
    }


response = requests.get(BASE, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")

first = soup.select("article.product_pod")[0]
print(parse_book(first))
```

Run `python extract.py`. You'll get a dictionary - but a slightly grubby one.
The stock text is wrapped in whitespace and newlines, and the price has a stray
character on the front. Let's clean it.

## Clean the text as you pull it

Raw HTML text is full of indentation, newlines, and the odd encoding artifact.
Clean it at the moment of extraction, so every record downstream is already
tidy. A few moves cover almost everything:

| Problem | Fix |
|---------|-----|
| Leading/trailing whitespace, newlines | `.text.strip()` |
| A currency symbol you want as a number | `.replace("£", "")` then `float(...)` |
| Internal double spaces | `" ".join(text.split())` |

Here's `parse_book` with cleaning built in, turning the price into a real number
we can sort and total later:

```python
def parse_book(book):
    link = book.select_one("h3 a")
    title = link["title"].strip()

    raw_price = book.select_one("p.price_color").text   # e.g. "£51.77"
    price = float(raw_price.replace("£", "").strip())

    rating_classes = book.select_one("p.star-rating")["class"]
    rating = rating_classes[1]

    in_stock = book.select_one("p.instock.availability").text.strip()

    url = BASE + link["href"]

    return {
        "title": title,
        "price": price,
        "rating": rating,
        "in_stock": in_stock,
        "url": url,
    }
```

Now `price` is `51.77`, a float, not a string. Decide on the *type* you want for
each field at extraction time - a scraper that emits clean, typed records is
worth ten that emit strings someone has to scrub later.

## Survive a missing field

Here's the thing that separates a script that works once from a scraper you can
trust: real pages are inconsistent. One book is missing a rating. Another has no
price because it's out of stock. The moment you call `.text` on a
`select_one` that found nothing, you get `AttributeError: 'NoneType' object has
no attribute 'text'` and the whole run dies on item 47 of 1000.

The fix is to check before you reach in. A small helper keeps the main function
readable.

## Your turn: get_text

You already have everything this needs: `select_one` from last phase, and the
fact that it returns `None` when nothing matches. Write a helper that looks up
`selector` inside `element` and returns its stripped text - or `default` if the
selector found nothing.

```python
def get_text(element, selector, default=""):
    # your turn
    return default
```

You need `select_one`, `.text`, `.strip()`, and a plain `if` - nothing new.
Check it against `parse_book` below once you're done: swap it in and a missing
field should come back as `""` instead of crashing the whole run.

```python
def get_text(element, selector, default=""):
    found = element.select_one(selector)
    return found.text.strip() if found else default
```

`select_one` returns `None` when nothing matches, and `None` is falsy, so the
`if found` guard catches it. When the element is missing you get your default
instead of a crash. Wire it in:

```python
def parse_book(book):
    link = book.select_one("h3 a")
    title = link["title"].strip() if link else "Unknown"

    raw_price = get_text(book, "p.price_color")
    price = float(raw_price.replace("£", "")) if raw_price else None

    rating_el = book.select_one("p.star-rating")
    rating = rating_el["class"][1] if rating_el else None

    in_stock = get_text(book, "p.instock.availability")

    url = BASE + link["href"] if link else None

    return {
        "title": title,
        "price": price,
        "rating": rating,
        "in_stock": in_stock,
        "url": url,
    }
```

Now a missing price becomes `None`, not a stack trace. `None` is a deliberate
"we looked and it wasn't there" - far more useful than an empty string, because
later you can ask "which records are missing a price?" and get a real answer.

## Pull the whole page into records

Put it together: loop every book, build a list of dicts, and report.

```python
import requests
from bs4 import BeautifulSoup

BASE = "https://books.toscrape.com/"


def get_text(element, selector, default=""):
    found = element.select_one(selector)
    return found.text.strip() if found else default


def parse_book(book):
    link = book.select_one("h3 a")
    raw_price = get_text(book, "p.price_color")
    rating_el = book.select_one("p.star-rating")
    return {
        "title": link["title"].strip() if link else "Unknown",
        "price": float(raw_price.replace("£", "")) if raw_price else None,
        "rating": rating_el["class"][1] if rating_el else None,
        "in_stock": get_text(book, "p.instock.availability"),
        "url": BASE + link["href"] if link else None,
    }


response = requests.get(BASE, timeout=10)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")

records = [parse_book(b) for b in soup.select("article.product_pod")]

print(f"Extracted {len(records)} records")
for r in records[:3]:
    print(r)

cheapest = min(records, key=lambda r: r["price"])
print("Cheapest:", cheapest["title"], "at £", cheapest["price"])
```

Run it. Twenty clean dictionaries, and because the price is a real number, you
can find the cheapest book with one line. That's the payoff of typing your data
as you extract it.

## Where we are

You have a list of clean, structured, typed records, and code that won't fall
over when a page leaves a field blank. One problem remains: this is only the
first 20 books. There are a thousand. Next phase we follow the "next" link
through every page - and we do it without being a nuisance to the server.


---

# Pagination and Being Polite

We've been scraping one page. The catalog has fifty. This phase teaches the
scraper to walk from page to page until there are no more - and, equally important, to do it in a way that doesn't hammer the server or get you blocked.
By the end you'll have a loop that collects the whole catalog at a respectful
pace.

I'm putting the manners and the law in the same phase as the loop on purpose.
The mechanics of pagination take ten minutes. Knowing how *not* to be a problem
is the part that keeps you out of trouble, and it's the part a lot of tutorials
skip. We won't.

## Find the "next" link

Open the site and scroll to the bottom - there's a "next" button. Inspect it.
On `books.toscrape.com` it's a `<li class="next">` containing an `<a>` whose
`href` points at the following page. When you're on the last page, that element
isn't there at all. That presence-or-absence is exactly the signal we need: keep
going while there's a next link, stop when there isn't.

```python
next_link = soup.select_one("li.next a")
if next_link:
    next_href = next_link["href"]   # e.g. "catalogue/page-2.html"
else:
    next_href = None                # we're on the last page
```

One wrinkle: that `href` is *relative*. From the catalogue pages it might be
`page-2.html`, meaning "relative to where I am now." Python's standard library
has the right tool so you don't guess: `urljoin` combines the current page's URL
with a relative link and gives you a correct absolute URL.

```python
from urllib.parse import urljoin

current = "https://books.toscrape.com/catalogue/page-1.html"
print(urljoin(current, "page-2.html"))
# -> https://books.toscrape.com/catalogue/page-2.html
```

## Set a User-Agent first

Before we loop, one courtesy and one practicality. Every request carries a
`User-Agent` header that says who's calling. By default `requests` sends
something like `python-requests/2.x`, which is genuine but anonymous. Many servers
treat the default Python agent with suspicion or block it outright.

Set a real one that identifies you and, ideally, how to reach you. This is both
polite (the site owner can see who you are) and effective (you're less likely to
be filtered). A shared session attaches the header to every request:

```python
import requests

session = requests.Session()
session.headers.update({
    "User-Agent": "weekend-book-scraper/1.0 (you@example.com)"
})
```

Don't impersonate a real browser to sneak past defenses. If a site plainly
doesn't want bots, that's information, not a challenge - more on that below.

## The pagination loop

Now the full walk. Start at page one, parse it, find the next link, sleep a
moment, repeat - until there's no next link.

```python
import time
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

START = "https://books.toscrape.com/catalogue/page-1.html"
DELAY = 1.0   # seconds between requests

session = requests.Session()
session.headers.update({
    "User-Agent": "weekend-book-scraper/1.0 (you@example.com)"
})


def parse_book(book):
    link = book.select_one("h3 a")
    raw_price = book.select_one("p.price_color")
    rating_el = book.select_one("p.star-rating")
    return {
        "title": link["title"].strip() if link else "Unknown",
        "price": float(raw_price.text.replace("£", "")) if raw_price else None,
        "rating": rating_el["class"][1] if rating_el else None,
    }


def scrape_all(start_url):
    records = []
    url = start_url
    page = 1
    while url:
        print(f"Fetching page {page}: {url}")
        response = session.get(url, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")

        for book in soup.select("article.product_pod"):
            records.append(parse_book(book))

        next_link = soup.select_one("li.next a")
        url = urljoin(url, next_link["href"]) if next_link else None
        page += 1
        time.sleep(DELAY)      # be a good guest

    return records


all_books = scrape_all(START)
print(f"Collected {len(all_books)} books across all pages")
```

Run it. It walks every page, printing as it goes, pausing a second between
fetches, and ends with all thousand books in one list. The `while url:` loop is
the engine: `url` becomes `None` the moment there's no next link, and the loop
ends on its own.

## Why the delay matters

That `time.sleep(DELAY)` isn't decoration. Without it, a fast loop fires
requests as quickly as Python can manage - dozens per second - and to the server
that's indistinguishable from an attack. You can knock a small site over, and
you will absolutely get your IP blocked. A one-second pause makes you a normal
visitor. Slow is polite, and polite is what keeps you scraping.

A second knob worth knowing for bigger jobs: retries with backoff. If a request
fails, wait a bit and try again - and wait *longer* each time rather than
pounding a struggling server. For this project a fixed delay is enough; keep
backoff in your back pocket for production.

## robots.txt - read it, respect it

Most sites publish a file at `/robots.txt` saying which paths automated clients
may and may not touch. It's a request, not a wall, but ignoring it is rude and,
depending on where you are and what you're doing, can be a factor against you
legally. Python's standard library reads it for you:

```python
from urllib.robotparser import RobotFileParser

rp = RobotFileParser()
rp.set_url("https://books.toscrape.com/robots.txt")
rp.read()

agent = "weekend-book-scraper/1.0"
print(rp.can_fetch(agent, "https://books.toscrape.com/catalogue/page-1.html"))
```

`can_fetch` returns `True` or `False`. The straightforward move is to check it before you
scrape a path and skip what's disallowed. (Our practice site allows everything;
real sites often disallow `/search`, `/cart`, login areas, and the like.)

## The ethics and the law, plainly

I'm not your lawyer, and this isn't legal advice - but here's the working
understanding a careful person operates with:

- **Public, factual data is the safe ground.** Facts (prices, titles, public
  listings) generally aren't copyrightable. Re-publishing someone's *creative*
  content wholesale is a different matter.
- **Terms of Service exist.** Many sites' ToS forbid automated access. Violating
  them can be a breach of contract even when the data itself is public. Read them
  for anything you'll do more than once.
- **Don't scrape personal data carelessly.** Names, emails, profiles - privacy
  law (GDPR, CCPA, and friends) applies regardless of whether data is "public."
- **Don't degrade the service.** Hammering a server can cross from rude into
  unlawful (computer-misuse statutes). The delay isn't only manners.
- **Logins and paywalls change everything.** Scraping behind authentication, or
  bypassing access controls, is a sharp escalation. Don't, unless you have clear
  permission.
- **Prefer an API.** If the site offers one, use it. It's faster, more stable,
  and it's them *inviting* you in.

The short version: scrape public facts, slowly, with a genuine User-Agent, while
respecting robots.txt and ToS, and never for personal data or behind a login.
That posture covers the vast majority of legitimate scraping.

```mermaid
graph TD
  A[Want some data] --> B{Is there an API?}
  B -->|yes| C[Use the API]
  B -->|no| D{robots.txt + ToS allow it?}
  D -->|no| E[Stop or ask permission]
  D -->|yes| F[Scrape slowly, genuine UA]
```

## Where we are

The scraper now collects an entire catalog, page by page, at a pace that won't
get you blocked or sued - with a real User-Agent and a robots.txt check in the
toolkit. The data's all in memory, though, and memory vanishes when the program
exits. Last phase: write it to disk, then look at where you'd take this next.


---

# Saving the Data, and Where to Take It

We've got a thousand clean records sitting in memory. The moment the program
ends, they're gone. This phase fixes that - we write them to CSV and JSON - and
then turns the finished scraper into one complete script. After that, the fun
part: a tour of where you'd take this when a weekend project meets a real need.

Both file formats ship with Python. No installs. `csv` for the spreadsheet
people, `json` for the program-talking-to-program people. We'll write both,
because they answer different questions and cost nothing extra.

## Write to CSV

CSV opens in Excel, Numbers, Google Sheets, and every data tool on earth. Our
records are a list of dicts with identical keys, which is exactly what
`csv.DictWriter` is built for.

```python
import csv


def save_csv(records, filename="books.csv"):
    if not records:
        print("Nothing to save")
        return
    fieldnames = records[0].keys()
    with open(filename, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=fieldnames)
        writer.writeheader()
        writer.writerows(records)
    print(f"Wrote {len(records)} rows to {filename}")
```

Two details that save you grief. `newline=""` stops Python from adding blank
lines between every row on Windows - leave it off and your CSV looks
double-spaced in Excel. And `encoding="utf-8"` makes sure titles with accents or
symbols survive instead of turning into garbage. Always pass both when writing
CSV.

## Write to JSON

JSON keeps your data's *shape* - nested structures, real numbers, `None` as
`null`. It's the format you'd hand to another program or a web front-end.

```python
import json


def save_json(records, filename="books.json"):
    with open(filename, "w", encoding="utf-8") as f:
        json.dump(records, f, indent=2, ensure_ascii=False)
    print(f"Wrote {len(records)} records to {filename}")
```

`indent=2` makes the file human-readable instead of one giant line.
`ensure_ascii=False` lets real characters (£, é, -) appear as themselves rather
than `\u` escapes. Drop both and JSON still works, but you'll thank yourself for
the readable version when you open it to debug.

## The whole thing, in one file

Here's the complete scraper - fetch, parse, extract defensively, paginate
politely, save both formats. This is the program the project was building toward.
Save it as `scraper.py`.

```python
import csv
import json
import time
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

START = "https://books.toscrape.com/catalogue/page-1.html"
DELAY = 1.0

session = requests.Session()
session.headers.update({
    "User-Agent": "weekend-book-scraper/1.0 (you@example.com)"
})


def parse_book(book):
    link = book.select_one("h3 a")
    raw_price = book.select_one("p.price_color")
    rating_el = book.select_one("p.star-rating")
    stock_el = book.select_one("p.instock.availability")
    return {
        "title": link["title"].strip() if link else "Unknown",
        "price": float(raw_price.text.replace("£", "")) if raw_price else None,
        "rating": rating_el["class"][1] if rating_el else None,
        "in_stock": stock_el.text.strip() if stock_el else "",
        "url": urljoin(START, link["href"]) if link else None,
    }


def scrape_all(start_url):
    records = []
    url = start_url
    page = 1
    while url:
        print(f"Fetching page {page}: {url}")
        response = session.get(url, timeout=10)
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")
        for book in soup.select("article.product_pod"):
            records.append(parse_book(book))
        next_link = soup.select_one("li.next a")
        url = urljoin(url, next_link["href"]) if next_link else None
        page += 1
        time.sleep(DELAY)
    return records


def save_csv(records, filename="books.csv"):
    if not records:
        return
    with open(filename, "w", newline="", encoding="utf-8") as f:
        writer = csv.DictWriter(f, fieldnames=records[0].keys())
        writer.writeheader()
        writer.writerows(records)


def save_json(records, filename="books.json"):
    with open(filename, "w", encoding="utf-8") as f:
        json.dump(records, f, indent=2, ensure_ascii=False)


if __name__ == "__main__":
    books = scrape_all(START)
    save_csv(books)
    save_json(books)
    print(f"Done. Saved {len(books)} books to books.csv and books.json")
```

Run it:

```bash
python scraper.py
```

Watch it walk the pages, then open `books.csv` in a spreadsheet. There's your
weekend's work: a thousand books with titles, prices, ratings, stock, and links -
sortable, filterable, yours. That's a finished, working scraper.

## Where to take it next

You've got the core skill. Here's the map of what's past the edge of this
project, roughly in order of effort.

| Upgrade | What it buys you | First tool to look at |
|---------|------------------|------------------------|
| A database | Query, dedupe, update over time | `sqlite3` (built in) |
| Scheduling | Runs itself on a timer | cron, Task Scheduler |
| Concurrency | Many pages at once, faster | `httpx` + `asyncio` |
| Headless browser | Scrape JS-built pages | Playwright |
| A framework | Big crawls, built-in plumbing | Scrapy |

A few of those deserve a sentence.

**A database.** When you scrape the same site repeatedly, a CSV per run gets
messy fast. SQLite - which ships with Python as `sqlite3` - lets you store
records in a real table, ask questions with SQL, and update yesterday's data
instead of duplicating it. It's the natural next step when "save a file" stops
being enough.

**Scheduling.** A scraper that runs itself is worth ten you have to remember to
run. On macOS or Linux, cron fires your script on a schedule; on Windows, Task
Scheduler does the same. Point it at `python scraper.py` nightly and wake up to
fresh data.

**Headless browsers, for the sites that fight back.** Here's the wall you'll hit
eventually: some pages arrive nearly empty and build their content with
JavaScript *after* loading. `requests` only sees that empty shell - it doesn't
run JavaScript. When `response.text` is missing data you can plainly see in your
browser, that's the symptom. The cure is a headless browser like Playwright,
which drives a real (invisible) browser, lets the JavaScript run, and *then*
hands you the finished HTML to feed into the very same BeautifulSoup code you
wrote this weekend. Everything you learned still applies - you've upgraded the
fetch step, nothing else.

**Scrapy.** When a one-file script grows into a serious crawler - many sites,
retries, pipelines, politeness baked in - Scrapy is the framework built for it.
It's more to learn, so reach for it when you've outgrown a script, not before.

## Where we are

You built a real web scraper this weekend. It fetches pages, parses messy HTML,
extracts clean and typed records, walks an entire catalog at a respectful pace,
and saves the results to formats you can actually use. Every piece is code you
understand, because you wrote it one phase at a time.

The same five-box loop - fetch, parse, extract, next, save - scales from this
practice site to almost anything you'll want to point it at. Swap the selectors
for a new site's HTML, keep the politeness, and you're scraping. Go find some
public data worth having, and treat the servers kindly while you get it.
