# Build a CSV Summary Report (Python)

> Turn raw CSV data into a summary report with Python’s csv module - parse rows, aggregate, group, and format - all runnable in your browser from an embedded sample.


---

# Build a CSV Summary Report (Python)

You have a CSV file. Maybe it's sales, maybe it's expenses, maybe it's a log of who did what. Opening it in a spreadsheet works once, but the moment someone asks "what's the total per region?" every week, you want a script that answers it in a second.

That's what we're building this weekend: a small Python program that reads a CSV, adds up the numbers, breaks them out by category, and prints a clean report you'd be happy to paste into an email. No libraries to install - Python ships with everything we need in its `csv` module.

## What you'll build

A self-contained report generator. Feed it rows like this:

```
date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
```

and it prints something like this:

```
SALES SUMMARY
=============
Rows:          12
Total amount:  1,884.50
Average:       157.04
Largest sale:  410.00

By region:
  North        642.50
  South        531.00
  East         711.00
```

We'll get there in four steps, each one a working piece you can run on its own.

## Runs in your browser

Every code block in this project is **run-along** - it executes right here in the page, no setup, no install. Each block is self-contained: it carries its own sample data and prints its own result, so you can run them in any order and tweak the numbers to see what changes. The last phase also shows you how to point the same code at a real file on your own machine.

## The stack

| Piece | What it does |
|-------|--------------|
| `csv.DictReader` | Reads each row into a dict keyed by the header |
| `io.StringIO` | Lets us treat a string as a file (so the sample lives in the code) |
| `collections.defaultdict` | Groups rows by category without messy setup |
| `str.format` / f-strings | Lines up the numbers into a tidy report |

All four are in the standard library. Nothing else.

## The shape of the build

```mermaid
graph LR
  A[CSV text] --> B[Parse rows]
  B --> C[Totals & counts]
  C --> D[Group by category]
  D --> E[Formatted report]
```

## Roughly how long

A focused afternoon - maybe two to three hours if you stop to experiment, which you should. Each phase stands alone, so you can do one, walk away, and come back.

## What you'll learn

- Reading CSV safely with `DictReader` instead of splitting strings by hand
- Converting text columns to numbers (and why the CSV gives you strings)
- Aggregating: sum, count, min, max, average
- Grouping rows with `defaultdict` - the pattern you'll reuse forever
- Formatting numbers into aligned, readable columns
- How to swap the embedded sample for a real file when you run it locally

By the end you'll have one script that does the whole job, and the muscle memory to write the next one without looking it up. Let's start by getting the rows out of the CSV.


---

# Parsing the CSV

The first job is getting the rows out of the file and into something Python can work with. We'll turn the raw CSV text into a list of dictionaries - one dict per row, keyed by the column header.

## Why not split on commas yourself?

It's tempting to do `line.split(",")` and call it a day. Don't. Real CSV data has commas inside quoted fields (`"Smith, John"`), quotes inside values, and blank cells. The `csv` module handles all of that correctly, and it's already in the standard library. Splitting strings by hand is the kind of shortcut that works on your test data and breaks the day someone has a comma in a product name.

So we reach for `csv`.

## The file that isn't a file

Normally you'd open a CSV with `open("sales.csv")`. But this code runs in your browser, where there's no file to open. The trick: `io.StringIO` wraps a string so it behaves exactly like an open file. The `csv` module can't tell the difference.

That means we can keep our sample data right inside the code, and the same `DictReader` call you'd use on a real file works here unchanged. (In the last phase, you'll swap `StringIO` for `open()` and nothing else changes.)

## DictReader

`csv.DictReader` reads the first line as the header, then hands you each following row as a dict. So this CSV:

```
region,amount
North,120.50
```

becomes `{"region": "North", "amount": "120.50"}`.

Notice the amount is a **string**, not a number. CSV is only text - every value comes back as a string, even the ones that look numeric. We'll deal with converting those in the next phase. For now, our goal is to get the rows parsed and look at them.

Here's the full thing. Run it.

```python runnable
import csv
import io

# Our sample data lives right here in the code.
CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

# StringIO makes the string behave like an open file.
f = io.StringIO(CSV_TEXT)
reader = csv.DictReader(f)

# DictReader gives one dict per row; collect them into a list.
rows = list(reader)

print(f"Parsed {len(rows)} rows.\n")
for row in rows:
    print(row)
```

You should see six dictionaries, each with `date`, `region`, `product`, and `amount` keys. The headers became the keys automatically - that's `DictReader` doing its job.

## Reading the output

Look closely at one row:

```
{'date': '2026-01-03', 'region': 'North', 'product': 'Widget', 'amount': '120.50'}
```

Every value is in quotes - they're all strings. `'120.50'` is text, not the number `120.50`. If you tried to add up the amounts right now, you'd get `'120.50' + '89.00'` which concatenates into `'120.5089.00'`. Not what we want. Holding onto that fact is the whole reason the next phase exists.

## Pull out one column

Since each row is a dict, grabbing a single column across all rows is one line. Let's list every region we saw.

Before you run this, guess how many *unique* regions show up in the six rows. Then check.

```python runnable
import csv
import io

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

regions = [row["region"] for row in rows]
print("All regions:", regions)
print("Unique regions:", sorted(set(regions)))
```

That `set()` trick collapses duplicates, and `sorted()` makes the order stable. You now have a clean list of the categories you'll group by later.

## What you've got

A reliable way to turn CSV text into a list of dicts you can loop over, index by column name, and pull values from. That's the foundation. The data's a little raw - everything's still a string - but it's structured, and that's the hard part done.

In the next phase we'll convert the `amount` column to real numbers and start adding things up. Keep this parsing pattern in your head: `StringIO` (or `open`) → `DictReader` → `list`. You'll write it at the top of every one of these scripts.


---

# Totals and Counts

We have rows. Now we want numbers about those rows: how much in total, how many sales, the smallest, the largest, the average. This is the part people actually ask for.

## First, fix the strings

Remember from the last phase: every value the CSV gives us is a string. `'120.50'` is text. Before we can do math, we convert the `amount` column to a float.

`float("120.50")` gives `120.5`. We do that for each row. The cleanest move is a list comprehension that pulls out only the numbers we care about:

```python runnable
import csv
import io

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

amounts = [float(row["amount"]) for row in rows]
print(amounts)
```

Now they're real numbers - no quotes - and we can add, compare, and average them.

## The five numbers

We want five things about the whole dataset: how many amounts there are, their total, the smallest, the largest, and the average. Python's built-ins can get you there without much code.

**Your turn.** This one's the point of the phase, so give it a shot before you read on. Fill in `summarize` and hit Run - the checks underneath tell you whether it works. My version is in the next block whenever you want it.

```python runnable
def summarize(amounts):
    # Given a list of numbers, return a dict with:
    #   "count"    - how many numbers
    #   "total"    - their sum
    #   "smallest" - the minimum
    #   "largest"  - the maximum
    #   "average"  - total / count
    pass


# --- checks: fix your function until this prints "All good." ---
stats = summarize([120.50, 89.00, 210.00, 55.50, 134.00, 410.00])
assert isinstance(stats, dict), f"summarize should return a dict, got: {stats!r}"
assert stats["count"] == 6, f"count should be 6, got: {stats.get('count')}"
assert stats["total"] == 1019.0, f"total should be 1019.0, got: {stats.get('total')}"
assert stats["smallest"] == 55.5, f"smallest should be 55.5, got: {stats.get('smallest')}"
assert stats["largest"] == 410.0, f"largest should be 410.0, got: {stats.get('largest')}"
assert abs(stats["average"] - 169.8333333333333) < 0.0001, f"average should be about 169.83, got: {stats.get('average')}"
print("All good.")
```

Stuck on one of the five? Each is a single built-in call away - the question is which built-in fits.

### One way to write it

There's no special "average" built-in - it's the sum divided by the count, which you already have. Python's built-ins do the rest:

| Question | Code |
|----------|------|
| How many? | `len(amounts)` |
| Total? | `sum(amounts)` |
| Smallest? | `min(amounts)` |
| Largest? | `max(amounts)` |
| Average? | `sum(amounts) / len(amounts)` |

```python runnable
import csv
import io

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))
amounts = [float(row["amount"]) for row in rows]


def summarize(amounts):
    return {
        "count": len(amounts),
        "total": sum(amounts),
        "smallest": min(amounts),
        "largest": max(amounts),
        "average": sum(amounts) / len(amounts),
    }


stats = summarize(amounts)
print(f"Count:    {stats['count']}")
print(f"Total:    {stats['total']:.2f}")
print(f"Smallest: {stats['smallest']:.2f}")
print(f"Largest:  {stats['largest']:.2f}")
print(f"Average:  {stats['average']:.2f}")
```

The `:.2f` in the f-string rounds to two decimal places - so `157.041666...` prints as `157.04`. Money never wants fifteen digits after the dot.

## The one bug waiting to happen

Look at `average = total / count`. If `count` is zero - an empty CSV, or one with only a header - that line blows up with `ZeroDivisionError`. It's worth one guard.

Before you run this, guess what the average line prints when there are no data rows at all. Then check.

```python runnable
import csv
import io

# An empty file: header only, no data rows.
CSV_TEXT = "date,region,product,amount"

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))
amounts = [float(row["amount"]) for row in rows]

count = len(amounts)
total = sum(amounts)
average = total / count if count else 0.0

print(f"Count:   {count}")
print(f"Total:   {total:.2f}")
print(f"Average: {average:.2f}")
```

`total / count if count else 0.0` reads as "divide if there's anything, otherwise zero." `sum([])` is already `0` and `len([])` is `0`, so those two are fine on their own - it's the division that needs the guard. One small condition saves you a crash on the inevitable empty file.

## Which row was the biggest?

`max(amounts)` tells you the biggest number, but often you want the whole row - what product, what region. For that, give `max` a `key` so it compares rows by their amount and hands back the row itself.

Before you run it, guess which row prints as the biggest sale - the region and product, not just the number.

```python runnable
import csv
import io

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

biggest = max(rows, key=lambda r: float(r["amount"]))
print("Biggest sale:")
print(f"  {biggest['date']}  {biggest['region']}  {biggest['product']}  {biggest['amount']}")
```

`key=lambda r: float(r["amount"])` tells `max` how to rank the rows. Without the `float`, it would compare the amounts as strings - and `'89.00'` sorts higher than `'410.00'` alphabetically, which would be wrong. The conversion matters everywhere you compare.

## What you've got

The headline numbers for the whole dataset: count, total, min, max, average - plus the single biggest row. That's already a useful summary. But a real report breaks the totals out by category - total per region, per product - so you can see where the money's actually coming from.

That's grouping, and it's the next phase. The pattern there builds straight on what you have: convert to a number, then accumulate. We'll do it per group instead of all at once.


---

# Grouping with a Dict

A single grand total is fine, but the question people really have is "where's it coming from?" Total per region. Sales per product. That's grouping: split the rows into buckets by a category column, then total each bucket.

## The idea

We walk through every row, look at its `region`, and add its `amount` to a running total for that region. We need somewhere to keep those running totals - one slot per region - and that's a dict:

```
{"North": 330.50, "South": 223.00, "East": 465.50}
```

The catch is the first time we see a region, there's no slot yet. We have to create it before we can add to it.

## The clumsy way first

So you can see what `defaultdict` saves you, here's grouping with a plain dict. Every loop has to check "have I seen this region before?":

```python runnable
import csv
import io

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

totals = {}
for row in rows:
    region = row["region"]
    amount = float(row["amount"])
    if region not in totals:      # first time? make the slot.
        totals[region] = 0.0
    totals[region] += amount

for region, total in totals.items():
    print(f"{region}: {total:.2f}")
```

It works, but that `if region not in totals` line is noise we repeat in every grouping script. There's a cleaner tool.

## defaultdict does the check for you

`collections.defaultdict(float)` is a dict that, when you ask for a key it's never seen, quietly creates it with a default value first. Pass `float` and missing keys start at `0.0`. Pass `int` and they start at `0`. Pass `list` and they start at `[]`.

**Your turn.** This is the technique the rest of the guide leans on, so try it yourself first. Write `total_by_group` using `defaultdict` instead of the `if region not in totals` check from the clumsy version above. The checks underneath tell you if it works. My version is right after.

```python runnable
from collections import defaultdict


def total_by_group(rows, group_col, value_col):
    # Return a dict mapping each distinct value in `group_col`
    # to the sum of `value_col` for rows with that value.
    # Use defaultdict(float) so you don't need to check
    # "have I seen this key before?"
    pass


# --- checks: fix your function until this prints "All good." ---
sample = [
    {"region": "North", "amount": "10"},
    {"region": "South", "amount": "5"},
    {"region": "North", "amount": "3"},
]
totals = total_by_group(sample, "region", "amount")
assert isinstance(totals, dict), f"total_by_group should return a dict, got: {totals!r}"
assert totals["North"] == 13.0, f"North should total 13.0, got: {totals.get('North')}"
assert totals["South"] == 5.0, f"South should total 5.0, got: {totals.get('South')}"
assert len(totals) == 2, f"expected 2 groups, got: {len(totals)}"
print("All good.")
```

Stuck? Look at the clumsy version above - `total_by_group` is that same loop, minus the `if`.

### One way to write it

`totals[region] += amount` works even the first time with a defaultdict - the slot gets created as `0.0` on the spot, then the amount is added. The `if` disappears:

```python runnable
import csv
import io
from collections import defaultdict

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))


def total_by_group(rows, group_col, value_col):
    totals = defaultdict(float)
    for row in rows:
        totals[row[group_col]] += float(row[value_col])
    return totals


totals = total_by_group(rows, "region", "amount")
for region in sorted(totals):
    print(f"{region}: {totals[region]:.2f}")
```

Same result, less ceremony. `sorted(totals)` prints the regions alphabetically so the output is stable instead of in whatever order the rows happened to arrive.

## Count and total per group

Usually you want more than the total - you also want how many sales made up that total. Keep two defaultdicts side by side, or keep one dict of small lists. Two is the readable choice.

Before you run this, guess which region has the highest *average* sale - not the highest total.

```python runnable
import csv
import io
from collections import defaultdict

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

totals = defaultdict(float)
counts = defaultdict(int)
for row in rows:
    region = row["region"]
    totals[region] += float(row["amount"])
    counts[region] += 1

for region in sorted(totals):
    avg = totals[region] / counts[region]
    print(f"{region:6}  {counts[region]} sales  total {totals[region]:7.2f}  avg {avg:6.2f}")
```

Now each line shows the region, how many sales, the total, and the per-region average. Notice the format codes are doing the alignment: `{region:6}` pads the name to six characters, `{totals[region]:7.2f}` reserves seven characters for the number. That's a preview of the report formatting coming in the last phase.

## Group by a different column

The grouping key is only a column name. Swap `region` for `product` and the same loop totals sales per product instead - no other change.

Before you run it, guess whether Widget or Gadget has the higher total.

```python runnable
import csv
import io
from collections import defaultdict

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

GROUP_BY = "product"   # change to "region" and rerun

totals = defaultdict(float)
for row in rows:
    totals[row[GROUP_BY]] += float(row["amount"])

print(f"Totals by {GROUP_BY}:")
for key in sorted(totals):
    print(f"  {key:8} {totals[key]:.2f}")
```

Pulling the column into a `GROUP_BY` variable means you can repoint the whole script at a different breakdown by editing one line. That's the kind of small lever that makes a script reusable.

## What you've got

The per-group breakdown - total and count for each region (or product, or anything). Combined with the grand totals from the last phase, you now have every number the report needs. The only thing left is presentation: lining it all up so it reads like a report instead of debug output.

That's the final phase, where the pieces become one script - and where you'll learn to point it at a real file on your machine.


---

# Formatting the Report

Time to put it all together. We have the grand totals from phase 2 and the per-group breakdown from phase 3. This phase turns those numbers into a single report that lines up cleanly, then wires the whole thing into one function - and shows you how to run it on a real CSV on your own machine.

## Lining numbers up

Raw `print()` output looks ragged because names and numbers are different widths. Format specs fix that. Two you'll use constantly:

| Spec | Effect | Example |
|------|--------|---------|
| `{name:<10}` | left-align in 10 chars | `North     ` |
| `{value:>10.2f}` | right-align in 10 chars, 2 decimals | `    642.50` |

Left-align text, right-align numbers - that's the whole secret to a column that reads well. Numbers line up on their right edge so the decimal points stack.

A thousands separator helps too: `{value:,.2f}` prints `1884.5` as `1,884.50`. Easier to read large totals at a glance.

## The whole script

Everything from the last three phases - parse the rows, get the five numbers, group by category, line it up with the format specs above - becomes one function that returns the finished report as a string.

**Your turn.** This is the capstone: the function the rest of the project ships. Fill in `build_report` and hit Run. The checks underneath tell you if it works. My version, in full, is right after.

```python runnable
def build_report(rows, group_col="region", value_col="amount"):
    # Return the finished report as one string ("\n".join(lines)).
    # It must contain, in order:
    #   - a "SALES SUMMARY" header
    #   - "Rows:" followed by the row count
    #   - "Total amount:" followed by the sum, 2 decimals
    #   - "Average:" followed by the average, 2 decimals
    #   - "Largest sale:" followed by the biggest single value, 2 decimals
    #   - "By <group_col>:" followed by one line per group (sorted by
    #     key), each showing the group's name and its total
    pass


# --- checks: fix your function until this prints "All good." ---
sample_rows = [
    {"region": "North", "amount": "120.50"},
    {"region": "South", "amount": "89.00"},
    {"region": "North", "amount": "210.00"},
]
report = build_report(sample_rows)
assert isinstance(report, str), f"build_report should return a string, got: {type(report)}"
assert "SALES SUMMARY" in report, f"missing the header, got:\n{report}"
assert "Rows:" in report and "3" in report, f"row count (3) should show up, got:\n{report}"
assert "419.50" in report, f"total amount (419.50) should show up, got:\n{report}"
assert "139.83" in report, f"average (139.83) should show up, got:\n{report}"
assert "210.00" in report, f"largest sale (210.00) should show up, got:\n{report}"
lines = report.splitlines()
north_line = next((l for l in lines if "North" in l), "")
south_line = next((l for l in lines if "South" in l), "")
assert "330.50" in north_line, f"North's total (330.50) should be on its own line, got: {north_line!r}"
assert "89.00" in south_line, f"South's total (89.00) should be on its own line, got: {south_line!r}"
assert lines.index(north_line) < lines.index(south_line), "groups should be sorted alphabetically (North before South)"
print("All good.")
```

Nothing here is new work - every piece was built in a previous phase. Stuck? Build it in stages: get the five numbers printing first, then add the group loop, then worry about the exact wording of each line.

### One way to write it

```python runnable
import csv
import io
from collections import defaultdict

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50
2026-01-14,South,Widget,134.00
2026-01-19,East,Gadget,410.00
2026-01-22,North,Widget,98.00
2026-01-25,South,Gadget,175.50"""


def build_report(rows, group_col="region", value_col="amount"):
    amounts = [float(r[value_col]) for r in rows]
    count = len(amounts)
    total = sum(amounts)
    average = total / count if count else 0.0
    largest = max(amounts) if amounts else 0.0

    group_totals = defaultdict(float)
    for r in rows:
        group_totals[r[group_col]] += float(r[value_col])

    lines = []
    lines.append("SALES SUMMARY")
    lines.append("=" * 13)
    lines.append(f"Rows:          {count}")
    lines.append(f"Total amount:  {total:>12,.2f}")
    lines.append(f"Average:       {average:>12,.2f}")
    lines.append(f"Largest sale:  {largest:>12,.2f}")
    lines.append("")
    lines.append(f"By {group_col}:")
    for key in sorted(group_totals):
        lines.append(f"  {key:<10} {group_totals[key]:>12,.2f}")

    return "\n".join(lines)


rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))
report = build_report(rows)
print(report)
```

That's the project. One function takes the rows and gives you the formatted report as a string. Building the report into a list of lines and joining at the end (instead of printing as you go) means you can return it, write it to a file, or print it - your choice, decided by the caller.

Try changing `build_report(rows)` to `build_report(rows, group_col="product")` and rerun. Same report, broken out by product instead. That flexibility came free from passing the column name as an argument.

## A quick sanity check

When you write aggregation code, check it against a number you can verify by hand. The per-group totals must add up to the grand total - if they don't, something's double-counting or dropping rows. Let's assert exactly that:

```python runnable
import csv
import io
from collections import defaultdict

CSV_TEXT = """date,region,product,amount
2026-01-03,North,Widget,120.50
2026-01-05,South,Gadget,89.00
2026-01-08,North,Gadget,210.00
2026-01-11,East,Widget,55.50"""

rows = list(csv.DictReader(io.StringIO(CSV_TEXT)))

grand = sum(float(r["amount"]) for r in rows)

group_totals = defaultdict(float)
for r in rows:
    group_totals[r["region"]] += float(r["amount"])
group_sum = sum(group_totals.values())

# Floats don't compare exactly; allow a tiny tolerance.
assert abs(grand - group_sum) < 0.001, "groups don't sum to the total!"
print(f"Grand total:     {grand:.2f}")
print(f"Sum of groups:   {group_sum:.2f}")
print("Check passed: the groups add up.")
```

The `abs(...) < 0.001` is there because floats don't always compare exactly - `0.1 + 0.2` famously isn't `0.3` to a computer. Comparing with a small tolerance instead of `==` is the right habit for any money math.

## Running it on a real file

Everything so far ran in the browser using `StringIO`. On your own machine you read an actual file instead - and that's the only line that changes. The rest of `build_report` is identical.

You don't need to install anything; `csv` ships with Python. Save this as `report.py` next to your CSV, then run `python report.py sales.csv`:

```python
import csv
import sys
from collections import defaultdict


def build_report(rows, group_col="region", value_col="amount"):
    amounts = [float(r[value_col]) for r in rows]
    count = len(amounts)
    total = sum(amounts)
    average = total / count if count else 0.0
    largest = max(amounts) if amounts else 0.0

    group_totals = defaultdict(float)
    for r in rows:
        group_totals[r[group_col]] += float(r[value_col])

    lines = ["SALES SUMMARY", "=" * 13]
    lines.append(f"Rows:          {count}")
    lines.append(f"Total amount:  {total:>12,.2f}")
    lines.append(f"Average:       {average:>12,.2f}")
    lines.append(f"Largest sale:  {largest:>12,.2f}")
    lines.append("")
    lines.append(f"By {group_col}:")
    for key in sorted(group_totals):
        lines.append(f"  {key:<10} {group_totals[key]:>12,.2f}")
    return "\n".join(lines)


def main():
    path = sys.argv[1] if len(sys.argv) > 1 else "sales.csv"
    with open(path, newline="", encoding="utf-8") as f:
        rows = list(csv.DictReader(f))

    report = build_report(rows)
    print(report)

    # Also save it next to the input.
    with open("report.txt", "w", encoding="utf-8") as out:
        out.write(report)
    print("\nSaved to report.txt")


if __name__ == "__main__":
    main()
```

Run it from a terminal:

```bash
python report.py sales.csv
```

Three things to notice about the file version:

- `open(path, newline="", encoding="utf-8")` - the `newline=""` is the one `csv` quirk worth remembering. It stops Python from mangling line endings on Windows; the `csv` docs ask for it every time you open a CSV file. `encoding="utf-8"` handles names with accents.
- `sys.argv[1]` lets you pass the filename on the command line, with `sales.csv` as a fallback. No argument-parsing library needed.
- We write the report to `report.txt` as well as printing it, so you've got a file to email. Same string, two destinations.

## What you built

A working CSV summary tool. It reads rows, totals and averages a numeric column, breaks the total out by any category, and prints a clean aligned report - from an embedded sample in the browser, or from a real file on your machine with one changed line.

The shape you learned here - parse to dicts, convert to numbers, accumulate into a `defaultdict`, format with alignment specs - is the backbone of nearly every quick data script you'll write. Next time someone asks "what's the total per X?", you won't open a spreadsheet. You'll write ten lines and have the answer.
