# Dockerize an App

> Take a small web app and containerize it - a real Dockerfile, image caching, and a compose file with a database - built and run on your machine.


---

# Dockerize an App

You have an app that runs on your laptop. It works. Then a teammate clones it, runs the same commands, and gets a wall of errors - a missing library, the wrong language version, a database that isn't there. You spend an afternoon over a screen-share fixing their machine instead of building anything.

This weekend you're going to fix that for good. You'll take a small web app, put it inside a container, and end up with something anyone can run with one command - same versions, same dependencies, same database, every time. No "works on my machine."

## What you'll build

A small Python web app (Flask) running inside a Docker container, talking to a Postgres database in a second container, the two wired together with a single compose file. By the end you'll bring the whole stack up with `docker compose up` and shut it down with `docker compose down`, and you'll have an image you can tag and push to a registry for someone else to pull and run.

## The stack

| Piece | What it is |
| --- | --- |
| Docker Engine | Runs containers on your machine |
| Flask | A tiny Python web app (the thing we containerize) |
| Postgres | The database, run as its own container |
| Docker Compose | Wires the app and database together |

## Built on your machine

This is a **build-on-your-machine** project, not a run-in-the-browser one. Every command here you run in your own terminal. You need Docker installed before you start - get **Docker Desktop** (macOS, Windows) or **Docker Engine** (Linux) from docker.com, then confirm it's alive:

```bash
docker --version
docker run hello-world
```

If `hello-world` prints a friendly message, you're set. If it can't connect to the Docker daemon, start Docker Desktop (or run `sudo systemctl start docker` on Linux) and try again.

You'll also want Python 3 on hand for the first couple of phases, though by the end the container holds its own Python and your machine's version stops mattering - which is the whole point.

## Rough time

About a weekend, in four sittings:

- **Phase 1** - the "works on my machine" problem, and your first Dockerfile that builds and runs.
- **Phase 2** - a real Dockerfile: caching, `.dockerignore`, a slim base, a non-root user.
- **Phase 3** - `docker compose` to run the app and a Postgres database together.
- **Phase 4** - ship it: environment variables, secrets, a healthcheck, and pushing to a registry.

## What you'll learn

- The difference between an **image** (the recipe) and a **container** (a running instance of it).
- How Docker's layer cache makes the difference between a 2-second rebuild and a 2-minute one.
- Why the order of lines in a Dockerfile matters more than almost anything else.
- How containers find and talk to each other on a network.
- How to pass configuration in without baking secrets into the image.
- How to hand your image to someone else so they run it with no setup at all.

Open a terminal, make a folder for this project, and turn to Phase 1.


---

# Why Containers, and a First Dockerfile

Here's the problem we're solving, stated plainly. Your app depends on things that aren't in the app: a specific Python version, a web framework, a database driver, maybe a system library or two. When you run it, all of those happen to be on your machine. When someone else runs it, some of them aren't - or they're the wrong version. The app doesn't care that it's "the same code." It breaks.

A container fixes this by shipping the app **and** everything it depends on as one sealed unit. Inside the container is a tiny, predictable Linux system with exactly the Python and libraries your app needs. Outside, the host machine's setup stops mattering. Whoever runs the container gets your environment, not theirs.

Two words you'll see constantly:

- An **image** is the recipe - a frozen, read-only snapshot of the filesystem with your app and its dependencies baked in.
- A **container** is a running instance of an image. You can start many containers from one image, like many objects from one class.

You write a **Dockerfile** to describe the image. Docker reads it top to bottom and builds the image. Let's make one.

## The app

Make a project folder and drop two files in it. First the app - a Flask server with a single route:

```python
# app.py
from flask import Flask

app = Flask(__name__)

@app.route("/")
def home():
    return "Hello from inside a container!\n"

if __name__ == "__main__":
    app.run(host="0.0.0.0", port=5000)
```

One detail that trips everyone up: `host="0.0.0.0"`. Inside a container, the default `127.0.0.1` means "only listen to traffic from inside this container" - which is nobody, since you're connecting from outside. `0.0.0.0` tells Flask to accept connections on all interfaces so your host can reach it.

Now declare the dependency:

```text
# requirements.txt
flask==3.0.3
```

You could run this locally right now (`pip install -r requirements.txt && python app.py`), and if you want to confirm the app itself works before containerizing, go ahead. But the point is to stop depending on your machine's Python. So let's containerize it.

## The Dockerfile

Create a file named exactly `Dockerfile` (no extension) next to `app.py`:

```dockerfile
# Dockerfile
FROM python:3.12

WORKDIR /app

COPY requirements.txt .
RUN pip install -r requirements.txt

COPY . .

EXPOSE 5000

CMD ["python", "app.py"]
```

Read it line by line, because every line is doing real work:

- `FROM python:3.12` - start from an official image that already has Python 3.12 installed. This is your base layer; you build on top of it.
- `WORKDIR /app` - set the working directory inside the image. Commands after this run here, and it's created if it doesn't exist.
- `COPY requirements.txt .` - copy the requirements file from your folder into the image's `/app`.
- `RUN pip install -r requirements.txt` - install Flask inside the image. This runs at *build* time, so the dependency is baked in.
- `COPY . .` - copy the rest of your project (including `app.py`) into the image.
- `EXPOSE 5000` - document that the app listens on port 5000. (This is a label; it doesn't actually publish the port - you do that when you run.)
- `CMD [...]` - the default command to run when a container starts. This one launches your app.

## Build it

From the project folder, build the image and give it a name with `-t` (for "tag"):

```bash
docker build -t myapp .
```

The `.` at the end is the **build context** - the folder Docker sends to the build, and the root for those `COPY` lines. You'll watch Docker pull the base image (once - it's cached after that) and run each step. At the end you'll see something like `naming to docker.io/library/myapp`.

Confirm the image exists:

```bash
docker images
```

You should see `myapp` in the list with a size - probably around a gigabyte, because `python:3.12` is a full image. We'll shrink that dramatically in Phase 2.

## Run it

Start a container from the image:

```bash
docker run -p 8080:5000 myapp
```

The `-p 8080:5000` is the bridge between worlds: it maps port 8080 on your machine to port 5000 inside the container (the order is `host:container`). The `EXPOSE` line alone wouldn't do this - `-p` is what actually opens the door.

Open a browser to `http://localhost:8080` (or run `curl http://localhost:8080`). You should see:

```text
Hello from inside a container!
```

That response came from a Python that may not even be installed on your machine, running inside an isolated Linux environment, reachable because you mapped a port. The container is running in your terminal - press `Ctrl+C` to stop it.

To run it in the background instead, add `-d` (detached) and name it:

```bash
docker run -d -p 8080:5000 --name myapp-run myapp
docker ps          # see it running
docker logs myapp-run
docker stop myapp-run && docker rm myapp-run
```

## Where you are

You have a working image and a container serving HTTP. Anyone with Docker can now run your exact app - no Python setup, no dependency install - with `docker build` and `docker run`. That's the "works on my machine" problem solved in its crudest form.

It's also wasteful: a gigabyte for a Hello World, and a full reinstall of Flask every time you change one line of `app.py`. In Phase 2 we turn this rough Dockerfile into one you'd be happy to put your name on.


---

# A Real Dockerfile

The Phase 1 Dockerfile works, but it has three problems you'll feel the moment you use it for real: it rebuilds slowly, it's enormous, and it runs your app as root. Each has a fix, and each fix teaches you something about how Docker actually works. Let's take them one at a time and end with a Dockerfile you'd ship.

## How the layer cache works

Every instruction in a Dockerfile creates a **layer** - a saved snapshot of the filesystem after that step. When you rebuild, Docker walks the instructions in order and reuses a cached layer as long as nothing that feeds it has changed. The first instruction whose inputs changed busts the cache, and every instruction after it reruns.

That last sentence is the whole game. Look at the Phase 1 order:

```dockerfile
COPY . .
RUN pip install -r requirements.txt
```

If you wrote it this way, every edit to `app.py` changes the `COPY . .` layer, which busts the cache for `pip install` right below it - so Docker reinstalls Flask on every single code change. You changed a string and paid for a full dependency install.

We already avoided that in Phase 1 by copying `requirements.txt` *before* the rest of the code:

```dockerfile
COPY requirements.txt .
RUN pip install -r requirements.txt
COPY . .
```

Now `pip install` only reruns when `requirements.txt` changes. Edit `app.py` a hundred times and the install layer stays cached. The rule that falls out of this:

> Put the things that change rarely (dependencies) **above** the things that change often (your code).

Here's the cache decision as a picture:

```mermaid
flowchart TD
    A[Start build] --> B{requirements.txt<br/>changed?}
    B -- no --> C[Reuse cached<br/>pip install layer]
    B -- yes --> D[Rerun pip install]
    C --> E{app code<br/>changed?}
    D --> E
    E -- no --> F[Reuse cached<br/>COPY layer]
    E -- yes --> G[Rerun COPY]
    F --> H[Image ready]
    G --> H
```

## A .dockerignore

When you run `docker build .`, Docker sends the entire folder - the build context - to the build engine. If that folder contains a `.git` directory, a virtualenv, caches, or logs, all of it gets shipped, slowing the build and sometimes sneaking junk into your image through `COPY . .`.

Fix it the same way you'd fix `.gitignore` - list what to leave out. Create `.dockerignore` next to your Dockerfile:

```text
# .dockerignore
.git
.gitignore
__pycache__/
*.pyc
.venv/
venv/
.env
*.log
Dockerfile
.dockerignore
```

Excluding `.env` here matters for more than speed: it keeps local secrets from being copied into the image. We'll handle configuration properly in Phase 4.

## A slim base image

`python:3.12` is the full image - it includes compilers, build tools, and a complete Debian userland you mostly don't need to *run* a Flask app. The `-slim` variant strips that down to a minimal Debian plus Python:

| Base image | Rough size |
| --- | --- |
| `python:3.12` | ~1 GB |
| `python:3.12-slim` | ~150 MB |
| `python:3.12-alpine` | ~70 MB |

Swap one word and you've cut the image by roughly 850 MB:

```dockerfile
FROM python:3.12-slim
```

A note on `alpine`: it's even smaller, but it uses a different C library (musl, not glibc), which occasionally breaks Python packages that ship compiled code and expect glibc. For this project `-slim` is the sweet spot - small, and no surprises. Reach for `alpine` only when you've measured that the extra savings is worth the risk.

## Don't run as root

By default, the process inside your container runs as `root`. If someone finds a way to break out of your app, they're root inside the container - a worse starting point for them than an unprivileged user. Create a normal user and switch to it before the app runs:

```dockerfile
RUN useradd --create-home appuser
USER appuser
```

Everything after `USER appuser` runs as that user. Put this line *after* your `pip install` (which may need to write to system locations) but *before* `CMD`, so the app itself runs unprivileged.

## The real Dockerfile

Put it all together:

```dockerfile
# Dockerfile
FROM python:3.12-slim

WORKDIR /app

# Dependencies first - this layer is cached until requirements.txt changes
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# App code second - changes here don't bust the install layer
COPY . .

# Run as a non-root user
RUN useradd --create-home appuser
USER appuser

EXPOSE 5000

CMD ["python", "app.py"]
```

Two small additions worth calling out. `pip install --no-cache-dir` tells pip not to keep its download cache inside the image - you don't need it at runtime, and it only adds weight. And the comments document *why* the order is what it is, which the next person (often future you) will thank you for.

## Prove it's better

Rebuild:

```bash
docker build -t myapp .
docker images myapp
```

The size should now be a few hundred megabytes instead of a gigabyte. Now test the cache. Change the greeting in `app.py` - edit the string in the `home()` route - and rebuild:

```bash
docker build -t myapp .
```

Watch the output. The `pip install` step prints `CACHED` and finishes instantly; only the `COPY . .` and the steps after it rerun. The rebuild that took a minute now takes a couple of seconds. Run it again to confirm it still serves traffic:

```bash
docker run -d -p 8080:5000 --name myapp-run myapp
curl http://localhost:8080
docker stop myapp-run && docker rm myapp-run
```

Same `Hello`, smaller image, faster builds, no root. This is a Dockerfile you can be proud of. Next we give the app something to talk to - a real database, in its own container, wired up with compose.


---

# docker compose: App Plus Database

Your app runs in a container. Now it needs a database - and a database is its own piece of software, with its own version and config and storage. You could install Postgres on your machine, but that drags you right back to "works on my machine." Instead you run Postgres as a second container.

Two containers means two `docker run` commands with the right ports, names, networks, and environment variables - and you'd type them in the right order every time. That gets old fast. **Docker Compose** is the fix: you describe the whole stack once in a file, and bring it all up or down with one command.

## Give the app something to store

First, make the app actually use a database so we can see it working. Update `requirements.txt`:

```text
# requirements.txt
flask==3.0.3
psycopg[binary]==3.2.1
```

`psycopg` is the Postgres driver for Python; the `[binary]` extra pulls a prebuilt version so nothing needs compiling. Now update `app.py` to count visits in Postgres:

```python
# app.py
import os
import psycopg
from flask import Flask

app = Flask(__name__)

DB_URL = os.environ["DATABASE_URL"]

def init_db():
    with psycopg.connect(DB_URL) as conn:
        conn.execute(
            "CREATE TABLE IF NOT EXISTS visits (id serial PRIMARY KEY)"
        )

@app.route("/")
def home():
    with psycopg.connect(DB_URL) as conn:
        conn.execute("INSERT INTO visits DEFAULT VALUES")
        count = conn.execute("SELECT count(*) FROM visits").fetchone()[0]
    return f"You are visit number {count}\n"

if __name__ == "__main__":
    init_db()
    app.run(host="0.0.0.0", port=5000)
```

Notice the app reads its database connection from `os.environ["DATABASE_URL"]`. It doesn't hard-code where the database is - that comes from the environment, which is exactly what lets the same image run against different databases. Compose will set that variable for us.

## The compose file

Create `compose.yaml` next to your Dockerfile:

```yaml
# compose.yaml
services:
  web:
    build: .
    ports:
      - "8080:5000"
    environment:
      DATABASE_URL: postgresql://appuser:secret@db:5432/appdb
    depends_on:
      - db

  db:
    image: postgres:16
    environment:
      POSTGRES_USER: appuser
      POSTGRES_PASSWORD: secret
      POSTGRES_DB: appdb
    volumes:
      - dbdata:/var/lib/postgresql/data

volumes:
  dbdata:
```

Walk through it:

- **Two services**, `web` and `db`. Each becomes a container.
- `build: .` - `web` is built from your Dockerfile in the current folder. `db` instead pulls the official `postgres:16` image - no Dockerfile needed.
- `ports` on `web` does what `-p` did: maps host 8080 to container 5000.
- `environment` sets variables inside each container. The `db` service uses Postgres's own variables to create a user, password, and database on first start. The `web` service gets the `DATABASE_URL` its code reads.
- `depends_on: [db]` tells Compose to start `db` before `web`.
- `volumes` on `db` maps a named volume `dbdata` to where Postgres stores its files, so data survives a container restart. More on that below.

## The magic word: `db`

Look closely at the connection string: `postgresql://appuser:secret@db:5432/appdb`. The host is `db` - the name of the service. Compose puts every service on a shared private network and lets them find each other by service name. Inside the `web` container, `db` resolves to the Postgres container's address. You never deal with IP addresses.

```mermaid
flowchart LR
    H[Your browser<br/>localhost:8080] --> W[web container<br/>Flask :5000]
    W -- "connects to host 'db'" --> D[db container<br/>Postgres :5432]
    D --- V[(dbdata volume)]
```

Note that 5432 is *not* published to your host - only `web` can reach Postgres, over the private network. That's a sensible default: your database shouldn't be exposed to the outside world. (If you want to inspect it with a desktop tool, you can add a `ports` entry to `db`, but you don't need to.)

## Volumes: why your data survives

Containers are disposable. Delete one and everything written inside it is gone. That's fine for your stateless app, but a database that forgets everything on restart is useless.

A **volume** is storage that lives outside the container's lifecycle. The line `dbdata:/var/lib/postgresql/data` says "keep the contents of Postgres's data directory in a named volume called `dbdata`." Destroy and recreate the `db` container and the data is still there, because it never lived inside the container - it lived in the volume.

## Bring it up

From the project folder:

```bash
docker compose up --build
```

`up` starts the whole stack; `--build` rebuilds your `web` image first so code changes are picked up. You'll see logs from both services interleaved, color-coded by service. Wait for Postgres to report it's ready and Flask to say it's serving, then hit it:

```bash
curl http://localhost:8080
# You are visit number 1
curl http://localhost:8080
# You are visit number 2
```

The count climbs because each request writes a row to Postgres. Stop the stack with `Ctrl+C`, or run detached and manage it separately:

```bash
docker compose up --build -d   # background
docker compose ps              # what's running
docker compose logs web        # logs for one service
docker compose logs -f         # follow all logs
```

## Prove the volume works

Here's the satisfying test. Bring the stack down - but keep the volume:

```bash
docker compose down
docker compose up -d
curl http://localhost:8080
# You are visit number 3
```

The count picked up where it left off. `down` removed the containers, but the `dbdata` volume persisted, so Postgres still had your rows. Now wipe it for real with `-v`, which deletes the named volumes too:

```bash
docker compose down -v
docker compose up -d
curl http://localhost:8080
# You are visit number 1
```

Back to one - the data is gone because you destroyed the volume. That `-v` flag is the difference between "restart the stack" and "start fresh," and knowing which is which will save you a panic someday.

You now have a real two-container application defined in one file, brought up and down with one command, with data that persists exactly as long as you want it to. The only thing left is making it fit to hand to someone else - secrets, health, and a registry. That's Phase 4.


---

# Ship It

You have a working stack. Now make it fit to leave your laptop. Three things stand between you and that: the password is sitting in plain text in your compose file, nothing checks whether the database is actually *ready* before the app tries to use it, and the image only exists on your machine. Let's close all three, then push the image somewhere a teammate or a server can pull it.

## Get secrets out of the compose file

Right now `compose.yaml` has `POSTGRES_PASSWORD: secret` written in it. That file is in version control, which means your password is in version control - for everyone, forever, even after you "delete" it from a later commit.

The first step up is an **`.env` file** that Compose reads automatically. Create `.env` next to `compose.yaml`:

```text
# .env  -- do NOT commit this
POSTGRES_USER=appuser
POSTGRES_PASSWORD=a-better-password-than-secret
POSTGRES_DB=appdb
```

Then reference those variables in `compose.yaml` with `${...}`:

```yaml
# compose.yaml
services:
  web:
    build: .
    ports:
      - "8080:5000"
    environment:
      DATABASE_URL: postgresql://${POSTGRES_USER}:${POSTGRES_PASSWORD}@db:5432/${POSTGRES_DB}
    depends_on:
      db:
        condition: service_healthy

  db:
    image: postgres:16
    environment:
      POSTGRES_USER: ${POSTGRES_USER}
      POSTGRES_PASSWORD: ${POSTGRES_PASSWORD}
      POSTGRES_DB: ${POSTGRES_DB}
    volumes:
      - dbdata:/var/lib/postgresql/data
    healthcheck:
      test: ["CMD-SHELL", "pg_isready -U ${POSTGRES_USER} -d ${POSTGRES_DB}"]
      interval: 5s
      timeout: 3s
      retries: 5

volumes:
  dbdata:
```

Now add `.env` to both `.gitignore` and `.dockerignore` so it never lands in your repo or your image. Commit a `.env.example` with dummy values instead, so the next person knows which variables to set:

```text
# .env.example  -- safe to commit
POSTGRES_USER=appuser
POSTGRES_PASSWORD=change-me
POSTGRES_DB=appdb
```

A word on what `.env` is and isn't. It keeps secrets out of your committed files, which is the big win. It is *not* a vault. For real production you'd graduate to your platform's secret manager (Docker secrets, Kubernetes secrets, a cloud KMS) - but the contract stays identical: **the app reads its config from the environment**, and where that environment comes from is someone else's problem. Because your app already reads `DATABASE_URL` from `os.environ`, you change nothing in the code to move up the ladder.

## Add a healthcheck

You may have noticed a race in Phase 3: Compose starts `db` before `web` because of `depends_on`, but "started" isn't "ready." Postgres takes a second or two to accept connections, and if `web` connects in that window, it crashes.

The compose file above already fixes this. The `healthcheck` on `db` runs `pg_isready` every few seconds until Postgres answers, marking the container *healthy*. Then `depends_on` with `condition: service_healthy` makes `web` wait for that healthy state - not start, but *healthy* - before it launches.

A healthcheck is also useful on its own. It tells Docker (and orchestrators like Kubernetes) whether a container is alive, so a wedged container gets noticed and restarted instead of silently failing. Bring it up and watch the states:

```bash
docker compose up --build -d
docker compose ps
```

You'll see `db` go from `starting` to `healthy`, and `web` only comes up after. No more startup race.

## Tag the image for a registry

Your image is called `myapp` and lives only on your machine. To share it, you push it to a **registry** - a server that stores images. Docker Hub is the default and has a free tier; cloud providers and `ghcr.io` (GitHub) work the same way.

Registry image names follow a pattern: `registry/username/name:tag`. For Docker Hub the registry part is implied, so it's only `username/name:tag`. The `tag` is a version label - `latest` is the default, but a real version is better.

Tag your existing image (replace `yourname` with your Docker Hub username):

```bash
docker tag myapp yourname/dockerize-demo:1.0.0
docker tag myapp yourname/dockerize-demo:latest
```

Tagging doesn't copy anything - it only adds names pointing at the same image. Confirm:

```bash
docker images yourname/dockerize-demo
```

## Push it

Log in, then push:

```bash
docker login
docker push yourname/dockerize-demo:1.0.0
docker push yourname/dockerize-demo:latest
```

Docker uploads the layers. Layers you've pushed before are skipped, so the slim base you chose in Phase 2 pays off again here - smaller image, faster push. When it finishes, your image is on the registry.

Now the proof. From any machine with Docker - or after deleting your local copy with `docker rmi yourname/dockerize-demo:1.0.0` - anyone can run it:

```bash
docker run -p 8080:5000 \
  -e DATABASE_URL="postgresql://appuser:pass@somehost:5432/appdb" \
  yourname/dockerize-demo:1.0.0
```

No clone, no Python, no pip. They pull the image and run it, passing in their own `DATABASE_URL`. That's the "works on my machine" problem fully closed - it now works on *any* machine.

## A few production tips

You've built the real thing. A handful of habits separate a demo from something that runs in production:

| Habit | Why it matters |
| --- | --- |
| Pin versions (`postgres:16`, `python:3.12-slim`) | `latest` changes under you and breaks reproducibility |
| Tag images with real versions, not only `latest` | You can roll back to a known image when a deploy goes wrong |
| Use a production WSGI server (e.g. gunicorn), not `flask run` | Flask's built-in server is for development, not load |
| Set `restart: unless-stopped` on services | The container comes back after a crash or reboot |
| Keep secrets in a real secret store for prod | `.env` is fine locally; production wants more |
| Run as non-root (you already do) | Smaller blast radius if the app is compromised |

To swap in gunicorn, add it to `requirements.txt` and change the Dockerfile's last line:

```dockerfile
CMD ["gunicorn", "--bind", "0.0.0.0:5000", "app:app"]
```

That's the whole arc. You took an app that ran on exactly one machine and turned it into an image anyone can pull and run, wired to a database, with config kept out of source control, a healthcheck guarding startup, and a version you can roll back to. The next time someone says "it works on my machine," you can hand them an image and say: now it works on yours too.
