# Testcontainers, From Zero

> Integration tests against the real thing: spin up a throwaway Postgres, Kafka, or Redis in Docker for each test run, then tear it down automatically.


---

# Testcontainers, From Zero

You wrote a beautiful mock of your database, every test is green, and then production falls over on a query the mock happily accepted. The mock told you what you wanted to hear. Testcontainers fixes that by booting a real Postgres, a real Redis, a real Kafka in a throwaway Docker container for the duration of your test run, then deleting it the moment you're done. Your tests talk to the actual engine, not your guess about how it behaves. No shared staging database to fight over, no leftover state between runs.

## How to read this

Read the three phases in order. Phase 1 builds the mental model: why mocks lie and what a "throwaway container" really is. Phase 2 is the everyday loop: start a container, get its real port, point your code at it, clean up. Phase 3 is where it gets real: the Docker requirement, slow first runs, port collisions, and CI. Type the commands as you go; the muscle memory matters more than the prose.

## The phases

1. [Why mocks lie and what a container gives you](01-why-mocks-lie.md)
2. [The everyday loop: start, connect, tear down](02-the-everyday-loop.md)
3. [Production reality: Docker, speed, and CI](03-production-reality.md)


---

# Why mocks lie and what a container gives you

Here's the moment that sends most people looking for Testcontainers. You have a function that runs a SQL query. You don't want your unit test to need a database, so you mock the database client: when the code calls `query()`, the mock returns a hand-written row. Green checkmark. Ship it. Then a real query hits a real Postgres and dies on a `JSON` cast, a unique-constraint violation, a timezone the mock never modeled, or a migration that didn't run. The mock didn't test your SQL. It tested your *belief* about your SQL.

That gap has a name, and it's the whole reason this tool exists.

## A mock is a recording of your assumptions

When you mock a dependency, you write down what you think it will do. The mock then plays that recording back, flawlessly, forever. That's genuinely useful for the layers you own and want to isolate. But your database is not your assumption. Postgres has its own type system, its own constraint engine, its own SQL dialect, its own quirks about ordering and nulls. Redis has its own eviction and expiry semantics. Kafka has partitions and offsets and consumer-group rebalancing. None of that lives in your mock unless you re-implement it, badly, by hand.

```text
What you mocked:          What Postgres actually does:
  query() -> [{"id":1}]     - enforces the UNIQUE index
                            - rejects the bad ENUM value
                            - applies the timezone
                            - runs your real migration DDL
```

*What just happened:* the mock returns a clean row and never exercises the engine's rules, so any bug that lives in those rules sails straight past your test suite and into production.

The deeper you go into integration territory, the more this matters. For where mocks legitimately belong versus where you need the real thing, see /guides/unit-integration-e2e.

## What a "throwaway container" actually is

A container is a real running copy of the software, isolated from your machine, started from a published image. `postgres:16` is a complete, real Postgres server. When Testcontainers runs, it asks Docker to start that image, waits until the database is genuinely ready to accept connections, hands your test the connection details, runs your tests, and then destroys the container. Throwaway is the key word: the container exists only for the test run. Next run gets a brand-new one with zero leftover state.

```text
test run starts
   |
   v
docker pulls postgres:16 (first time only) -> starts container
   |
   v
Testcontainers WAITS until the DB accepts connections
   |
   v
your tests run against the REAL Postgres
   |
   v
container is stopped and DELETED  (state gone)
```

*What just happened:* you got a real database for the lifetime of one test run and nothing to clean up afterward, because the container and everything in it is discarded.

This is the mental shift. You are not faking the dependency and you are not borrowing a long-lived shared one. You are renting a real, private, disposable copy for a few seconds.

## Why not a shared staging database?

The other common move is to point integration tests at one shared test database that lives on a server somewhere. It works until two people run tests at once and stomp on each other's rows, or someone leaves the schema in a half-migrated state, or a test fails and leaves garbage data that breaks the *next* test. Shared mutable state plus concurrency is the classic recipe for flaky tests, and flaky tests are worse than no tests because they teach the team to ignore red.

A throwaway container sidesteps the whole problem. Every run is isolated. Two developers, or twenty CI jobs, can run simultaneously, each with its own private container that nobody else can see or corrupt.

> The one real cost: Testcontainers needs a working Docker (or compatible) runtime on whatever machine runs the tests. That's the trade. We confront it head-on in Phase 3, because it's the single thing that surprises people most.

## For builders

The pattern is language-agnostic. There are official Testcontainers libraries for Java, Go, Python, .NET, Node.js, Rust, and more, and they all share the same shape: declare a container, start it, read back the connection details, use them, let it clean up. Learn the model once and it transfers. The image names (`postgres:16`, `redis:7`, `confluentinc/cp-kafka`) are the same images you'd run anywhere else, so what your tests exercise is exactly what runs in production.

```quiz
[
  {
    "q": "Why does a mocked database let real bugs through?",
    "choices": [
      "Mocks are too slow to catch them",
      "A mock replays your assumptions and never runs the real engine's rules",
      "Mocks always return null",
      "Docker isn't involved"
    ],
    "answer": 1,
    "explain": "A mock plays back what you told it to; it doesn't enforce constraints, types, or dialect the way the real engine does."
  },
  {
    "q": "What does 'throwaway' mean for a Testcontainers container?",
    "choices": [
      "It runs forever in the background",
      "It is shared across the whole team",
      "It exists only for the test run and is deleted afterward, with no leftover state",
      "It only holds mocked data"
    ],
    "answer": 2,
    "explain": "The container is created for the run and destroyed after, so each run starts from a clean, isolated state."
  },
  {
    "q": "What is the main downside of a single shared staging database for integration tests?",
    "choices": [
      "It is too realistic",
      "Concurrent runs corrupt each other's state, causing flaky tests",
      "It cannot run migrations",
      "It is incompatible with Docker"
    ],
    "answer": 1,
    "explain": "Shared mutable state plus concurrent test runs produces leftover data and collisions, which makes tests flaky."
  }
]
```


---

# The everyday loop: start, connect, tear down

Once the mental model clicks, the day-to-day is a tight loop you'll repeat for every integration test file: declare a container, start it, ask it for the connection details, point your code at those details, and let the library tear it down. The single thing beginners get wrong is hardcoding a port. Let's kill that habit first, because it's the heart of how Testcontainers works.

## The dynamic port is the whole trick

Postgres listens on 5432 *inside* the container. But Testcontainers does not publish it on your host's 5432. It maps the container's 5432 to a random free port on your machine, something like 49173, different on every run. This is deliberate: it's why you can run ten containers at once without collisions and why CI never trips over a port that's already taken. You never hardcode the port. You ask the container what port it got.

```python
from testcontainers.postgres import PostgresContainer

with PostgresContainer("postgres:16") as pg:
    url = pg.get_connection_url()
    print(url)
    # postgresql+psycopg2://test:test@localhost:49173/test
```

*What just happened:* the container started, mapped Postgres's internal 5432 to a random host port, and `get_connection_url()` handed back the full address including that port, so your code connects to the right place without you ever naming a number.

The shape is identical across languages. In Java you call `container.getJdbcUrl()`; in Go you call `container.ConnectionString(ctx)`; in Node you read `container.getMappedPort(5432)`. Same idea every time: the library knows the real port, so you ask it.

> Hardcoding `localhost:5432` is the number-one Testcontainers mistake. It accidentally works on your laptop if you happen to have a local Postgres on 5432, then fails everywhere else, or worse, your test silently talks to your real local database. Always read the mapped port.

## The full loop, start to finish

Here's a complete, realistic test in Python so you can see every step. The `with` block is doing the lifecycle work: start on enter, stop and delete on exit, even if the test throws.

```python
import psycopg2
from testcontainers.postgres import PostgresContainer

def test_user_is_persisted():
    with PostgresContainer("postgres:16") as pg:
        conn = psycopg2.connect(pg.get_connection_url().replace("+psycopg2", ""))
        cur = conn.cursor()

        # run your real schema, not a mock
        cur.execute("CREATE TABLE users (id SERIAL PRIMARY KEY, email TEXT UNIQUE)")
        cur.execute("INSERT INTO users (email) VALUES ('a@b.com')")
        conn.commit()

        cur.execute("SELECT email FROM users WHERE id = 1")
        assert cur.fetchone()[0] == "a@b.com"
```

*What just happened:* a real Postgres started, you created a real table with a real `UNIQUE` constraint, inserted and read back a real row, and when the `with` block ended the container was destroyed, leaving nothing behind.

Notice what this test would now catch that a mock wouldn't: insert two rows with the same email and Postgres rejects the second one for real. That's the constraint actually firing, not your imagination of it.

## Start the container once per suite, not per test

Booting a container takes a second or two. If you start a fresh one for every single test, your suite crawls. The standard move is to start the container once for the whole test file (or session), and reset *data* between tests instead of restarting the *container*. Reset is cheap; restart is not.

```text
SLOW (don't):                  FAST (do):
test_a -> start container       start container ONCE
test_a -> stop                    test_a -> TRUNCATE tables
test_b -> start container         test_b -> TRUNCATE tables
test_b -> stop                    test_c -> TRUNCATE tables
test_c -> start container       stop container ONCE
...
```

*What just happened:* the fast version pays the startup cost a single time and clears data with a quick `TRUNCATE` between tests, so the suite stays fast while every test still starts from a clean slate.

In pytest you'd put the container in a session- or module-scoped fixture; in JUnit you'd mark the container `static` with `@Container`; in Go you start it in `TestMain`. Same goal: one slow startup, many fast tests.

## Containers other than databases

The same loop works for anything with a Docker image. Redis, Kafka, RabbitMQ, Elasticsearch, even a real HTTP service, all follow declare-start-read-port-teardown. When there's no dedicated module for your image, there's a generic container you point at any image and tell which port to wait for.

```python
from testcontainers.core.container import DockerContainer
from testcontainers.core.waiting_utils import wait_for_logs

with DockerContainer("redis:7").with_exposed_ports(6379) as redis:
    wait_for_logs(redis, "Ready to accept connections")
    host = redis.get_container_host_ip()
    port = redis.get_exposed_port(6379)
    print(host, port)  # 127.0.0.1 49201
```

*What just happened:* a generic Redis container started, you waited until its log said it was ready (so you don't connect too early), and then read the mapped host and port to connect, the same pattern as the Postgres module but spelled out by hand.

That `wait_for_logs` line is important: a container being *started* is not the same as the service inside it being *ready*. The dedicated modules bake in a sensible wait strategy for you; with the generic container you specify your own. More on readiness traps in Phase 3.

## In the wild

Most teams split fast unit tests from slower Testcontainers-backed integration tests so the quick feedback loop stays quick, then run both in CI. For how those layers fit together in a pipeline, see /guides/testing-in-ci.

```quiz
[
  {
    "q": "Why should you never hardcode localhost:5432 in a Testcontainers test?",
    "choices": [
      "Postgres doesn't use 5432",
      "The container maps the internal port to a random host port, so you must read the mapped port",
      "Docker blocks port 5432",
      "It's slower than a random port"
    ],
    "answer": 1,
    "explain": "Testcontainers maps the container port to a random free host port to avoid collisions; you ask the container for the actual port."
  },
  {
    "q": "What's the recommended way to keep a Testcontainers suite fast?",
    "choices": [
      "Start a fresh container for every test",
      "Mock the container",
      "Start the container once per suite and reset data (e.g. TRUNCATE) between tests",
      "Disable the wait strategy"
    ],
    "answer": 2,
    "explain": "Container startup is the slow part; start once and clear data between tests so each test is still isolated but the suite stays fast."
  },
  {
    "q": "When using the generic container for an image without a dedicated module, what extra step matters most?",
    "choices": [
      "Hardcoding the port",
      "Specifying a wait strategy so you connect only after the service is actually ready",
      "Disabling Docker",
      "Running it as root"
    ],
    "answer": 1,
    "explain": "A started container isn't necessarily a ready service; the generic container needs you to define when it's ready (e.g. wait_for_logs)."
  }
]
```


---

# Production reality: Docker, speed, and CI

Testcontainers is reliable, but it has one hard requirement and a handful of failure modes that look mysterious the first time you hit them. None are hard once you know what you're looking at. This phase is the list of things that will bite you, and the fix for each.

## The Docker requirement is non-negotiable

Testcontainers starts real containers, so it needs a container runtime it can talk to. On a dev laptop that's usually Docker Desktop, Colima, Rancher Desktop, or Podman; in CI it's a Docker daemon the job can reach. No runtime, no containers, and the error is blunt about it.

```text
Could not find a valid Docker environment.
Please check:
  - Docker is installed and the daemon is running
  - the current user can access the Docker socket
```

*What just happened:* Testcontainers tried to reach a container runtime, found nothing it could talk to, and stopped before any test ran, because there is no fallback. Real containers need a real daemon.

This is the trade you accept for testing against the real thing. If a machine genuinely cannot run Docker, Testcontainers is the wrong tool there. Plan for it: developers install a runtime as part of onboarding, and CI uses an image or service that provides Docker.

## The first run is slow because of the image pull

The first time you reference `postgres:16`, Docker has to download it. That can take a noticeable while and makes a fresh checkout's first test run feel broken when it's actually still pulling. Every run after that uses the cached image and is fast.

```text
first run:  pull postgres:16 (downloads layers) ... then start  -> slow
later runs: image already cached locally ......... then start  -> fast
```

*What just happened:* the slowness was a one-time download, not the test, and once the image is cached locally subsequent runs skip straight to starting the container.

In CI, cache the Docker layers between runs (most CI providers support this) or pre-pull the images in a setup step, otherwise every pipeline pays the download tax. Pin image tags (`postgres:16`, not `postgres:latest`) so a surprise upstream change can't silently alter what your tests run against.

## Readiness, not only "started"

The bug that produces flaky tests more than any other: connecting before the service is actually ready. Docker reports the container as running the instant the process starts, but Postgres needs a moment more before it accepts connections, and a database that's mid-startup will refuse you with a connection error that looks random.

```text
container state: RUNNING   <- Docker says go
postgres state:  starting  <- not ready yet!
your test:       connect   -> "connection refused"  (flaky failure)
```

*What just happened:* the test connected during the gap between the process launching and the database being ready to serve, producing an intermittent failure that has nothing to do with your code.

The dedicated modules (`PostgresContainer`, etc.) ship with a correct wait strategy, so prefer them. With the generic container, always set an explicit wait, on a log line, on a port, or on an HTTP health check, so the library blocks until the service is truly serving.

## The leftover-container fear, and Ryuk

A reasonable worry: if my test crashes hard, do containers pile up forever? Testcontainers guards against this with a companion container, commonly called Ryuk, that watches your test session and reaps the containers it started if your process dies without cleaning up. Normal teardown still happens through the lifecycle (the `with` block, `@Container`, `t.Cleanup`); Ryuk is the safety net for the crash case.

> Some locked-down CI environments block Ryuk. You can disable it, but then you must ensure cleanup yourself, otherwise abandoned containers accumulate on the runner. Check your platform's docs before turning it off, and prefer leaving it on where allowed.

## Resource use and parallelism

Each container is a real running service eating real memory and CPU. Spin up Postgres plus Kafka plus Elasticsearch for one test and you're running three real servers at once. That's fine, until you also run the suite in high parallelism on a small CI runner and it starts swapping or getting killed for memory.

```text
1 test, 3 services:  postgres + kafka + elasticsearch  = real RAM x3
x8 parallel workers: 24 real services at once           = OOM on a small box
```

*What just happened:* parallelism multiplies real resource use because each worker starts its own real containers, so a runner that's fine serially can run out of memory under heavy parallel load.

Tune parallelism to the runner's size, give CI integration jobs a box with enough memory, and don't start heavyweight services you don't need for a given test. Keep the truly fast unit tests separate from the container-backed integration tests so the quick loop stays quick.

## For builders

A solid setup looks like this: dedicated container modules with their built-in wait strategies, the container started once per suite with data reset between tests, image tags pinned, Ryuk left on, and CI configured with Docker available plus image-layer caching. Get those right and Testcontainers fades into the background, exactly what you want from a test dependency.

```quiz
[
  {
    "q": "What happens if Testcontainers can't reach a container runtime?",
    "choices": [
      "It falls back to mocks automatically",
      "It runs the tests against an in-memory fake",
      "It fails before any test runs, because real containers need a real daemon",
      "It downloads Docker for you"
    ],
    "answer": 2,
    "explain": "Testcontainers has no fallback; without a reachable Docker-compatible runtime it errors out immediately."
  },
  {
    "q": "Why does the first test run feel slow but later runs are fast?",
    "choices": [
      "The first run compiles the image",
      "The first run pulls (downloads) the image; later runs use the cached copy",
      "Ryuk slows the first run",
      "The wait strategy only runs once"
    ],
    "answer": 1,
    "explain": "The initial image download is a one-time cost; cached images make subsequent runs fast."
  },
  {
    "q": "What is the Ryuk companion container for?",
    "choices": [
      "Speeding up image pulls",
      "Reaping leftover containers if your test process dies without cleaning up",
      "Mapping ports",
      "Replacing the database with a mock"
    ],
    "answer": 1,
    "explain": "Ryuk is the safety net that removes containers Testcontainers started if the test session crashes before normal teardown."
  }
]
```
