# Ansible, From Zero

> Configuration management without agents: SSH into your servers and run idempotent playbooks that bring them to a desired state, every time.


---

# Ansible, From Zero

You have a handful of servers, and right now they drift. You SSH in, run a few commands, edit a config by hand, forget which box got the fix and which didn't. Three months later nobody knows what's actually installed where. Ansible kills that whole class of problem: you write down the state you want once, run it, and every server matches. Run it again next week and nothing changes, because everything is already correct.

## How to read this

Go in order. Phase 1 builds the mental model: agentless push over SSH, and the one idea that makes Ansible click - idempotency. Phase 2 is the everyday loop: inventory, playbooks, variables, roles, handlers. Phase 3 is production reality - what breaks, where Ansible is the wrong tool, and how it sits next to provisioning tools. You can run every example on one Linux box or a cheap VM; you don't need a fleet.

## The phases

1. [What Ansible Actually Is](01-the-mental-model.md) - agentless config management, push over SSH, and why idempotency changes everything.
2. [The Everyday Loop](02-the-everyday-loop.md) - inventory, playbooks, modules, variables, roles, and handlers in the order you'll actually use them.
3. [Production Reality](03-production-reality.md) - gotchas, scaling, secrets, and where Ansible ends and Terraform begins.


---

# What Ansible Actually Is

Picture the manual version of the job. You SSH into a web server, install nginx, copy a config file, start the service. Then you do it again on the second server. And the third. By the fourth you've made a typo, and now one box runs a slightly different config than the others. Nobody notices until 2am, when that one box behaves differently under load and you're trying to remember what you typed three weeks ago.

That gap - between "what I think is on the server" and "what is actually on the server" - is the whole problem Ansible exists to close. It's a configuration management tool: you describe the state a machine should be in, and Ansible makes the machine match that description. Not the steps to get there. The destination.

## Agentless: it's SSH underneath

Most of Ansible's competitors (Puppet, Chef, Salt in its classic mode) need an agent - a daemon running on every managed server, phoning home to a central master. That's another service to install, secure, monitor, and upgrade on every box. It's infrastructure to manage your infrastructure.

Ansible threw that out. There is no agent. The machine you run Ansible from (your laptop, a CI runner, a bastion host - the "control node") connects to each target over plain SSH, the same protocol you already use to log in. It pushes over a little Python, runs it, collects the result, and disconnects. The target needs two things it almost certainly already has: an SSH server and a Python interpreter.

```text
  control node                          managed nodes
 ┌────────────┐      SSH (push)        ┌────────────┐
 │  ansible   │ ─────────────────────▶ │  web-01     │
 │  + your    │ ─────────────────────▶ │  web-02     │
 │  playbooks │ ─────────────────────▶ │  db-01      │
 └────────────┘    no agent installed  └────────────┘
```

*What just happened:* The control node does all the thinking. The managed nodes run nothing special between runs - Ansible connects, acts, leaves. That's why it's called "push" (the control node initiates) and "agentless" (nothing persistent lives on the targets).

This is a real architectural choice with consequences. Push-over-SSH means you control exactly when changes happen - Ansible only does something when you run it. There's no background daemon drifting servers on its own schedule. The tradeoff: scaling to thousands of nodes means opening thousands of SSH connections from one place, which we'll deal with in Phase 3.

> If your SSH keys and access aren't already sorted, that's the real prerequisite here. See [/guides/ssh-and-keys](/guides/ssh-and-keys) - Ansible inherits whatever SSH setup you already have, including agents, jump hosts, and key auth.

## Idempotency: the idea that makes it click

Here's the concept that separates Ansible from a glorified shell-script runner. Read this carefully, because everything else depends on it.

A bash script says *do these steps*. An Ansible task says *ensure this state*. The difference shows up the second time you run it.

```bash
# A shell script - imperative, describes STEPS
useradd deploy
mkdir /opt/app
echo "PORT=8080" >> /etc/app.conf
```

*What just happened:* Run this once, it works. Run it twice, `useradd` errors because the user already exists, and the `echo >>` appends a *second* `PORT=8080` line to the config. The script isn't safe to re-run. Imperative scripts accumulate damage.

Now the Ansible way:

```yaml
# An Ansible task - declarative, describes STATE
- name: Ensure deploy user exists
  ansible.builtin.user:
    name: deploy
    state: present
```

*What just happened:* The `user` module checks whether a user named `deploy` already exists. First run: it doesn't, so Ansible creates it and reports **changed**. Second run: it does, so Ansible does nothing and reports **ok**. Same task, run a hundred times, and the system ends up identical every time. That property - running it again is always safe and converges to the same state - is **idempotency**.

This is why Ansible runs report counts like `changed=3 ok=12`. The `ok` items were already correct and left alone. You're not firing commands blind; you're asserting facts about the system and letting Ansible reconcile the difference. A run where everything is already correct shows `changed=0`, and that's the goal: a system that matches its definition.

## Why "modules" instead of raw commands

You might wonder why you'd write `ansible.builtin.user` instead of running `useradd`. Because raw commands aren't idempotent - that's the whole point. Modules are the idempotent building blocks. There's one for users, one for packages (`apt`, `dnf`, the generic `package`), one for files and templates, one for services, one for git checkouts, and hundreds more. Each one knows how to check the current state and change it only if needed.

You *can* run raw commands when you have to (the `command` and `shell` modules exist), but every time you reach for them you give up idempotency and take responsibility for it yourself. The skill of writing good Ansible is largely the skill of finding the right module instead of shelling out.

> For builders: think of a module as a tiny program that takes "desired state" as arguments, inspects reality, and makes the smallest change to close the gap. That's the same loop a reconciler runs in Kubernetes or Terraform - Ansible runs it once, on demand, over SSH, instead of continuously.

```quiz
[
  {
    "q": "What does 'agentless' mean for Ansible?",
    "choices": [
      "It runs without any configuration files",
      "There is no persistent daemon on managed nodes; the control node connects over SSH on demand",
      "It does not require SSH keys",
      "It manages only the local machine"
    ],
    "answer": 1,
    "explain": "Ansible pushes over SSH from a control node. Managed nodes need only an SSH server and Python - nothing persistent runs between executions."
  },
  {
    "q": "You run an Ansible task to ensure a user exists, twice. What happens on the second run?",
    "choices": [
      "It errors because the user already exists",
      "It creates a duplicate user",
      "It checks, sees the user exists, makes no change, and reports 'ok'",
      "It deletes and recreates the user"
    ],
    "answer": 2,
    "explain": "That's idempotency: the module checks current state and only changes what's needed. An already-correct system reports changed=0."
  },
  {
    "q": "Why prefer a module like 'user' over the 'command' module running 'useradd'?",
    "choices": [
      "Modules run faster than commands",
      "Modules are idempotent and check state; raw commands are not, so you'd own that logic yourself",
      "The command module is deprecated",
      "Modules don't need SSH"
    ],
    "answer": 1,
    "explain": "Modules know how to inspect current state and change only what's needed. Reaching for command/shell gives up that idempotency."
  }
]
```


---

# The Everyday Loop

Now the hands-on part. Day to day, Ansible is four things in a stack: an **inventory** that lists your machines, **playbooks** that say what to do to them, **variables** that let one playbook serve many machines, and **roles** that package it all up so you can reuse it. We'll build them in that order, because that's the order they depend on each other.

## Inventory: who am I talking to

Ansible needs to know what machines exist and how to reach them. That's the inventory. The simplest form is an INI-style file, usually named `hosts` or `inventory.ini`.

```ini
# inventory.ini
[web]
web-01 ansible_host=10.0.1.11
web-02 ansible_host=10.0.1.12

[db]
db-01 ansible_host=10.0.1.20

[production:children]
web
db

[web:vars]
ansible_user=deploy
```

*What just happened:* We defined two groups, `web` and `db`, with hosts in each. `production:children` makes a parent group containing both, so `production` means all five-ish boxes at once. `[web:vars]` sets `ansible_user=deploy` for every web host. Groups are how you target a slice of your fleet - "run this on `web`" - without listing machines one by one.

You can sanity-check what Ansible thinks your inventory looks like before touching anything:

```console
$ ansible-inventory -i inventory.ini --graph
@all:
  |--@production:
  |  |--@web:
  |  |  |--web-01
  |  |  |--web-02
  |  |--@db:
  |  |  |--db-01
  |--@ungrouped:
```

*What just happened:* `ansible-inventory --graph` renders the group tree so you can confirm grouping before you run a playbook. Cheap insurance against accidentally targeting the wrong set of hosts.

A quick connectivity check uses the `ping` module - which isn't ICMP ping, it's "can I SSH in and run Python here":

```console
$ ansible -i inventory.ini web -m ping
web-01 | SUCCESS => { "ping": "pong" }
web-02 | SUCCESS => { "ping": "pong" }
```

*What just happened:* The ad-hoc `ansible` command (not `ansible-playbook`) ran the `ping` module against the `web` group. `pong` back from both means SSH auth and Python are working. If this fails, fix it before writing playbooks - every playbook depends on this working.

## Playbooks: plays, tasks, modules

A playbook is a YAML file describing what to do. The structure has three nested layers, and the names matter:

- A **play** maps a group of hosts to a list of tasks (`hosts: web` plus the tasks for web servers).
- A **task** is one step - it calls one module with some arguments and gets a `name` for the run log.
- A **module** is the idempotent unit of work from Phase 1 (`apt`, `copy`, `service`).

```yaml
# site.yml
- name: Configure web servers       # this is a PLAY
  hosts: web
  become: true                      # run tasks with sudo
  tasks:
    - name: Install nginx           # this is a TASK
      ansible.builtin.apt:          # this is a MODULE
        name: nginx
        state: present
        update_cache: true

    - name: Start and enable nginx
      ansible.builtin.service:
        name: nginx
        state: started
        enabled: true
```

*What just happened:* One play targets the `web` group, escalates to root with `become: true`, then runs two tasks. `apt` ensures nginx is installed; `service` ensures it's running now (`started`) and set to start on boot (`enabled`). Run it:

```console
$ ansible-playbook -i inventory.ini site.yml

PLAY [Configure web servers] ***************************
TASK [Install nginx] **********************************
changed: [web-01]
changed: [web-02]
TASK [Start and enable nginx] *************************
changed: [web-01]
changed: [web-02]

PLAY RECAP ********************************************
web-01  : ok=3  changed=2  unreachable=0  failed=0
web-02  : ok=3  changed=2  unreachable=0  failed=0
```

*What just happened:* Both tasks reported `changed` on the first run because nginx wasn't there yet. The `ok=3` includes an implicit fact-gathering step Ansible runs first. Run the exact same command again and you'll get `changed=0` - everything is already in the desired state. That `changed=0` on a re-run is your proof the playbook is idempotent.

> Before a risky change, run with `--check` (a dry run that reports what *would* change without doing it) and `--diff` (shows the actual file/line differences). Together they're your "show me what you're about to do" button.

## Variables: one playbook, many machines

Hardcoding `nginx` and version numbers into tasks doesn't scale. Variables let you parameterize. They can live in the playbook, in the inventory, or - the clean way - in `group_vars/` and `host_vars/` directories that Ansible loads automatically by group or host name.

```yaml
# group_vars/web.yml  - applies to every host in the 'web' group
app_port: 8080
worker_count: 4
```

```yaml
# in a task, used via Jinja2 templating
- name: Deploy nginx config
  ansible.builtin.template:
    src: nginx.conf.j2
    dest: /etc/nginx/nginx.conf
  notify: reload nginx
```

```text
# templates/nginx.conf.j2
worker_processes {{ worker_count }};
server {
    listen {{ app_port }};
}
```

*What just happened:* The `template` module renders `nginx.conf.j2` through Jinja2, substituting `{{ worker_count }}` and `{{ app_port }}` from `group_vars/web.yml`, and writes the result to the target. Change the variable, re-run, and the config updates everywhere - one source of truth, many machines. The `notify:` line is a trigger we'll explain next.

## Handlers: do something only when something changed

Restarting nginx on every run is wasteful and disruptive. You only want to reload it when its config actually changed. That's what handlers are for: a handler is a task that runs *only if* it was notified, and only *once* at the end of the play, no matter how many tasks notified it.

```yaml
  handlers:
    - name: reload nginx
      ansible.builtin.service:
        name: nginx
        state: reloaded
```

*What just happened:* The `template` task above had `notify: reload nginx`. If the template task reports `changed` (the config differed), it queues the `reload nginx` handler. If the config was already correct, nothing is notified and nginx is left running undisturbed. The handler fires once at the end even if three different tasks notified it. This is how you get "reload the service, but only when its config actually moved."

## Roles: packaging it so you can reuse it

Once a playbook grows past a screen or two, you'll want structure. A **role** is a standard directory layout that bundles tasks, handlers, templates, defaults, and files for one responsibility - say, "set up an nginx web server" - so you can drop it into any playbook.

```text
roles/
  nginx/
    tasks/main.yml        # the tasks (auto-loaded)
    handlers/main.yml     # the handlers (auto-loaded)
    templates/nginx.conf.j2
    defaults/main.yml     # default variable values
```

```yaml
# site.yml - now it just composes roles
- name: Configure web servers
  hosts: web
  become: true
  roles:
    - nginx
    - app_deploy
```

*What just happened:* Ansible auto-discovers `tasks/main.yml`, `handlers/main.yml`, and `templates/` inside a role by convention - no paths to wire up. The playbook shrinks to a list of roles, and each role is independently reusable across projects. This is also how you consume other people's work: the public Ansible Galaxy registry is full of community roles for common software, installable with `ansible-galaxy`.

> In the wild: most real Ansible repos are a thin top-level playbook plus a `roles/` tree and `group_vars/`. The playbook reads like a table of contents; the roles hold the actual logic. When you inherit an Ansible codebase, read `site.yml` first, then the roles it lists.

```quiz
[
  {
    "q": "In Ansible's structure, what is a 'play'?",
    "choices": [
      "A single call to one module",
      "A mapping of a group of hosts to a list of tasks",
      "A directory of reusable templates",
      "The output recap of a run"
    ],
    "answer": 1,
    "explain": "A play binds hosts (like 'web') to the tasks that should run on them. Tasks call modules; a playbook is a list of plays."
  },
  {
    "q": "When does a handler actually run?",
    "choices": [
      "On every playbook run, always",
      "Before any tasks, at the start of the play",
      "Only if a task notified it, once, at the end of the play",
      "Once per task that notifies it"
    ],
    "answer": 2,
    "explain": "Handlers run only when notified by a changed task, and run a single time at the play's end no matter how many tasks notified them."
  },
  {
    "q": "What is the main benefit of organizing work into a role?",
    "choices": [
      "Roles run faster than plain playbooks",
      "Roles bundle tasks, handlers, templates, and defaults in a conventional layout you can reuse across playbooks",
      "Roles remove the need for an inventory",
      "Roles make tasks non-idempotent"
    ],
    "answer": 1,
    "explain": "A role is a standard directory structure that packages one responsibility so it's reusable and auto-discovered by Ansible."
  }
]
```


---

# Production Reality

The first time Ansible bites you, it's usually one of a few well-worn ways: a playbook that *looks* idempotent but isn't, secrets sitting in plain text, a run that hangs on three hundred hosts, or the slow realization that you're trying to use Ansible to do a job Terraform should be doing. Let's walk through each before it walks through you.

## The idempotency trap: command and shell

Phase 1 said modules are idempotent. The escape hatches - `command` and `shell` - are not, and they're the single most common source of "why does this report changed every time."

```yaml
# NOT idempotent - runs every time, always reports 'changed'
- name: Build the app
  ansible.builtin.shell: make build
```

*What just happened:* The `shell` module has no idea what "build the app" means or whether it's already done, so it runs `make build` on every single playbook execution and always reports `changed`. Your run never converges to `changed=0`.

You fix this by telling Ansible how to check:

```yaml
- name: Build the app
  ansible.builtin.shell: make build
  args:
    creates: /opt/app/dist/bundle.js   # skip if this file exists
```

*What just happened:* `creates:` tells the task "if this file already exists, you're done - skip me." Now it runs once and reports `ok` thereafter. For the inverse there's `removes:`, and for arbitrary conditions you pair a check task with `changed_when:` and `when:`. The rule of thumb: every `command`/`shell` task needs a guard, or it's a lie about idempotency.

## Secrets: never commit them in plain text

Database passwords, API tokens, TLS private keys - they end up in variables, and variables end up in git. Ansible ships a built-in answer: **Ansible Vault**, which encrypts files (or individual values) with a passphrase, so the ciphertext is safe to commit.

```console
$ ansible-vault encrypt group_vars/production/secrets.yml
New Vault password:
Confirm New Vault password:
Encryption successful

$ ansible-playbook -i inventory.ini site.yml --ask-vault-pass
Vault password:
```

*What just happened:* `ansible-vault encrypt` turned the secrets file into an encrypted blob - open it and you'll see ciphertext, not your password. At run time, `--ask-vault-pass` (or a password file referenced by `--vault-password-file`) decrypts it in memory. The plaintext never touches disk in the repo. Commit the encrypted file freely; guard the passphrase like any other credential.

## Scaling: SSH fan-out has limits

Agentless is elegant, but "one control node opens an SSH connection to every target" has a ceiling. By default Ansible runs against a batch of hosts at a time (the `forks` setting), so a run across hundreds of hosts proceeds in waves, not all at once.

```console
# crank up parallelism for a big fleet
$ ansible-playbook -i inventory.ini site.yml --forks 50
```

*What just happened:* `--forks 50` lets Ansible work 50 hosts in parallel instead of the conservative default. Higher fan-out finishes faster but loads the control node's CPU, memory, and network harder. For genuinely large fleets people add SSH pipelining, mitogen, or pull-mode (`ansible-pull`, where each node pulls and runs its own config from git on a schedule - flipping the push model). The straight summary: Ansible is superb for tens to low hundreds of hosts and needs care beyond that.

## Order, failure, and not breaking everything at once

A naive run hits every host in a group at full speed. For a rolling deploy you want the opposite - update a few, verify, move on - so one bad change doesn't take down the whole fleet simultaneously.

```yaml
- name: Rolling update
  hosts: web
  serial: 2            # two hosts at a time
  max_fail_percentage: 25
  tasks:
    - name: Deploy new release
      ansible.builtin.git:
        repo: https://example.com/app.git
        dest: /opt/app
        version: v2.1.0
```

*What just happened:* `serial: 2` updates the web group two hosts at a time instead of all at once, so a broken release fails on a small batch first. `max_fail_percentage: 25` aborts the whole run if more than a quarter of hosts fail - a circuit breaker that stops a bad change before it reaches everything. This is how you turn "config management" into a safe deploy.

## The big one: config management vs provisioning

This is the distinction that decides whether you've picked the right tool at all.

Ansible configures machines that **already exist**. It assumes there's a server at an IP, reachable over SSH, and it installs packages, writes configs, and starts services on it. It does not natively create the server, the network, the load balancer, or the DNS record. It's weak at managing the *lifecycle* of cloud resources, because it has no real state model of "what should exist."

Terraform is the mirror image. It **provisions** infrastructure - it creates and destroys cloud resources (VMs, networks, DNS, IAM) and tracks them in a state file so it knows exactly what exists and can reconcile drift in the resource graph. It's weak at the inside-the-box config that Ansible is great at.

```text
  Terraform                          Ansible
 ┌──────────────────────┐          ┌──────────────────────┐
 │ CREATE the servers,   │  hand    │ CONFIGURE the servers:│
 │ network, DNS, LB.     │   off →  │ packages, files,      │
 │ Tracks state of what  │          │ services, deploys.    │
 │ EXISTS.               │          │ Assumes they exist.   │
 └──────────────────────┘          └──────────────────────┘
```

*What just happened:* The common production pattern is both, in sequence: Terraform stands up the boxes and outputs their IPs, then Ansible configures what's inside them. They're complementary, not competing. If you find yourself fighting Ansible to manage cloud resources, that's the signal you want a provisioning tool. See [/guides/infrastructure-as-code-terraform](/guides/infrastructure-as-code-terraform) for the other half of this pairing.

> The shorthand that's stuck with me: Terraform owns the *shape* of your infrastructure (what exists and how it's wired); Ansible owns the *state inside* each piece. Reach for the one whose job matches your problem, and don't make either do the other's work.

```quiz
[
  {
    "q": "Why does a task using the 'shell' module to run 'make build' report 'changed' on every run?",
    "choices": [
      "shell is slower than other modules",
      "shell has no knowledge of whether the work is already done, so it always executes",
      "make is not supported by Ansible",
      "It needs become: true to be idempotent"
    ],
    "answer": 1,
    "explain": "command/shell aren't idempotent. Add a guard like creates:/removes: or changed_when: so the task can skip when already done."
  },
  {
    "q": "What does Ansible Vault give you?",
    "choices": [
      "A way to store inventory in the cloud",
      "Encryption of secret files/values with a passphrase so ciphertext is safe to commit to git",
      "Faster SSH connections",
      "Automatic provisioning of cloud servers"
    ],
    "answer": 1,
    "explain": "Vault encrypts secrets at rest; the passphrase decrypts them in memory at run time, so plaintext never lands in the repo."
  },
  {
    "q": "Which split between Terraform and Ansible is correct?",
    "choices": [
      "Terraform configures software inside servers; Ansible provisions cloud resources",
      "Terraform provisions/creates infrastructure and tracks its state; Ansible configures servers that already exist",
      "They do the same thing; pick either one",
      "Ansible tracks infrastructure state in a state file like Terraform"
    ],
    "answer": 1,
    "explain": "Terraform owns the shape of infra (what exists, tracked in state); Ansible owns the config inside existing machines. They're complementary."
  }
]
```
