# GitLab CI/CD, From Zero

> Pipelines defined in .gitlab-ci.yml: stages, jobs, and runners that build, test, and deploy on every push - with artifacts, caching, and environments.


---

# GitLab CI/CD, From Zero

You pushed a branch and someone said "wait for the pipeline." A wall of green and red dots appeared, a YAML file you didn't write decided whether your code was allowed to merge, and nobody could explain why the deploy button was grayed out. That confusion is normal, and it clears up fast once you see the shape underneath it. GitLab CI/CD is one file, a handful of ideas, and a machine that runs your commands for you. This guide gets you to the point where you can read any `.gitlab-ci.yml`, write one from scratch, and know what to do when it goes red.

## How to read this

Go in order. Phase 1 builds the mental model: stages, jobs, runners, and the single file that wires them together - read this even if you've copied a pipeline before, because the model is what makes the rest stick. Phase 2 is the everyday toolkit you'll actually type: passing files between jobs, caching dependencies, and controlling when each job runs. Phase 3 is production reality: environments, manual deploy gates, secrets, and the failures that wake people up. Each phase ends with a short quiz so you can check yourself before moving on.

## The phases

1. [Phase 1: The Mental Model - One File, A Pipeline, A Machine](01-the-mental-model.md)
2. [Phase 2: The Everyday Core - Artifacts, Cache, and Rules](02-artifacts-cache-rules.md)
3. [Phase 3: Production Reality - Environments, Gates, and Secrets](03-environments-gates-secrets.md)


---

# Phase 1: The Mental Model - One File, A Pipeline, A Machine

Here's the reality you're starting from: you push a commit, GitLab shows a little pipeline icon next to it, and a bunch of dots turn green (or one turns red and blocks your merge). It feels like magic happening on a server you've never logged into. It isn't magic. There are exactly three moving parts, and once you can name them, every pipeline you'll ever read becomes legible.

## The whole system in three words

GitLab CI/CD runs on three ideas:

- **A job** is a list of shell commands plus the context they run in. "Install dependencies and run the tests" is a job. A job either passes (exit code 0) or fails (anything else).
- **A stage** is a named group of jobs that run *together, in parallel*. The classic stages are `build`, `test`, `deploy`. All the jobs in `test` run at the same time; the pipeline doesn't move to `deploy` until every `test` job has passed.
- **A runner** is the actual machine (or container) that picks up a job and executes its commands. GitLab the website doesn't run your code - it hands the job to a runner, the runner runs it and reports back.

Put those together and you get a **pipeline**: stages run in order, jobs inside a stage run in parallel, runners do the work.

```text
push commit
   │
   ▼
┌─────────┐   ┌──────────────────────┐   ┌─────────┐
│  build  │ → │  test (3 in parallel)│ → │ deploy  │
└─────────┘   └──────────────────────┘   └─────────┘
  stage 1            stage 2               stage 3
```

*What just happened:* the commit triggered a pipeline with three stages. `build` runs first and must pass before `test` starts; the three test jobs run at once; only if all of them pass does `deploy` get its turn.

If you want the broader "why does any of this exist" framing - why teams automate build/test/deploy at all - see [/guides/what-cicd-does](/guides/what-cicd-does). This guide assumes you're sold on the idea and want to drive GitLab's version of it.

## The one file that controls everything

Everything lives in a file named `.gitlab-ci.yml` at the root of your repository. GitLab reads it on every push. There is no separate dashboard where the "real" config hides - the file *is* the config, it's version-controlled with your code, and changing the pipeline means editing this file and committing it.

A minimal-but-real pipeline looks like this:

```yaml
stages:
  - build
  - test

build-app:
  stage: build
  image: node:20
  script:
    - npm ci
    - npm run build

run-tests:
  stage: test
  image: node:20
  script:
    - npm ci
    - npm test
```

*What just happened:* you declared two stages and two jobs. `build-app` belongs to the `build` stage, `run-tests` belongs to `test`. Each job names a Docker `image` (the environment it runs in) and a `script` (the commands the runner executes). On a push, the runner spins up a `node:20` container, runs the `build-app` commands, then - only if that passed - spins up a fresh container for `run-tests`.

The two names you see at the top level (`build-app`, `run-tests`) are job names - you choose them, and they show up as the dots in the pipeline view. The reserved keywords (`stages`, `stage`, `image`, `script`) are GitLab's vocabulary; the job names are yours.

> A job always starts from a clean checkout in a fresh container. Nothing carries over from a previous job unless you explicitly tell it to - that "explicitly tell it to" is what artifacts and cache are for, and that's Phase 2. For now, hold the idea that jobs are isolated by default.

## What a runner actually is

The word "runner" trips people up because it's invisible. A runner is a small agent program installed on some machine - a cloud VM, a beefy server in a closet, GitLab's own shared fleet. It connects to your GitLab instance and says "I'm available." When a pipeline has a job ready, GitLab assigns it to a runner, which clones your repo, runs the `script`, captures the output and exit code, and reports back.

On GitLab.com you usually get **shared runners** for free (with a quota of minutes), so things work out of the box. In a company you'll often see **specific runners** the team installed - maybe to get more memory, a GPU, or access to an internal network. You rarely manage runners yourself early on; you need to know that the green dots cost real compute on a real machine somewhere.

```yaml
deploy-staging:
  stage: deploy
  tags:
    - linux-large
  script:
    - ./deploy.sh staging
```

*What just happened:* the `tags` key tells GitLab "only a runner that advertises the `linux-large` tag may take this job." Tags are how you route a heavy job to a beefy machine or a deploy job to a runner that has the right network access. No matching runner means the job sits pending - a classic "why is my pipeline stuck" cause.

## Reading a pipeline result

When a pipeline runs you'll see each job as a dot: green = passed, red = failed, gray = didn't run, blue/spinning = running, orange clock = pending (waiting for a runner). Click any job to see its full console log - every command and its output, exactly as the runner saw it. That log is your single source of truth when something breaks. Don't guess at why a job failed; open the log and read the last few lines.

```console
$ npm test
> jest

FAIL  src/auth.test.js
  ✕ rejects an expired token (12 ms)

Tests: 1 failed, 41 passed, 42 total
ERROR: Job failed: exit code 1
```

*What just happened:* the test job ran `npm test`, one test failed, `jest` exited with code 1, and GitLab marked the job (and the pipeline) red. The fix isn't in GitLab - it's in your code. CI didn't break; it did its job and told you the truth.

**For builders:** the fastest way to learn this file is to add a throwaway job that runs `echo` and `env`, push it, and read the log. You'll see the working directory, the branch name, the commit SHA, and dozens of `CI_*` variables GitLab injects automatically - the same variables you'll lean on in Phase 2 and Phase 3.

```quiz
[
  {
    "q": "In a pipeline with stages build, test, deploy, when do the jobs in the test stage run?",
    "choices": [
      "One at a time, in the order they appear in the file",
      "All at once, but only after every build-stage job has passed",
      "Before the build stage, to fail fast",
      "Only if you click a button to start them"
    ],
    "answer": 1,
    "explain": "Stages run in order; jobs within a stage run in parallel. The test stage starts only once all build jobs pass."
  },
  {
    "q": "What actually executes the commands in a job's script?",
    "choices": [
      "The GitLab web server itself",
      "Your local machine when you push",
      "A runner - an agent on some machine that picks up the job",
      "The .gitlab-ci.yml file"
    ],
    "answer": 2,
    "explain": "GitLab assigns the job to a runner, which clones the repo, runs the script, and reports the result back."
  },
  {
    "q": "Where does the pipeline configuration live?",
    "choices": [
      "In a hidden settings dashboard on GitLab.com",
      "In a file named .gitlab-ci.yml at the repo root, version-controlled with the code",
      "In a database only admins can edit",
      "In each runner's local config"
    ],
    "answer": 1,
    "explain": "The .gitlab-ci.yml file at the repository root is the config. It's committed with your code, so pipeline changes are reviewable like any other change."
  }
]
```


---

# Phase 2: The Everyday Core - Artifacts, Cache, and Rules

Phase 1 left you with a clean mental model and one nagging problem: jobs are isolated. Your `build` job compiled the app into a `dist/` folder, and then your `deploy` job started in a fresh container with no `dist/` in sight. Your tests reinstall every dependency from scratch and take four minutes when they could take forty seconds. And your deploy job runs on every single branch, including the doc-typo fix on someone's feature branch. This phase fixes all three. These three keywords - `artifacts`, `cache`, and `rules` - are the difference between a toy pipeline and one you'd actually ship behind.

## Artifacts: passing files forward

An **artifact** is a file (or folder) a job produces that you want to keep and hand to later jobs. You declare what to save; GitLab uploads it when the job finishes and downloads it into the right place for any later job that needs it.

```yaml
build-app:
  stage: build
  script:
    - npm ci
    - npm run build
  artifacts:
    paths:
      - dist/
    expire_in: 1 week

deploy-app:
  stage: deploy
  script:
    - ./deploy.sh dist/
```

*What just happened:* `build-app` declared `dist/` as an artifact. When it finishes, GitLab uploads that folder. `deploy-app` runs in a later stage, and because it's downstream, GitLab automatically downloads `dist/` into its workspace before the script runs - so `./deploy.sh dist/` finds real files. `expire_in: 1 week` tells GitLab to delete the stored copy after a week so storage doesn't grow forever.

By default, artifacts flow to *all* later jobs. Artifacts are also how you publish things humans want: a built binary, a coverage report, a `.zip` you download from the pipeline page. There's one specially-handled flavor - test reports:

```yaml
run-tests:
  stage: test
  script:
    - npm test -- --reporters=jest-junit
  artifacts:
    when: always
    reports:
      junit: junit.xml
```

*What just happened:* `reports: junit` tells GitLab "this artifact is a test report - parse it and show pass/fail counts in the merge request." `when: always` uploads the report even when the job fails (the default is `on_success`), which matters here because a *failing* test run is exactly when you want to see the report.

## Cache: making jobs fast

Artifacts move *outputs forward* between stages. **Cache** is different: it stores files you want to *reuse across pipeline runs* to avoid redoing slow work - almost always your dependency folder. The mental split that keeps people sane:

- **Artifact** = "I built this, the next job needs it." Tied to one pipeline. Correctness.
- **Cache** = "Downloading this again is slow, reuse last time's copy if you can." Spans pipelines. Speed, best-effort.

```yaml
run-tests:
  stage: test
  cache:
    key:
      files:
        - package-lock.json
    paths:
      - node_modules/
  script:
    - npm ci
    - npm test
```

*What just happened:* the runner restores `node_modules/` from the cache before the script, keyed on the contents of `package-lock.json`. If the lockfile hasn't changed, you get the same cache and `npm ci` is fast. When `package-lock.json` changes, the `key` changes, so you get a fresh cache and rebuild dependencies - exactly the behavior you want.

> Never put correctness on cache. Cache can be empty (first run), stale, or evicted at any time - treat it as a maybe. If your `deploy` job *needs* the `dist/` from `build`, that's an artifact, not a cache. A common painful bug is a deploy that "works on my machine and sometimes in CI" because it secretly depended on a warm cache.

## Rules: deciding when a job runs

By default, every job runs in every pipeline. That's rarely what you want - you don't deploy to production from a feature branch, and you might skip the slow integration suite on draft merge requests. The `rules` keyword controls when a job is created.

```yaml
deploy-prod:
  stage: deploy
  script:
    - ./deploy.sh production
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'
```

*What just happened:* `deploy-prod` is only added to the pipeline when the commit is on the `main` branch. On any other branch the job simply doesn't exist - no gray dot, no skipped step. `$CI_COMMIT_BRANCH` is one of the built-in variables GitLab injects into every pipeline.

You can combine conditions and change the job's behavior per rule:

```yaml
integration-tests:
  stage: test
  script:
    - ./run-integration.sh
  rules:
    - if: '$CI_PIPELINE_SOURCE == "merge_request_event"'
    - if: '$CI_COMMIT_BRANCH == "main"'
    - when: never
```

*What just happened:* this job runs in two situations - when the pipeline comes from a merge request, or when the commit lands on `main`. The final `when: never` is the catch-all: any other case falls through and the job is skipped. Rules are evaluated top to bottom, and the first match wins.

You may still see older pipelines using `only:` and `except:` instead of `rules:`:

```yaml
legacy-deploy:
  stage: deploy
  script:
    - ./deploy.sh staging
  only:
    - main
```

*What just happened:* `only: [main]` is the older syntax for "run this only on the main branch." It still works, but `rules` is the current, more expressive replacement - `only`/`except` can't express "run on MRs *and* main but nothing else" cleanly. Read `only`/`except` fluently; reach for `rules` when you write new jobs.

## Putting it together

Here's a small but realistic pipeline using all three ideas at once:

```yaml
stages:
  - build
  - test
  - deploy

build:
  stage: build
  script:
    - npm ci
    - npm run build
  artifacts:
    paths: [dist/]

test:
  stage: test
  cache:
    key:
      files: [package-lock.json]
    paths: [node_modules/]
  script:
    - npm ci
    - npm test

deploy:
  stage: deploy
  script:
    - ./deploy.sh dist/
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'
```

*What just happened:* `build` produces `dist/` as an artifact that flows to `deploy`; `test` caches `node_modules/` so it's fast on repeat runs; `deploy` only runs on `main` and consumes the artifact. Every branch gets a build and tests; only `main` gets a deploy. That's the everyday shape of a working pipeline.

**In the wild:** the single most common pipeline speedup is caching the dependency folder with a lockfile-based key, and the single most common pipeline *bug* is depending on a cache for correctness. Get those two right and you've avoided most of the pain teams hit in their first month.

```quiz
[
  {
    "q": "Your deploy job needs the dist/ folder that the build job produced. What should you use?",
    "choices": [
      "cache, keyed on the branch name",
      "artifacts in the build job, which flow to the deploy job",
      "Nothing - files automatically persist between jobs",
      "A shared cache with when: always"
    ],
    "answer": 1,
    "explain": "Artifacts pass a job's outputs forward within a pipeline and are about correctness. Cache is best-effort speed and can be empty, so it must never carry required files."
  },
  {
    "q": "What is the safest way to key a cache for node_modules?",
    "choices": [
      "On the branch name, so each branch gets its own cache",
      "On the contents of package-lock.json, so it refreshes when dependencies change",
      "On the commit SHA, so it's unique every push",
      "On a fixed string, so it never changes"
    ],
    "answer": 1,
    "explain": "Keying on the lockfile reuses the cache while dependencies are unchanged and rebuilds it only when they actually change - the right trade between speed and freshness."
  },
  {
    "q": "How do you make a job run only on the main branch?",
    "choices": [
      "Put it in a stage named main",
      "Add rules with if: '$CI_COMMIT_BRANCH == \"main\"'",
      "Give it a runner tag called main",
      "Set artifacts: paths: [main]"
    ],
    "answer": 1,
    "explain": "rules with a condition on $CI_COMMIT_BRANCH decides when the job is created. On other branches the job simply won't be part of the pipeline."
  }
]
```


---

# Phase 3: Production Reality - Environments, Gates, and Secrets

You can now write a pipeline that builds, tests, and deploys. The moment it touches real servers, three new questions appear, and they're the ones that decide whether your pipeline is trustworthy. Where did this deploy go, and what's running there right now? Who gets to push the button that ships to production? And how do you give your deploy job a database password without writing it into a file the whole company can read? This phase answers all three, then walks through the failures that actually page people at 2am.

## Environments: naming where you deploy

An **environment** is GitLab's record of a deploy target - `staging`, `production`, a per-branch review app. It's not infrastructure; it's a label plus a history. When a job declares `environment:`, GitLab tracks every deploy to that target: which commit, when, by whom, and a one-click way to see what's live and to roll back.

```yaml
deploy-staging:
  stage: deploy
  script:
    - ./deploy.sh staging
  environment:
    name: staging
    url: https://staging.example.com
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'
```

*What just happened:* this job deploys to `staging` and tells GitLab the live URL. In the **Deployments → Environments** view you now get a `staging` entry showing the current commit, a link straight to the running site, and a deployment history. The `url` becomes a clickable button - small thing, but it turns "is the deploy actually up?" from a Slack question into one click.

## Manual gates: the button you have to press

Deploying to staging on every `main` push is fine. Deploying to *production* automatically is how teams ship a Friday-night outage. The fix is a **manual job**: it appears in the pipeline but waits, doing nothing, until a human clicks play.

```yaml
deploy-production:
  stage: deploy
  script:
    - ./deploy.sh production
  environment:
    name: production
    url: https://example.com
  rules:
    - if: '$CI_COMMIT_BRANCH == "main"'
      when: manual
```

*What just happened:* on a `main` pipeline, `deploy-production` shows up as a play button instead of running. The pipeline can finish "successfully" with this job still pending - the deploy happens only when someone clicks it. That click is recorded against the production environment, so you have an audit trail of who shipped what.

> A manual job is a *gate*, not a guarantee of safety. Anyone with the right project role can press it. For real protection on production, combine the manual gate with **protected branches** and **protected environments** in GitLab's project settings, which restrict who may run that deploy. The YAML expresses intent; the settings enforce it.

You can make the gate block the pipeline if you'd rather force a decision, by adding `allow_failure: false` - then the pipeline stays "blocked" until someone acts, instead of going green with the deploy still waiting.

## Secrets: CI/CD variables

Your deploy job needs credentials - an API token, a database URL, an SSH key. These must never live in `.gitlab-ci.yml`, because that file is in the repo and visible to everyone with read access. Instead you store them as **CI/CD variables** in the project (or group) settings, and GitLab injects them into the job's environment at runtime.

```yaml
deploy-production:
  stage: deploy
  script:
    - echo "Deploying with token..."
    - curl -H "Authorization: Bearer $DEPLOY_TOKEN" https://api.example.com/release
  environment:
    name: production
```

*What just happened:* `$DEPLOY_TOKEN` isn't defined anywhere in the file. It's set in **Settings → CI/CD → Variables**, and the runner injects it as an environment variable when the job runs. The repo stays clean; the secret stays out of version control.

When you add a CI/CD variable, two checkboxes matter:

- **Masked** - GitLab replaces the value with `[masked]` in job logs, so an accidental `echo` doesn't leak it. Turn this on for every secret. (The value must meet GitLab's masking rules - long enough, no problematic characters - or masking silently won't apply.)
- **Protected** - the variable is exposed only to jobs running on protected branches or tags. This stops a feature branch from reading your production credentials. Turn it on for anything production-grade.

```console
$ echo "Token is $DEPLOY_TOKEN"
Token is [masked]
```

*What just happened:* even though the script echoed the variable, masking replaced it in the log. Masking is a safety net, not permission to print secrets - but it saves you the day someone forgets.

## When it breaks: the failures that page you

Real pipelines fail in a handful of recognizable ways. Knowing the shape saves you from staring at a red dot in confusion.

**The job is stuck pending forever.** No runner matches it. Either there are no runners available, or the job has a `tags:` value no runner advertises. Check the job - GitLab tells you "This job is stuck because there are no active runners." Fix the tag or the runner, don't blame the YAML.

**The deploy "passed" but nothing changed.** Classic masked-failure: a command in the middle of the script failed but a later command returned 0, so the job went green. By default the shell stops on the first failing command, but pipes and subshells can swallow errors. Make failures loud:

```yaml
deploy:
  stage: deploy
  script:
    - set -euo pipefail
    - ./build.sh | tee build.log
    - ./deploy.sh production
```

*What just happened:* `set -euo pipefail` makes the shell exit on any failed command (`-e`), treat unset variables as errors (`-u`), and - crucially - fail if any command in a pipe fails (`pipefail`), not only the last one. Without it, `./build.sh | tee build.log` would report success as long as `tee` succeeded, hiding a broken build.

**A secret is empty in the job.** Nearly always the variable is **protected** and the job is running on a non-protected branch, so GitLab refused to inject it. Either protect the branch or, for non-production targets, drop the protected flag. The log won't say "permission denied" - the variable is blank, so guard against it:

```yaml
deploy:
  stage: deploy
  script:
    - test -n "$DEPLOY_TOKEN" || { echo "DEPLOY_TOKEN is empty - check protected branch settings"; exit 1; }
    - ./deploy.sh production
```

*What just happened:* the job fails fast with a clear message instead of making a credential-less API call that fails confusingly downstream. A two-line guard like this turns a 20-minute head-scratch into an instant diagnosis.

**For builders:** the muscle to build here is reading the job log first and the YAML second. CI failures feel mysterious because the pipeline is "out there," but the log is a complete, accurate transcript of exactly what ran. If you came from GitHub Actions, the concepts map closely - same build/test/deploy shape, different file format and vocabulary; see [/guides/your-first-pipeline-github-actions](/guides/your-first-pipeline-github-actions) for the comparison.

```quiz
[
  {
    "q": "How do you make a production deploy require a human to click before it runs?",
    "choices": [
      "Put it in its own stage",
      "Add when: manual to the job's rule (or as a top-level key)",
      "Mark its artifacts as protected",
      "Set expire_in: never"
    ],
    "answer": 1,
    "explain": "when: manual turns the job into a play button that waits for someone to trigger it, giving you a deliberate gate before production."
  },
  {
    "q": "Where should a production database password live so it's not in the repo?",
    "choices": [
      "Hardcoded in .gitlab-ci.yml under a comment",
      "In a CI/CD variable in project settings, marked masked and protected",
      "In an artifact that expires quickly",
      "In the cache, keyed on the branch"
    ],
    "answer": 1,
    "explain": "CI/CD variables keep secrets out of version control; masked hides them in logs and protected restricts them to protected branches."
  },
  {
    "q": "A deploy job shows green but the site didn't change. What's the most likely cause and a good guard?",
    "choices": [
      "The runner was too fast; add a sleep",
      "A mid-script command failed silently; add set -euo pipefail so failures stop the job",
      "The artifact expired; set expire_in longer",
      "The environment URL was wrong; remove it"
    ],
    "answer": 1,
    "explain": "Errors inside pipes or subshells can be swallowed so the job exits 0. set -euo pipefail makes any failed command fail the job loudly."
  }
]
```
