# Argo CD and GitOps

> Deployment by pull, not push: GitOps makes a Git repo the source of truth for your cluster, and Argo CD continuously reconciles reality to match it.


---

# Argo CD and GitOps

You ship a change to Kubernetes, and a week later nobody can say for sure what's actually running. Someone hot-patched a replica count by hand, a config map drifted, and the cluster no longer matches anything you can point at. GitOps fixes the trust problem: the desired state of your cluster lives in Git, and a controller named Argo CD watches that repo and quietly drags reality back into line whenever it wanders. This guide gives you the mental model and the muscle memory to run it without surprises.

## How to read this

Read the three phases in order. Phase 1 builds the mental model so the rest stops feeling like magic. Phase 2 is the everyday loop: defining an app, syncing, rolling back. Phase 3 is where it bites - drift, sync waves, secrets, and the failures you'll actually be paged for. If you've never run Kubernetes, skim /guides/kubernetes-without-the-hype first; if "CI vs CD" is fuzzy, /guides/what-cicd-does sets the frame.

## The phases

1. [The pull model: Git as the source of truth](01-the-pull-model.md)
2. [Your daily loop: apps, sync, and rollback](02-daily-loop.md)
3. [When reconciliation bites: drift, waves, and secrets](03-when-it-bites.md)


---

# The pull model: Git as the source of truth

Think about how most teams deploy before GitOps. A pipeline runs, and at the end it reaches into the cluster and *pushes*: `kubectl apply`, a `helm upgrade`, maybe a hand-typed command at 2am during an incident. The cluster does whatever the last push told it. Now ask the uncomfortable question: what's running right now, and who decided that? Often nobody can answer with certainty. The cluster's state and your repo's state have quietly diverged, and the only way to know is to go poke at the live thing.

GitOps flips the direction. Instead of CI pushing into the cluster, a controller *inside* the cluster pulls from Git and makes the cluster match. The repo stops being a record of intentions and becomes the literal, enforced definition of what runs. That one inversion - pull instead of push - is the whole idea. Everything else in this guide is a consequence of it.

## Desired state vs actual state

Kubernetes is already declarative. You don't tell it "start three pods" - you write a Deployment that says `replicas: 3`, and a controller works to make that true and keep it true. If a pod dies, Kubernetes notices the gap between desired (3) and actual (2) and starts a replacement. That gap-closing behavior has a name: a **reconciliation loop**.

GitOps extends that loop one level up. The desired state is now a set of YAML files in a Git repo. The actual state is the live cluster. Argo CD is the controller that watches both and works to close any gap between them.

```text
   Git repo  ──────────►  Argo CD  ──────────►  Kubernetes cluster
   (desired)              (compares,            (actual)
                           reconciles)
        ▲                                              │
        └──────────────  reads back actual  ◄──────────┘
```

*What just happened:* desired state flows one way (Git → Argo CD → cluster), and Argo CD reads the cluster back to compare. Nothing pushes into the cluster from outside; the controller pulls.

The two states get compared constantly, and the comparison produces a status you'll live by:

- **Synced** - the cluster matches Git. All is well.
- **OutOfSync** - the cluster has drifted from Git, or Git has new commits the cluster hasn't applied yet.

## Push vs pull, side by side

The difference sounds academic until you see what each makes easy.

```text
PUSH (classic CI/CD)
  CI pipeline  ──has cluster credentials──►  kubectl apply  ──►  cluster
  Truth: "whatever the last successful pipeline pushed"

PULL (GitOps)
  Git repo  ◄──watches──  Argo CD (in cluster)  ──reconciles──►  cluster
  Truth: "whatever Git says, enforced continuously"
```

*What just happened:* in push, your CI system holds cluster credentials and is the actor. In pull, the cluster holds a read-only token to Git and is its own actor - credentials never leave the cluster's blast radius.

That credential point matters more than it looks. In a push world, every CI runner that can deploy has keys to prod. In a pull world, Argo CD reads Git (often read-only) and acts from *inside* the cluster, so your CI never needs cluster access at all. CI's job shrinks to "build the image, write the new tag into Git." Deployment becomes a Git commit.

> Mental model: GitOps doesn't add a new way to deploy. It removes deploying as a verb. You change Git; the cluster catches up. "Deploy" becomes "merge."

## Why this is worth the trouble

Three properties fall out of the pull model for free, and they're the reason teams adopt it.

**Auditability.** Every change to production is a Git commit - authored, timestamped, reviewed in a pull request. Want to know who changed the replica count and why? It's `git log`, not a forensic dig through cluster history.

**Rollback is `git revert`.** Because Git is the source of truth, undoing a bad deploy is undoing a commit. You don't hunt for the previous image tag or the old config - you revert the commit, and Argo CD reconciles the cluster back to that state. Same mechanism as deploying, run backwards.

**Drift detection and self-healing.** If someone runs `kubectl edit` by hand, the cluster no longer matches Git. Argo CD sees that as OutOfSync and can flag it - or, if you turn on self-heal, quietly undo it. The cluster can't silently drift away from what's reviewed and recorded.

```text
$ argocd app get payments
Name:        payments
Health:      Healthy
Sync Status: OutOfSync   (someone edited the live Deployment)
```

*What just happened:* a manual edit to the live cluster shows up immediately as OutOfSync. Git still says one thing, the cluster says another, and Argo CD refuses to pretend they agree.

## For builders

You don't need GitOps for a hobby cluster you alone touch. It earns its keep the moment more than one person can change production, or the moment "what's actually running?" becomes a question with a scary answer. The payoff is a single, reviewable, revertible record of your infrastructure - and a controller that won't let reality drift away from it behind your back.

```quiz
[
  {
    "q": "What is the defining inversion that GitOps introduces?",
    "choices": [
      "Manifests are written in JSON instead of YAML",
      "A controller inside the cluster pulls desired state from Git, instead of CI pushing into the cluster",
      "Deployments run twice as fast",
      "Kubernetes is replaced by a simpler scheduler"
    ],
    "answer": 1,
    "explain": "GitOps reverses the direction: an in-cluster controller pulls from Git and reconciles, rather than an external pipeline pushing in."
  },
  {
    "q": "In a GitOps setup, how do you roll back a bad deployment?",
    "choices": [
      "SSH into each node and restart services",
      "Run kubectl rollback against the live cluster",
      "Revert the Git commit; Argo CD reconciles the cluster back to that state",
      "Delete the namespace and recreate it by hand"
    ],
    "answer": 2,
    "explain": "Because Git is the source of truth, rollback is git revert - the same reconciliation mechanism, run against an earlier state."
  },
  {
    "q": "Why does the pull model reduce your production credential exposure?",
    "choices": [
      "It encrypts all YAML automatically",
      "CI no longer needs cluster credentials; the controller acts from inside the cluster and only needs to read Git",
      "It disables kubectl entirely",
      "It moves all secrets into the Git history"
    ],
    "answer": 1,
    "explain": "In pull mode, deployment is a Git commit. The cluster's controller reads Git and acts internally, so external CI never needs cluster keys."
  }
]
```


---

# Your daily loop: apps, sync, and rollback

You've got the mental model. Now the practical question: what do you actually touch day to day? Almost less than you'd expect. The whole point of GitOps is that your normal workflow is editing files and merging pull requests - Argo CD does the rest. But you'll lean on a handful of concepts and commands often enough that they should become reflex. This phase walks the loop you'll run dozens of times a week.

## The Application is the unit of work

In Argo CD, the thing you manage is an **Application**. It's a small piece of config that answers three questions: *where does the desired state live, what's in it, and where does it go?*

- **source** - the Git repo, a path inside it, and a revision (a branch, tag, or commit).
- **destination** - the cluster and namespace to deploy into.
- **syncPolicy** - whether Argo CD applies changes automatically or waits for you.

Here's a minimal one as YAML, which itself usually lives in Git:

```yaml
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
  name: payments
  namespace: argocd
spec:
  project: default
  source:
    repoURL: https://github.com/acme/infra.git
    targetRevision: main          # track this branch
    path: apps/payments           # folder of manifests in the repo
  destination:
    server: https://kubernetes.default.svc
    namespace: payments
  syncPolicy:
    automated:
      prune: true                 # delete resources removed from Git
      selfHeal: true              # undo manual cluster edits
```

*What just happened:* this tells Argo CD to watch `apps/payments` on the `main` branch and keep the `payments` namespace matching it - automatically, including deleting things you remove from Git (`prune`) and reverting manual edits (`selfHeal`).

Notice the Application is *itself* a Kubernetes resource. That leads to a common pattern: one root Application that points at a folder of other Applications, so adding a new app to the cluster is - you guessed it - a Git commit. That's the "app of apps" pattern, and it's how teams scale to dozens of services without clicking around a UI.

## The sync: making the cluster match Git

A **sync** is the act of applying Git's desired state to the cluster. With `automated` syncPolicy, Argo CD syncs on its own whenever it detects a new commit (it polls the repo on an interval, or reacts instantly if you wire up a webhook). With manual policy, you trigger it yourself - useful when you want a human to press the button on prod.

The everyday CLI rhythm:

```console
$ argocd app get payments
Name:        payments
Health:      Healthy
Sync Status: OutOfSync  (main is 1 commit ahead)

$ argocd app sync payments
Syncing... apps/payments
  Deployment/payments  configured
  Service/payments     unchanged

$ argocd app get payments
Sync Status: Synced
Health:      Healthy
```

*What just happened:* `get` showed the cluster was a commit behind Git. `sync` applied that commit, changing only what differed (the Deployment), and left untouched what already matched (the Service). The status flipped to Synced.

Two status fields ride along and you read both constantly:

- **Sync Status** - does the cluster match Git? (Synced / OutOfSync)
- **Health** - are the running resources actually OK? (Healthy / Progressing / Degraded)

They're independent, and the difference matters. A freshly-applied Deployment can be **Synced** (Git applied successfully) but **Progressing** or **Degraded** because the new pods are crash-looping. Synced means "Git was applied"; Health means "the result is working." Watch both - Synced-but-Degraded is exactly the state a broken deploy lives in.

## The deploy flow, end to end

Putting it together, here's how a real change rides to production:

```text
1. Build:   CI builds image  acme/payments:v1.4.2
2. Promote: CI commits to infra repo, bumping the image tag in apps/payments
3. Detect:  Argo CD sees the new commit on main  → OutOfSync
4. Sync:    Argo CD applies it (auto, or you click)  → rolling update
5. Verify:  Health goes Progressing → Healthy
```

*What just happened:* CI's only job is steps 1–2 - build the image and write the new tag into Git. From step 3 on, Argo CD owns the deploy. Your CI system never touched the cluster; it touched a file.

> Keep image tags immutable and specific. Pin `acme/payments:v1.4.2` (or a digest), never `:latest`. If Git says `:latest`, two syncs can produce two different clusters from the same commit - which quietly destroys the "Git is the source of truth" guarantee. See /guides/what-cicd-does for where tag promotion fits in the pipeline.

## Rollback: deploy in reverse

Something's wrong in prod. In a push world you'd scramble for the previous tag. Here, rollback is the same loop pointed backward - and you have two clean ways to do it.

The GitOps-pure way is to revert the commit. Argo CD reconciles the cluster to the reverted state, and your Git history stays accurate about what happened:

```console
$ git revert a1b2c3d        # undo the bad deploy commit
$ git push
# Argo CD detects the new commit and syncs the cluster back
```

*What just happened:* the revert creates a *new* commit that undoes the bad one. The cluster follows Git back to the good state, and the revert is recorded - your history shows both the break and the fix.

Argo CD also keeps a deploy history and can roll back directly, handy when prod is on fire and you want speed before tidiness:

```console
$ argocd app history payments
ID  DATE                 REVISION
3   2026-06-30 14:02     a1b2c3d (bad)
2   2026-06-29 09:10     9f8e7d6 (good)

$ argocd app rollback payments 2
Rolled back to revision 9f8e7d6
```

*What just happened:* Argo CD re-applied the manifests from history entry 2, instantly. But note the catch: the live cluster now matches an *old* commit while `main` still points at the bad one, so the app reads OutOfSync. Use this to stop the bleeding, then revert in Git to make the source of truth agree again. The Git revert is the durable fix; the CLI rollback is the fire extinguisher.

## In the wild

Most teams structure the infra repo so each environment is a folder or branch (`apps/payments` for staging, a separate path or repo for prod), and "promoting" a release means copying the tested image tag from one to the other in a reviewed PR. The deploy, the promotion, and the rollback all become the one operation you already know how to review: a change to a file in Git.

```quiz
[
  {
    "q": "What does an Argo CD Application define?",
    "choices": [
      "The CI pipeline steps for building an image",
      "A source (repo/path/revision), a destination (cluster/namespace), and a sync policy",
      "The Dockerfile for a service",
      "The list of developers allowed to deploy"
    ],
    "answer": 1,
    "explain": "An Application maps a Git source to a cluster destination and says how to sync them - it's the unit Argo CD reconciles."
  },
  {
    "q": "An app shows Sync Status: Synced but Health: Degraded. What does that mean?",
    "choices": [
      "Git failed to apply to the cluster",
      "Git was applied successfully, but the running resources are unhealthy (e.g. pods crash-looping)",
      "The repo URL is wrong",
      "Argo CD is offline"
    ],
    "answer": 1,
    "explain": "Sync Status and Health are independent. Synced means Git was applied; Degraded means the result isn't working."
  },
  {
    "q": "Why is pinning an immutable image tag (not :latest) important in GitOps?",
    "choices": [
      "It makes pulls faster",
      "It's required by Kubernetes",
      "With :latest, the same Git commit can produce different clusters on different syncs, breaking the source-of-truth guarantee",
      "It reduces the size of the manifest"
    ],
    "answer": 2,
    "explain": "If Git says :latest, two syncs of the same commit may run different images - so Git no longer fully determines cluster state."
  }
]
```


---

# When reconciliation bites: drift, waves, and secrets

GitOps is calm right up until it isn't. The same reconciliation loop that quietly keeps your cluster correct will, in the wrong situation, fight you, loop forever, or refuse to deploy in the order you needed. None of these are bugs - they're the loop doing exactly what you told it, when what you told it was incomplete. This phase is the set of gotchas that turn a 3am page into a 30-second fix once you've seen them before.

## Self-heal: the feature that fights you

`selfHeal: true` is wonderful - until you're mid-incident and trying to hand-patch the live cluster to test a theory. You `kubectl edit` the Deployment, and seconds later Argo CD reverts your change, because to it, your edit *is* drift. You and the controller are now in a quiet tug-of-war, and the controller never tires.

```console
$ kubectl scale deploy/payments --replicas=10   # emergency scale-up by hand
deployment.apps/payments scaled
# ...moments later, Argo CD reconciles...
$ kubectl get deploy payments
NAME       READY
payments   3/3                                   # back to what Git says
```

*What just happened:* self-heal saw the live replica count (10) diverge from Git (3) and pulled it back. Your manual change evaporated. The fix isn't to disable self-heal in a panic - it's to make the change *in Git*, or temporarily disable auto-sync for that one app:

```console
$ argocd app set payments --sync-policy none   # pause automation
# ...do your manual experiment, then re-enable...
$ argocd app set payments --sync-policy automated
```

The lesson: in a self-heal world, the cluster is read-only to humans. The keyboard you type production changes on is the one that edits Git. Treat any live edit as temporary at best.

## Sync waves: order is not free

A naive sync applies everything in a commit at once. Usually fine. But some resources *must* come up before others: a database migration Job before the app that needs the new schema, a namespace before the things inside it, a CRD before the custom resources that use it. Apply them all at once and the dependents fail because what they need isn't ready yet.

Argo CD orders a sync with **sync waves** - an annotation that sorts resources into ordered groups. Lower waves apply (and become healthy) before higher waves start.

```yaml
metadata:
  annotations:
    argocd.argoproj.io/sync-wave: "1"   # runs before wave 2
```

```text
Wave 0:  Namespace, CRDs            (foundations)
Wave 1:  ConfigMap, Secret, DB Job  (prerequisites)
Wave 2:  Deployment, Service        (the app itself)
```

*What just happened:* Argo CD applies wave 0, waits for it to be healthy, then wave 1, then wave 2. The migration Job (wave 1) finishes before the Deployment (wave 2) starts, so the app never boots against an un-migrated database.

There's a sharp variant: **hooks**, like a `PreSync` Job that must succeed before the rest applies. Use them for things that should run *every* deploy - a migration, a smoke test - not only on first creation. The trap is leaving a hook Job around so it can't recreate; annotate hooks with a delete policy so they clean up, or the next sync stalls on a leftover.

## Pruning: the blast radius of deleting a file

`prune: true` means "if it's gone from Git, delete it from the cluster." That's the correct, tidy behavior - and also a foot-gun. Delete the wrong folder in a refactor, merge it, and Argo CD will faithfully delete the corresponding live resources. The source of truth said remove them, so it removed them.

Two guardrails are worth knowing:

- **Review prune carefully.** A diff that *removes* manifests is a diff that *deletes production resources*. Treat deletions in infra PRs with the same care as a `DROP TABLE`.
- **`Prune=false` on the resource you can't afford to lose.** A PersistentVolumeClaim, for instance, can carry a `Prune=false` annotation so an accidental removal from Git leaves the live volume alone.

```yaml
metadata:
  annotations:
    argocd.argoproj.io/sync-options: Prune=false   # never auto-delete this
```

*What just happened:* even if this resource vanishes from Git, Argo CD will report it as out of sync rather than deleting it - buying you a human decision before data disappears.

## Secrets: the one thing you can't commit raw

Here's the contradiction at the center of GitOps: everything lives in Git, but database passwords and API keys absolutely cannot live in Git in plaintext. A public repo would leak them; even a private one bakes them into history forever. So how do secrets fit the "Git is the source of truth" model?

The answer is to commit secrets *encrypted*, so what's in Git is useless without a key the cluster holds. The common approaches:

- **Sealed Secrets** - you encrypt a secret with a public key into a `SealedSecret` that's safe to commit; only the controller in the cluster can decrypt it. The encrypted blob is your Git-tracked desired state.
- **External Secrets Operator** - Git holds only a *reference* ("fetch `prod/db-password` from Vault/AWS Secrets Manager"), and an operator pulls the real value into the cluster at runtime.

```text
Git (committed, safe):  SealedSecret { encryptedData: AgB7x9...== }
                                   │  Argo CD applies it
                                   ▼
Cluster controller decrypts ──►  Secret { password: <plaintext> }  (never in Git)
```

*What just happened:* the encrypted form is the source of truth in Git and reconciles like any other resource. Decryption happens only inside the cluster, so the secret's plaintext never appears in a commit, a PR, or a repo's history.

> The wrong fix is to apply secrets out-of-band with `kubectl` and leave them out of Git. Now you've reintroduced exactly the drift GitOps exists to kill - a piece of production state that isn't in the source of truth and that self-heal can't protect. Encrypt and commit; don't smuggle.

## Stuck OutOfSync forever

A maddening failure mode: an app that's perpetually OutOfSync no matter how often it syncs. Usually it's not a sync problem at all - it's that something *else* is mutating the resource after Argo CD applies it. A mutating admission webhook injects a sidecar, an autoscaler rewrites the replica count, a defaulting controller adds fields. Argo CD applies Git's version, the other actor changes it back, and the loop never converges.

The fix is to tell Argo CD to *ignore* the field that legitimately differs, so it stops counting a managed difference as drift:

```yaml
spec:
  ignoreDifferences:
    - group: apps
      kind: Deployment
      jsonPointers:
        - /spec/replicas        # an autoscaler owns this; don't fight it
```

*What just happened:* Argo CD now reconciles everything *except* `replicas`, leaving that field to the autoscaler. The endless OutOfSync tug-of-war stops because the two actors no longer claim ownership of the same field.

This is the deeper lesson of running GitOps in production: the reconciliation loop assumes Argo CD is the *only* writer. Wherever something else also writes - autoscalers, webhooks, operators - you have to draw an explicit boundary, or the loop spins forever trying to win a fight it was never meant to be in. If your cluster's foundations feel shaky underneath all this, /guides/kubernetes-without-the-hype is the place to firm them up.

```quiz
[
  {
    "q": "You hand-edit a live Deployment during an incident and Argo CD keeps reverting it. Why?",
    "choices": [
      "The cluster is out of disk space",
      "selfHeal treats your manual edit as drift and reconciles the cluster back to Git",
      "Your kubectl context is wrong",
      "Argo CD has crashed and is restarting"
    ],
    "answer": 1,
    "explain": "With selfHeal on, any divergence from Git - including your manual edit - is drift, so Argo CD pulls it back. Change Git, or pause auto-sync."
  },
  {
    "q": "A migration Job must finish before the app Deployment starts. What ensures that order?",
    "choices": [
      "Listing the Job first in the YAML file",
      "Sync waves (or a PreSync hook) - lower waves apply and become healthy before higher waves",
      "Running two separate Applications",
      "Setting prune: false on the Deployment"
    ],
    "answer": 1,
    "explain": "Sync waves sort resources into ordered groups; Argo CD waits for each wave to be healthy before starting the next."
  },
  {
    "q": "How do secrets fit the 'everything in Git' model without leaking?",
    "choices": [
      "Commit them in plaintext but only to a private repo",
      "Apply them with kubectl and keep them out of Git",
      "Commit them encrypted (e.g. Sealed Secrets), or commit only a reference the cluster resolves at runtime",
      "Base64-encode them, which is encryption"
    ],
    "answer": 2,
    "explain": "Git holds the encrypted form or a reference; the cluster decrypts or fetches the real value, so plaintext never lands in a commit."
  }
]
```
