# AWS CloudFormation

> AWS's native infrastructure as code: declare resources in a template, and CloudFormation creates, updates, and rolls back the whole stack as one unit.


---

# AWS CloudFormation

You have a pile of AWS resources that belong together: a bucket, a queue, a role, a function. Today they live in your head and your click history in the console. Tomorrow you need the same thing in staging, or you need to tear it all down cleanly, or a teammate needs to know what exists and why. CloudFormation lets you write that pile down as one file, hand it to AWS, and get back a managed group of resources that creates, updates, and deletes as a single unit. No more guessing what's running. No more orphaned resources nobody dares touch.

## How to read this

Three phases, in order. Phase 1 builds the mental model: what a template and a stack actually are, and why a managed unit beats a folder of scripts. Phase 2 is the everyday work: writing templates with parameters and outputs, previewing changes before they land, and wiring resources together. Phase 3 is production reality: rollbacks, drift, the failure modes that bite, and a clear-eyed look at where Terraform wins instead.

If you have never touched infrastructure as code, the broader idea is worth a detour: see [/guides/infrastructure-as-code-terraform](/guides/infrastructure-as-code-terraform) for the tool-agnostic concept. And if AWS itself is new, [/guides/cloud-platforms-explained](/guides/cloud-platforms-explained) sets the scene.

## The phases

1. [The mental model: templates and stacks](01-templates-and-stacks.md) - what you're actually declaring, and why a stack is one unit.
2. [Writing and changing stacks for real](02-writing-and-changing-stacks.md) - parameters, outputs, intrinsic functions, and change sets.
3. [When it breaks: rollback, drift, and Terraform](03-rollback-drift-and-reality.md) - the gotchas, the recovery moves, and where each tool wins.


---

# The mental model: templates and stacks

Picture how most AWS environments actually start. Someone opens the console, clicks through a wizard to make an S3 bucket, clicks again for an SQS queue, attaches an IAM role by hand, and wires a Lambda to it. It works. Then three weeks later you need the same setup in a second region, or a new hire asks "what's actually running in this account?", and the real answer is "nobody fully knows." The knowledge lives in click history and tribal memory. That's the pain CloudFormation exists to kill.

The core idea is two words: **template** and **stack**. Get those two clear and everything else is detail.

## A template is the recipe

A template is a text file - YAML or JSON - that *describes* the AWS resources you want. Not the clicks to make them. The end state. You say "I want a bucket named like this, a queue with this retention, a role with these permissions," and CloudFormation figures out the order to create them in and the calls to make.

Here's a small but complete template that creates one S3 bucket:

```yaml
# template.yaml
AWSTemplateFormatVersion: "2010-09-09"
Description: A single private bucket

Resources:
  AppBucket:
    Type: AWS::S3::Bucket
    Properties:
      BucketName: missing-manual-demo-bucket
      PublicAccessBlockConfiguration:
        BlockPublicAcls: true
        BlockPublicPolicy: true
        IgnorePublicAcls: true
        RestrictPublicBuckets: true
```

*What just happened:* you declared one resource. `AppBucket` is a **logical ID** - a name *you* pick that's local to this template. `Type` says what AWS thing it is. `Properties` are the settings. Notice you never wrote "create" or "make an API call" anywhere - the file is a description, not a script.

That's the whole shift in thinking. A shell script that calls the AWS CLI is a list of *steps*. A CloudFormation template is a *declaration of the result*. If you run a script twice you get two buckets (or an error). If you apply a template twice, CloudFormation sees the bucket already matches and does nothing. This property - same input, same end state, no matter how many times you run it - is **idempotence**, and it's the foundation of every infrastructure-as-code tool. The broader idea is covered tool-agnostically in [/guides/infrastructure-as-code-terraform](/guides/infrastructure-as-code-terraform).

## A stack is the running unit

When you hand that template to CloudFormation, what comes back is a **stack**. The stack is the live, managed group of resources that the template produced. This is the part people underestimate, so let it land: the stack is a *unit*. CloudFormation tracks every resource it created and the relationships between them.

That unit-ness buys you three things you cannot get from a folder of scripts:

- **Delete the stack, delete everything in it.** One command tears down the bucket, the queue, the role, the function - in the right order, cleaning up dependencies. No orphans.
- **CloudFormation knows what it owns.** It keeps a record of every resource and its current state. Ask it and it tells you exactly what's in the stack.
- **An update is a diff, not a redo.** Change the template and re-apply, and CloudFormation works out what actually changed and touches only that.

```mermaid
flowchart LR
  T[template.yaml] -->|create-stack| CF[CloudFormation]
  CF -->|provisions| S[Stack]
  S --> B[S3 Bucket]
  S --> Q[SQS Queue]
  S --> R[IAM Role]
  CF -.->|delete-stack| X[All removed in order]
```

*What just happened:* the template flows in once, and from then on the stack is the thing you operate on. You don't manage the bucket and the queue separately - you manage the stack, and the stack manages them.

## Why declare it at all

If you've only ever clicked through consoles, the payoff might still feel abstract. Three concrete wins:

**Repeatability.** The same template makes the same stack in `us-east-1`, in `eu-west-1`, in a teammate's sandbox account. Environments stop drifting apart because they come from one source.

**Review.** The template is a file. It lives in Git. A change to your infrastructure becomes a pull request someone can read, comment on, and approve - *before* it touches a real account. Infrastructure changes get the same safety net as code changes.

**Truth.** The template is the source of truth for what exists. Six months from now, the answer to "what's running?" is "read the template," not "spelunk through the console and hope."

> [!NOTE]
> CloudFormation is AWS's *native* tool - it's part of AWS, free to use (you pay only for the resources it creates), and it knows every AWS service the moment that service launches. That nativeness is its biggest strength and, as you'll see in Phase 3, the root of its main limitation too.

## For builders

You don't have to write a template to read one. When you adopt an AWS feature, search for its CloudFormation resource type - `AWS::Lambda::Function`, `AWS::DynamoDB::Table`, `AWS::EC2::Instance`. Every property maps to something you'd otherwise set by clicking. Reading the resource reference for a service is one of the fastest ways to learn what that service can actually do, because the template surface *is* the configuration surface.

```quiz
[
  {
    "q": "What is the relationship between a template and a stack?",
    "choices": [
      "They are two names for the same file",
      "A template describes the desired resources; a stack is the live managed group CloudFormation creates from it",
      "A stack is the YAML file and a template is the running infrastructure",
      "A template is for staging and a stack is for production"
    ],
    "answer": 1,
    "explain": "The template is the declarative recipe (text); the stack is the resulting managed unit of real resources."
  },
  {
    "q": "What does the logical ID (e.g. AppBucket) refer to?",
    "choices": [
      "The real AWS resource name, which must be globally unique",
      "A name local to the template that you use to reference the resource within it",
      "The AWS account ID the resource belongs to",
      "The region the resource is created in"
    ],
    "answer": 1,
    "explain": "Logical IDs are template-local names. The real (physical) name is either set in Properties or auto-generated by AWS."
  },
  {
    "q": "Why is a stack being a single unit useful?",
    "choices": [
      "It makes each resource cheaper to run",
      "Deleting the stack cleanly removes all its resources in dependency order, with no orphans",
      "It lets you skip IAM permissions",
      "It automatically encrypts every resource"
    ],
    "answer": 1,
    "explain": "Because CloudFormation tracks the whole group, it can create, update, and tear it down as one coherent unit."
  }
]
```


---

# Writing and changing stacks for real

The single-bucket template in Phase 1 was plain but lonely. Real stacks have parts that talk to each other, values that change between environments, and outputs other systems need to read. This phase is the day-to-day craft: making a template flexible with parameters, wiring resources together with intrinsic functions, handing results out with outputs, and - the habit that will save you most often - previewing every change before it lands.

## Parameters: one template, many environments

You don't want a separate template for staging and production that differ by one bucket name. You want one template with knobs. Those knobs are **parameters**.

```yaml
AWSTemplateFormatVersion: "2010-09-09"
Description: Bucket with an environment knob

Parameters:
  EnvName:
    Type: String
    Default: staging
    AllowedValues: [staging, production]
    Description: Which environment this stack is for

Resources:
  AppBucket:
    Type: AWS::S3::Bucket
    Properties:
      BucketName: !Sub "missing-manual-${EnvName}-data"
```

*What just happened:* `Parameters` declares an input named `EnvName` that defaults to `staging` and refuses anything but the two allowed values. Down in the bucket, `!Sub` substitutes the parameter into a string, so the bucket comes out named `missing-manual-staging-data` or `missing-manual-production-data` depending on what you pass.

You pass parameters when you create or update the stack:

```bash
aws cloudformation deploy \
  --template-file template.yaml \
  --stack-name app-prod \
  --parameter-overrides EnvName=production
```

*What just happened:* the `deploy` command sends the template plus your parameter values to AWS. `deploy` is the friendly wrapper - it creates the stack if it doesn't exist and updates it if it does, so you run the same command every time.

## Intrinsic functions: how resources reference each other

`!Sub` is one of CloudFormation's **intrinsic functions** - the small built-in operations that let a static text file express relationships. You'll reach for a handful constantly:

- **`!Ref`** - get the value of a parameter, or the physical name/ID of a resource.
- **`!GetAtt`** - get a specific *attribute* of a resource (an ARN, a URL, an endpoint).
- **`!Sub`** - substitute variables into a string.
- **`!Join`** - glue a list of strings together with a separator.

Here's a queue and a role where the role's policy points at the queue:

```yaml
Resources:
  Jobs:
    Type: AWS::SQS::Queue
    Properties:
      MessageRetentionPeriod: 345600   # 4 days, in seconds

  WorkerRole:
    Type: AWS::IAM::Role
    Properties:
      AssumeRolePolicyDocument:
        Version: "2012-10-17"
        Statement:
          - Effect: Allow
            Principal: { Service: lambda.amazonaws.com }
            Action: sts:AssumeRole
      Policies:
        - PolicyName: read-jobs
          PolicyDocument:
            Version: "2012-10-17"
            Statement:
              - Effect: Allow
                Action: [sqs:ReceiveMessage, sqs:DeleteMessage]
                Resource: !GetAtt Jobs.Arn   # <-- the link
```

*What just happened:* `!GetAtt Jobs.Arn` pulls the queue's ARN - a value that doesn't exist until AWS creates the queue. CloudFormation reads that reference, realizes the role *depends on* the queue, and creates the queue first. **Dependencies are inferred from references.** You almost never order resources by hand; you reference them and CloudFormation works out the graph.

> [!TIP]
> If two resources genuinely depend on each other but don't reference one another, use `DependsOn:` to state the order explicitly. Reach for it only when an implicit reference can't express the relationship - overusing it turns a clean dependency graph into hand-maintained ordering.

## Outputs: handing values back out

A stack often produces values other people or systems need - a bucket name, a queue URL, an endpoint. Don't make them dig through the console. Declare **outputs**.

```yaml
Outputs:
  JobsQueueUrl:
    Description: URL of the jobs queue
    Value: !Ref Jobs
  JobsQueueArn:
    Value: !GetAtt Jobs.Arn
    Export:
      Name: shared-jobs-queue-arn
```

*What just happened:* after the stack settles, `JobsQueueUrl` and `JobsQueueArn` show up in the stack's Outputs, queryable by CLI or console. The `Export` on the second one publishes it account-wide so a *different* stack can import it - that's how you share a value across stacks without copy-pasting an ARN.

## Change sets: look before you leap

This is the habit that separates calm operators from people who break production on a Friday. When you change a template, you don't have to apply it blind. A **change set** is a dry run: CloudFormation computes exactly what it *would* do and shows you, and nothing happens until you say go.

```bash
# 1. Create a change set from your edited template
aws cloudformation create-change-set \
  --stack-name app-prod \
  --change-set-name bump-retention \
  --template-body file://template.yaml \
  --parameters ParameterKey=EnvName,ParameterValue=production

# 2. See what it plans to do
aws cloudformation describe-change-set \
  --stack-name app-prod --change-set-name bump-retention
```

A trimmed view of what comes back:

```text
Changes:
  - ResourceChange:
      Action: Modify
      LogicalResourceId: Jobs
      ResourceType: AWS::SQS::Queue
      Replacement: False        # <-- in-place edit, not a rebuild
      Details: [ MessageRetentionPeriod ]
```

*What just happened:* CloudFormation tells you it will *modify* the `Jobs` queue and, critically, `Replacement: False` - it'll edit in place rather than destroy and recreate. That `Replacement` flag is the one to read every time. Some property changes force a replacement, which for a database or a queue can mean data loss or a new endpoint. The change set surfaces that *before* you commit.

When the plan looks right, execute it:

```bash
aws cloudformation execute-change-set \
  --stack-name app-prod --change-set-name bump-retention
```

*What just happened:* now - and only now - does CloudFormation make the change. The console's "preview changes" button does the same thing under the hood. Treat the change set as your seatbelt: cheap to create, free to throw away, and the only reliable preview of what an update will really do.

## In the wild

A common, sane workflow: keep templates in Git, run `create-change-set` in CI on every pull request, and post the described changes as a comment so reviewers see the blast radius before approving. The merge then runs `execute-change-set`. Infrastructure changes get reviewed exactly like code, with a real diff attached. This is the same review-before-apply loop you'd build around any IaC tool - see [/guides/infrastructure-as-code-terraform](/guides/infrastructure-as-code-terraform) for how the pattern looks elsewhere.

```quiz
[
  {
    "q": "How does CloudFormation usually decide the order to create resources?",
    "choices": [
      "Alphabetically by logical ID",
      "Top to bottom as written in the template",
      "From references between resources (e.g. !GetAtt and !Ref build a dependency graph)",
      "It creates everything simultaneously regardless of dependencies"
    ],
    "answer": 2,
    "explain": "References imply dependencies, so CloudFormation infers the order. DependsOn is only for cases a reference can't express."
  },
  {
    "q": "What is a change set?",
    "choices": [
      "A backup of the current stack",
      "A dry-run preview of what an update would do, which you must execute separately to apply",
      "A list of parameters with default values",
      "A way to delete several stacks at once"
    ],
    "answer": 1,
    "explain": "A change set computes and shows the planned changes; nothing happens until you execute it."
  },
  {
    "q": "In a change set, why does the Replacement field matter most?",
    "choices": [
      "It shows how much the change will cost",
      "Replacement: True means the resource will be destroyed and recreated, which can cause data loss or a new endpoint",
      "It indicates whether the template is valid YAML",
      "It controls which region the change applies to"
    ],
    "answer": 1,
    "explain": "A replacement rebuilds the resource. For stateful resources that can mean lost data or changed identifiers, so always check it."
  }
]
```


---

# When it breaks: rollback, drift, and Terraform

Everything so far assumed the happy path. Production is where templates meet reality: an update half-applies, someone hotfixes a resource by hand, a stack gets wedged in a state you can't escape. This phase is the survival kit - what CloudFormation does automatically when things go wrong, how to spot the changes you didn't make, the stuck states and how to get out of them, and a clear-eyed comparison with Terraform so you pick the right tool instead of defending the one you know.

## Automatic rollback: the safety net you'll meet first

CloudFormation treats a stack operation as a transaction. If creating or updating a stack fails partway, it doesn't leave you with a half-built mess - it **rolls back**, undoing what it did to return the stack to the last known-good state.

```text
Events (most recent first):
  UPDATE_ROLLBACK_COMPLETE      app-prod
  UPDATE_ROLLBACK_IN_PROGRESS   app-prod   The following resource(s) failed to update: [Jobs]
  UPDATE_FAILED                 Jobs       Resource handler returned message: "Invalid retention period"
  UPDATE_IN_PROGRESS            Jobs
```

*What just happened:* the update to `Jobs` failed on a bad property, so CloudFormation reverted the whole update and the stack landed back at `UPDATE_ROLLBACK_COMPLETE` - its previous working configuration. You read these stack **events** bottom-to-top to find the *first* failure; everything above it is just the cleanup. The first `UPDATE_FAILED` line is almost always your real error message.

This is genuinely great default behavior. It also has a sharp edge: a rolled-back stack is back to working, but your *change* didn't apply. Don't celebrate the green `ROLLBACK_COMPLETE` as success - it means "I undid your change safely," not "your change worked."

## Drift: the gap between template and reality

The template is supposed to be the source of truth. Then someone opens the console at 2 a.m. during an incident, flips a setting on a resource by hand, and forgets to put it back in the template. Now the live resource and the template disagree. That gap is **drift**, and it quietly erodes the whole promise of infrastructure as code - because the template no longer describes what's actually running.

CloudFormation can detect it for you:

```bash
# Kick off a drift check (asynchronous)
aws cloudformation detect-stack-drift --stack-name app-prod

# Then read the per-resource results
aws cloudformation describe-stack-resource-drifts \
  --stack-name app-prod --stack-resource-drift-status-filters MODIFIED
```

A trimmed result:

```text
StackResourceDriftStatus: MODIFIED
LogicalResourceId: AppBucket
PropertyDifferences:
  - PropertyPath: /VersioningConfiguration/Status
    ExpectedValue: Suspended
    ActualValue: Enabled
    DifferenceType: NOT_EQUAL
```

*What just happened:* CloudFormation compared the live bucket to the template and found versioning was turned on by hand. The template still says `Suspended`. The fix is not to update the console again - it's to decide which side is right, put that truth into the template, and re-apply so the two agree. Run drift detection on a schedule; finding drift early is the difference between a one-line correction and an archaeology project.

> [!WARNING]
> CloudFormation detects drift but does not automatically *fix* it. And it doesn't see everything - some resource types and some property kinds aren't covered by drift detection. Treat a clean drift report as "no detected drift," not a guarantee of none.

## Stuck stacks and the moves that free them

A few failure states are common enough to keep in your back pocket:

- **`ROLLBACK_COMPLETE` after a failed *create*.** A brand-new stack that failed to create lands here and can't be updated - only deleted. Delete it, fix the template, create again.
- **`UPDATE_ROLLBACK_FAILED`.** The rollback itself couldn't finish (often a resource it can't revert). Use `continue-update-rollback`, optionally skipping the resources it's stuck on, to get back to a stable state.
- **Stuck `*_IN_PROGRESS` for a long time.** Usually a resource waiting on something that will never happen (a wait condition, a missing dependency, an IAM permission). Read the events for the resource that's hanging - the answer is almost always there.

```bash
# Nudge a wedged rollback past resources it can't revert
aws cloudformation continue-update-rollback \
  --stack-name app-prod \
  --resources-to-skip StuckResourceLogicalId
```

*What just happened:* you told CloudFormation to finish the rollback while skipping the resource it couldn't handle, returning the stack to a stable `UPDATE_ROLLBACK_COMPLETE` so you can try again. Skipping isn't free - that resource may now be inconsistent with the template, so reconcile it afterward.

## A deletion guard worth knowing

Because deleting a stack deletes everything in it, one habit prevents the worst day: protect the resources you can't afford to lose.

```yaml
Resources:
  CriticalDB:
    Type: AWS::RDS::DBInstance
    DeletionPolicy: Retain          # keep the DB even if the stack is deleted
    UpdateReplacePolicy: Retain     # keep the old one if an update forces replacement
    Properties:
      Engine: postgres
      AllocatedStorage: 20
      DBInstanceClass: db.t3.micro
```

*What just happened:* `DeletionPolicy: Retain` tells CloudFormation to leave this database alone if the stack is deleted, and `UpdateReplacePolicy: Retain` keeps the old instance if a change would otherwise rebuild it. For databases and anything holding state, set these on purpose. The default is to delete, and the default has ruined weekends.

## CloudFormation vs Terraform: where each wins

You can do infrastructure as code on AWS with CloudFormation (native) or [/guides/infrastructure-as-code-terraform](/guides/infrastructure-as-code-terraform) (third-party, from HashiCorp). Both are good. They win in different places, and choosing well matters more than loyalty.

**CloudFormation wins when:**

- **You're all-in on AWS.** It's native, free to use, needs no separate state file to host or lock, and supports new AWS features the day they ship.
- **You want managed rollback and drift built in.** Both are first-class, no extra tooling.
- **You're deep in the AWS ecosystem** - service catalog, StackSets across accounts, organizations integration all assume CloudFormation.

**Terraform wins when:**

- **You're multi-cloud or use third-party services.** One language and workflow covers AWS, other clouds, Cloudflare, Datadog, GitHub, and more. CloudFormation only speaks AWS.
- **You value the language and ecosystem.** HCL with modules, a huge registry of reusable modules, and `terraform plan` as a fast, readable preview many engineers find friendlier than change sets.
- **You want explicit state you control** - though that state file is also a thing you must store, lock, and secure, which CloudFormation spares you.

The plain summary: if your world is one AWS account or org and likely to stay that way, CloudFormation's nativeness and zero state-management overhead are a real edge. The moment a second provider enters the picture, Terraform's single workflow usually wins. Plenty of shops run both - CloudFormation for AWS-only foundations, Terraform where the estate spans providers. Pick for the estate you actually have, not the one on a slide.

## In the wild

A pragmatic split many teams settle on: CloudFormation (often via the higher-level AWS SAM or CDK, which both compile down to CloudFormation templates) for tightly AWS-coupled stacks like serverless apps, and Terraform for the cross-cutting platform layer. The shared discipline underneath both is the same - template in Git, preview every change, detect drift on a schedule, protect stateful resources from accidental deletion.

```quiz
[
  {
    "q": "Your stack update fails and ends at UPDATE_ROLLBACK_COMPLETE. What does that mean?",
    "choices": [
      "Your change applied successfully",
      "CloudFormation safely reverted the failed update, so your change did NOT apply",
      "The stack was deleted",
      "Drift was detected and fixed"
    ],
    "answer": 1,
    "explain": "Rollback returns the stack to its last working state. It's a safe undo, not a successful change."
  },
  {
    "q": "What is drift in CloudFormation?",
    "choices": [
      "A gradual increase in stack cost over time",
      "When live resources no longer match the template, usually from manual console changes",
      "A slow region-to-region replication delay",
      "The time it takes a stack to finish updating"
    ],
    "answer": 1,
    "explain": "Drift is the gap between the template (source of truth) and the actual resource state. CloudFormation can detect it but won't auto-fix it."
  },
  {
    "q": "Which scenario most favors Terraform over native CloudFormation?",
    "choices": [
      "An AWS-only serverless app that uses brand-new AWS features",
      "An estate spanning AWS plus Cloudflare, Datadog, and GitHub under one workflow",
      "A team that wants managed rollback with no extra tooling",
      "A team that wants to avoid hosting and locking a state file"
    ],
    "answer": 1,
    "explain": "Terraform's single workflow across many providers is its key edge. The other options are CloudFormation's strengths."
  }
]
```
