# Zero-Downtime Deploys

> Ship a new version without a maintenance window: rolling, blue-green, and canary deploys, plus the health checks and migrations that make them safe.


---

# Zero-Downtime Deploys

You've felt the dread: it's release night, you put up a "back in 30 minutes" page, and you cross your fingers while the new version goes live. Maybe it works. Maybe it doesn't, and now you're rolling back at 11pm with users watching the maintenance page. There's a better way, and it's not magic - it's a handful of patterns that let you swap a running system out from under live traffic without anyone noticing.

By the end of this guide, deploys stop being a held-breath event. You'll understand *why* the naive approach drops requests, the three strategies that fix it, and the one part everyone gets wrong - the database - so your next release is something you do on a Tuesday afternoon instead of a Saturday night.

## How to read this

- **Want it to finally click?** Read in order. Phase 1 shows why the obvious deploy breaks; Phase 2 covers the three strategies that fix it; Phase 3 tackles the hard part - migrations and health checks - that make all three actually safe.
- **Already deploying and only need the migration trick?** Jump to [Phase 3: The Hard Part - Migrations and Health](03-migrations-and-health.md). That's where most real outages hide.

## The phases

1. **[Why Naive Deploys Hurt](01-why-naive-deploys-hurt.md)** - the mental model: a deploy means *two versions of your code wanting to exist at the same moment*. Stop-and-replace drops that moment on the floor, and your users feel it.
2. **[The Three Strategies](02-the-three-strategies.md)** - rolling, blue-green, and canary. How each one moves traffic off the old version and onto the new without a gap, and when to reach for which.
3. **[The Hard Part - Migrations and Health](03-migrations-and-health.md)** - why the database is what actually bites you, the expand-then-contract pattern, and the health checks and connection draining that let a load balancer do its job.


---

# Why Naive Deploys Hurt

Here's the deploy you probably learned first, even if nobody named it. You SSH into the server, pull the new code, stop the app, and start it again. For the few seconds between "stop" and "ready," the server answers nothing. If a user clicks during those seconds, they get an error or a hang. Scale that to a busy site and a "few seconds" of nothing is a wall of failed requests.

The trap is that on your laptop, this *looks* fine. You restart, you refresh, the new version is there. You were the only user, and you weren't clicking during the gap. Production is different in exactly one way that matters: there is never a moment when nobody is clicking.

## The deploy is a moment when two versions want to exist

Strip away the tooling and every deploy is the same shape: right now, version 1 of your code is serving requests. You want version 2 to serve them instead. There is some instant where the switch happens - and the entire problem of zero-downtime deploys is *what occurs during that instant*.

The naive deploy handles it like this:

```text
v1 running ──► [STOP] ──► (nothing answers) ──► [START v2] ──► v2 ready
                            ^^^^^^^^^^^^^^^^
                            requests die here
```

*What just happened:* between stop and ready, there's a window where the port is closed or the app is still booting. Every request that arrives in that window fails. The window might be two seconds or thirty, depending on how slow your app is to start - and a slow-starting app makes the wound bigger.

## Why "it starts in a second" is a lie under load

Booting an app is not instant, and it gets slower exactly when you can least afford it. A real app, on startup, has to:

- Load and parse its code.
- Open a connection pool to the database.
- Warm caches, compile templates, or JIT hot paths.
- Run any startup checks before it's truly ready to answer.

```console
$ systemctl restart myapp
$ curl localhost:8080/health
curl: (7) Failed to connect to localhost port 8080: Connection refused
$ # ...wait...
$ curl localhost:8080/health
curl: (7) Failed to connect to localhost port 8080: Connection refused
$ # ...wait some more...
$ curl localhost:8080/health
{"status":"ok"}
```

*What just happened:* the process was "restarted" the instant the command returned, but the app wasn't *ready* for several more seconds. "Process is up" and "app can serve a request" are two different events, and the gap between them is pure downtime. This distinction - running versus ready - is the seed of everything in Phase 3.

## The fix is never "make the gap smaller" - it's "have no gap"

The instinct is to optimize the restart: faster boot, quicker swap, shave the window down. That's a losing game. Even a one-second gap drops requests under real traffic, and you can't get to zero by subtraction.

The actual fix is structural: **keep the old version serving until the new version is fully ready, then move traffic across - never tear down what's working before the replacement can take over.** To do that, you need something sitting *in front* of your app that decides where requests go.

## The piece that makes it possible: something in front

Every zero-downtime strategy depends on a layer between users and your app instances - a **load balancer** (or reverse proxy, or service mesh; same idea). Users talk to it; it forwards each request to one of your backend instances.

```text
                 ┌──────────────┐
   users ──────► │ load balancer│
                 └──────┬───────┘
                  ┌─────┴─────┐
                  ▼           ▼
              instance A   instance B   (your app, more than one copy)
```

*What just happened:* because the load balancer chooses the target per request, you can change *which* instances are healthy and in-rotation without users ever addressing your app directly. Take instance A out, upgrade it, put it back - the balancer routes around the gap. This indirection is the hinge the next phase turns on: you can't roll, flip, or canary anything if every user is wired straight to a single process.

> The implication people miss: zero-downtime deploys assume **more than one instance** and a router in front. A single box with one process can't be upgraded without *some* gap. If you're on one server, the first real step toward safe deploys is running at least two instances behind a balancer.

**For builders:** look at how your app is reachable right now. If users hit a single process directly (one `node server.js`, one container, one port), there's no seam to deploy through. The cheapest upgrade isn't a fancy tool - it's a second instance and a proxy (nginx, a cloud load balancer, your platform's built-in one) so that "take one down" stops meaning "take the site down."

```quiz
[
  {
    "q": "Why does a stop-and-replace deploy drop requests in production but seem fine in local testing?",
    "choices": [
      "Local machines are faster, so the gap is zero",
      "In production there is never a moment when no one is sending requests, so the stop-to-ready gap always catches some",
      "Production code is buggier than local code",
      "The database is slower in production"
    ],
    "answer": 1,
    "explain": "The gap exists in both places; locally you just aren't sending traffic during it. Under real load, requests always arrive in that window and fail."
  },
  {
    "q": "What is the difference the health gap exposes between 'process is up' and 'app is ready'?",
    "choices": [
      "There is no difference; they happen at the same instant",
      "A process can be running while still booting - opening pools, warming caches - and unable to serve requests yet",
      "'Ready' means the process has exited cleanly",
      "'Up' refers to the database, 'ready' refers to the cache"
    ],
    "answer": 1,
    "explain": "Startup work (pools, caches, compilation) happens after the process launches. Treating 'up' as 'ready' routes traffic to an app that can't answer yet."
  },
  {
    "q": "What structural piece must exist before any zero-downtime strategy can work?",
    "choices": [
      "A faster CPU to shrink the restart window",
      "A single very reliable server",
      "A load balancer (or proxy) in front of more than one instance, so traffic can be routed around an instance being upgraded",
      "A larger database connection pool"
    ],
    "answer": 2,
    "explain": "Zero-downtime deploys rely on indirection: a router in front of multiple instances lets you take one out, upgrade it, and return it without users hitting a gap."
  }
]
```


---

# The Three Strategies

Now that something sits in front of your instances, you have choices about *how* to move traffic from the old version to the new. There are three patterns worth knowing, and they're not really competitors - they trade off speed, cost, and how much risk you want to take in one swing. Most teams use more than one over time.

The thread running through all three: at no point is there a moment where the only thing that can serve a request is mid-restart. You always have a healthy version answering while the change happens.

## Rolling: replace a few at a time

A rolling update keeps your fleet mostly intact and swaps instances out in small batches. Take one (or a few) out of rotation, upgrade them, wait until they're healthy, put them back, then move to the next batch. Repeat until the whole fleet runs the new version.

```text
start:   [v1] [v1] [v1] [v1]      all old

step 1:  [v2] [v1] [v1] [v1]      one upgraded, rest still serving
step 2:  [v2] [v2] [v1] [v1]
step 3:  [v2] [v2] [v2] [v1]
done:    [v2] [v2] [v2] [v2]      all new, never fewer than 3 serving
```

*What just happened:* at every step, a majority of instances stay healthy and in-rotation, so there's always capacity to serve traffic. The new version proves itself healthy on one instance before the next one is touched. This is the default for Kubernetes and most container platforms because it needs no extra hardware - you reuse the instances you already have.

The catch with rolling: for a while, **both versions serve real traffic at the same time.** A user might hit v2 on one request and v1 on the next. If those versions disagree about the shape of the data or an API response, that user sees something inconsistent. Hold that thought - it's the whole reason Phase 3 exists.

## Blue-green: two full environments, flip the switch

Blue-green runs two complete copies of your environment. "Blue" is live and serving everyone. You deploy the new version to "green" - an idle full-size copy - and let it warm up and pass its checks with zero real traffic on it. When green looks good, you point the load balancer at green. One change, all traffic moves.

```text
before flip:   users ──► [ BLUE  v1 ]  (live)
                         [ GREEN v2 ]  (ready, no traffic)

after flip:    users ──► [ GREEN v2 ]  (live)
                         [ BLUE  v1 ]  (idle - kept warm for rollback)
```

*What just happened:* the flip is near-instant and atomic - traffic goes from all-blue to all-green in one routing change, so there's never a mix of versions serving at once. The superpower is rollback: if green misbehaves, you flip straight back to blue, which is still sitting there fully running. Recovery is seconds, not a redeploy.

The cost is in the name: you're paying for **two full environments** during the deploy. For a large fleet that's real money, even if green only exists for a short window. Blue-green also doesn't give you a gentle "try it on a few users first" - it's all or nothing.

## Canary: a small slice first, watch the numbers

A canary deploy sends a *small percentage* of traffic to the new version and keeps the rest on the old one. You watch your metrics - error rate, latency, the business numbers that matter - on that small slice. If it stays healthy, you raise the percentage in steps until it's serving everyone. If it goes bad, you've only exposed a few percent of users, and you pull it back.

```text
phase 1:   95% ──► [v1]      5% ──► [v2]   watch error rate, latency
phase 2:   75% ──► [v1]     25% ──► [v2]   still healthy? continue
phase 3:    0% ──► [v1]    100% ──► [v2]   full rollout
```

*What just happened:* the new version is tested against *real production traffic* but with a blast radius you control. The name comes from the canary in a coal mine - a small, early warning. The trade-off is that canary needs the most machinery: traffic-splitting at a percentage, and good enough metrics to actually tell "this canary is sick" from normal noise. Without solid observability, a canary is just a slower rollout you can't read.

## Which one, when

| Strategy   | Extra cost          | Rollback speed     | Both versions live at once? | Reach for it when |
|------------|---------------------|--------------------|-----------------------------|-------------------|
| Rolling    | None (reuse fleet)  | Roll back forward  | Yes, during the roll        | Default; you have many instances and limited budget |
| Blue-green | A second full env   | Instant (flip back)| No (atomic flip)            | Rollback speed matters most; you can afford the double |
| Canary     | Traffic-split + metrics | Fast (small slice) | Yes, by design          | Risky change; you want real-traffic proof on a few users first |

*What just happened:* there's no winner - there's a fit. Rolling is the cheap default. Blue-green buys you instant rollback at the price of a second environment. Canary buys you a tiny blast radius at the price of needing real observability. Teams often combine them: canary the risky releases, blue-green the ones where rollback speed is everything, roll the routine stuff.

> Notice what blue-green and canary lean on that rolling tries to skip: a clean way to **not mix versions**, or to mix them *on purpose and carefully*. Whenever two versions touch the same database, the strategy alone isn't enough - the data has to be ready for both. That's the hard part, and it's next.

**For builders:** start with what your platform already does. Kubernetes, ECS, and most PaaS offer rolling out of the box - you may already be doing zero-downtime deploys and not know it. Reach for blue-green or canary when a specific pain shows up: "rollback takes too long" points at blue-green; "that last release broke things for everyone before we noticed" points at canary. Don't build canary infrastructure you don't yet need - see [What CI/CD Does](/guides/what-cicd-does) for where these fit in the bigger release picture.

```quiz
[
  {
    "q": "During a rolling update, what is true about the versions serving traffic?",
    "choices": [
      "Only the old version serves until the very last second",
      "Both the old and new versions serve real traffic simultaneously for part of the rollout",
      "Neither version serves; there is a planned gap",
      "Only the new version serves from the first instant"
    ],
    "answer": 1,
    "explain": "Rolling swaps instances in batches, so for a stretch some instances run v1 and some run v2, both taking live requests - which is why version compatibility matters."
  },
  {
    "q": "What is the defining advantage of blue-green over rolling?",
    "choices": [
      "It uses no extra infrastructure",
      "It splits traffic by percentage for gradual testing",
      "Rollback is near-instant because the old environment is still fully running, ready to flip back to",
      "It guarantees the database never needs changes"
    ],
    "answer": 2,
    "explain": "Blue-green keeps the previous environment alive and idle, so a bad release is reverted by flipping the router back - seconds, not a redeploy."
  },
  {
    "q": "What does a canary deploy most depend on to be useful?",
    "choices": [
      "A second full copy of the environment",
      "Good observability - metrics that can tell a sick canary from normal noise on a small traffic slice",
      "The ability to stop all traffic during the deploy",
      "A single instance running the app"
    ],
    "answer": 1,
    "explain": "Canary exposes the new version to a small percentage of real traffic; without solid metrics you can't read whether that slice is healthy, so it's just a slow rollout."
  }
]
```

Watch it animated: [zero-downtime deployment strategies](/explainers/Deployments.dc.html)


---

# The Hard Part - Migrations and Health

You picked a strategy. Traffic moves smoothly from old to new. And then a release still takes the site down - because the *code* deployed cleanly but the *database* didn't get the memo. This is where most real outages live, and it's the part the shiny deploy tools can't do for you. Two topics: making schema changes that two code versions can survive, and making sure the balancer only ever sends traffic to instances that can actually answer.

## The trap: your migration and your code can't both win

Remember from Phase 2 that rolling and canary have old and new code running *at the same time*. Now add a schema change to the mix. Say the new version renames a column `name` to `full_name`. You write one migration: `ALTER TABLE users RENAME COLUMN name TO full_name`.

Watch what happens during the rollout:

```text
migration runs ──► column is now `full_name`
                   but old (v1) code still running, still doing SELECT name ...
                   v1 query: ERROR - column "name" does not exist
```

*What just happened:* the instant the rename lands, every still-running v1 instance starts throwing errors, because it's asking for a column that no longer exists. You didn't deploy zero-downtime; you deployed an outage with extra steps. Reverse it (migrate after the code) and now *new* code asks for `full_name` before it exists. There is no single ordering of "one big migration + swap the code" that avoids a broken window. The premise is wrong, not the order.

## The fix: expand, then contract

The way out is to stop thinking of a schema change as one step. Split it so that **at every moment, the database supports both the old code and the new code at once.** This is the expand-contract pattern (also called parallel change), and the rename above becomes a sequence of *separately deployed* steps:

1. **Expand** - add the new thing without removing the old. Add a `full_name` column; keep `name`. Deploy this migration alone. Old code still uses `name` and is perfectly happy.
2. **Migrate code (and backfill)** - deploy code that *writes to both* columns and reads the new one. Backfill `full_name` from existing `name` values. Now both columns are populated and current.
3. **Contract** - once no running code reads or writes `name`, deploy a final migration that drops it.

```sql
-- Step 1 (deploy alone): expand. Old code untouched, still uses `name`.
ALTER TABLE users ADD COLUMN full_name TEXT;

-- Step 2 (with the new code release): backfill existing rows.
UPDATE users SET full_name = name WHERE full_name IS NULL;
-- ...and the new code writes BOTH name and full_name on every save,
--    while reading full_name. Old code, if any is still up, still reads name.

-- Step 3 (a later, separate deploy): contract, once nothing reads `name`.
ALTER TABLE users DROP COLUMN name;
```

*What just happened:* no single step ever breaks a running version. During step 1 only `name` is read; during step 2 both columns are valid and kept in sync; by step 3 nothing touches `name` anymore, so dropping it harms no one. The cost is being upfront about the calendar: a "rename" is now three deploys spread over time, not one. That patience is the entire price of zero downtime on a schema change.

> The rule to carry everywhere: **a migration must be backward-compatible with the code already running.** Adding a column, adding a nullable field, adding a new table - safe, because old code ignores what it doesn't know about. Removing or renaming a column, making a column `NOT NULL`, narrowing a type - destructive, because old code breaks. Destructive changes always become expand → migrate → contract.

A few more moves in the same spirit:

- **Adding a `NOT NULL` column?** Add it nullable with a default first, backfill, *then* add the constraint. A bare `NOT NULL` add fails on existing rows.
- **Renaming a table?** Same as a column: new table, dual-write, backfill, switch reads, drop old.
- **Big backfills?** Do them in batches so you don't lock the table and cause the very downtime you're avoiding.

## Health checks: how the balancer knows who's ready

Back in Phase 1 we saw the gap between "process is up" and "app is ready." Health checks are how the load balancer learns the difference. Your app exposes an endpoint the balancer polls; only instances that answer healthy get traffic. There are two flavors, and conflating them causes its own outages:

- **Liveness** - "is this process alive, or wedged and needing a restart?" If liveness fails, the orchestrator kills and restarts the instance.
- **Readiness** - "can this instance serve a request *right now*?" If readiness fails, the balancer stops sending it traffic but leaves it running.

```text
new instance boots ──► readiness = NO  (still opening DB pool, warming)
                       balancer sends it ZERO traffic
pool open, warm ─────► readiness = YES
                       balancer adds it to rotation
```

*What just happened:* readiness is what closes the Phase 1 gap. The instance starts, but the balancer holds traffic back until the app says "I'm actually ready," so no request lands on a half-booted process. A rolling deploy *relies* on this: it won't move to the next batch until the new instances report ready. A readiness check that returns OK too early - before the DB pool is open - reintroduces the exact downtime you're trying to kill.

> Make readiness mean it. A check that always returns `200 OK` regardless of whether the app can reach its database is theater - it tells the balancer "send traffic" to an instance that will then fail every request. Check the things you actually need to serve (DB reachable, critical deps up), but keep it cheap; the balancer hits it constantly.

## Draining: let the old instance finish what it started

The last gap is on the *way out*. When you take an instance down, it's probably in the middle of serving requests. Kill it instantly and those in-flight requests die. **Connection draining** (graceful shutdown) fixes this: the instance is removed from the balancer's rotation so it gets *no new* requests, but it's given a window to finish the ones already in progress before it actually stops.

```console
$ # orchestrator wants to stop instance A
1. mark A "not ready"     -> balancer stops routing NEW requests to A
2. A keeps serving the requests already in flight
3. wait for them to finish (up to a drain timeout)
4. only now: send SIGTERM, A exits cleanly
```

*What just happened:* nothing in flight got dropped. New traffic moved to other instances the moment A left rotation, and A's existing requests got to complete. Skip draining and even a perfect rolling deploy sheds a burst of errors on every instance you cycle - small, easy to miss in testing, very real to the users who hit it. Most orchestrators do this for you *if* your app handles `SIGTERM` by finishing work and closing cleanly instead of dying on the spot.

## Putting it together

A genuinely zero-downtime release is the strategy from Phase 2 *plus* this phase's discipline:

1. Ship schema changes expand-first, backward-compatible - never break the running version.
2. Roll/flip/canary the new code, with readiness checks gating each instance into rotation.
3. Drain old instances so in-flight requests finish.
4. Later, in a separate deploy, contract - drop what nothing uses anymore.

*What just happened:* every gap from Phase 1 is now closed - boot gap by readiness, shutdown gap by draining, and the data gap by expand-contract. The deploy stops being an event. It becomes a Tuesday.

**For builders:** before your next deploy, ask one question of every migration - "would the code currently in production survive this change?" If the answer is no, it's a destructive change and needs the expand → migrate → contract split. Then confirm your readiness check actually pings the database, and that your app finishes in-flight work on `SIGTERM`. Those three habits prevent the large majority of "but the deploy tool said it was zero-downtime" outages. For where this sits in the pipeline, see [Your First Pipeline (GitHub Actions)](/guides/your-first-pipeline-github-actions).

```quiz
[
  {
    "q": "Why is a single 'rename column name to full_name' migration unsafe during a rolling deploy?",
    "choices": [
      "Renames are always slow",
      "While both code versions run, one of them will reference a column that no longer exists (or doesn't exist yet), causing errors",
      "Databases cannot rename columns at all",
      "It locks the table forever"
    ],
    "answer": 1,
    "explain": "With old and new code live at once, a hard rename breaks whichever version expects the other name. The fix is expand-contract, where both names coexist during the transition."
  },
  {
    "q": "What is the correct order of the expand-contract pattern?",
    "choices": [
      "Drop the old column, then add the new one, then deploy code",
      "Add the new column (keep old), deploy code that dual-writes and backfills, then later drop the old column",
      "Rename the column, then deploy code, then backfill",
      "Add a NOT NULL column immediately, then deploy code"
    ],
    "answer": 1,
    "explain": "Expand (add new, keep old) → migrate (dual-write, backfill, read new) → contract (drop old once nothing uses it). Every step stays compatible with the code that's running."
  },
  {
    "q": "What does connection draining (graceful shutdown) accomplish when removing an instance?",
    "choices": [
      "It speeds up the database",
      "It stops new requests from going to the instance while letting in-flight requests finish before the process exits",
      "It returns 200 OK from the health check no matter what",
      "It makes the instance boot faster"
    ],
    "answer": 1,
    "explain": "Draining takes the instance out of rotation so it gets no new traffic, then gives it a window to complete requests already in progress before SIGTERM - so nothing in flight is dropped."
  }
]
```
