# Feature Flags and Rollbacks

> Ship code dark, flip it on for a few users, and turn it off instantly when it breaks: feature flags decouple deploy from release.


---

# Feature Flags and Rollbacks

You merged the feature. It deployed clean. And then the support tickets start, and your only lever is a frantic redeploy of last week's build while real users hit the broken thing. That whole panic exists because deploying your code and releasing your feature got welded into one irreversible act. This guide pries them apart. You'll learn to ship code that's switched off, turn it on for a handful of people, watch, and kill it in seconds when something's wrong - no rebuild, no scramble.

## How to read this

Read the three phases in order; each one builds the mental model the next assumes. Phase 1 reshapes how you think about "released." Phase 2 is the day-to-day mechanics. Phase 3 is the part nobody warns you about: the cost of flags and how to keep them from rotting your codebase. If you only have five minutes, read Phase 1 - the mental model is the whole game.

## The phases

1. [The switch in your code](01-deploy-is-not-release.md) - why deploy and release are two different things, and what a flag actually is.
2. [Living with flags](02-rollouts-and-kill-switches.md) - gradual rollouts, kill switches, A/B tests, and trunk-based development.
3. [The bill comes due](03-flag-debt-and-rollback.md) - flag debt, combinatorial testing, expiry, and rollback versus roll-forward.


---

# The switch in your code

Picture the worst version of a release. You merge a feature, the pipeline turns green, the new code goes live, and within minutes something is on fire for real users. Your only move is to rebuild the previous version and push it through the whole pipeline again while the clock runs. That panic isn't bad luck. It's the consequence of one quiet assumption baked into how most teams work: that putting code on the server and turning a feature on for users are the same event.

They don't have to be. Pulling them apart is the entire idea behind feature flags.

## Deploy and release are two different verbs

It's worth slowing down on these two words, because the whole guide hinges on the difference.

- **Deploy** means the new code is running on your servers. The binary is there. The functions exist in memory. It's *present*.
- **Release** means users can actually experience the feature. It's *active*.

Most teams treat these as one moment because, by default, they are. You merge code that adds a checkout button, you deploy, and the button is right there for everyone the instant the deploy finishes. Deploy *is* release. That coupling is exactly what turns a small bug into a five-alarm incident - the only way to un-release is to un-deploy.

A feature flag breaks the weld. It's a runtime switch wrapped around the new behavior, so you can deploy the code in the "off" position and decide *separately*, later, whether and when to flip it on.

> The one-line version: **deploy is a thing you do to servers; release is a thing you do to users.** Flags let you do them at different times.

## What a flag actually is

Strip away the marketing and a feature flag is humble: a named boolean your code checks before doing the new thing.

```js
// The new checkout flow is deployed, but gated behind a flag.
if (flags.isEnabled("new-checkout-flow")) {
  renderNewCheckout();   // the new code path
} else {
  renderOldCheckout();   // the old, known-good path
}
```

*What just happened:* both code paths are deployed and sitting on the server right now. Which one a user sees is decided at runtime by `isEnabled`, not at deploy time. Flip the flag's value somewhere else and behavior changes with no new deploy.

That's the mechanical heart of it. The new code shipped, but it's *dark* - present and inert. The old path is still doing its job. Nobody noticed the deploy because, as far as users are concerned, nothing released.

The value behind `isEnabled` doesn't live in the code. It lives somewhere you can change at runtime: a config file the app re-reads, a row in a database, an environment variable, or a hosted flag service. The key property is that changing it does **not** require a rebuild or a redeploy. That's the difference between a feature flag and an `if (false)` you have to edit and ship.

## Why this changes everything

Once deploy and release are separate, a pile of previously-scary things get calm.

- **The dreaded bug becomes a switch, not a rebuild.** Feature breaks in production? Flip the flag off. Users are back on the old path in seconds. No pipeline, no rollback build, no waiting. (Phase 2 calls this a kill switch.)
- **You can release to *some* people.** On for your own account first. Then 1% of users. Then 10%. Watch the dashboards between each step. (Phase 2: gradual rollout.)
- **Unfinished code can live on the main branch safely.** A half-built feature behind an off flag deploys with everything else and hurts nobody, because no one can reach it. That's what makes trunk-based development practical instead of terrifying.

```text
WITHOUT FLAGS:   merge → deploy → live for everyone   (one irreversible step)
WITH FLAGS:      merge → deploy (dark) → flip on for 1% → 10% → 100%
                                          ↑ any step reversible instantly
```

*What just happened:* the bottom row turned a single all-or-nothing leap into a series of small, reversible steps. Every arrow after "deploy" is a flag flip you can undo, not a release you have to chase down.

> Worth saying plainly: a flag is risk control, not a feature. Users never see it. Its only job is to give *you* a dial where there used to be a cliff.

## For builders

If your team runs CI/CD, flags are the missing safety layer on top of it. A green pipeline tells you the code *builds and tests pass* - it says nothing about whether the feature behaves under real traffic. (For how that pipeline works, see [/guides/what-cicd-does](/guides/what-cicd-does).) Flags let your pipeline deploy aggressively and often, while the release decision stays a small, human, reversible flip made *after* the code is safely in production. Fast deploys plus slow, controlled releases is the combination you're after.

```quiz
[
  {
    "q": "What is the core distinction a feature flag exploits?",
    "choices": [
      "The difference between writing code and reviewing it",
      "The difference between deploying code to servers and releasing a feature to users",
      "The difference between unit tests and integration tests",
      "The difference between staging and production environments"
    ],
    "answer": 1,
    "explain": "Deploy means the code is present on servers; release means users can experience it. Flags let those two happen at different times."
  },
  {
    "q": "Where does the value behind a feature flag need to live?",
    "choices": [
      "Hardcoded in the source, like if (false)",
      "Somewhere you can change at runtime without a rebuild or redeploy",
      "Only inside compiled binaries",
      "In the git commit message"
    ],
    "answer": 1,
    "explain": "If changing the flag required editing code and shipping, it would be no better than an if (false). The whole point is runtime control."
  },
  {
    "q": "What does it mean for code to be deployed 'dark'?",
    "choices": [
      "The code is on the server but gated off, so it runs nowhere users can reach",
      "The code is encrypted on disk",
      "The code is deployed only at night",
      "The code is deleted after deploy"
    ],
    "answer": 0,
    "explain": "Dark code is present and inert - deployed but switched off, hurting no one until you flip the flag."
  }
]
```


---

# Living with flags

You've got the mental model: code can be deployed and still dark. Now the practical question - what do you actually *do* with that switch? It turns out a single boolean, once you can target *who* sees it, covers four very different jobs that teams used to handle with four very different (and more painful) tools. Let's walk each one the way you'd hit them in a real week.

## The gradual rollout: don't release to everyone at once

You finished a feature. The tests pass. The instinct is to flip it on for all users and move on. Resist it. The flag lets you turn it on for a *slice* and grow the slice only if the dashboards stay calm.

A rollout is a sequence of flag states, not one flip:

```text
1. Enable for: my own account            → click around, sanity check
2. Enable for: internal team             → real humans, real data
3. Enable for: 1% of users               → watch error rate, latency, support
4. Enable for: 10%                        → still calm? keep going
5. Enable for: 100%                       → fully released
```

*What just happened:* you converted a binary launch into a dimmer switch. At each step the blast radius of a bug is capped - a problem at 1% touches 1% of users, and you caught it before it reached the other 99%. This is what people mean by a *canary release*: a small group goes first, like the canary in the mine.

The targeting logic usually lives in the flag service, but conceptually it's plain:

```js
function isEnabled(flagName, user) {
  const flag = lookupFlag(flagName);        // current config, fetched at runtime
  if (flag.allowList.includes(user.id)) return true;   // explicit early access
  return hashUserId(user.id) % 100 < flag.percentage;  // stable % bucketing
}
```

*What just happened:* hashing the user id and comparing against a percentage means the *same* user lands in the same bucket every call. That stability matters - without it, a user at 10% rollout would flicker between old and new on every page load, which feels broken even when nothing is wrong.

> Roll out by percentage of *users*, not percentage of *requests*. A single confused user seeing the new feature half the time is a worse experience than a clean 10% who consistently get it.

## The kill switch: the reason flags earn their keep

This is the payoff Phase 1 promised. The new feature is at 10%, the error rate spikes, and you don't reach for a rollback build. You flip one value.

```bash
# Feature is misbehaving in production. One command, no redeploy.
$ flagctl disable new-checkout-flow
✓ new-checkout-flow → OFF (propagating to all instances ~10s)
```

*What just happened:* every server re-reads the flag within seconds and routes all traffic back to the old, known-good path. No pipeline ran. No artifact was rebuilt. The mean-time-to-recovery dropped from "however long a deploy takes" to "however long a config propagation takes" - typically seconds.

A kill switch is most valuable around the things most likely to go wrong: a new third-party integration, an expensive query, a risky algorithm. Wrap those in a flag *specifically* so you have an off button, even if you never plan a gradual rollout. The flag exists to fail safely, not to release slowly.

## A/B tests: the flag as a measuring tool

The same targeting machinery answers a different question. Instead of "is this safe?" you ask "which version performs better?" Split users into groups, show each group a variant, and compare a metric.

```text
Group A (50%):  old button copy   → measured conversion: baseline
Group B (50%):  new button copy   → measured conversion: +4.1%
```

*What just happened:* the flag stopped being a pure on/off and became a *which-variant* selector. The mechanism is identical - stable per-user bucketing - but the goal shifted from risk control to learning. (One plain caveat: a real A/B test needs enough traffic and a proper significance check before you trust a difference. The flag delivers the split; it doesn't do the statistics for you.)

## Trunk-based development: flags as the unlock

Here's the workflow consequence that surprises people. When unfinished features can be deployed dark, long-lived feature branches stop being necessary. Everyone commits to the main branch ("trunk") in small pieces; half-built work hides behind an off flag and ships with every deploy, harming nobody because no one can reach it.

```text
Branch-heavy:  feature lives on a branch for 3 weeks → painful merge → big-bang release
Trunk-based:   small commits to main daily, feature behind an off flag → flip on when ready
```

*What just happened:* the bottom row eliminated the giant merge and the big-bang release. The work integrated continuously, in tiny increments, while staying invisible to users. The flag is what makes "merge unfinished code to main" sane instead of reckless.

## For builders

These four uses share one engine: a runtime-changeable value plus per-user targeting. You can start with the crudest possible version - a JSON file your app re-reads, or a single environment variable - and it'll genuinely work for a small team. Reach for a hosted flag service when you need per-user targeting, an audit trail of who flipped what, and instant propagation across many instances. Don't buy that complexity on day one; grow into it when a config file stops being enough.

```quiz
[
  {
    "q": "Why should percentage rollout bucket users with a stable hash rather than randomly per request?",
    "choices": [
      "Random is slower to compute",
      "So the same user consistently gets the same version instead of flickering between old and new",
      "Hashing encrypts the user id for privacy",
      "Random rollouts are not allowed by most flag services"
    ],
    "answer": 1,
    "explain": "Without stable bucketing a user at 10% would flip between old and new on every page load - that feels broken even when nothing is wrong."
  },
  {
    "q": "What is the main thing a kill switch reduces compared to a traditional rollback?",
    "choices": [
      "The number of tests you need to write",
      "The size of the codebase",
      "Mean-time-to-recovery - from 'a deploy length' down to a config propagation",
      "The cost of cloud hosting"
    ],
    "answer": 2,
    "explain": "A kill switch routes traffic back to the safe path via a config change in seconds, instead of waiting for a full rebuild-and-redeploy."
  },
  {
    "q": "How do feature flags make trunk-based development practical?",
    "choices": [
      "They prevent anyone from committing buggy code",
      "They let unfinished features deploy to main behind an off flag, invisible to users",
      "They automatically merge long-lived branches",
      "They remove the need for code review"
    ],
    "answer": 1,
    "explain": "Half-built work hides behind an off flag and ships with every deploy harming nobody, which removes the need for long-lived feature branches."
  }
]
```


---

# The bill comes due

Flags feel like free safety. They aren't. Every one you add is a fork in the road that lives in your code, and forks that nobody removes pile up into a quiet, compounding mess. This phase is the clear-eyed part: what flags cost, why old ones are dangerous, how to keep them from rotting your codebase, and the one decision you'll face in every incident - roll back or roll forward.

## Flag debt: the switches nobody turned off

A flag you added for a rollout three months ago is fully released - everyone's on the new path. But the `if (flags.isEnabled(...))` is still in the code, and so is the old `else` branch underneath it. That dead fork is *flag debt*: a switch that's done its job but never got removed.

```js
// Shipped 4 months ago. Flag is at 100% and will never be turned off.
if (flags.isEnabled("new-checkout-flow")) {
  renderNewCheckout();
} else {
  renderOldCheckout();   // dead code - but is it? nobody's sure anymore
}
```

*What just happened:* the `else` is unreachable in practice, but no one's certain enough to delete it, so it lingers. Multiply that by every rollout your team has ever done and the codebase fills with forks that obscure what the code actually does. Worse, an old forgotten flag is a live hazard - flip the wrong stale toggle and you've just re-enabled code that hasn't run in months and was never maintained.

## Combinatorial testing: flags multiply states

Here's the cost that scales worst. Each independent flag doubles the number of states your system can be in. Two flags is four combinations. Ten flags is over a thousand. You cannot test them all.

```text
1 flag   →  2 states     (on, off)
2 flags  →  4 states
3 flags  →  8 states
10 flags →  1024 states   ← you are not testing all of these
```

*What just happened:* the state space grows as 2 to the power of the flag count. In reality most flags aren't truly independent, so it's not quite that bad - but the direction is unforgiving. A combination you never tested *will* eventually occur in production, because some user's account hits exactly that mix of toggles. Fewer live flags is fewer untested combinations.

> The lesson isn't "avoid flags." It's that **every live flag is permanent inventory you have to carry.** The carrying cost is real, so the inventory must be actively managed down.

## Give every flag an expiry

The fix is a discipline, not a tool: a flag is born with a death date. The moment you add one, you decide when it dies.

- **Release flags** (gradual rollouts) are temporary by design. Once at 100% and stable, the flag *and the old branch* get deleted. That's the cleanup task, not an optional nicety.
- **Kill switches** for risky dependencies may be long-lived on purpose - but they're a deliberate, documented exception, not the default.
- Track expiry dates somewhere visible. Some teams make a stale flag fail the build; others run a recurring review. The mechanism matters less than the rule: **no flag lives forever by accident.**

```text
GOOD:  add flag → roll out → 100% stable → DELETE flag + old branch  (debt: 0)
BAD:   add flag → roll out → 100% → move on → ...repeat 40 times      (debt: 40 forks)
```

*What just happened:* the top row treats removal as part of the feature's lifecycle, so debt stays near zero. The bottom row treats the flag as done at 100%, and the forks accumulate until the codebase is a minefield of stale toggles.

## Rollback versus roll-forward

When something breaks in production, you face one fork in the road, and a flag changes which one is even available.

**Rollback** means going back to a known-good state. With a flag, this is the kill switch - flip it off, you're instantly back on the old path. Fast, low-risk, reversible. It's almost always the right *first* move in an incident: stop the bleeding now, diagnose later.

**Roll-forward** means fixing the problem with a *new* deploy that goes forward, not back. You'd choose it when there's nothing safe to roll back *to* - for example, a database migration already changed the data shape, so the old code can no longer run correctly. You can't un-migrate cleanly, so the only way out is through: ship a fix.

```text
Bug in production?
  ├─ Is there a safe state to return to?
  │     YES → ROLL BACK  (flip the kill switch - fast, reversible)
  │     NO  → ROLL FORWARD (ship a fix - the only path when going back is unsafe)
```

*What just happened:* the deciding question isn't "which is better" - it's "is going back actually safe?" Flags make rollback cheap and instant *when it's available*, which is exactly why you wrap risky changes in them. But a flag can't undo a data migration, so some incidents force you forward no matter how good your toggles are.

> Decide the rollback plan *before* you ship, not during the incident. "If this breaks, do we flip a flag or do we ship a fix?" is a five-minute conversation calm, and an agonizing one at 3am.

## For builders

The mature version of all this is a habit, not a platform. When you add a flag, write down three things: who it's for, when it dies, and how you'd back out the change it gates. Flags are part of your release machinery, so treat them with the same care as the rest of your pipeline - see [/guides/build-and-release-basics](/guides/build-and-release-basics) for where they sit in the broader release flow. The teams that get burned aren't the ones using flags; they're the ones who added forty and removed zero.

```quiz
[
  {
    "q": "What is 'flag debt'?",
    "choices": [
      "The licensing cost of a hosted flag service",
      "Flags that have done their job but were never removed, leaving dead forks in the code",
      "The latency added by checking a flag",
      "Flags that are turned on for too many users"
    ],
    "answer": 1,
    "explain": "A released flag whose if/else lingers in the code is debt - dead, confusing, and a live hazard if someone flips the stale toggle."
  },
  {
    "q": "Why does the number of live flags create a testing problem?",
    "choices": [
      "Each flag requires its own separate server",
      "Each independent flag roughly doubles the number of system states, and you can't test them all",
      "Flags slow down the test suite linearly",
      "Tests cannot read flag values at all"
    ],
    "answer": 1,
    "explain": "State space grows roughly as 2^(flag count). Untested combinations will eventually occur in production, so fewer live flags is safer."
  },
  {
    "q": "When would you choose roll-forward instead of rolling back?",
    "choices": [
      "Whenever a flag exists for the feature",
      "When rolling forward is always faster than a flag flip",
      "When there's no safe state to return to - e.g. a migration already changed the data shape",
      "Roll-forward should never be used in production"
    ],
    "answer": 2,
    "explain": "Rollback needs a safe state to return to. If a migration changed data so the old code can't run, going back is unsafe and you must ship a fix forward."
  }
]
```
