# Backups and Disaster Recovery

> The 3-2-1 rule, RPO and RTO, and the one practice that matters most: actually testing the restore before the disaster forces you to.


---

# Backups and Disaster Recovery

Everyone agrees backups matter, right up until the morning a drive dies, a `DROP TABLE` runs against
prod instead of staging, or ransomware encrypts every file you own and leaves a note. That's the moment
you discover what your backups are actually worth - and for an uncomfortable number of teams, the answer
is "nothing," because the job had been silently failing for months, or the backup itself got encrypted too.

This guide is here so that morning is survivable. You'll get the mental model that keeps the two halves
of this problem straight - getting your *data* back versus getting your *service* back - plus the
handful of rules and numbers the whole field runs on, and the one habit (testing the restore) that
separates teams who recover in an hour from teams who never recover at all.

## How to read this

- **Want the gist fast?** Read [Phase 1](01-backup-vs-disaster-recovery.md). It installs the one
  distinction everything else hangs on, then the 3-2-1 rule in a sentence.
- **Want it to actually stick?** Read in order. Each phase builds: what these words mean, the numbers
  that drive every decision and its cost, then what happens when it all goes wrong on purpose and on a
  bad day.

## The phases

1. **[Backup vs Disaster Recovery](01-backup-vs-disaster-recovery.md)** - the two different problems
   hiding under one word, and the 3-2-1 rule that makes a backup trustworthy.
2. **[RPO, RTO, and the Cost Dial](02-rpo-rto-and-cost.md)** - how much data you can lose and how long
   you can be down, and how those two numbers set your budget.
3. **[The Untested Backup, and Ransomware](03-testing-and-ransomware.md)** - why the restore drill is
   the only proof that counts, and the offline, immutable copy that survives an attacker.


---

# Backup vs Disaster Recovery

Picture two bad mornings. In the first, a teammate runs a query against the wrong database and deletes a
table of customer orders. The servers are fine, the app is up, but a chunk of *data* is gone. In the
second, the data center your VPS lives in catches fire. Nothing is corrupt - it's *unreachable*,
and your whole service is dark.

Those are two different problems, and the word "backup" only solves the first one cleanly - which is why
"we have backups" never quite answers "are we safe?"

## The two problems, kept separate

A **backup** answers: *can I get my data back?* It's a copy of your bytes - the database, the user
uploads, the config - stored somewhere you can pull from later. Backups are about **data loss**.

**Disaster recovery (DR)** answers a bigger question: *can I get my whole service running again?* That
includes the data, but also the servers, the network, the DNS, the certificates, the deploy process - 
everything between "I have my bytes" and "customers can log in." DR is about **downtime**.

```text
BACKUP                         DISASTER RECOVERY
--------------------           ----------------------------------
"I have my data back."         "Customers can use the service again."
copy of the bytes              data + servers + network + DNS + deploy
fixes: deletion, corruption    fixes: fire, region outage, ransomware
the input to recovery          the whole play, start to finish
```

*What just happened:* a backup is one ingredient; disaster recovery is the finished meal. You can have a
perfect backup and still be down for three days because nobody knew how to rebuild the server it
restores onto. Holding these apart is the whole mental model - most "our backups failed us" stories are
actually "we had backups but no recovery plan."

## Why a single copy is not a backup

Here's the trap that catches people: they copy the database to a second folder on the *same machine* and
call it a backup. Then the machine dies, and both copies die together. A copy that shares a fate with the
original isn't protecting you - it's using more disk.

A real backup has to survive the thing that kills the original. That means a *separate failure domain*:
different disk, different machine, different building. The question to ask of any copy is blunt: **what
single event takes out both this copy and the thing it's backing up?** If you can name one, it's not a
backup yet.

> The first rule of backups: a backup you've never restored from is a *hope*, not a backup. We'll come
> back to this hard in Phase 3 - it's the single most expensive mistake in this whole topic.

## The 3-2-1 rule

The industry boiled "separate failure domain" down to a rule you can recite from memory. It's old, it
predates the cloud, and it still holds:

```text
3  copies of your data        (the live one + two backups)
2  different media / systems   (so one storage failure can't kill both copies)
1  copy kept offsite           (so one building can't kill everything)
```

*What just happened:* each number kills a different disaster. **Three copies** means one corrupt copy
doesn't leave you at zero. **Two media** means a single storage technology failing (a bad disk batch, a
filesystem bug, one cloud bucket misconfigured) can't take all your copies at once. **One offsite** means
a fire, flood, or region-wide outage in one location doesn't end you. Miss any one number and you've left
a specific disaster un-handled.

A concrete, modest version for a small project:

```text
copy 1:  the live database on your VPS         (the original)
copy 2:  nightly dump to a cloud object store   (different system, offsite)
copy 3:  weekly pull of that dump to your laptop or a home NAS (different media)
```

*What just happened:* that's 3-2-1 without buying anything exotic. Three copies, more than one kind of
storage, and at least one copy that isn't in the same building as your server. A side project can hit
this; there's no excuse rooted in scale.

## For builders

If you're running anything on your own box - see [/guides/what-a-server-is](/guides/what-a-server-is) and
[/guides/deploying-to-a-vps](/guides/deploying-to-a-vps) - the cheapest correct first move is a nightly
database dump shipped to object storage. That single step takes you from "one copy that dies with the
server" to genuinely offsite. It's not the whole 3-2-1, but it's the rung that matters most, and it's a
cron job and a bucket away.

```bash
# nightly: dump the DB and push it offsite (sketch, not production-hardened)
pg_dump mydb | gzip > /tmp/mydb-$(date +%F).sql.gz
# then upload /tmp/mydb-*.sql.gz to an offsite bucket and rotate old ones
```

*What just happened:* this is the second copy in the 3-2-1 list - a point-in-time snapshot living
somewhere your server can't take down with it. Note what it does *not* do: prove it restores. That proof
is Phase 3.

```quiz
[
  {
    "q": "What's the core difference between a backup and disaster recovery?",
    "choices": [
      "Backups are for databases; DR is for files",
      "A backup gets your data back; DR gets your whole service running again",
      "DR is just a backup stored in the cloud",
      "They're two words for the same thing"
    ],
    "answer": 1,
    "explain": "A backup is a copy of your bytes (fixes data loss). DR is the full play to restore the running service - data plus servers, network, DNS, and deploy (fixes downtime)."
  },
  {
    "q": "You copy your database to a second folder on the same server. Why isn't that a backup?",
    "choices": [
      "It uses too much disk space",
      "Folders can't hold database files",
      "It shares a failure domain - one dead machine kills both copies",
      "Backups must always be encrypted"
    ],
    "answer": 2,
    "explain": "A real backup has to survive the event that destroys the original. A copy on the same machine dies with it, so it protects against nothing structural."
  },
  {
    "q": "In the 3-2-1 rule, what does the '1' stand for?",
    "choices": [
      "One copy kept offsite",
      "One backup per day",
      "One person responsible for backups",
      "One cloud provider only"
    ],
    "answer": 0,
    "explain": "3 copies, 2 different media/systems, and 1 copy kept offsite - so a single building, fire, or region outage can't take out everything at once."
  }
]
```


---

# RPO, RTO, and the Cost Dial

"How often should we back up?" feels technical. It isn't - it's a *business* question wearing a technical
hat, and you can't answer it sensibly until you know two things: how much recent data the business can
afford to lose, and how long it can afford to be down.

Those two numbers have names, and once you have them, every other decision - schedule, storage, spend - 
falls out almost automatically. This is where backups stop being a checkbox and start being a deliberate
trade: the dial you're turning is cost.

## The two numbers: RPO and RTO

**RPO - Recovery Point Objective** - is how much data you're willing to lose, measured in *time*. It
answers: "when we recover, how far back is our most recent good copy?" An RPO of one hour means: in the
worst case, you lose up to the last hour of changes. RPO is set by your **backup frequency** - you can't
recover to a point you never captured.

**RTO - Recovery Time Objective** - is how long you're willing to be down, measured from "disaster
strikes" to "service is back." An RTO of four hours means: from the fire alarm to customers logging in
again, you've promised four hours. RTO is set by how fast your **recovery process** runs.

```text
        disaster
           │
   ...─────●───────────────────────●─────►  time
        ▲  │                       ▲
     last  │                    service
     good  │◄──── RTO ─────────► back up
     backup│   (downtime you accept)
        │
        │◄ RPO ►│  (data you accept losing:
                   the gap from last backup to disaster)
```

*What just happened:* RPO looks *backward* from the disaster - how much recent work vanished. RTO looks
*forward* from the disaster - how long the lights stay off. They're independent: you can lose only five
minutes of data (tiny RPO) but still take two days to get running (huge RTO), or vice versa. People mix
them up constantly; the fix is to remember RP**O** = recovery *point* (a moment in the past), RT**O** =
recovery *time* (a duration of downtime).

## How the numbers drive cost

Here's the part nobody tells you up front: **smaller numbers cost exponentially more.** Each one has its
own cost curve.

Tightening **RPO** means backing up more often. Going from nightly to hourly is cheap-ish. Going from
hourly to "near-zero" means continuous replication or streaming the database's write-ahead log to a
standby - a permanent second system, always running, always costing money.

Tightening **RTO** means recovering faster. Restoring from a cold backup might take hours of copying and
rebuilding. Getting RTO to minutes means a warm or hot standby already running and ready to take over - 
again, a second system you pay for around the clock.

```text
RPO target     typical mechanism                  relative cost
-----------    --------------------------------   -------------
24 hours       nightly dump to object storage     $
1 hour         hourly snapshots                    $$
minutes        continuous WAL / log shipping       $$$
near-zero      synchronous replication             $$$$

RTO target     typical mechanism                  relative cost
-----------    --------------------------------   -------------
1 day          restore from cold backup            $
hours          scripted rebuild + restore          $$
minutes        warm standby, ready to promote      $$$
seconds        hot standby / active-active         $$$$
```

*What just happened:* both dials run from "cheap and slow" to "expensive and instant," and the bottom of
each table is a permanently-running duplicate of your system. That's why you don't set these numbers by
asking engineers "how good can we make it" - the answer is always "infinitely good, for infinite money."
You set them by asking the business "what does an hour of downtime, or an hour of lost data, actually cost
us?" and buying down to where the cost of protection meets the cost of the loss.

## Different data deserves different numbers

A common mistake is picking one RPO/RTO for everything. Your customer database and your cache of
thumbnail images do not deserve the same protection. The database is irreplaceable; the thumbnails
regenerate from the originals. Paying for near-zero RPO on regenerable data is lighting money on fire.

```text
data                 RPO          RTO          why
-------------------  -----------  -----------  ------------------------------
orders / payments    minutes      minutes      losing it = losing money + trust
user accounts        ~1 hour      hours        important, changes slower
app logs             ~1 day       days         useful, not load-bearing
rendered thumbnails  "who cares"  on rebuild   regenerated from source images
```

*What just happened:* you tier your data and spend the tight (expensive) numbers only where loss actually
hurts. This single move - refusing to protect everything at the highest tier - is what keeps a backup
budget sane.

> A tempting shortcut: "we'll figure out RTO during the incident." You won't. During an incident you'll
> be improvising under pressure, and improvised recovery is slow recovery. The number is a promise you
> make *now* so you can build (and rehearse) the process that keeps it - which is exactly where Phase 3
> goes.

## For builders

If you've deployed something to a VPS (see [/guides/deploying-to-a-vps](/guides/deploying-to-a-vps)),
write your RPO and RTO down in plain words *before* you size any backup tooling: "We can lose at most one
hour of data. We need to be back within four hours." Now the schedule and the storage choice aren't
guesses - hourly backups satisfy the RPO, and a tested rebuild script that runs in well under four hours
satisfies the RTO. The numbers turn a vague worry into a spec you can actually verify.

```quiz
[
  {
    "q": "Your RPO is 1 hour. What does that promise?",
    "choices": [
      "You'll be back online within 1 hour of a disaster",
      "You back up exactly once per hour, no more",
      "In the worst case you lose at most the last hour of data changes",
      "Recovery takes no longer than 1 hour"
    ],
    "answer": 2,
    "explain": "RPO = Recovery Point Objective = how much data you can lose, measured in time. A 1-hour RPO means your most recent good copy is never more than an hour behind. Downtime is RTO, a different number."
  },
  {
    "q": "Why do tighter RPO and RTO targets cost so much more?",
    "choices": [
      "Cloud providers charge premium rates for small numbers",
      "Near-zero data loss and near-instant recovery require a second system running all the time",
      "Faster backups need faster internet, which is rare",
      "Smaller numbers require more engineers to watch the dashboards"
    ],
    "answer": 1,
    "explain": "The bottom of both cost curves is a permanently-running duplicate - continuous replication for RPO, a hot standby for RTO - that you pay for around the clock."
  },
  {
    "q": "Which is the sound way to handle backups for regenerable data like rendered thumbnails?",
    "choices": [
      "Give it the same tight RPO/RTO as the orders database",
      "Don't back it up at the highest tier - it regenerates from source, so a loose RPO is fine",
      "Never back up anything that can be regenerated",
      "Back it up more often than the database since there's more of it"
    ],
    "answer": 1,
    "explain": "Tier your data. Spend the expensive, tight numbers only where loss actually hurts. Paying for near-zero RPO on data you can regenerate is wasted money."
  }
]
```


---

# The Untested Backup, and Ransomware

You've got the model, the rule, and the numbers. Now the two ways teams with all of that *still* lose
everything - both avoidable. The first is quiet and self-inflicted: a backup that was never tested and
turns out not to work. The second is loud and adversarial: ransomware that reaches for your backups on
purpose. This phase turns a backup *plan* into something that actually saves you.

## The untested backup is the classic trap

Here is the most expensive sentence in this entire topic: **a backup you have never restored from is not
a backup - it's a hope.**

It feels paranoid until you've lived it. Backup jobs fail silently all the time: a path changed, a
credential expired, a disk filled, a config typo meant the job dumped an empty file every night for six
months. The job kept reporting "success" because it ran without crashing - it wasn't backing up
anything. Nobody noticed, because nobody ever tried to *use* the output. The dashboard was green the
entire time.

```text
What the dashboard shows:   backup job: ✅ SUCCESS (ran every night)
What's actually in the file: 0 rows. Empty dump. Six months of nothing.
When you find out:          during the disaster, when you go to restore.
```

*What just happened:* "the job ran" and "the backup works" are different claims, and only one of them
gets checked by default. The gap between them is where companies die. A successful run proves the script
executed; it proves nothing about whether the bytes it produced can rebuild your system.

## The only proof: the restore drill

There is exactly one way to know a backup works: **restore from it and check the result.** Not read the
logs. Not confirm the file exists and has a plausible size. Actually pull the backup into a clean
environment, bring the data up, and verify it's real.

```bash
# the restore drill, in spirit:
# 1. take last night's backup (the real artifact you'd use in a disaster)
# 2. restore it into a fresh, isolated environment
createdb restore_test
gunzip -c mydb-2026-06-29.sql.gz | psql restore_test

# 3. verify it's actually your data, not an empty shell
psql restore_test -c "SELECT count(*) FROM orders;"   # expect a real, sane number
psql restore_test -c "SELECT max(created_at) FROM orders;"  # expect ~last night
```

*What just happened:* you proved the backup is restorable *and* recent, in a sandbox, on a calm Tuesday - 
not at 3am during a real outage. The row count catches the empty-dump failure; the latest timestamp
catches a stale or wrong-source backup. Do this on a schedule (a quarterly drill at minimum, automated if
you can), and treat a failed drill as a real incident, because it is one - you've discovered you were
unprotected.

A restore drill also quietly validates your **RTO**: time the drill. If "promise: back in 4 hours" but
the drill takes 9, your RTO is fiction and now you know - while it's cheap to fix.

> Schrödinger's backup: the condition of any backup is unknown until you attempt a restore. Don't leave
> the box closed.

## Ransomware changes the threat model

Classic backups assume *accidents* - a dead disk, a fat-fingered delete. Ransomware is different: it's an
**intelligent adversary** who *wants* your recovery to fail. Modern ransomware doesn't immediately
encrypt and announce itself. It often sits quietly, finds your backups, and encrypts or deletes *those
first* - because a victim who can restore doesn't pay the ransom.

This breaks an assumption hiding inside ordinary 3-2-1. If your "offsite" backup is a cloud bucket your
production server can write to and delete from, then anyone - or any malware - with your server's
credentials can wipe that backup too. Reachable and deletable means *destroyable*. Same failure domain,
a logical one instead of a physical one.

## Immutable and offline: the copy they can't reach

The defense is a copy the attacker *cannot* alter or delete, even with full control of your servers. Two
forms, often combined:

```text
IMMUTABLE - write-once, read-many. The storage itself refuses deletion or
             overwrite until a retention period expires. (Object-lock /
             "WORM" on cloud object storage.) Even with your keys, an
             attacker can't erase what's locked.

OFFLINE - physically disconnected. Tapes in a vault, or a drive that's
             unplugged between backup runs. An air gap is unhackable over the
             network because there's no network to it.
```

*What just happened:* both close the loophole. Immutability means even a fully compromised account can't
delete the locked copy until its retention window passes; offline (the "air gap") means there's no live
path to the copy at all. This is the modern upgrade to the "1" in 3-2-1: not enough that the offsite copy
is *elsewhere* - at least one copy must be *unreachable* by a live, compromised system.

A practical small-team version: keep your normal offsite backups, *and* enable object-lock on the bucket
(or pull periodic copies to a drive you disconnect) so there's always one copy that today's compromised
credentials cannot touch.

## For builders

If you run your own infrastructure (see [/guides/what-a-server-is](/guides/what-a-server-is)), bake two
non-negotiables into the plan from day one. One: a recurring restore drill, with a row-count-and-timestamp
check, treated as an incident when it fails. Two: at least one immutable or offline copy, so a stolen
credential or a ransomware run can't take your recovery down with your production. Everything earlier in
this guide is preparation; these two are what make it *real* when the bad morning comes - and it does
come.

```quiz
[
  {
    "q": "Why is a backup job that reports 'success' every night still not proof you're protected?",
    "choices": [
      "Success messages are often delayed by an hour",
      "'The job ran' only proves the script executed, not that the bytes can rebuild your system",
      "Nightly is too infrequent to count as a real backup",
      "Success only counts if the job runs on a weekend"
    ],
    "answer": 1,
    "explain": "Jobs fail silently - an expired credential or config typo can dump an empty file while still 'succeeding.' Only an actual restore proves the output is usable."
  },
  {
    "q": "What is the only reliable way to know a backup actually works?",
    "choices": [
      "Confirm the backup file exists and has a plausible size",
      "Read the backup job logs for errors",
      "Restore it into a clean environment and verify the data is real and recent",
      "Check that the job has run successfully 30 days in a row"
    ],
    "answer": 2,
    "explain": "The restore drill is the only proof: pull the real artifact into a sandbox, bring the data up, and check it (e.g. row counts and the latest timestamp). It also validates your RTO if you time it."
  },
  {
    "q": "Why can ordinary 3-2-1 backups still fall to ransomware?",
    "choices": [
      "Ransomware encrypts faster than backups can run",
      "If your offsite copy is reachable and deletable by your server, compromised credentials can wipe it too",
      "3-2-1 was never designed for cloud storage",
      "Ransomware only targets databases, which 3-2-1 ignores"
    ],
    "answer": 1,
    "explain": "An attacker with your server's credentials can destroy any backup that server can write to or delete. The fix is an immutable (object-locked) or offline (air-gapped) copy they can't reach."
  }
]
```
