# Database Backups and Restores

> A backup you have never restored is a hope, not a backup. Logical versus physical dumps, point-in-time recovery, and testing the restore.


---

# Database Backups and Restores

Everyone backs up their database. Almost nobody restores it - until the morning a bad `DELETE`, a
dropped table, or a dead disk forces the question, and the plain answer is *we think we have backups*.
That gap between "we have backups" and "we have proven we can get the data back" is where companies
quietly die. A backup you have never restored is not a safety net; it is a hope you wrote to disk.

This guide closes that gap: the mental model that the *restore* is the real product and the backup is
only its raw material, the three ways data actually gets backed up and when each fits, and the failures
that turn a backup strategy into a false sense of security - so next time it matters, you're running a
rehearsed procedure instead of improvising in front of an audience.

## How to read this

- **In an incident right now and need the data back?** Jump to
  [Phase 3: When It Breaks](03-when-it-breaks.md) for the restore-day checklist and the failure modes
  to rule out fast.
- **Want backups to stop being a black box?** Read in order. Phase 1 reframes what a backup is *for*,
  Phase 2 covers the three backup types and point-in-time recovery, Phase 3 covers testing, drills, and
  the ways it all goes wrong.

## The phases

1. **[The Restore Is the Real Thing](01-the-restore-is-the-real-thing.md)** - the mental model: a
   backup exists only to be restored, so the restore is the thing you measure and test. RPO and RTO in
   plain terms, and why "we have backups" is a sentence with no information in it.
2. **[The Three Kinds of Backup](02-the-three-kinds-of-backup.md)** - logical dumps versus physical
   snapshots versus the write-ahead log, what each gives you, and how the WAL unlocks point-in-time
   recovery so you can rewind to the second before the bad command.
3. **[When It Breaks](03-when-it-breaks.md)** - the cautionary tale of the backup job that wrote empty
   files for months, the 3-2-1 rule, and how to automate, verify, and *drill* the restore so it works
   the day you need it.

**Related:** [What a Database Is](/guides/what-a-database-is) ·
[Transactions and ACID](/guides/transactions-and-acid)


---

# The Restore Is the Real Thing

Picture the moment this guide is really about. Someone ran `DELETE FROM orders` without a `WHERE`
clause, or a migration dropped the wrong table, or a disk gave up. The room goes quiet. Then comes the
sentence everyone says and nobody has verified: *"It's fine, we have backups."*

Here's the uncomfortable truth that the rest of this guide builds on. Nobody has ever needed a backup.
What people need is a *restore* - the data, back in the database, with the application working again. The
backup is the raw material. And raw material you've never actually used is an assumption, not a
capability.

## The asymmetry nobody plans for

Backups and restores feel like two sides of one coin. They are not. They get wildly unequal attention,
and that imbalance is the root of most backup disasters.

- **The backup runs constantly.** It's a cron job, a managed snapshot schedule, a checkbox in a console.
  It runs every night for years, silently, and produces a green checkmark.
- **The restore runs almost never.** It runs in the worst hour of someone's career, under pressure, often
  for the first time, often by someone who didn't write the backup.

So the thing you've practiced thousands of times is the half that doesn't save you. The thing that
actually saves you, you've practiced zero times. That's the trap. A green "backup succeeded" light tells
you a *file got written*. It tells you nothing about whether that file can become a working database
again.

💡 **Reframe.** Stop thinking "do we have backups?" Start thinking "when did we last *restore* one, and
how long did it take?" The first question has a comforting answer that means nothing. The second has a
real answer that's worth everything.

## "We have backups" is a sentence with no information

When someone says "we have backups," ask three questions and watch the confidence drain:

```text
Q: When was the last backup taken?
A: ...last night? I think the job runs nightly.

Q: Have you ever restored one into a clean database?
A: ...not the production one, no.

Q: How long would a full restore take - minutes, hours, a day?
A: ...not sure. We've never timed it.
```
*What just happened:* Three plain questions exposed that "we have backups" was three separate untested
assumptions wearing one confident coat: that the job runs, that the output is usable, and that recovery
fits inside the time the business can survive being down. None of those is verified by the backup
succeeding. Each one is only verified by an actual restore.

This is why the entire discipline reframes around the restore. The backup is a *means*. The restore is
the *goal*. You measure, budget, and test the goal - not the means.

## RPO and RTO: the two numbers that define "good enough"

Before you can say a restore is acceptable, you need to define *acceptable*. Two numbers do that, and
they're refreshingly concrete once you strip the jargon.

📝 **Terminology.**
- **RPO - Recovery Point Objective.** How much *data* can you afford to lose, measured in time. "We can
  lose at most 5 minutes of data" is an RPO of 5 minutes. It answers: *how far back does the restore
  rewind us?*
- **RTO - Recovery Time Objective.** How much *time* the recovery itself can take. "We must be back up
  within 1 hour" is an RTO of 1 hour. It answers: *how long are we down while we restore?*

A simple way to feel the difference: RPO is the gap *before* the disaster (data between your last good
backup and the moment things broke). RTO is the gap *after* (the wall-clock time to get running again).

```text
        last good backup            disaster              back online
  ───────────●───────────────────────●────────────────────────●────────►  time
              └──── RPO: data lost ───┘
                                       └──── RTO: downtime ─────┘
```
*What just happened:* The timeline makes the two numbers physical. The stretch *before* the disaster is
data you'll lose - shrink it with more frequent backups (or a continuous log, which Phase 2 covers). The
stretch *after* is downtime - shrink it with a faster, rehearsed restore. They're different problems with
different fixes, and conflating them is how teams over-invest in one while the other quietly fails them.

The point of naming these numbers is that they turn a vague fear into an engineering target. "Don't lose
data" is unachievable and untestable. "RPO 5 minutes, RTO 1 hour" is something you can design for and,
crucially, *measure against in a drill*.

⚠️ **Gotcha - your RPO is set by your backup frequency, not your hopes.** If you back up once a night and
the disaster strikes at 5pm, you've lost a full day of data, full stop. Your real RPO is "up to 24
hours," no matter what the slide deck says. Want a smaller RPO? You need more frequent backups or a
continuous mechanism - there's no other lever.

## For builders

When you size a backup strategy, write the RPO and RTO down *first*, as a business decision, before you
pick any tooling. They're not technical preferences - they're answers to "how much data can the company
lose?" and "how long can it be down?", which only the business can answer. A storefront that loses an
hour of orders has a real problem; an internal analytics warehouse rebuilt from source data nightly
might happily tolerate an RPO of a day. The tooling in Phase 2 is chosen *to hit those numbers* - pick
the numbers first, or you'll buy machinery for a target you never defined.

## Recap

1. Nobody needs a backup; everyone needs a **restore**. The backup is raw material - the restore is the
   capability that saves you.
2. Backups run constantly and restores run almost never, so the half that actually rescues you is the
   half you've never practiced. That asymmetry is the core danger.
3. **"We have backups"** is three untested assumptions in a trench coat: the job ran, the output is
   usable, and recovery fits the time you have.
4. **RPO** = how much data you can lose (the gap before the disaster). **RTO** = how long recovery can
   take (the gap after). Name both as targets before choosing any tool.

```quiz
[
  {
    "q": "Why does this guide insist the restore - not the backup - is the thing you actually test?",
    "choices": [
      "Restores are cheaper to run than backups",
      "A successful backup only proves a file was written, not that the data can be brought back into a working database",
      "Backups are always reliable, so only restores can fail",
      "Testing backups is impossible"
    ],
    "answer": 1,
    "explain": "A green backup light means a file got written. Whether that file can become a working database again is only proven by an actual restore."
  },
  {
    "q": "Your team backs up once every 24 hours. A disaster hits 23 hours after the last backup. What is your real RPO in this event?",
    "choices": [
      "5 minutes",
      "1 hour",
      "Up to 24 hours of data lost",
      "Zero, because you have backups"
    ],
    "answer": 2,
    "explain": "RPO is set by backup frequency. Once-a-night backups mean you can lose up to a full day of data, regardless of intentions."
  },
  {
    "q": "Which pair correctly describes RPO and RTO?",
    "choices": [
      "RPO = how long recovery takes; RTO = how much data is lost",
      "RPO = how much data can be lost; RTO = how long recovery can take",
      "Both measure downtime in different units",
      "RPO is for physical backups, RTO is for logical backups"
    ],
    "answer": 1,
    "explain": "RPO (Recovery Point Objective) is the data-loss budget before the disaster; RTO (Recovery Time Objective) is the downtime budget after it."
  }
]
```


---

# The Three Kinds of Backup

Phase 1 set two targets: RPO (how much data you can lose) and RTO (how long recovery can take). Now we
pick the machinery that hits them. There are three fundamentally different ways to capture a database's
state, trading off speed, portability, and how little data you lose. Know the three, and you stop
cargo-culting a `pg_dump` cron job and start choosing the right tool for your RPO.

We'll use Postgres-flavored commands because they're concrete and widely known, but the *categories* are
universal - MySQL, SQL Server, and the rest all have a logical export, a physical copy, and a transaction
log.

## Kind 1: the logical dump - a recipe to rebuild the database

A **logical backup** doesn't copy the database's files. It reads the data out and writes down the
*instructions* to recreate it: the `CREATE TABLE` statements and the `INSERT`s (or a compact equivalent)
that would rebuild everything from an empty database.

```console
$ pg_dump --format=custom --file=shop_2026-06-30.dump shopdb
$ ls -lh shop_2026-06-30.dump
-rw-r--r--  1 you  staff   84M  Jun 30 02:00 shop_2026-06-30.dump
```
*What just happened:* `pg_dump` connected to `shopdb` and produced a single self-contained file
describing how to rebuild it. That file is portable - you can restore it into a different Postgres
version, a differently-sized machine, or a fresh empty database, and you can even restore one table out
of it. It's a recipe, not a photograph.

You restore a logical dump by *replaying the recipe* into a target database:

```console
$ createdb shopdb_restored
$ pg_restore --dbname=shopdb_restored shop_2026-06-30.dump
$ psql shopdb_restored -c "SELECT count(*) FROM orders;"
 count
-------
 41902
(1 row)
```
*What just happened:* `pg_restore` rebuilt the schema and re-inserted every row into a brand-new
database, and the row count confirms the data arrived. This is a genuine, observed restore - the only
thing Phase 1 said actually proves anything. Logical dumps are the most portable and easiest to
spot-check, which is why they're the friendliest for *testing*.

The cost: logical dumps are slow to take and slow to restore on large databases, because rebuilding from
`INSERT`s is far more work than copying files. A terabyte database can take hours to dump and longer to
restore - which can blow your RTO.

## Kind 2: the physical backup - a photograph of the files

A **physical backup** copies the actual data files the database lives in (or takes a storage-level
snapshot of the disk). It's a photograph of the bytes on disk at a moment in time, not a recipe.

```console
# Conceptually: a consistent copy of the data directory / a disk snapshot
$ pg_basebackup --pgdata=/backups/base-2026-06-30 --format=tar --gzip
```
*What just happened:* Instead of reading rows and writing `INSERT`s, this copied the database's files
wholesale into a backup directory. Restoring is correspondingly fast - put the files back and start the
engine, no row-by-row rebuild. Big databases recover dramatically faster this way, which is how you hit a
tight RTO at scale. Cloud "snapshot" backups (RDS snapshots, disk snapshots) are this category.

The trade-off is the mirror image of logical dumps. Physical backups are *fast* but *rigid*: they're tied
to the same database engine version and often the same platform, and you generally restore the *whole*
thing - you can't cherry-pick one table out of a raw file copy. Photograph, not recipe.

| | Logical dump | Physical backup |
|---|---|---|
| What it stores | Instructions to rebuild (`CREATE`/`INSERT`) | A copy of the data files |
| Portable across versions? | Yes | No (version/platform-locked) |
| Restore one table? | Yes | No (all-or-nothing) |
| Speed on big data | Slow | Fast |
| Best for | Small/medium DBs, migrations, spot-checks | Large DBs, tight RTO |

## Kind 3: the write-ahead log - the stream that fills the gap

Here's the limitation both of the above share: each is a *point in time*. Whether you dump nightly or
snapshot nightly, a disaster at 5pm still loses everything since the last capture. That's your RPO ceiling
from Phase 1, and neither full-backup type can break through it alone.

The fix is a mechanism most databases already have for durability: the **write-ahead log** (WAL). Before
the database changes any data, it first appends a record of the change to a sequential log. (If you've read
[Transactions and ACID](/guides/transactions-and-acid), this is the same log that makes "durable" mean
durable - the change is safe on disk the instant the log entry is written.) It's a continuous recording of
*every change*, in order.

Keep (archive) the WAL continuously and you have not just nightly snapshots but the complete ordered
stream of everything that happened *between* them.

```text
  full backup        WAL stream (every change, continuously archived)
       ●─────▶ w w w w w w w w w w w w w w w w w w w w ─────▶  now
    (Sun 02:00)        each w = one logged change
```
*What just happened:* The full backup gives you a starting point; the archived WAL gives you every change
since. Together they don't represent one moment - they represent *every* moment from the backup forward.
That's what shrinks RPO from "since last night" toward "the last few seconds."

## Point-in-time recovery: rewinding to the second before the mistake

Combine a physical base backup with the archived WAL and you unlock a technique that feels like a time
machine: **point-in-time recovery (PITR)**. Restore the base backup, then *replay the WAL up to a chosen
instant* - and stop.

```console
# Restore the base backup, then tell the engine: replay WAL, but stop just before the disaster
$ recovery_target_time = '2026-06-30 16:59:30'   # the bad DELETE ran at 17:00:00
```
*What just happened:* The engine restored the base files, then replayed every logged change in order up
to 16:59:30 and stopped - half a minute before someone ran `DELETE FROM orders` with no `WHERE`. The
table comes back exactly as it was the moment before the mistake. You didn't lose a day of orders, just
thirty seconds, and you chose where to stop. That's the payoff of the WAL: recovery to a *moment*, not
just last night's snapshot.

💡 **Key point.** The three kinds aren't competitors - the strong setups *combine* them. A common shape:
periodic physical base backups for fast bulk recovery, continuous WAL archiving for a tiny RPO and PITR,
and occasional logical dumps for portability and easy spot-checking. You pick the mix that hits the RPO
and RTO you wrote down in Phase 1.

## For builders

If your RPO is "minutes," a nightly dump alone will never get you there - no amount of polishing the dump
job changes that it's a once-a-day snapshot. The lever that breaks the once-a-day ceiling is continuous
WAL archiving, the only mechanism that captures the gaps *between* full backups. When someone asks for
near-zero data loss, the answer isn't "back up more often" - it's "archive the log continuously and set
up PITR." Match the mechanism to the number; don't run the backup tool harder.

## Recap

1. A **logical dump** is a recipe (`CREATE`/`INSERT`) to rebuild the database - portable, table-granular,
   easy to spot-check, but slow on large data.
2. A **physical backup** is a photograph of the data files - fast to restore at scale, but locked to the
   engine version/platform and usually all-or-nothing.
3. The **write-ahead log** is the continuous, ordered record of every change. Archiving it captures
   everything *between* full backups, which is what shrinks your RPO.
4. **Point-in-time recovery** = base backup + replayed WAL, stopped at a chosen instant - recovery to the
   second before the mistake. Strong setups combine all three kinds to hit their RPO/RTO.

```quiz
[
  {
    "q": "You need to restore a single accidentally-dropped table from a 2-terabyte database, into a different Postgres version for inspection. Which backup type fits best?",
    "choices": [
      "A physical backup, because it restores fast",
      "A logical dump, because it's portable across versions and lets you restore one table",
      "The write-ahead log alone",
      "None - single-table restores are impossible"
    ],
    "answer": 1,
    "explain": "Logical dumps are portable across versions and table-granular. Physical backups are version-locked and usually all-or-nothing."
  },
  {
    "q": "What does archiving the write-ahead log give you that nightly full backups alone cannot?",
    "choices": [
      "Faster nightly backups",
      "A continuous record of every change between backups, shrinking RPO toward seconds",
      "Smaller backup files",
      "Automatic restore testing"
    ],
    "answer": 1,
    "explain": "The WAL is the ordered stream of every change. Archiving it captures the gaps between full backups, which is what lets RPO drop from 'since last night' to seconds."
  },
  {
    "q": "Point-in-time recovery (PITR) works by...",
    "choices": [
      "Taking a backup every second",
      "Restoring a base backup, then replaying the archived WAL and stopping at a chosen instant",
      "Keeping the database in read-only mode",
      "Copying the data files twice for redundancy"
    ],
    "answer": 1,
    "explain": "PITR combines a base backup with the archived WAL: restore the base, replay logged changes in order, and stop at the target time - e.g. just before a bad command."
  }
]
```


---

# When It Breaks

You know what a backup is *for* (the restore) and the three ways to take one. This phase is about the
part that decides whether any of it saves you: the unglamorous discipline of making sure the restore
actually works on the day it counts. The most dangerous backup isn't the one that's missing - it's the
one that's there, green, and quietly useless. Here's the story that haunts everyone who's lived it.

## The cautionary tale: the backup job that wrote empty files

A team set up a nightly `pg_dump`. It ran. The job exited cleanly, the monitoring stayed green, the file
landed in storage. Every morning, for months, a backup file appeared. Everyone slept fine.

Then they needed it. And when they opened the backups, every file was a few kilobytes - empty. A
credential had changed early on; the dump connected, failed to read the data, wrote a near-empty file,
and *still exited zero* because the wrapper script only checked "did a file get created," not "is there a
database inside it." Months of green checkmarks, and not one usable backup among them.

```console
$ ls -lh /backups/ | tail -3
-rw-r--r--  1 db  db   2.1K  Jun 28 02:00 shop_2026-06-28.dump
-rw-r--r--  1 db  db   2.1K  Jun 29 02:00 shop_2026-06-29.dump
-rw-r--r--  1 db  db   2.1K  Jun 30 02:00 shop_2026-06-30.dump
```
*What just happened:* Every nightly file is 2.1K - for a database that should dump to tens of megabytes.
The job "succeeded" every night and produced nothing. This is the signature failure of backups: not a
loud error, but a *silent* one that the success light never noticed. The lesson isn't "check the file
size." It's deeper: **a backup is only verified by restoring it.** Nothing short of that would have
caught this.

## The 3-2-1 rule: don't keep your eggs near the fire

Even a perfect backup is worthless if it burns in the same fire as the database. The **3-2-1 rule** is the
old, boring, undefeated answer to "where do I keep these?":

- **3** copies of your data (the live database counts as one).
- on **2** different kinds of media or storage (not all on the same disk/array).
- with **1** copy off-site (a different region, a different provider, somewhere a single disaster can't
  reach both it and production).

```text
  [ live DB ]  ──▶  [ backup on different storage ]  ──▶  [ off-site copy ]
     copy 1               copy 2 (2nd medium)               copy 3 (1 off-site)
   ┕━━━━━━━━━━━━━━ one regional outage / ransomware must not take all three ━━━━━━━━━━━━━━┙
```
*What just happened:* The rule spreads your copies so that no single event - a failed disk, a deleted
bucket, a region outage, a ransomware run that encrypts everything it can reach - can destroy every copy
at once. The off-site one is the clause people skip and regret: a backup sitting in the same account or
region as production shares its fate.

⚠️ **Gotcha - a backup the attacker can also delete isn't off-site enough.** Ransomware and compromised
credentials often reach the backups precisely because they're in the same account with the same access.
The strongest off-site copy is one that production's credentials *cannot* delete or overwrite - separate
account, write-once/immutable storage, or a separate retention lock. Off-site means "out of the blast
radius," not "another folder."

## Automate, verify, drill - the three habits that make it real

Three habits turn a pile of files into an actual recovery capability, and they build on each other.

**1. Automate the backup.** A backup that depends on someone remembering to run it will be skipped the
week it matters. Schedule it, and alert on the job *not running* - silence should be loud.

**2. Verify automatically - and prove there's a database in the file.** Don't trust the exit code (the
empty-files team did). At minimum, check the backup is plausibly sized and structurally valid. Far
better: do a real automated restore into a throwaway database and run a sanity query.

```console
# Nightly verify: restore last night's dump into a scratch DB and sanity-check it
$ createdb verify_scratch
$ pg_restore --dbname=verify_scratch /backups/shop_2026-06-30.dump
$ psql verify_scratch -tAc "SELECT count(*) FROM orders;"
41902
$ dropdb verify_scratch
```
*What just happened:* The verifier didn't inspect the file - it *restored* it and asked the rebuilt
database a question. A non-zero, sensible order count proves the backup contains a real, queryable
database. The empty-files disaster cannot survive this check: a 2.1K file would fail the restore or
return zero rows, and the alert fires while it's still a Tuesday, not a catastrophe.

**3. Drill the full restore - like a fire drill.** Verification proves the *file* is good. A drill proves
the *whole procedure* is good: that a human can find the right backup, run the restore end to end, and
get the app working - within your RTO. Schedule it (quarterly is a common cadence), do it under realistic
conditions, and *time it*.

```text
RESTORE DRILL  (run it before you need it)
  1. Pick a target moment.       → e.g. "restore to 16:59:30 yesterday"
  2. Provision a clean target.   → fresh instance, NOT production
  3. Restore base + replay WAL.  → follow the written runbook, step by step
  4. Bring the app up against it.→ does it actually work end to end?
  5. Run sanity queries.         → row counts, recent records present?
  6. Record the wall-clock time. → did you beat your RTO?  if not, fix it now
```
*What just happened:* The drill converts every Phase-1 assumption into an observed fact: the backup
restores, the runbook is correct, a human can execute it, and the whole thing fits inside the RTO you
promised. Each drill also surfaces the gaps - a missing step, a permission you lacked, a restore that's
slower than your RTO - while they're cheap to fix instead of discovering them mid-incident.

💡 **Key point.** Automate, verify, drill map straight onto the failures they prevent: automation stops
the *missing* backup, verification stops the *empty/corrupt* backup, and the drill stops the *unrestorable
or too-slow* backup. Skip any one and you've left a hole exactly where disasters walk in.

## The restore-day checklist

When it's real and the pressure is on, don't improvise - run the procedure you rehearsed:

1. **Stop the bleeding.** Pause whatever is corrupting or deleting data (the runaway job, the bad deploy)
   so the damage doesn't grow while you work.
2. **Decide the target moment.** With PITR, identify the instant *before* the damage. Without it,
   identify the last good full backup.
3. **Restore to a *new* target, not over production.** Never restore on top of the damaged database - 
   you may need it for forensics, and a failed restore mustn't destroy your last evidence.
4. **Verify before cutting over.** Run sanity queries against the restored copy. Confirm the data you
   expect is present and the damage is gone.
5. **Cut over, then write it down.** Point the app at the restored database, then record what happened
   and what the drill should change next time.

## For builders

Put the restore time on a dashboard, not the backup time. The metric that predicts whether you survive an
incident is "how long did our last *restore drill* take, and did it beat RTO" - not "did last night's
backup succeed." Tracking the restore makes the team invest in the half that actually saves them, and it
kills the empty-files class of failure, because you can't fake a drill that has to produce a working
database at the end.

## Recap

1. The deadliest backup is **silently useless** - green, present, and empty. Only an actual restore would
   have caught the empty-files job; a success exit code proves nothing.
2. The **3-2-1 rule**: 3 copies, on 2 kinds of storage, with 1 truly off-site (out of the blast radius,
   ideally where production's own credentials can't delete it).
3. **Automate** (stop the missing backup), **verify** by restoring into a scratch DB (stop the empty/corrupt
   backup), and **drill** the full restore against your RTO (stop the unrestorable or too-slow backup).
4. On restore day, **stop the bleeding, pick the target moment, restore to a fresh target, verify, then
   cut over** - run the rehearsed procedure, never improvise.

```quiz
[
  {
    "q": "The 'backup job that wrote empty files' ran green for months. What single practice would have caught it earliest?",
    "choices": [
      "Adding more storage",
      "Automatically restoring each backup into a throwaway database and running a sanity query",
      "Backing up more frequently",
      "Switching from logical to physical backups"
    ],
    "answer": 1,
    "explain": "A backup is only verified by restoring it. An automated restore + sanity query would have failed (or returned zero rows) the first night the files were empty."
  },
  {
    "q": "Which copy best satisfies the '1' in the 3-2-1 rule?",
    "choices": [
      "A second backup file in the same storage bucket as production",
      "A copy in a separate account/region that production's credentials cannot delete or overwrite",
      "A copy on the same disk, compressed",
      "A copy kept only in the database's own WAL"
    ],
    "answer": 1,
    "explain": "Off-site means out of the blast radius. A copy that a regional outage, account compromise, or ransomware run can also reach shares production's fate."
  },
  {
    "q": "Verification proves the backup file is restorable. What does a full restore drill additionally prove?",
    "choices": [
      "That the backup file is smaller",
      "That the whole procedure works - a human can restore end to end, the app comes up, and it all fits within RTO",
      "That the database engine is up to date",
      "Nothing beyond what verification proves"
    ],
    "answer": 1,
    "explain": "A drill tests the entire procedure under realistic conditions and times it against RTO, surfacing runbook gaps and slow restores while they're cheap to fix."
  }
]
```
