# HashiCorp Vault

> Stop hardcoding secrets: Vault stores them encrypted, gates access by policy, issues short-lived dynamic credentials, and keeps an audit trail.


---

# HashiCorp Vault

Right now there's a password in a config file, an API key in an environment variable that got committed once, and a database credential that hasn't changed in three years because nobody remembers everywhere it's pasted. That's not a bug, that's how most teams ship - until one leak turns it into a breach. Vault exists to make that sprawl stop: one encrypted place for secrets, access decided by written policy, credentials that expire on their own, and a log of every read. This guide builds the mental model first, then the daily moves, then the parts that bite in production.

## How to read this

Read phase 1 slowly. Vault makes a lot more sense once you understand that it spends its life *sealed* and that everything inside is encrypted by a key Vault itself can't reach without help. Get that, and unseal, policies, and dynamic secrets stop feeling like ceremony. Phase 2 is the everyday loop: logging in, reading and writing secrets, and the magic trick that is dynamic database credentials. Phase 3 is production reality - token leases, revocation, the audit log, and the failure modes that page you at 3am.

## The phases

1. [Phase 1: Sealed by Default](01-sealed-by-default.md) - what Vault actually is and why it exists
2. [Phase 2: The Daily Loop](02-the-daily-loop.md) - auth, policies, static and dynamic secrets
3. [Phase 3: Leases, Revocation, and Reality](03-leases-revocation-and-reality.md) - the gotchas and production concerns


---

# Sealed by Default

Picture the state your secrets are in today. A database password lives in `application.yml`, which lives in a repo, which lives on every laptop that ever cloned it. An AWS key is in a CI variable that ten people can read. None of these secrets has an expiry date, none of them is logged when read, and revoking one means a frantic grep across every service that might use it. This is *secrets sprawl*, and it's not laziness - it's the default. Every shortcut was reasonable in the moment. The problem is that the moments add up into a blast radius nobody can measure.

Vault's whole pitch is to collapse that blast radius. Instead of secrets scattered as plaintext, there's one service that holds them encrypted, hands them out only to identities that policy allows, can make many of them short-lived, and writes down every access. Before any of that makes sense, you need the one idea the rest hangs from: Vault spends most of its life **sealed**.

## What "sealed" actually means

When Vault stores a secret, it doesn't write it to disk in the clear. It encrypts it with an internal *encryption key*. That encryption key is itself encrypted by a *root key* (sometimes called the master key). So far that's normal at-rest encryption. The twist is where the root key lives: it does **not** live on disk in usable form.

When the Vault process starts, it's *sealed*. It has the encrypted data and the encrypted root key, but it cannot decrypt anything, because it's missing the piece that unlocks the root key. In that state Vault can do almost nothing - it can't read secrets, can't issue tokens, can't serve requests. It's a locked safe whose combination isn't in the building.

```text
                sealed                          unsealed
        ┌─────────────────┐             ┌─────────────────────┐
        │ encrypted data  │  unseal     │ encrypted data      │
        │ encrypted root  │ ─────────►  │ root key in memory  │
        │ (no root key)   │  keys       │ can decrypt + serve │
        └─────────────────┘             └─────────────────────┘
```

*What just happened:* a sealed Vault is inert by design. The root key only ever exists in memory, only after an explicit unseal, and it's gone the moment the process restarts.

## Unsealing, and why it's split into pieces

You might expect one unseal password. Vault deliberately avoids that, because one password is one person who can be coerced, phished, or fired. Instead, the classic setup uses **Shamir's Secret Sharing**: at initialization, the root key is split into several *unseal keys* (shares), and you set a *threshold* of how many are needed to reconstruct it.

```console
$ vault operator init -key-shares=5 -key-threshold=3
Unseal Key 1: hf2k...    Unseal Key 2: 9xQp...
Unseal Key 3: Lm4v...    Unseal Key 4: bN7c...
Unseal Key 5: Tz0a...
Initial Root Token: hvs.AbCdEf...
```

*What just happened:* Vault generated 5 unseal keys and decided that any 3 of them, together, can reconstruct the root key. No single key does anything alone. The threshold means no one person can unseal Vault, and losing one or two keys doesn't lock you out forever.

To unseal, you feed in keys one at a time until you hit the threshold:

```console
$ vault operator unseal hf2k...
Sealed   true    Unseal Progress   1/3
$ vault operator unseal Lm4v...
Sealed   true    Unseal Progress   2/3
$ vault operator unseal Tz0a...
Sealed   false   Unseal Progress   0/3
```

*What just happened:* three of the five share-holders showed up, the root key got reassembled in memory, and Vault is now serving requests. Restart the process and you're back to sealed - you unseal again.

> In real production, almost nobody types unseal keys by hand. Vault is configured with **auto-unseal**, where a cloud KMS or HSM holds the unlocking key and Vault asks it on startup. Shamir is the model to understand; auto-unseal is the model you run. Phase 3 returns to this.

## The pieces Vault is made of

Once unsealed, Vault is a set of pluggable parts. You don't need them all on day one, but the names will keep coming up:

- **Secrets engines** - mounted at a path, each one *does* something with secrets. The `kv` engine stores static key-value secrets. The `database` engine *generates* credentials on demand. The `transit` engine encrypts data without ever storing it. Different engines, mounted at different paths, same Vault.
- **Auth methods** - how a human or machine proves who they are *before* Vault gives them a token. A person might log in with their company SSO; a Kubernetes pod proves itself with its service-account token; a CI job uses AppRole. Each method maps an external identity to Vault policies.
- **Policies** - written rules saying which paths an identity may read, write, or delete. No policy, no access. This is how Vault turns "who are you" into "what may you touch."
- **Tokens** - the currency of every request. After you authenticate, you get a token; every later call carries it. Tokens have a lease and expire.
- **Audit devices** - the tamper-evident log of every request and response (with secrets hashed, not printed).

```text
   identity ──auth method──► token ──(carries)──► request
                                                    │
                                              policy check
                                                    │
                                          secrets engine at a path
```

*What just happened:* every interaction follows the same spine - prove identity, get a token, make a request, policy decides, an engine serves it, the audit device records it. Hold that shape and the rest of Vault is detail.

## Why this beats a config file

The payoff of all this structure is concrete. Secrets are encrypted at rest under a key that isn't sitting on disk. Access is governed by policy you can read and review, not by file permissions scattered across hosts. Many secrets can be *dynamic* - generated when asked and revoked automatically - so a leaked credential is worth little because it expires. And the audit log means "who read the prod database password, and when" has an actual answer.

For the bigger picture of why teams move off plaintext and what other tools play in this space, see [/guides/secrets-management](/guides/secrets-management). Vault is one strong answer to the problem that guide frames.

```quiz
[
  {
    "q": "What does it mean that Vault is 'sealed' when it starts?",
    "choices": ["It refuses network connections", "It can't decrypt its data because the root key isn't reconstructed yet", "Its data is deleted until you log in", "It runs in read-only mode"],
    "answer": 1,
    "explain": "Sealed Vault holds encrypted data and an encrypted root key, but lacks the means to unlock the root key, so it can decrypt nothing until unsealed."
  },
  {
    "q": "With -key-shares=5 -key-threshold=3, how many unseal keys are needed to unseal?",
    "choices": ["All 5", "Any 3", "Just 1", "Exactly 2"],
    "answer": 1,
    "explain": "Shamir's Secret Sharing splits the root key into 5 shares, any 3 of which reconstruct it. No single key works alone."
  },
  {
    "q": "Which part of Vault decides whether an identity may read a given path?",
    "choices": ["The auth method", "A policy", "The audit device", "The unseal key"],
    "answer": 1,
    "explain": "Auth methods establish identity; policies map that identity to which paths it may read, write, or delete."
  }
]
```


---

# The Daily Loop

Most days you won't unseal anything or split a root key. You'll log in, read a secret or two, maybe write one, and your apps will fetch credentials at startup. This phase walks the loop you'll actually live in. To follow along without touching real infrastructure, run a dev server - it starts unsealed, in memory, with a known root token, and it prints a warning because it's only for learning.

```console
$ vault server -dev
==> Vault server started! Root Token: hvs.devroot...
WARNING: dev mode is enabled! In-memory, unsealed, NOT for production.
```

*What just happened:* a throwaway Vault is running on `http://127.0.0.1:8200`. Open a second terminal, point the CLI at it, and you're ready.

```console
$ export VAULT_ADDR='http://127.0.0.1:8200'
$ export VAULT_TOKEN='hvs.devroot...'
$ vault status
Sealed     false
Version    1.x.x
```

*What just happened:* the CLI knows where Vault is and who you are. Every command below rides on that token.

## Static secrets: the KV engine

The simplest engine is `kv` - a versioned key-value store for secrets you hold and rotate yourself, like a third-party API key. Write one, read it back:

```console
$ vault kv put secret/myapp/stripe api_key=sk_live_abc123 webhook_secret=whsec_xyz
$ vault kv get secret/myapp/stripe
====== Data ======
Key               Value
api_key           sk_live_abc123
webhook_secret    whsec_xyz
```

*What just happened:* you stored two fields under one path and read them back. In a dev server, `secret/` is a KV v2 mount, so it keeps version history - overwrite `api_key` and the old value is still retrievable by version until you delete it.

To pull only one field, ask for it directly - handy in scripts:

```console
$ vault kv get -field=api_key secret/myapp/stripe
sk_live_abc123
```

*What just happened:* Vault returned the raw value with no formatting, ready to pipe into an env var or config at deploy time instead of baking it into the image.

## Identity first: auth methods and policies

The root token can do anything, which is exactly why apps and people should never use it. Real access starts with an **auth method** that maps an identity to a **policy**. Let's build the smallest real example: a policy that allows reading one path, and an AppRole identity bound to it.

First the policy - written in HCL, granting capabilities on paths:

```hcl
# myapp-policy.hcl
path "secret/data/myapp/*" {
  capabilities = ["read"]
}
```

*What just happened:* this policy says the bearer may `read` anything under `secret/data/myapp/` (the `data/` segment is how KV v2 paths look under the hood) and nothing else. Default-deny means everything not listed is forbidden.

```console
$ vault policy write myapp myapp-policy.hcl
$ vault auth enable approle
$ vault write auth/approle/role/myapp token_policies=myapp token_ttl=1h
```

*What just happened:* you registered the policy, turned on the AppRole auth method, and created a role `myapp` whose logins receive the `myapp` policy and a token that lives one hour. AppRole is built for machines: a service presents a `role_id` (like a username) and a `secret_id` (like a password) and gets a scoped token back.

```console
$ vault read auth/approle/role/myapp/role-id
role_id    7c8f...
$ vault write -f auth/approle/role/myapp/secret-id
secret_id  b41a...
$ vault write auth/approle/login role_id=7c8f... secret_id=b41a...
token             hvs.scoped...
token_policies    ["default" "myapp"]
token_ttl         1h
```

*What just happened:* the app authenticated and got a token scoped to exactly what the `myapp` policy allows. That token can read `secret/myapp/*` and nothing else - a leaked copy can't touch the rest of Vault, and it expires in an hour.

> Pick the auth method that matches the identity you already have. A pod in Kubernetes should use the `kubernetes` method (it proves itself with its service-account token, no extra secret to manage). Humans should use OIDC/SSO. AppRole is the fallback when nothing better fits. The goal is always: stop inventing new credentials to protect other credentials.

## The magic trick: dynamic secrets

Here's where Vault stops being a fancier password file. The `database` secrets engine doesn't *store* a database password - it *creates a fresh one on demand* and deletes it later. Set it up once by telling Vault how to connect as an admin and what kind of user to mint:

```console
$ vault secrets enable database
$ vault write database/config/appdb \
    plugin_name=postgresql-database-plugin \
    connection_url="postgresql://{{username}}:{{password}}@db:5432/app" \
    allowed_roles="readonly" \
    username="vault_admin" password="..."
$ vault write database/roles/readonly \
    db_name=appdb \
    creation_statements="CREATE ROLE \"{{name}}\" WITH LOGIN PASSWORD '{{password}}' VALID UNTIL '{{expiration}}'; GRANT SELECT ON ALL TABLES IN SCHEMA public TO \"{{name}}\";" \
    default_ttl="1h" max_ttl="24h"
```

*What just happened:* you taught Vault to log into Postgres as an admin and defined a `readonly` role whose users get SELECT and self-destruct. Nothing's been issued yet - this is the recipe, not a meal.

Now any app with permission asks for credentials, and Vault generates a brand-new database user each time:

```console
$ vault read database/creds/readonly
username    v-approle-readonly-x7k2p
password    A1a-9zQ...random...
lease_id    database/creds/readonly/abc123
lease_duration   1h
```

*What just happened:* Vault created a real Postgres user that didn't exist a second ago, valid for one hour. When the lease ends, Vault logs back in as admin and drops the user. There's no shared password to leak, no rotation cron to maintain, and a stolen credential is worthless within the hour.

```mermaid
sequenceDiagram
  participant App
  participant Vault
  participant DB
  App->>Vault: read database/creds/readonly
  Vault->>DB: CREATE ROLE (admin)
  Vault-->>App: username + password (1h lease)
  Note over Vault,DB: at lease expiry
  Vault->>DB: DROP ROLE
```

*What just happened:* the secret's whole lifecycle - birth, hand-off, death - is owned by Vault. The app never holds a long-lived credential, and the database never has a stale account lying around.

## Encryption as a service: transit

One more engine worth knowing, because it solves a different problem. Sometimes you don't want Vault to *hold* a secret - you want it to encrypt *your* data without your app ever touching a key. That's the `transit` engine:

```console
$ vault secrets enable transit
$ vault write -f transit/keys/orders
$ vault write transit/encrypt/orders plaintext=$(echo -n "card-4242" | base64)
ciphertext    vault:v1:abcDEF123...
```

*What just happened:* Vault encrypted your data and handed back ciphertext tagged with a key version. Your app stores that ciphertext in its own database. The encryption key never leaves Vault, so a dump of your database is useless without a `transit/decrypt` call - which is policy-gated and audited like everything else.

For builders: transit is how you get strong encryption and key rotation without becoming a cryptographer or shipping keys in your binary. Rotate the key in Vault and old ciphertext (`vault:v1:`) still decrypts while new writes use `vault:v2:`.

```quiz
[
  {
    "q": "What is fundamentally different about a dynamic database secret versus a KV secret?",
    "choices": ["It is encrypted and the KV one is not", "Vault generates a brand-new credential on demand and revokes it at lease end", "It can only be read once", "It is stored in a faster engine"],
    "answer": 1,
    "explain": "The database engine creates a fresh DB user per request and drops it when the lease expires, so there's no shared long-lived password to leak."
  },
  {
    "q": "Why should an app authenticate with AppRole or Kubernetes instead of the root token?",
    "choices": ["The root token is slower", "Scoped tokens are limited by policy and expire, shrinking the blast radius", "Root tokens don't work over the network", "AppRole secrets never expire"],
    "answer": 1,
    "explain": "The root token can do anything. A scoped token is bound to a policy and a TTL, so a leak is contained and self-limiting."
  },
  {
    "q": "What does the transit secrets engine store?",
    "choices": ["Your encrypted application data", "Nothing of yours - it encrypts/decrypts data you keep, the key stays in Vault", "Database credentials", "Unseal keys"],
    "answer": 1,
    "explain": "Transit is encryption-as-a-service: the key never leaves Vault, and you store the resulting ciphertext yourself."
  }
]
```


---

# Leases, Revocation, and Reality

The dev server lied to you a little. It's unsealed forever, it never restarts, and nothing it issues ever expires under load. Production is the opposite, and the gap between the two is where people get paged. This phase is the part you'll wish you'd read before the incident: leases, revocation, the audit log, and the handful of failure modes that define life with Vault.

## Leases: everything has a clock

Almost everything Vault hands out - tokens, dynamic database creds, AWS credentials - comes with a **lease**: an ID and a duration. When the lease expires, Vault revokes the thing automatically. This is the feature, not a nuisance. A credential that dies on its own is a credential a thief can't keep.

The catch is that long-running apps outlive their leases. If your service grabbed a 1-hour database credential at boot and you do nothing, hour two is an outage. The fix is to **renew** before expiry:

```console
$ vault lease renew database/creds/readonly/abc123
lease_id          database/creds/readonly/abc123
lease_duration    1h
```

*What just happened:* the lease clock reset for another hour. Real apps use a Vault client library or the Vault Agent sidecar to renew on a timer, so the credential stays fresh as long as the app runs - and stops being renewed the moment it dies.

> The number that bites: a lease can be renewed only up to its `max_ttl`. A creds role with `default_ttl=1h max_ttl=24h` can renew for at most a day, then the credential is gone no matter how often you ask. Long-lived services must be able to fetch a *new* credential, not renew one forever. Design for the credential disappearing, because eventually it will.

## Revocation: the reason any of this matters

Leasing gives you scheduled cleanup. Revocation gives you the panic button. If a credential leaks, you don't grep your fleet - you tell Vault to kill it, and Vault undoes whatever it created (drops the DB user, deletes the AWS key):

```console
$ vault lease revoke database/creds/readonly/abc123
Success! Revoked lease
```

*What just happened:* that one credential is dead and the underlying database user is gone. During a real incident you can go wider - `vault lease revoke -prefix database/creds/readonly` kills every credential that role ever issued. This is the difference between Vault and a password vault: Vault knows what it issued and can take it all back.

## The audit log: who touched what

Turn on an audit device and Vault writes every request and response to an append-only log. This is non-negotiable for production - it's how you answer "who read the prod secret and when," and it's required for most compliance.

```console
$ vault audit enable file file_path=/var/log/vault/audit.log
```

A log line (trimmed) looks like this:

```json
{"time":"2026-06-30T09:12:04Z","type":"response",
 "auth":{"display_name":"approle-myapp","policies":["default","myapp"]},
 "request":{"operation":"read","path":"secret/data/myapp/stripe"},
 "response":{"data":{"api_key":"hmac-sha256:9f3c..."}}}
```

*What just happened:* Vault recorded who (the AppRole identity), what (read), where (the path), and when - but the secret value is HMAC'd, not printed. You can prove a specific value was accessed without the log itself becoming a place secrets leak. One sharp edge: if *all* configured audit devices fail to write, Vault blocks the request rather than serving it unlogged. Auditing is treated as more important than availability, so monitor your audit sink.

## Auto-unseal, and the seal you'll actually run

Phase 1 used Shamir unseal keys typed by hand. In production that means a credential outage every time Vault restarts, at 3am, requiring three humans. So real deployments use **auto-unseal**: Vault stores its root key wrapped by an external KMS (AWS KMS, GCP KMS, Azure Key Vault, or an HSM) and unwraps it automatically on startup.

```hcl
seal "awskms" {
  region     = "us-east-1"
  kms_key_id = "alias/vault-unseal"
}
```

*What just happened:* Vault now unseals itself by calling the cloud KMS on boot. You've moved the "who can unlock Vault" question to your cloud's IAM, which is auditable and doesn't require waking people up. The Shamir model still matters - it's how recovery keys work even under auto-unseal - but you don't run it by hand.

## The failure modes that define daily life

A few realities that no tutorial mentions until they hurt:

- **Vault is now a hard dependency.** If Vault is down or sealed, apps can't fetch secrets. Run it highly available (a cluster with integrated storage / Raft), and cache or template secrets to disk via Vault Agent so a brief outage doesn't instantly cascade.
- **Sealed-on-restart is normal.** A node that reboots comes back *sealed*. Without auto-unseal, it stays useless until someone unseals it. This surprises people every single time.
- **Tokens and leases pile up.** Apps that authenticate on every request instead of reusing a token can flood Vault with short-lived leases. Reuse tokens, renew, and set sane TTLs.
- **Policy is default-deny and exact.** A path that's off by one segment (`secret/myapp` vs `secret/data/myapp` under KV v2) returns permission denied, not a helpful error. Most "Vault is broken" tickets are a policy path typo.
- **Recovery is a real plan, not a hope.** Losing the unseal/recovery keys *and* the storage backend means the data is unrecoverable by design - that's the security guarantee working against you. Back up storage and store recovery keys the way you'd store the keys to the building.

```mermaid
flowchart LR
  A[node restarts] --> B{auto-unseal?}
  B -- yes --> C[KMS unwraps key, serving]
  B -- no --> D[sealed, outage until human unseals]
```

*What just happened:* the single most common production surprise in one picture - restart equals sealed unless you've set up auto-unseal.

## Where Vault sits in the bigger picture

Vault is the runtime heart of secrets, but it's one layer of a larger discipline. The leases-and-revocation model only protects you if secrets aren't also leaking through your build pipeline, your dependencies, or a committed `.env`. For how secrets fit into the wider problem of protecting what your software depends on, see [/guides/supply-chain-security](/guides/supply-chain-security), and for the broader landscape of secret storage choices, [/guides/secrets-management](/guides/secrets-management).

In the wild: the teams who get the most from Vault are the ones who lean into dynamic, short-lived secrets and let leases do the cleaning. The teams who fight it are the ones who use it as a slightly fancier KV store with eternal static secrets - they get the operational cost of Vault without the payoff. The whole point was to make secrets boring, brief, and accountable. Aim for that.

```quiz
[
  {
    "q": "A long-running service got a 1h database credential with max_ttl=24h. What must it eventually do?",
    "choices": ["Renew the same lease forever", "Fetch a new credential, because the lease can't be renewed past max_ttl", "Restart Vault", "Switch to the root token"],
    "answer": 1,
    "explain": "Renewal works only up to max_ttl. After 24h the credential is gone, so the app must be able to request a fresh one."
  },
  {
    "q": "What does Vault do if every configured audit device fails to write?",
    "choices": ["Serves the request and logs later", "Blocks the request rather than serving it unaudited", "Disables auditing automatically", "Silently drops the log line"],
    "answer": 1,
    "explain": "Vault treats auditing as more important than availability: if it can't record a request, it refuses to serve it."
  },
  {
    "q": "Why do production Vault deployments use auto-unseal?",
    "choices": ["It's faster at encryption", "A restarted node comes back sealed; auto-unseal avoids a 3am manual unseal outage", "It removes the need for policies", "It makes secrets last forever"],
    "answer": 1,
    "explain": "Any restart leaves a node sealed. Auto-unseal lets a KMS/HSM unwrap the root key on boot so the node returns to service without humans."
  }
]
```
