# Rate Limits and Retries

> Calling a flaky or throttled API without melting it or yourself: 429s, exponential backoff with jitter, idempotency, and circuit breakers.


---

# Rate Limits and Retries

You wired up an API call, it worked, you shipped it. Then traffic grew, or the upstream had a bad
minute, and suddenly your logs are full of `429` and `503` - and your "fix" of retrying in a tight loop
somehow made everything *worse*. That sinking feeling, where the thing meant to help is now the thing
hurting you, is exactly what this guide clears up.

Here's the relief: calling a flaky or throttled API well is a small, learnable set of habits, not a dark
art. APIs rate-limit you on purpose, and they tell you how to behave - you mostly need to listen. Retrying
safely is a handful of rules: wait longer each time, add a little randomness, cap how hard you try, and
never blindly retry something that moves money. Learn those, and a wobbly dependency becomes a minor
annoyance instead of a 2am page.

## How to read this

- **Mid-incident, need the rule right now?** Phase 2 has the backoff-and-jitter recipe and the
  "is this safe to retry?" checklist; Phase 3 has the circuit-breaker idea for when a dependency is
  truly down.
- **Want it to actually make sense?** Read in order. Phase 1 explains *why* limits exist (so you stop
  taking them personally), Phase 2 turns retrying into a safe reflex, and Phase 3 covers the failure
  modes that bite teams who only learned the happy path.

## The phases

1. **[Why APIs Push Back](01-why-apis-push-back.md)** - the mental model: rate limits aren't an insult,
   they're how a shared service protects itself. Token buckets, the `429` status, and the `Retry-After`
   header that tells you exactly how long to wait.
2. **[Retrying Without Making It Worse](02-retrying-without-making-it-worse.md)** - the everyday core:
   exponential backoff, why you *must* add jitter, a retry budget so you give up gracefully, and the
   golden rule of only retrying requests that are safe to repeat.
3. **[When Retrying Isn't Enough](03-when-retrying-isnt-enough.md)** - the deeper payoff: the
   thundering-herd problem, the circuit breaker that stops you from hammering a downed service, and what
   it means to be a genuinely good client.

> This guide stays at the level you need to write a resilient client. The provider's side - designing
> and enforcing limits, sizing buckets, fairness across tenants - is its own topic; for the broader API
> picture see [REST APIs, Explained](/guides/rest-apis-explained) and [What an API Is](/guides/what-an-api-is).


---

# Why APIs Push Back

The first time you see a `429 Too Many Requests` come back from an API you've been calling happily for
weeks, it feels personal - like the service singled you out and slammed a door. It didn't. Rate limiting
is one of the most boringly reasonable things a service does, and once you see it from the *server's*
side, you'll stop fighting it and start working with it. That shift is the whole point of this phase.

## The reader's reality: a shared restaurant kitchen

Picture an API as one kitchen serving thousands of customers at once. If a single customer could fire
ten thousand orders a second, the kitchen would grind to a halt and *everyone* - including you - would
wait forever or get nothing. So the kitchen sets a pace: each customer gets a fair number of orders per
minute. That's a rate limit. It exists to keep the kitchen *up*, which is the only way you get served at
all.

A rate limit protects three things at once:

- **The service's stability** - no single caller can overwhelm it and take it down for everyone.
- **Fairness** - your noisy neighbor can't starve you of capacity.
- **Cost and abuse control** - it caps runaway scripts, scrapers, and accidental infinite loops.

> The mental flip that helps: a rate limit is not a punishment for *you*. It's a promise to *everyone*
> that the service will still be there in a minute. You're one of the "everyone" it's protecting.

## How limits are usually counted: the token bucket

You'll hear "100 requests per minute" and imagine a strict odometer that resets on the minute. Real
systems are usually gentler and smarter than that, and the most common model is the **token bucket**.

Picture a bucket that holds, say, 100 tokens. Every request you make takes one token out. The bucket
refills steadily - for example, a token or two every second - up to its maximum of 100. As long as the
bucket has tokens, your request goes through. When it's empty, you're throttled until it refills a bit.

```text
bucket capacity: 100 tokens        refill: ~2 tokens/second

[||||||||||||||||||||] 100   <- full: you can burst 100 requests right now
   you fire 100 fast
[                    ] 0     <- empty: the next request gets a 429
   wait ~5 seconds
[||||||||||          ] ~10   <- refilled a little: ~10 requests allowed again
```

The bucket let you **burst** - fire a clump of requests quickly using saved-up tokens - but it caps your
*sustained* rate at the refill speed. Burst for the spike, then settle into the steady pace. This is why
"100 per minute" doesn't mean "exactly one every 0.6 seconds": you can spend fast for a moment, not
*forever*.

📝 **Terminology.** *Burst* = a short clump of requests above your steady rate, allowed because tokens
accumulated while you were idle. *Sustained rate* = the long-run average the bucket refills at. Other
models exist (fixed windows, sliding windows, leaky buckets), but token bucket is the one to picture
first - it explains both the bursting and the throttling you'll actually see.

## What the server tells you: the 429 and friends

When you've run out, a well-behaved API doesn't go silent - it answers with a clear status and, often,
instructions. The signal you'll see most is:

```http
HTTP/1.1 429 Too Many Requests
Retry-After: 30
Content-Type: application/json

{ "error": "rate_limited", "message": "Too many requests. Try again in 30 seconds." }
```

The server said **429 Too Many Requests** - "I heard you, but you're going too fast" - and the
`Retry-After: 30` header is it *telling you exactly how long to wait*: 30 seconds. This is the single most
useful header in this whole guide. When the server hands you a number, you don't guess your backoff - you
obey the number.

⚠️ **`429` is not an error you caused by being wrong.** Your request was *valid*; you were merely too
frequent. That's different from a `400` (your request is malformed) or a `403` (you're not allowed). A
`429` means "good request, bad timing" - so retrying it later is the *right* move, where retrying a `400`
would be pointless.

Two more things you'll commonly meet:

- **`Retry-After` can be a date instead of seconds.** Some servers send `Retry-After: Wed, 21 Oct 2026
  07:28:00 GMT` - an absolute time to wait until. Same meaning, different format; handle both.
- **`X-RateLimit-*` headers (a widespread convention, not an official standard).** Many APIs include
  hints on *every* response, not only the `429`:

```http
HTTP/1.1 200 OK
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 7
X-RateLimit-Reset: 1730531280
```

Even on a successful `200`, the server quietly told you the bucket holds **100**, you have **7** left, and
it resets at that Unix timestamp. A thoughtful client watches `Remaining` and *slows down on its own* as
it approaches zero - getting throttled is something you can often see coming and avoid.

## A status-code cheat sheet for "should I even retry?"

Not every failure is a rate limit, and not every failure is worth retrying. Here's the quick read:

| Status | Means | Retry? |
|---|---|---|
| `429 Too Many Requests` | You're going too fast | **Yes** - wait (use `Retry-After`), then retry |
| `503 Service Unavailable` | Server is overloaded or down briefly | **Yes** - back off and retry |
| `502 / 504` Gateway errors | A proxy couldn't reach the real server in time | **Usually** - transient; retry with backoff |
| `500 Internal Server Error` | Something broke server-side | **Maybe** - could be transient; retry cautiously |
| `400 Bad Request` | Your request is malformed | **No** - retrying sends the same bad request |
| `401 / 403` | Not authenticated / not allowed | **No** - fix credentials/permissions instead |
| `404 Not Found` | The thing isn't there | **No** - it won't appear by asking again |

💡 **Key point.** Retry the *transient* failures (`429`, `5xx`) and leave the *permanent* ones (`4xx`
except `429`) alone. Retrying a `400` in a loop is the classic way to turn one bug into a storm of
identical, doomed requests.

## For builders

If you're building the *client*, the cheapest reliability win is reading what the server already tells
you: honor `Retry-After`, and if `X-RateLimit-Remaining` is getting low, ease off before you hit the
wall instead of bouncing off it. The server has done the hard work of telling you how to behave well - 
most rate-limit pain comes from clients that ignore those signals and retry blindly. Phase 2 turns
"retry later" into a precise, safe recipe.

```quiz
[
  {
    "q": "What does an HTTP 429 status code mean?",
    "choices": [
      "Your request was malformed and should be fixed before retrying",
      "Your request was valid but you're sending too many too quickly",
      "You're not authorized to access the resource",
      "The server has permanently shut down the endpoint"
    ],
    "answer": 1,
    "explain": "429 Too Many Requests means a valid request arrived too frequently - good request, bad timing. That's why retrying later is the correct response."
  },
  {
    "q": "In a token-bucket rate limiter, why can you sometimes fire a burst of requests faster than the stated steady rate?",
    "choices": [
      "The limit only applies to POST requests, not GET",
      "Tokens accumulate up to the bucket's capacity while you're idle, so you can spend a saved-up clump at once",
      "The server ignores the first minute of any session",
      "Bursting is a bug that providers haven't fixed yet"
    ],
    "answer": 1,
    "explain": "The bucket fills up to its capacity when you're not using it; a burst spends those accumulated tokens. Your sustained rate is still capped at the refill speed."
  },
  {
    "q": "A 429 response includes a `Retry-After: 30` header. What's the right thing to do?",
    "choices": [
      "Retry immediately - the header is only a suggestion",
      "Treat it as a permanent failure and stop calling the API",
      "Wait 30 seconds before retrying, because the server told you exactly how long",
      "Double the value and wait 60 seconds to be safe"
    ],
    "answer": 2,
    "explain": "Retry-After is the server explicitly telling you how long to wait. When you get a concrete number, honor it instead of guessing."
  }
]
```

Watch it animated: [API rate limiting](/explainers/RateLimiting.dc.html)


---

# Retrying Without Making It Worse

You know the failure is probably temporary, so retrying feels obvious - and it is, right up until the
naive version bites you. Two failure stories haunt this phase. One: you retry in a tight `while` loop and
turn a momentary blip into a self-inflicted flood. Two: you retry a "create charge" call that *actually
succeeded* but timed out on the reply - and bill the customer twice. This phase gives you the recipe that
avoids both: wait longer each time, add randomness, cap your effort, and only retry what's safe to repeat.

## Don't retry instantly: exponential backoff

The first instinct - retry immediately, maybe in a loop - is the worst one. If the service is struggling,
a wall of instant retries is more load, which makes it struggle more. The fix is to **wait longer after
each failure**, and the standard pattern is **exponential backoff**: each wait roughly doubles.

```text
attempt 1  ->  fails  ->  wait ~1s
attempt 2  ->  fails  ->  wait ~2s
attempt 3  ->  fails  ->  wait ~4s
attempt 4  ->  fails  ->  wait ~8s
attempt 5  ->  give up (out of budget)
```

Instead of hammering, you backed off - 1, 2, 4, 8 seconds - giving the service room to recover. A short
hiccup gets caught by the early, quick retries; a longer outage doesn't get a flood from you, because your
waits stretch out fast. The doubling means you make only a handful of attempts even across many seconds.

📝 **Terminology.** *Backoff* = the wait you insert before a retry. *Exponential* = that wait grows
multiplicatively (×2 each time here), not by a fixed step. The base delay (1s here) and the multiplier
(2×) are yours to tune; doubling from ~1s is a sane default.

## The bit everyone forgets: jitter

Here's the trap that catches teams who "did backoff right." Imagine the service blips and a thousand of
your clients all fail at the same instant. With pure exponential backoff, all thousand wait *exactly* 1
second, then all thousand retry *at the same moment*. You've synchronized a thousand clients into a
drumbeat of simultaneous spikes - every retry round lands as one giant wave. That's a self-made
**thundering herd** (much more on this in Phase 3).

The fix is **jitter**: add randomness to each wait so the retries spread out instead of clumping.

```text
without jitter (everyone waits exactly 2s):
   all clients ->|             |<- one big spike at t=2s

with jitter (each waits a random slice up to 2s):
   clients spread ->| . . . . . |<- load smeared across the window
```

Jitter smeared the retry wave across the whole window instead of stacking it into one spike. The service
sees a gentle trickle it can actually absorb, rather than a synchronized wall.

A simple, widely used recipe is **"full jitter"**: instead of waiting the full computed backoff, wait a
*random* amount between zero and that backoff.

```python runnable
import random

base = 1.0          # base delay in seconds
cap = 30.0          # never wait longer than this
for attempt in range(5):
    backoff = min(cap, base * (2 ** attempt))   # 1, 2, 4, 8, 16
    wait = random.uniform(0, backoff)           # full jitter: random slice of it
    print(f"attempt {attempt + 1}: backoff window={backoff:.0f}s, actually wait={wait:.2f}s")
```

The backoff *window* still doubles (1, 2, 4, 8, 16s), but the actual wait is a random point inside that
window - so two clients running this same code almost never wait the same amount, and their retries don't
collide. The `cap` keeps a long outage from producing absurd 10-minute waits.

💡 **Key point.** Exponential backoff *without* jitter is a known foot-gun: it turns many clients into a
synchronized herd. Backoff decides *how long*; jitter decides *who goes when*. You need both.

## Know when to stop: a retry budget

Retrying forever is its own bug. If the service is genuinely down for an hour, infinite retries pile up
work, hold connections open, and can take *your* service down too. So you give yourself a **budget** and
give up gracefully when it's spent.

A budget is usually one or both of:

- **Max attempts** - e.g. "try at most 5 times, then fail."
- **A deadline** - e.g. "keep retrying, but stop after 30 seconds total, no matter the attempt count."

```text
budget: 5 attempts OR 30s total, whichever comes first

if attempts >= 5            -> stop, surface the error
if elapsed_seconds >= 30    -> stop, surface the error
otherwise                   -> back off (with jitter) and try again
```

You bounded the effort in *both* dimensions. The deadline matters because a few exponential backoffs can
quietly add up to a long time - you want a hard ceiling so a stuck operation doesn't hang a user's request
for minutes. When the budget runs out, you stop and report a clean failure rather than retrying into the
void.

⚠️ **Beware nested retries.** If your client retries, and the library *it* calls also retries, and the
service behind *that* retries too, the attempts multiply: 3 × 3 × 3 = 27 real requests for one logical
call. This "retry amplification" is a notorious way to accidentally DDoS yourself. Pick *one* layer to
own retries and turn the others off.

## The golden rule: only retry what's safe to repeat

This is the rule that separates a safe retry from a money-losing one. Before retrying, ask: **if this
request already ran, is running it again harmless?**

That property has a name you met if you read the webhooks guide: **idempotency**. An operation is
idempotent if doing it twice has the same effect as doing it once.

- `GET /users/42` - reading. Run it a hundred times; nothing changes. **Safe to retry.**
- `DELETE /sessions/abc` - deleting. Already gone? Deleting again is still "gone." **Safe to retry.**
- `PUT /users/42 {name: "Sam"}` - setting to a value. Same result every time. **Safe to retry.**
- `POST /charges {amount: 5000}` - *creating* a charge. Retry a timed-out one and you might **bill twice.**
  **Not safe** - unless you make it safe.

The danger case is the timed-out write. Your `POST` to create a charge may have *succeeded* on the server
while the *response* got lost on the way back. From your side it looks like a failure. Retry it naively
and you've created a second charge.

```text
you: POST /charges  ----------------->  server: charge created
you: (no response - network dropped it)  <-- response lost here
you: "looks failed, retry!"  -------->  server: ANOTHER charge created  💸
```

The lost *response*, not a lost request, is what makes blind retries dangerous on writes. The work
happened; only your confirmation went missing. A naive retry does the work a second time.

## The fix for unsafe retries: idempotency keys

You can't make "create a charge" naturally idempotent - each call is meant to create something new. So
serious APIs give you a tool: an **idempotency key**. You generate a unique ID for the *logical*
operation and send it with the request. The server remembers that key. If it sees the same key twice, it
returns the *original* result instead of doing the work again.

```http
POST /v1/charges HTTP/1.1
Idempotency-Key: a1b2c3-charge-order-9921
Content-Type: application/json

{ "amount": 5000, "currency": "usd", "source": "tok_visa" }
```

You stamped this charge with a key tied to the *order*, not the attempt. The first time the server sees
`a1b2c3-charge-order-9921`, it creates the charge and records the result against the key. When your retry
arrives with the *same* key, the server recognizes it, skips creating a second charge, and returns the
first charge's response. One key, one charge - no matter how many times the network makes you retry.

📝 **Terminology.** *Idempotency key* = a caller-generated unique ID for one logical operation, sent so
the server can deduplicate retries. Generate it *once per operation* and reuse it across that operation's
retries - if you make a fresh key every attempt, you've defeated the whole point.

💡 **Key point.** Retry idempotent requests freely. For non-idempotent ones (most `POST`s that create or
charge), either don't retry, or use an idempotency key so the server can make the retry safe. "Could this
double-charge?" is the question to ask before every retry of a write.

## Putting it together

A safe retry loop, in plain terms:

```text
for each attempt, until the budget is spent:
    send the request
    if it succeeded                      -> return the result
    if it failed with a 4xx (not 429)    -> stop; retrying won't help
    if it failed with 429 or 5xx:
        if the response had Retry-After  -> wait that long
        else                             -> wait backoff(attempt) with jitter
        (writes carry the same idempotency key on every attempt)
budget spent -> surface a clean error
```

Every habit from this phase is in there - honor `Retry-After` when given, otherwise exponential backoff
*with jitter*, stop on permanent errors, respect a budget, and keep the idempotency key constant across
attempts. That loop turns a flaky dependency into a non-event.

## For builders

Reach for a battle-tested retry library before hand-rolling this - most ecosystems have one that gives you
backoff, jitter, and budgets in a few lines, and they've already fixed the bugs you'd hit. Your real job
is *configuration with judgment*: set sane caps, add jitter (confirm the default actually does), retry
only transient statuses, and decide per-endpoint whether a retry could double a side effect. Phase 3
covers what to do when even a perfect retry policy isn't enough - when a dependency is *down*.

```quiz
[
  {
    "q": "Why is jitter added on top of exponential backoff?",
    "choices": [
      "To make retries happen faster overall",
      "To spread retries out in time so many clients don't all retry at the exact same instant",
      "To encrypt the request so it can't be replayed",
      "Because the Retry-After header requires it"
    ],
    "answer": 1,
    "explain": "Pure exponential backoff makes many clients wait the same fixed amount and retry in sync - a self-made thundering herd. Jitter randomizes each wait so the load spreads out."
  },
  {
    "q": "A POST that creates a charge times out with no response. Why is blindly retrying it dangerous?",
    "choices": [
      "POST requests can never be retried under any circumstances",
      "The charge may have actually succeeded and only the response was lost, so a retry could create a second charge",
      "The server will reject any retry automatically",
      "Retrying a POST always corrupts the request body"
    ],
    "answer": 1,
    "explain": "A lost response, not a lost request, is the trap: the work may already be done. A naive retry repeats a non-idempotent operation and can double-charge - unless you use an idempotency key."
  },
  {
    "q": "What is the correct way to use an idempotency key across retries of the same charge?",
    "choices": [
      "Generate a brand-new key for every retry attempt",
      "Use the same key on every attempt of that one logical operation so the server deduplicates them",
      "Only send the key on the final attempt",
      "Let the server generate the key and ignore it on the client"
    ],
    "answer": 1,
    "explain": "The key identifies the logical operation, not the attempt. Keeping it constant across retries lets the server recognize duplicates and return the original result instead of redoing the work."
  }
]
```


---

# When Retrying Isn't Enough

You did everything right in Phase 2 - backoff, jitter, a budget, idempotency keys. And there's still a
failure mode that recipe can't fix: the dependency isn't *blipping*, it's *down*. When that happens, even
polite retries from enough clients become an attack, and retrying a corpse wastes your own resources
and makes recovery harder. This phase is about recognizing "stop trying for a bit" as the resilient move,
and the patterns that do it for you.

## The thundering herd, properly

You met the herd in Phase 2 as the reason for jitter. Now see its bigger, scarier form. A popular service
goes down for thirty seconds. During that window, every client's request fails. The instant it comes back
up, *all of those clients retry at once* - plus all the new traffic that arrived in the meantime. The
service, already fragile from rebooting, gets slammed by a spike far larger than its normal load and falls
over again. Now it's down longer, more clients pile up, and the next recovery attempt faces an even bigger
wall. That's a **retry storm**, and it can keep a service down long after the original problem is gone.

```mermaid
flowchart TD
  A[service goes down 30s] --> B[every client's calls fail]
  B --> C[service recovers]
  C --> D[all clients retry at once]
  D --> E[spike far above normal load]
  E --> F[service falls over again]
  F --> B
```

The diagram is a *loop* on purpose - the retries feed the very outage they're reacting to. Jitter softens
each wave, but past a certain scale, smearing the spike isn't enough. You need clients that recognize
"this thing is down" and *stop calling it entirely* for a while. That's the circuit breaker.

## The circuit breaker

The name comes from the electrical panel in your home. When a circuit draws too much current, the breaker
*trips* - it cuts the connection so the wiring doesn't catch fire. You don't keep jamming the switch back;
you wait, fix the cause, then reset it. A software **circuit breaker** does the same for a failing
dependency: after enough failures, it *opens* and makes your calls fail instantly - without even touching
the network - for a cooldown period. This protects two parties at once: the struggling service gets
breathing room, and *you* stop wasting time and threads waiting on calls that are going to fail anyway.

A breaker has three states:

```mermaid
stateDiagram-v2
  [*] --> Closed
  Closed --> Open: failures cross threshold
  Open --> HalfOpen: cooldown elapsed
  HalfOpen --> Closed: trial call succeeds
  HalfOpen --> Open: trial call fails
```

- **Closed** - normal. Calls flow through. The breaker counts failures.
- **Open** - tripped. Calls fail *instantly* (no network call) for a cooldown window. This is the herd
  protection: a thousand clients with open breakers send *zero* requests instead of a storm.
- **Half-open** - after the cooldown, the breaker lets *one* trial call through. If it succeeds, the
  service is back: close the breaker and resume. If it fails, re-open and wait again.

The half-open state is the clever bit - instead of all clients rushing back the instant the cooldown ends
(which would re-create the herd), the breaker probes with a single request and only fully reopens the
floodgates once that probe proves the service is healthy. It recovers *gently*.

📝 **Terminology.** *Trip / open* = the breaker has decided the dependency is unhealthy and is short-
circuiting calls. *Cooldown* = how long it stays open before testing again. *Half-open* = the cautious
"let one through and see" probe state. Most resilience libraries implement all of this for you; you mostly
tune the thresholds.

💡 **Key point.** Retries handle a service that's *struggling*; a circuit breaker handles a service that's
*down*. Retrying says "try again soon"; the breaker says "stop trying for now." A robust client uses both
 - retry the blips, trip the breaker on a real outage.

## When the breaker is open: degrade gracefully

A breaker that's open means a feature is unavailable *right now*. The question becomes: what does your app
do instead of the call? Failing instantly is good for the *dependency*, but your user still needs an
answer. The art is **graceful degradation** - a reduced but working experience instead of a hard crash.

Common fallbacks, roughly best to worst:

- **Serve a cached or stale value.** "Last known price" beats a spinner that never resolves.
- **Use a sensible default.** Recommendations service down? Show the popular-items list.
- **Queue the work for later.** Can't process the upload now? Accept it, enqueue it, confirm when done.
- **Fail with a clear, plain message.** "Search is temporarily unavailable, try again shortly" - not a
  blank page or a stack trace.

```text
checkout flow, fraud-check service is down (breaker open):

  bad:   call fraud-check -> hang 30s -> time out -> 500 -> customer lost
  good:  breaker open -> skip the call -> flag order for manual review -> complete checkout
```

The good path didn't pretend the dependency was up and didn't crash the whole checkout because one service
was down. It chose a safe fallback (review the order later) so the *core* flow - taking the order - still
worked. Deciding these fallbacks ahead of time is what separates an app that *bends* in an outage from one
that *breaks*.

## Being a genuinely good client

Pull the whole guide together into the posture that keeps you off everyone's incident reports - including
your own:

- **Listen to what the server tells you.** Honor `Retry-After`. Watch `X-RateLimit-Remaining` and slow
  down *before* you hit the wall, instead of bouncing off it repeatedly.
- **Never retry without backoff and jitter.** Instant or synchronized retries are how you turn one blip
  into an outage.
- **Bound your effort.** A retry budget and (where it fits) a circuit breaker keep a downstream failure
  from becoming *your* failure.
- **Retry only what's safe.** Idempotent calls freely; non-idempotent writes only with an idempotency key.
- **Cache and batch to need fewer calls.** The cheapest request is the one you didn't have to make. Cache
  responses that don't change often; batch many small calls into one where the API supports it.
- **Spread out scheduled work.** If a hundred of your machines all hit an API exactly on the minute via
  cron, you've built a thundering herd on a timer. Add a little randomness to scheduled jobs too.

> The throughline of this entire guide: a rate limit or an outage is the *system asking you to behave*.
> A good client *listens* - it slows down when asked, backs off when refused, and stops knocking when the
> door is clearly bolted. That's not only courtesy; it's what keeps the service (and you) alive.

## For builders

You don't have to hand-build breakers or backoff - mature resilience libraries exist in every major
ecosystem (think "retry + circuit-breaker" toolkits) and they've already handled the edge cases. Your
engineering judgment goes into the *policy*: which dependencies get a breaker, what the thresholds and
cooldowns are, and - most importantly - what each feature *falls back to* when its dependency is gone.
Decide those fallbacks during calm design time, not at 2am during the incident. For where this fits in the
wider picture of how services talk to each other, see [REST APIs, Explained](/guides/rest-apis-explained)
and [What an API Is](/guides/what-an-api-is).

## Recap

1. A **retry storm / thundering herd** is when many clients retry in sync and re-crash a service that was
   trying to recover. Jitter softens it; at scale you need more.
2. A **circuit breaker** *opens* after repeated failures and makes calls fail instantly during a cooldown,
   sparing both the downed service and your own resources.
3. The breaker's **half-open** state probes with a single call so recovery is gentle, not another stampede.
4. When the breaker is open, **degrade gracefully** - cache, default, queue, or a clear message - instead
   of crashing the whole flow.
5. A **good client** listens to the server's signals, always backs off with jitter, bounds its effort,
   retries only what's safe, and makes fewer calls by caching, batching, and de-syncing scheduled jobs.

You came in panicking at `429`s and a tight retry loop that made things worse. You leave knowing *why*
APIs push back, how to retry so it helps instead of harms, and what to do when retrying alone won't cut
it. That's the full toolkit for calling a flaky, throttled, or down dependency like someone who's been
paged for it before - and built so they won't be again.

```quiz
[
  {
    "q": "What problem does a circuit breaker solve that retries with backoff and jitter do not?",
    "choices": [
      "It encrypts requests to a failing service",
      "It stops a client from calling a service that is genuinely down, instead of repeatedly waiting on doomed calls",
      "It guarantees the request will eventually succeed",
      "It removes the need for a Retry-After header"
    ],
    "answer": 1,
    "explain": "Retries handle a struggling service; a breaker handles a down one. When open, it fails calls instantly - sparing the downed service and freeing your own resources instead of waiting on calls that will fail."
  },
  {
    "q": "Why does a circuit breaker use a 'half-open' state instead of fully closing as soon as the cooldown ends?",
    "choices": [
      "To make the code more complicated on purpose",
      "So it can probe with a single trial call and only fully reopen if the service is actually healthy, avoiding a new stampede",
      "Because half-open requests are faster than closed ones",
      "To bill the dependency for the downtime"
    ],
    "answer": 1,
    "explain": "If every client resumed at once when the cooldown ended, that's another thundering herd. Half-open lets one trial call test the waters, so recovery is gentle."
  },
  {
    "q": "A dependency's circuit breaker is open during checkout's optional fraud check. What is the most graceful behavior?",
    "choices": [
      "Keep calling the fraud service in a tight loop until it answers",
      "Crash the entire checkout with a 500 error",
      "Skip the call and flag the order for later manual review so checkout still completes",
      "Silently approve nothing and show the customer a blank page"
    ],
    "answer": 2,
    "explain": "Graceful degradation keeps the core flow working with a safe fallback. Flagging the order for review lets checkout complete instead of crashing the whole flow because one service is down."
  }
]
```
