# RabbitMQ, From Zero

> The classic message broker: exchanges, queues, and bindings route messages to workers, with acknowledgements and dead-letter queues for reliable delivery.


---

# RabbitMQ, From Zero

You have a web request that needs to send an email, resize an image, or charge a card, and you do not want the user staring at a spinner while it happens. So you reach for a queue. Then you open the RabbitMQ docs and meet exchanges, bindings, routing keys, vhosts, and a dozen acknowledgement modes, and the simple idea of "put work on a list" suddenly has more knobs than your stove. This guide gives you the mental model first, so every knob has a place to live.

By the end you will know what each piece does, how a message actually travels from your code to a worker, and how to make delivery reliable instead of hopeful.

## How to read this

Read the phases in order. Phase 1 builds the picture in your head: the broker, the post office, the difference between an exchange and a queue. Phase 2 is the daily work: publishing, consuming, the exchange types, prefetch. Phase 3 is what bites you in production: lost messages, poison messages, dead-letter queues, and how RabbitMQ differs from Kafka so you pick the right tool. If you have never touched a queue before, the broader idea is covered in [/guides/webhooks-and-message-queues](/guides/webhooks-and-message-queues).

## The phases

1. [Phase 1: The Smart Post Office](01-the-smart-post-office.md) - the mental model: broker, exchanges, queues, bindings, and why the producer never names a queue.
2. [Phase 2: Publishing and Consuming for Real](02-publishing-and-consuming.md) - the everyday loop: exchange types, routing keys, acknowledgements, and prefetch.
3. [Phase 3: When Delivery Goes Wrong](03-when-delivery-goes-wrong.md) - durability, redelivery, poison messages, dead-letter queues, and RabbitMQ versus Kafka.


---

# Phase 1: The Smart Post Office

You probably already have a picture in your head when someone says "queue": a list, items go in one end, workers pull them off the other. That picture is right, and it is also the source of most early confusion with RabbitMQ. Because in RabbitMQ, your code that sends a message does not put it on a queue. It hands it to something else entirely, and that something decides where it lands.

Hold onto one image for this whole phase: **a post office**. You drop a letter in the slot. You do not walk it to a specific mailbox. The post office reads the address and routes it. RabbitMQ is that post office, and it is a *smart* one - routing is its whole job.

## The three nouns you must keep separate

RabbitMQ speaks a protocol called AMQP (Advanced Message Queuing Protocol). The protocol gives you three building blocks, and the entire mental model is keeping them straight:

- **Exchange** - the mail slot. Producers publish *here*, never directly to a queue. The exchange holds nothing; it routes.
- **Queue** - the mailbox. This is the buffer that actually stores messages until a consumer takes them.
- **Binding** - the forwarding rule that connects an exchange to a queue. "Letters addressed like *this* go into *that* mailbox."

A producer publishes a message to an exchange with a **routing key** (think of it as the address on the envelope). The exchange looks at its bindings and copies the message into every queue whose binding matches. Consumers read from queues. That is the full circuit.

```text
producer ──publish(routing_key)──▶ EXCHANGE
                                      │  (checks bindings)
                       ┌──────────────┼──────────────┐
                       ▼              ▼               ▼
                    queue.A        queue.B        (no match → dropped)
                       │
                       ▼
                    consumer
```

*What just happened:* the producer never named a queue. It named an exchange and an address. Routing is the broker's decision, driven by bindings - which means you can add, remove, or re-point consumers without ever touching producer code.

## Why this indirection is the whole point

The first time you see it, the exchange feels like a pointless middleman. Why not let the producer write straight to the queue? Because that one layer of indirection is what makes the system flexible.

Picture an `order.placed` event. Today one service emails a receipt. Next month, billing wants the same event, and analytics wants it too. If the producer wrote directly to a queue, you would change the producer three times. With an exchange, the producer keeps publishing `order.placed` to the same exchange forever; you bind two new queues and the new consumers start receiving. The producer never learns they exist.

> This is the loose coupling that message brokers exist to give you. The sender knows *what happened*, not *who cares*. Adding a new listener is a config change on the consumer side, not a code change on the sender side.

## What lives where, physically

A RabbitMQ server is called a **broker** (or node). Inside it, everything is namespaced by a **virtual host** (vhost) - a logical partition, like a database name. Exchanges, queues, and bindings all belong to a vhost. A connection picks a vhost when it authenticates, and it can only see that vhost's resources. Most small setups use the default vhost `/` and never think about it again; you reach for separate vhosts when you want hard isolation between, say, staging and production traffic on one broker.

```bash
# Peek at what's actually defined on a running broker
rabbitmqctl list_exchanges name type
rabbitmqctl list_queues name messages consumers
rabbitmqctl list_bindings
```

*What just happened:* these admin commands show the three nouns live. `list_queues` with `messages consumers` is the one you will run most - it tells you how many messages are waiting and how many workers are attached, which is the first thing you check when a queue is backing up.

## A few exchanges already exist

When you connect, the broker already has some exchanges defined for you, including the **default exchange** - a nameless direct exchange (its name is the empty string `""`). It has a quiet special rule: every queue is automatically bound to it using the queue's own name as the routing key.

That rule is why the simplest possible RabbitMQ example looks like it breaks the model:

```text
publish to exchange="", routing_key="task_queue"
   → lands in the queue literally named "task_queue"
```

*What just happened:* publishing to the default exchange with a routing key equal to a queue name delivers straight to that queue. It looks like "publishing to a queue," but under the hood it is still exchange-then-binding - the binding was just created for you automatically. Useful for quick demos; in real systems you almost always declare your own exchange so routing is explicit.

## For builders

When you sketch a RabbitMQ design on a whiteboard, draw the exchange as a box and the queues as cylinders hanging off it, with labeled arrows (the bindings) between them. If you can draw that picture for your system, you understand it. If you cannot, you do not yet know where your messages go - and "I am not sure where this message ends up" is the bug you least want in production.

```quiz
[
  {
    "q": "In RabbitMQ, where does a producer send a message?",
    "choices": ["Directly to a queue", "To an exchange, with a routing key", "To a consumer", "To a binding"],
    "answer": 1,
    "explain": "Producers publish to an exchange with a routing key. The exchange uses its bindings to decide which queues receive the message."
  },
  {
    "q": "What is a binding?",
    "choices": ["A connection from a producer to the broker", "A worker process that reads messages", "A rule linking an exchange to a queue", "A copy of a message in storage"],
    "answer": 2,
    "explain": "A binding is the routing rule that connects an exchange to a queue, telling the exchange which messages to forward where."
  },
  {
    "q": "Why publish to an exchange instead of straight to a queue?",
    "choices": ["It is faster", "It encrypts the message", "New consumers can be added by binding new queues, with no change to the producer", "Queues cannot store messages"],
    "answer": 2,
    "explain": "The exchange layer decouples sender from receiver: you add listeners by adding bindings, never by editing producer code."
  }
]
```


---

# Phase 2: Publishing and Consuming for Real

You have the post-office picture. Now you actually send mail. This phase is the loop you will write a hundred times: declare your plumbing, publish, consume, acknowledge. The new decision is *which kind of exchange* to use, because that single choice changes how a routing key behaves and therefore how your messages spread.

The examples use the AMQP CLI shape and pseudocode that maps cleanly onto every client library (Python's `pika`, Node's `amqplib`, Go, Java). The names of the operations are the same everywhere because they come from the protocol, not the library.

## Declare before you use

Exchanges and queues do not spring into being. You **declare** them - an idempotent operation that creates the thing if missing and otherwise checks it matches. Declaring is safe to run every time your app starts; both producer and consumer typically declare what they need, so neither depends on the other booting first.

```text
exchange.declare  name="orders"  type="topic"  durable=true
queue.declare     name="email_q" durable=true
queue.bind        queue="email_q" exchange="orders" routing_key="order.placed"
```

*What just happened:* you created a durable topic exchange, a durable queue, and a binding that routes `order.placed` messages into `email_q`. `durable=true` means the broker remembers these definitions across a restart (we cover durability of the *messages* in Phase 3 - declaration durability and message durability are separate things).

## The exchange types - pick by routing shape

There are four exchange types. You will use three of them constantly; the fourth is a niche tool.

**Direct** - exact-match routing. A message goes to queues whose binding key equals the routing key, string for string.

```text
exchange type=direct
bind  queue=pdf_q   routing_key="pdf"
bind  queue=image_q routing_key="image"

publish routing_key="pdf"   → pdf_q only
publish routing_key="image" → image_q only
```

*What just happened:* direct is the workhorse for "route this job to the worker that handles its type." One routing key, one matching queue (or several if multiple queues share the same binding key).

**Topic** - pattern routing on dotted keys. Bindings use wildcards: `*` matches exactly one word, `#` matches zero or more words.

```text
exchange type=topic
bind queue=all_orders  routing_key="order.#"
bind queue=eu_orders   routing_key="order.eu.*"

publish routing_key="order.eu.placed" → all_orders AND eu_orders
publish routing_key="order.us.refunded" → all_orders only
```

*What just happened:* topic exchanges let one publish fan out by category. `order.#` catches everything under `order`; `order.eu.*` catches one more word after `order.eu`. This is the flexible default most teams reach for.

**Fanout** - ignore the routing key entirely; copy to *every* bound queue.

```text
exchange type=fanout
bind queue=cache_q
bind queue=audit_q
bind queue=search_q

publish (any routing key) → cache_q AND audit_q AND search_q
```

*What just happened:* fanout is broadcast. Every consumer gets its own copy. Use it for "this event happened, everyone who cares should react" - cache invalidation, live notifications.

The fourth type, **headers**, routes on message header attributes instead of the routing key. It is rarely worth the complexity; reach for topic first and only consider headers if you genuinely need to match on multiple non-hierarchical attributes.

> Mental shortcut: **direct** = "to this one", **topic** = "to whoever subscribed to this pattern", **fanout** = "to everyone".

## Work queues: one queue, many workers

The most common real pattern is not fancy routing - it is one queue with several identical consumers chewing through jobs in parallel. When multiple consumers subscribe to the same queue, RabbitMQ delivers each message to *one* of them. That is competing consumers, and it is how you scale a worker pool: start more processes, they share the load automatically.

```text
queue=task_q  ← consumer #1
              ← consumer #2
              ← consumer #3

100 messages → spread across the three, each message to exactly one worker
```

*What just happened:* a single queue with three consumers triples your throughput with zero code change. Each message is handled once. Contrast this with fanout, where each *queue* gets a copy - here it is one queue, and the *consumers* compete.

## Acknowledgements: the broker needs to know you finished

Here is the part people skip and regret. When a consumer receives a message, the broker does not consider it done. It waits for an **acknowledgement** (ack). Until the ack arrives, the broker considers the message "in flight" - delivered but unconfirmed.

```text
broker → deliver message → consumer
consumer does the work...
consumer → basic.ack → broker   (now the broker deletes it)
```

*What just happened:* the ack is your promise that the work succeeded. If your consumer crashes after receiving but before acking, the broker sees the connection drop, marks the message unacknowledged, and **redelivers** it to another consumer. No work is lost. This is the heart of reliable delivery.

There are two ack modes, and the default in raw AMQP matters:

- **Manual ack** (what you want): you call `basic.ack` after the work succeeds. Crash mid-work → redelivery. This is the safe choice for anything that matters.
- **Auto ack** (`auto_ack=true`): the broker treats the message as done the instant it ships it out. Fast, but if your consumer dies before finishing, the message is gone forever. Only use it for data you can afford to lose.

```text
# Safe consumer loop, in any language
consume(queue="task_q", auto_ack=false)
on_message(msg):
    do_the_work(msg)        # if this throws, no ack is sent
    basic.ack(msg)          # only reached on success
```

*What just happened:* by acking *after* the work and only on success, a crash leaves the message unacknowledged, and the broker hands it to someone else. Move the ack to the top and you have silently switched to "lose work on crash."

## Prefetch: stop one greedy worker from hoarding

By default a consumer will accept as many unacknowledged messages as the broker wants to push. So if you have three workers and 300 messages, one fast-connecting worker might grab 200 while the others sit idle - then if it is slow, your queue drains slowly even though two workers are bored.

The fix is **prefetch** (`basic.qos` with `prefetch_count`): cap how many unacknowledged messages one consumer may hold at once.

```text
basic.qos(prefetch_count=1)
```

*What just happened:* with prefetch 1, a consumer gets exactly one message, must ack it before getting the next, and slow workers stop starving fast ones. The work spreads by actual capacity, not by who grabbed first. For quick uneven jobs, 1 is a fine default; for high-throughput tiny messages, a small number like 10 to 50 reduces round-trips. Tune it; do not leave it unlimited.

## In the wild

A typical background-job setup: a topic exchange named after your domain, one durable queue per job type bound with a specific routing key, manual acks, `prefetch_count` tuned to your job size, and a worker pool of identical consumers per queue. That handful of decisions covers the vast majority of production RabbitMQ usage. The event-driven thinking behind *why* you publish events at all is worth the detour in [/guides/event-driven-architecture](/guides/event-driven-architecture).

```quiz
[
  {
    "q": "Which exchange type ignores the routing key and copies the message to every bound queue?",
    "choices": ["direct", "topic", "fanout", "headers"],
    "answer": 2,
    "explain": "Fanout broadcasts: every bound queue receives a copy, regardless of routing key. Direct and topic both route by key."
  },
  {
    "q": "With manual acknowledgement, when should a consumer call basic.ack?",
    "choices": ["Immediately on receiving the message", "After the work succeeds", "Before doing any work", "Never; the broker acks automatically"],
    "answer": 1,
    "explain": "Ack after success. If the consumer crashes before acking, the broker redelivers the message, so no work is lost."
  },
  {
    "q": "What does setting prefetch_count=1 accomplish?",
    "choices": ["Deletes the queue after one message", "Limits a consumer to one unacknowledged message at a time, spreading load fairly", "Makes the exchange a fanout", "Disables acknowledgements"],
    "answer": 1,
    "explain": "Prefetch caps in-flight unacked messages per consumer, preventing one worker from hoarding the queue while others sit idle."
  }
]
```


---

# Phase 3: When Delivery Goes Wrong

Everything in Phase 2 works on a good day. This phase is about the bad days: the broker restarts, a message can never be processed, a buggy consumer rejects the same message forever, or your queue quietly grows until memory runs out. RabbitMQ has answers for each, but they are opt-in. The defaults favor speed, and "I assumed it was durable" is a classic 3am lesson.

## Durability: surviving a broker restart

By default, queues and messages live in memory. Restart the broker and they vanish. Reliability needs three things turned on together - miss any one and you still lose data:

```text
1. queue.declare  durable=true        # the queue definition survives restart
2. publish with   delivery_mode=2     # the message is persisted to disk
3. exchange.declare durable=true      # the exchange definition survives restart
```

*What just happened:* a durable queue holding non-persistent messages still loses the messages on restart - the queue comes back empty. You need the queue durable *and* each message marked persistent (`delivery_mode=2`, sometimes exposed as a `persistent=true` flag). Persistence costs disk writes, so it is a deliberate trade: durability for throughput.

> Durable does not mean instant. There is a brief window where a confirmed-to-the-app message is still in the OS buffer, not yet on disk. If you need a hard guarantee that the broker has the message, use **publisher confirms** - the broker sends back a confirmation once it has taken responsibility for the message. Without confirms, a publish that "succeeds" only means it left your process.

## Redelivery and the poison message problem

Phase 2's safety net - unacked messages get redelivered - has a sharp edge. Suppose a message is malformed and your consumer throws every single time it tries. The flow becomes a loop:

```text
deliver → consumer throws → no ack → broker redelivers → consumer throws → ...forever
```

*What just happened:* a **poison message** jams the queue. Because the consumer never acks, the broker keeps handing the same message back. One bad message can stall an entire worker pool, burning CPU on a job that will never succeed. Redelivery alone is not enough; you need somewhere for hopeless messages to go.

## Reject vs nack: telling the broker "not this one"

A consumer is not limited to ack. It can also refuse a message:

- **`basic.nack`** (or `basic.reject`) with `requeue=true` - "I could not handle this, put it back." The message returns to the queue for another attempt.
- **`basic.nack`** with `requeue=false` - "I could not handle this, and do not give it back." The message is removed from the queue - and *this* is the hook that sends it to a dead-letter queue.

```text
on_message(msg):
    try:
        do_the_work(msg)
        basic.ack(msg)
    except PermanentError:
        basic.nack(msg, requeue=false)   # give up → dead-letter it
    except TransientError:
        basic.nack(msg, requeue=true)    # retry later
```

*What just happened:* you decide per-error whether a failure is worth retrying. A network blip is transient - requeue it. A message your code can never parse is permanent - `requeue=false` so it leaves the main queue instead of poisoning it. The difference between these two branches is the difference between self-healing and an infinite loop.

## Dead-letter queues: the hospital for failed messages

A **dead-letter exchange** (DLX) is a normal exchange that a queue forwards messages to when they are dead-lettered. A message gets dead-lettered when it is nacked with `requeue=false`, when it expires (TTL), or when the queue overflows its length limit.

You configure it as an argument on the *source* queue:

```text
queue.declare  name="task_q"
  arguments:
    x-dead-letter-exchange: "dlx"
    x-dead-letter-routing-key: "task.failed"

# then a normal queue catches the dead letters:
queue.declare  name="task_dead_q"
queue.bind     queue="task_dead_q" exchange="dlx" routing_key="task.failed"
```

*What just happened:* failed messages from `task_q` are not lost and do not loop - they are routed through the DLX into `task_dead_q`, where you can inspect them, alert on them, or replay them after a fix. The dead-letter queue is your evidence locker: it tells you *what* failed and lets you decide what to do, instead of silently dropping or silently retrying forever.

```text
task_q ──(nack requeue=false)──▶ DLX ──▶ task_dead_q ──▶ you, reading the failures
```

*What just happened:* the poison loop from earlier is broken. Bad messages exit the hot path after one decisive failure and land somewhere safe and visible. A common pattern adds a retry queue with a TTL in between, so a message bounces back for a few delayed attempts before finally giving up to the dead-letter queue.

## RabbitMQ vs Kafka: pick the right shape

People reach for Kafka and RabbitMQ for overlapping reasons, but they are built on opposite ideas, and knowing which you have saves you from forcing the wrong model.

| | RabbitMQ | Kafka |
| --- | --- | --- |
| Core model | Smart broker, dumb consumer | Dumb broker, smart consumer |
| What it is | A router that pushes messages to queues | A durable, ordered log readers pull from |
| After delivery | Message is acked and **deleted** | Message **stays** in the log (retention window) |
| Re-read old messages | No - once consumed and acked, it is gone | Yes - rewind your offset and replay |
| Routing | Rich: exchanges, topics, bindings | Minimal: topics and partitions |
| Best fit | Task queues, RPC, complex routing, per-message workflows | High-volume event streams, replay, many independent readers |

*What just happened:* the deciding question is *do consumers need to re-read history?* RabbitMQ treats a message as work to be done once and removed - perfect for "resize this image," "send this email," "charge this card." Kafka treats messages as a permanent log you can replay - perfect for "every page view, forever, read by analytics and billing and ML independently." If you want rich routing and fire-and-forget jobs, RabbitMQ. If you want a replayable stream consumed at each reader's own pace, Kafka. Using one where the other belongs is the most expensive RabbitMQ mistake there is.

## In the wild

A production-grade RabbitMQ queue almost never stands alone. It comes with: durable + persistent messages, manual acks, a dead-letter exchange catching permanent failures, and usually a delayed retry queue in front of the DLX. Set those up once as your default template and most reliability questions answer themselves. When you find yourself wanting infinite retention and replay, that is the signal you have outgrown the queue model - not a reason to fight RabbitMQ into being a log.

```quiz
[
  {
    "q": "To survive a broker restart, what does a message need beyond a durable queue?",
    "choices": ["Nothing; a durable queue is enough", "To be marked persistent (delivery_mode=2)", "A fanout exchange", "auto_ack enabled"],
    "answer": 1,
    "explain": "A durable queue with non-persistent messages comes back empty. The message itself must be persisted to disk (delivery_mode=2)."
  },
  {
    "q": "What sends a message to a dead-letter exchange?",
    "choices": ["Acking it successfully", "Nacking it with requeue=false (or TTL expiry, or queue overflow)", "Publishing to a topic exchange", "Setting prefetch_count=1"],
    "answer": 1,
    "explain": "Dead-lettering happens on nack/reject with requeue=false, on message TTL expiry, or on queue length overflow - routing the message out of the hot path."
  },
  {
    "q": "The biggest difference between RabbitMQ and Kafka is:",
    "choices": ["RabbitMQ is faster in all cases", "Kafka cannot do routing at all", "RabbitMQ deletes messages after they are acked; Kafka keeps them in a replayable log", "Kafka has no consumers"],
    "answer": 2,
    "explain": "RabbitMQ is a broker that routes and deletes; Kafka is a durable log you can rewind and replay. That shapes which workloads fit each."
  }
]
```
