# Event-Driven Architecture

> Systems that talk by emitting events instead of calling each other directly: queues, pub/sub, and choreography versus orchestration.


---

# Event-Driven Architecture

Right now, somewhere in your system, service A picks up the phone and calls service B directly: "Hey, a user signed up - go send the welcome email." It works. Then you add a service that gives out signup bonuses, and another that warms a cache, and another that pings analytics. Suddenly the signup code knows about five other services, breaks when any of them is down, and slows to the speed of its slowest dependency. There's a calmer way to build this, and it starts with a single shift: instead of *calling* the others, A announces "a user signed up" and walks away. Whoever cares, reacts.

This guide gives you the mental model - producers, a broker, consumers - and then makes you clear-eyed about the bill that comes with it: eventual consistency, harder debugging, and messages that sometimes arrive twice. You'll leave knowing not only how event-driven systems work, but when the trade is worth it and when it isn't.

## How to read this
- **Want the core idea fast?** Read [Phase 1](01-the-mental-model.md). It's the whole "announce, don't call" shift in one sitting.
- **Want it to actually stick?** Read in order. Phase 1 builds the model, Phase 2 shows how events really flow (queues, pub/sub, choreography vs orchestration), and Phase 3 is the hard-won reality: where this bites and how to survive it.

## The phases
1. **[Announce, Don't Call](01-the-mental-model.md)** - the core shift: a producer emits an event, a broker holds it, consumers react. Why this decouples your system, and what an "event" actually is.
2. **[How Events Really Flow](02-how-events-flow.md)** - queues vs pub/sub, choreography vs orchestration, and the delivery guarantees you actually get. The day-to-day mechanics.
3. **[The Bill Comes Due](03-the-bill-comes-due.md)** - eventual consistency, debugging across hops, ordering, and duplicate delivery. Why consumers must be idempotent, and when *not* to go event-driven at all.


---

# Announce, Don't Call

Picture the code that runs when someone signs up. In a lot of systems, it looks like a to-do list the signup handler is personally responsible for completing:

```text
on signup(user):
    db.save(user)
    email_service.send_welcome(user)        # call out
    billing_service.start_trial(user)        # call out
    analytics_service.track("signup", user)  # call out
    crm_service.create_contact(user)         # call out
    return "ok"
```

*What just happened:* the signup handler made four outbound calls. It now *knows about* four other services, *waits for* all four, and *fails* if any one of them is down. Add a fifth thing that should happen on signup, and you edit this file again. The signup code has quietly become the manager of the whole company.

## The shift: stop calling, start announcing

Event-driven architecture replaces that to-do list with a single announcement:

```text
on signup(user):
    db.save(user)
    broker.emit("user.signed_up", {id: user.id, email: user.email})
    return "ok"
```

*What just happened:* the handler did its own job (save the user), shouted one fact into the room - "a user signed up, here are the details" - and returned. It does not know who's listening. It does not wait. Sending the welcome email, starting the trial, tracking analytics - those are now somebody else's problem, and the signup code has no idea how many somebodies there are.

That's the entire idea. An **event** is a statement of fact about something that already happened, in the past tense: `user.signed_up`, `order.placed`, `payment.failed`. Not a command ("send the email"), not a question ("are you there?") - a fact. The producer announces it and moves on. Interested parties react on their own time.

> The past tense matters more than it looks. `send_welcome_email` is an instruction aimed at one specific service - it couples the sender to the receiver. `user.signed_up` is true; it doesn't care who acts on it. That difference is what buys you the decoupling.

## The three parts: producer, broker, consumer

Every event-driven system, no matter how fancy the tooling, is these three roles:

```text
  PRODUCER              BROKER                  CONSUMERS
  (emits)        (holds & routes the event)     (react)

  signup  ──►  ┌──────────────────────┐  ──►  email service
  handler      │   "user.signed_up"   │  ──►  billing service
               │   (a queue or log)   │  ──►  analytics service
               └──────────────────────┘  ──►  crm service
```

- **Producer** - the thing that emits the event. It knows nothing about consumers. Its only job is to publish the fact.
- **Broker** - the piece in the middle that receives events and holds them until consumers are ready. This is the part that's new. It's a queue, or a log, or a pub/sub topic - software like RabbitMQ, Kafka, SQS, NATS. We'll get into the flavors in Phase 2.
- **Consumer** - a service that subscribes to events it cares about and does work when one arrives. Consumers are independent: each one fails, retries, and scales on its own.

The broker is the hero of the story. Without it, "emit an event" would still mean A has to know where B is. *With* it, A talks only to the broker, and the broker is the only thing that knows the routing. That one indirection is where all the benefits come from.

## Why anyone bothers: three real wins

**1. Decoupling.** The producer and consumers don't know each other exist. You can add a sixth consumer - say, a service that sends a Slack alert to the sales team - without touching the signup code at all. You deploy the new consumer, it subscribes to `user.signed_up`, done. The blast radius of "add a feature" shrinks dramatically.

**2. Buffering.** Because the broker *holds* events, a slow or briefly-down consumer doesn't break the producer. If the email service is overwhelmed at peak, signups keep flowing - the welcome emails queue up and get sent as the email service catches up. The broker absorbs the spike. (This is the same backpressure idea covered in [Webhooks & Message Queues](/guides/webhooks-and-message-queues).)

**3. Replay.** Some brokers (the log-shaped ones) keep events around after they're consumed. That means a brand-new consumer can start at the beginning and process the entire history - rebuild a search index, populate a new analytics warehouse, recover from a bug - by re-reading events that were emitted months ago. The event stream becomes a kind of audit trail you can rewind.

## A quick gut-check on the difference

Direct calls are **synchronous and pointed**: A calls B, A waits, A is coupled to B, A is only as available as B. Events are **asynchronous and broadcast**: A emits, A doesn't wait, A is coupled to nothing, A's availability doesn't depend on anyone downstream.

That sounds strictly better, which is exactly the trap. You're trading a problem you can see (tight coupling, cascading slowness) for problems you *can't* see yet (the work happens "eventually," not now, and debugging means following a fact across services). Phase 3 is where we pay that bill in full. First, Phase 2 shows how events actually move.

**For builders:** before reaching for a broker, look at your existing call graph and ask which calls are *fire-and-forget* - the caller doesn't need the result to respond. Welcome emails, analytics, cache warming: classic fire-and-forget. Those are the calls that want to be events. A call whose result you need *in the same response* (charge this card, then tell the user it worked) usually wants to stay a direct call. Sorting your calls into those two buckets is most of the design work.

```quiz
[
  {
    "q": "In event-driven architecture, what is an 'event'?",
    "choices": ["A command telling a specific service what to do", "A statement of fact about something that already happened", "A question asking whether a service is available", "A scheduled task that runs on a timer"],
    "answer": 1,
    "explain": "An event is a past-tense fact (user.signed_up), not a command or a query. The producer announces what happened and doesn't dictate who acts on it."
  },
  {
    "q": "What role does the broker play that makes the whole pattern work?",
    "choices": ["It runs the business logic for every consumer", "It calls each consumer directly on the producer's behalf", "It receives events and holds them until consumers are ready, hiding the routing from the producer", "It guarantees every event is processed exactly once"],
    "answer": 2,
    "explain": "The broker is the indirection in the middle. The producer talks only to it, so producers and consumers never need to know about each other."
  },
  {
    "q": "Which of these is the BEST candidate to convert from a direct call into an event?",
    "choices": ["Charging a credit card before showing the user a success page", "Sending a welcome email after signup", "Reading a user's profile to render the current page", "Validating a password during login"],
    "answer": 1,
    "explain": "The welcome email is fire-and-forget: the caller doesn't need its result to respond. Calls whose results you need in the same response should usually stay direct."
  }
]
```

Watch it animated: [event-driven architecture](/explainers/EventDriven.dc.html)


---

# How Events Really Flow

You've got the model: producer, broker, consumer. Now the practical questions show up. When the broker holds an event, does *one* consumer get it or *everyone*? When several services have to cooperate on one workflow, who's in charge? And when the network hiccups - because it will - does the event get delivered once, or never, or twice? These three questions decide how your system actually behaves, so let's take them one at a time.

## Queue vs pub/sub: one taker, or all takers?

There are two delivery shapes, and the difference is who consumes a given event.

A **queue** is a line of work where each message is handled by exactly *one* consumer. You run several copies of the same worker for throughput, and the broker hands each message to whichever worker is free. This is for **work distribution** - "this job needs doing once, by somebody."

```text
QUEUE  (competing consumers - each message goes to ONE)

  producer ─► [ msg msg msg msg ] ─┬─► worker A   (takes msg 1, 3)
                                    └─► worker B   (takes msg 2, 4)
```

*What just happened:* four messages, two identical workers. The broker split the line between them - no message was handled twice. Add a third worker and the line drains faster. This is how you scale a queue: more workers, same queue.

**Pub/sub** (publish/subscribe) is broadcast. The producer publishes to a *topic*, and *every* subscriber gets its own copy of every event. This is for **fan-out** - "this fact happened, and several different services each need to know."

```text
PUB/SUB  (fan-out - each subscriber gets its OWN copy)

  producer ─► topic "user.signed_up" ─┬─► email service     (gets a copy)
                                       ├─► billing service   (gets a copy)
                                       └─► analytics service (gets a copy)
```

*What just happened:* one event, three independent copies. The email, billing, and analytics services each received `user.signed_up` and reacted in parallel. None of them competed for it; none of them blocked the others.

In real systems you combine these: pub/sub fans an event out to several services, and *within* each service a queue distributes the work across that service's worker pool. The two shapes aren't rivals - they stack.

## Choreography vs orchestration: who runs the workflow?

Single events are easy. The hard part is a *multi-step* workflow - say, placing an order: reserve inventory, charge payment, schedule shipping, send a confirmation. There are two ways to coordinate the dance.

**Choreography** - no conductor. Each service reacts to events and emits its own, like dancers who each know their cue. There's no central brain; the workflow *emerges* from the chain of reactions.

```text
CHOREOGRAPHY  (each service reacts and emits the next event)

  order.placed ─► [inventory] ─► inventory.reserved
                                  └─► [payment] ─► payment.charged
                                                   └─► [shipping] ─► order.shipped
```

*What just happened:* nobody is in charge. Inventory heard `order.placed`, did its bit, and emitted `inventory.reserved`. Payment heard *that* and continued the chain. The flow is the sum of independent reactions. This is maximally decoupled - but to understand the whole workflow, you have to trace events across every service, because no single place describes it.

**Orchestration** - a conductor. One component (an *orchestrator* or *saga manager*) owns the workflow and tells each service what to do next, usually still over events or commands.

```text
ORCHESTRATION  (one orchestrator drives every step)

         ┌──────────── ORCHESTRATOR ────────────┐
         │ 1.reserve   2.charge   3.ship   4.notify │
         └──┬──────────┬─────────┬────────┬───────┘
            ▼          ▼         ▼        ▼
       inventory   payment   shipping   email
```

*What just happened:* the orchestrator drove the order through four steps and waited for each to report back. The whole workflow lives in *one* readable place, and if step 2 fails the orchestrator can run compensating steps (un-reserve the inventory). The cost: the orchestrator is now a central thing that knows about every service - you've traded some decoupling for clarity and control.

The plain rule of thumb: **choreography** for simple, short chains where loose coupling matters most; **orchestration** once the workflow has branches, rollbacks, or more than a few steps and you need to *see* it in one place. Many teams start choreographed, watch it turn into spaghetti nobody can follow, and add an orchestrator for the gnarly workflows.

## Delivery guarantees: the fine print that will bite you

When you emit an event, how many times does each consumer receive it? There are three theoretical answers, and only one of them is what you almost always get in practice.

| Guarantee | What it means | Reality |
|---|---|---|
| **At-most-once** | Delivered 0 or 1 times - may be lost | Fast, but you can silently drop events. Rarely acceptable. |
| **At-least-once** | Delivered 1 or more times - never lost, may repeat | **The common default.** Safe against loss, but you *will* see duplicates. |
| **Exactly-once** | Delivered precisely 1 time | The dream. Genuinely hard end-to-end; often "exactly-once *processing*" faked atop at-least-once. |

Most brokers you'll actually use default to **at-least-once**. Here's *why* duplicates are nearly unavoidable: after a consumer processes an event, it sends an *acknowledgment* ("ack") back to the broker so the broker can stop tracking it. If the consumer crashes - or the ack gets lost on the network - *after* doing the work but *before* the ack lands, the broker never hears confirmation. So it does the safe thing: it redelivers. The work happens twice.

```text
1. broker delivers  "payment.charged"  ──►  consumer
2. consumer charges the card  ✅
3. consumer crashes before sending ack  💥
4. broker waited, heard nothing, REDELIVERS  ──►  consumer
5. consumer charges the card AGAIN  💸💸   ← the duplicate bites
```

*What just happened:* the broker did exactly what it promised - never lose a message - and the cost of that promise was a double charge. The broker can't tell "the consumer died" apart from "the consumer is just slow," so when in doubt it redelivers. This is not a bug; it's the guarantee working as designed.

Which is why the single most important habit in event-driven systems is making consumers **idempotent** - safe to run the same event twice with no extra effect. That's the heart of Phase 3, and the reason it deserves its own phase.

**For builders:** when you pick or configure a broker, find the delivery guarantee in its docs *before* you write a consumer, not after a double-charge in production. Assume at-least-once unless the docs prove otherwise, and design every consumer as if duplicates are guaranteed - because they effectively are. For the practical retry/backpressure side of running a queue, see [Webhooks & Message Queues](/guides/webhooks-and-message-queues).

```quiz
[
  {
    "q": "What's the core difference between a queue and pub/sub?",
    "choices": ["Queues are faster than pub/sub", "In a queue each message goes to exactly one consumer; in pub/sub every subscriber gets its own copy", "Pub/sub can't lose messages but queues can", "Queues are for events and pub/sub is for commands"],
    "answer": 1,
    "explain": "A queue distributes work (one taker per message) for throughput; pub/sub fans out (every subscriber gets a copy) so multiple services each learn about the same event."
  },
  {
    "q": "When does orchestration tend to beat choreography?",
    "choices": ["When you want maximum decoupling and the simplest possible services", "When the workflow has branches, rollbacks, or many steps and you need to see it in one place", "Whenever there is more than one consumer", "When the broker only supports at-most-once delivery"],
    "answer": 1,
    "explain": "Orchestration puts the whole workflow in one readable place and can run compensating steps on failure - worth the central coupling once the flow gets complex."
  },
  {
    "q": "Why do at-least-once brokers deliver duplicates?",
    "choices": ["Because the network always doubles every packet", "Because consumers ask for events twice on purpose", "Because if the ack is lost or the consumer crashes after doing the work, the broker can't tell and redelivers to avoid losing the event", "Because duplicates make processing faster"],
    "answer": 2,
    "explain": "The broker can't distinguish a dead consumer from a slow one. When it doesn't get an ack, it redelivers rather than risk losing the event - so the work can happen twice."
  }
]
```


---

# The Bill Comes Due

Every benefit in Phase 1 had a hidden cost, and now we pay it - straight, with no spin. Event-driven systems are genuinely powerful, and they are also genuinely harder to reason about than a stack of direct calls. The teams that succeed with them aren't the ones who avoided these problems; they're the ones who saw them coming. So here are the four that catch everyone, what each one feels like in production, and how to live with it.

## 1. Eventual consistency: "done" doesn't mean done yet

When the signup handler emits `user.signed_up` and returns "ok," the welcome email has *not* been sent. The trial has *not* started. Those happen moments later, when the consumers get around to it. The system is **eventually consistent**: it will be correct soon, but for a window after the producer returns, different parts disagree about the state of the world.

```text
t=0   signup handler returns "ok"  ──► user sees "Welcome!"
t=0   billing consumer hasn't run yet
t=1s  user clicks "My Plan"        ──► "No active trial"  😟
t=2s  billing consumer processes the event ──► trial now active
```

*What just happened:* for about two seconds the user existed but had no trial, and if they were quick they saw a wrong answer. Nothing is broken - the work just hadn't propagated yet. With direct calls this gap doesn't exist, because everything finishes before the response. With events, the gap is the *price of decoupling*, and you have to design the UI and the data flows to tolerate it (show "setting up your account…", read from the source of truth for critical checks, don't promise what hasn't happened).

> This is the trade in one sentence: synchronous calls give you *consistency now* at the cost of coupling; events give you *decoupling* at the cost of consistency *later*. Neither is free. Choose based on whether your use case can tolerate the gap.

## 2. Debugging gets harder: the request you can't follow

With direct calls, a failed request has a stack trace - one thread, one path, top to bottom. With events, a single user action fans out into a constellation of independent reactions across services, at different times, with no shared call stack. "Why didn't this user get their welcome email?" stops being a stack trace and becomes detective work across five log files.

The survival tools here are not optional in a serious system:
- **Correlation IDs** - stamp every event with an ID that follows the whole workflow, so you can grep one ID across every service and reconstruct the chain.
- **Distributed tracing** - tools that stitch those hops into a single timeline you can actually look at.
- **A dead-letter queue (DLQ)** - a side queue where events that keep failing get parked instead of looping forever, so you can inspect the poison message instead of losing it.

Without these, an event-driven system is a black box that *mostly* works and is *miserable* to debug the day it doesn't.

## 3. Ordering: events don't always arrive in the order they happened

You might assume `cart.item_added` always arrives before `cart.checked_out`. Across a distributed broker with multiple partitions and parallel consumers, that assumption can break - events can arrive out of order, or be processed in parallel out of order.

Most brokers only guarantee ordering within a **partition** or a single queue, and only if you route related events to the same partition (usually by a key like `cart_id`). So the fix is twofold: route events that must stay ordered to the same partition using a stable key, and where you can, **design events so order doesn't matter** - carry enough state in the event that a consumer can act correctly regardless of what it has or hasn't seen yet.

## 4. Duplicates, and the cure: idempotency

We saw in Phase 2 *why* at-least-once delivery hands you duplicates. Here's how you survive them. An operation is **idempotent** when doing it twice has the same effect as doing it once. You don't try to *prevent* duplicate delivery (you can't, reliably); you make duplicate *processing* harmless.

The workhorse pattern: give every event a unique ID, and have each consumer record which IDs it has already handled. Before doing the work, check.

```sql
-- consumer receives an event with id 'evt_8a3f' and a payment to apply
INSERT INTO processed_events (event_id) VALUES ('evt_8a3f')
ON CONFLICT (event_id) DO NOTHING;
-- only do the real work if THIS consumer hasn't seen evt_8a3f before
```

*What just happened:* the first time `evt_8a3f` arrives, the INSERT succeeds and the consumer proceeds to charge the card. If the broker redelivers `evt_8a3f` after a lost ack, the INSERT hits the conflict, changes nothing, and the consumer skips the charge. The double-charge from Phase 2 is now impossible - the duplicate still *arrives*, but processing it is a no-op. (Doing the dedupe insert and the real work in one transaction is what makes this airtight.)

This is why "make your consumers idempotent" is the most-repeated advice in this whole field. At-least-once is the delivery reality; idempotency is how you make it safe.

## So when should you NOT do this?

The most senior move in this guide is knowing when to skip it. Event-driven architecture is the wrong default for a small or new system: if your whole app is one service and a database, a broker buys you eventual-consistency bugs, distributed-debugging pain, and new infrastructure to operate - in exchange for decoupling between parts that aren't even separate yet. Reach for events when you have **real, separate services** that need to react to each other, when **fire-and-forget** work is clogging your request path, when you need **buffering** against spikes, or when **multiple independent consumers** genuinely need the same facts. Until then, a direct call is simpler, easier to debug, and consistent *now*. (Whether you even have separate services is exactly what [Monolith vs Microservices](/guides/monolith-vs-microservices) helps you decide *first*.)

**For builders:** if you take one habit from this guide, take this: **assume every event will be delivered more than once, and make every consumer safe under that assumption.** Add the event-ID dedupe table on day one, not after the first double-charge. It's a few lines, and it's the difference between an event-driven system you trust and one that quietly corrupts data every time the network sneezes.

```quiz
[
  {
    "q": "What does 'eventual consistency' mean for a user right after a producer emits an event and returns?",
    "choices": ["All downstream work is guaranteed finished before the response", "There's a window where different parts of the system disagree until the event is processed", "The event will never be processed", "The producer blocks until every consumer is done"],
    "answer": 1,
    "explain": "The producer returns immediately, but consumers process the event later. For a short window the state hasn't propagated, so different parts can disagree - the price of decoupling."
  },
  {
    "q": "Why must consumers be idempotent in an at-least-once system?",
    "choices": ["To make events arrive faster", "Because the broker guarantees exactly-once delivery", "Because duplicates are effectively unavoidable, so processing one twice must be harmless", "To enforce strict ordering of events"],
    "answer": 2,
    "explain": "At-least-once delivery means duplicates will happen (lost acks, crashes, redelivery). You can't prevent duplicate delivery reliably, so you make duplicate processing a no-op."
  },
  {
    "q": "Which situation is the BEST reason NOT to adopt event-driven architecture yet?",
    "choices": ["You have several independent services that need to react to the same facts", "Fire-and-forget work is clogging your request path", "Your whole app is one service and a database with no separate parts to decouple", "You need to buffer against traffic spikes"],
    "answer": 2,
    "explain": "A single service and database gains little from a broker but pays full price in eventual-consistency bugs and distributed-debugging pain. Events shine when you have real, separate services."
  }
]
```
