# Celery From Zero

> Learn the distributed task queue that runs background work for Python web apps: the broker/worker/result model, defining and calling tasks, results and state, retries and error handling, scheduled tasks with Celery Beat, and production scaling, monitoring, and the pitfalls. Move slow work off the request and run it reliably.


---

# Celery From Zero

Some work is too slow to do during a web request: sending an email, generating a report, processing an
upload, calling a slow third-party API. Make the user wait for it and your app feels broken. Celery is the
answer the Python world reaches for - a **distributed task queue** that lets your app hand a job off to
run *later*, on a separate pool of worker processes, while the request returns instantly. FastAPI,
Django, and Flask apps all lean on it for exactly this.

The mental model that makes Celery click is a four-part hand-off: your app puts a **task** message on a
**broker** (a queue, usually Redis or RabbitMQ); one of several **workers** picks it up and runs it; and
an optional **result backend** stores what it returned. Once you see those four pieces - and that they're
separate processes talking through a queue - Celery stops being intimidating configuration and becomes a
shape you can reason about (and debug).

> 📝 This assumes **Python** ([Python From Zero](/guides/python-from-zero)) and pairs naturally with the
> queue concepts in [Webhooks & Message Queues](/guides/webhooks-and-message-queues). It's the background
> worker that the framework guides ([FastAPI](/guides/fastapi-from-zero), [Django](/guides/django-from-zero),
> [Flask](/guides/flask-from-zero)) all hand heavy jobs to. Celery needs a broker + worker processes, so
> examples are shown with the commands to run them yourself.

## How to read this

Read in order - it builds from a single background task up to scheduled jobs and a monitored production
setup, using a running example of a web app offloading email and report work. Phases carry difficulty
badges.

## The phases

**Part 1 - The core model (🟢 Basic → 🟡)**
1. **[What Celery Is & Why](01-what-celery-is.md)** 🟢 - the problem, and the broker/worker/result mental model.
2. **[The Broker & Worker](02-the-broker-and-worker.md)** 🟢 - the queue (Redis/RabbitMQ) and the worker processes that drain it.
3. **[Defining & Calling Tasks](03-defining-and-calling-tasks.md)** 🟡 - `@task`, `.delay()`/`.apply_async()`, and how a job gets queued and run.

**Part 2 - Doing it right (🟡 → 🔴)**
4. **[Results & State](04-results-and-state.md)** 🟡 - the result backend, `AsyncResult`, and tracking task status.
5. **[Retries & Error Handling](05-retries-and-error-handling.md)** 🔴 - retrying failures, idempotency, and not losing work.
6. **[Scheduled Tasks with Celery Beat](06-scheduled-tasks-celery-beat.md)** 🟡 - periodic, cron-like jobs.

**Part 3 - Production (🔴 → 🟢)**
7. **[Production: Scaling, Monitoring & Pitfalls](07-production-scaling-monitoring.md)** 🔴 - concurrency, Flower, and the mistakes that bite everyone.
8. **[Where to Go Next](08-where-to-go-next.md)** 🟢 - Celery vs the alternatives, framework integration, and what to build.

> The whole thing is a hand-off: app → broker → worker → (result). Hold those four pieces and Celery's
> configuration and quirks all fall into place.


---

# What Celery Is & Why

Picture a user clicking "Sign up." Your code creates an account, then tries to send a welcome email
through a third-party mail service that takes four or five seconds to respond. The user sits there
watching a frozen page, waiting on an email *they aren't even reading yet*. Worse: the web worker
handling their request is stuck too, unable to serve anyone else until that mail call returns.

The fix is one idea: **do the slow part later, in the background, and answer the user right now.**
Celery is the tool most Python apps reach for to make that happen. By the end of this phase you'll have
the mental model - the four moving pieces and why they're separate - so the hands-on setup in later
phases makes sense instead of feeling like magic incantations.

## The problem - some work is too slow for a request

📝 **The core tension.** A web request should be *fast* - measured in milliseconds. But plenty of
real work isn't: sending an email, generating a PDF report, processing an uploaded video, calling a
sluggish third-party API. Do that work *inline* (right there in the request handler) and two bad
things happen at once: the user waits for the whole thing to finish, and your web worker is tied up
the entire time, unable to handle other requests.

Here's the painful version - work done inline:

```python
# views.py - the SLOW way
def signup(request):
    user = create_user(request.data)
    send_welcome_email(user)      # blocks ~5s waiting on the mail service
    return Response({"status": "ok"})   # user has been staring at a spinner the whole time
```

*What just happened:* `send_welcome_email(user)` runs *before* the `return`. Python won't move past
that line until the mail service answers, so the response is held hostage for five seconds. Multiply
that by a busy signup hour and your web workers spend most of their time waiting.

Now the fix - hand the slow work off and return immediately:

```python
# views.py - the BACKGROUND way
def signup(request):
    user = create_user(request.data)
    send_welcome_email.delay(user.id)   # enqueue it; returns instantly
    return Response({"status": "ok"})    # user gets a snappy response right away
```

*What just happened:* `.delay(...)` doesn't *run* the email code - it drops a "please send this email"
message onto a queue and returns in microseconds. The response goes back to the user immediately. The
actual sending happens moments later, in a different process, while your web worker has already moved on
to the next request. (Don't worry about `.delay` or why we pass `user.id` instead of `user` yet - those
are later phases. For now, notice the *shape*: enqueue, return, run-later.)

💡 **The reframe.** The slow work didn't get faster. You stopped *waiting* for it. That single shift - 
from "do it now while the user waits" to "schedule it and answer now" - is the entire point of a task
queue.

## What Celery is

📝 **Celery** - a distributed task queue for Python. You define ordinary Python functions as **tasks**,
your application **enqueues** them when it wants them run, and separate **worker** processes pick them up
and execute them asynchronously. You write the function once; Celery handles getting it onto a queue,
delivering it to a worker, running it, and (optionally) reporting back.

⚠️ **"Asynchronous" here is not asyncio.** This trips up nearly everyone coming from `async`/`await`. In
Celery, "asynchronous" means the work runs in a *different process* - often on a *different machine* - 
not on the same event loop a moment later. It's about handing work off to an entirely separate program
that runs it independently. If you've read [Python from Zero](/guides/python-from-zero), this is a
different axis from generators or coroutines: those keep one process busy efficiently, Celery spreads
work across many processes.

## The four pieces - the mental model

This is the model to burn into memory. Everything else in Celery is detail hanging off these four parts:

📝 **(1) The client (your app)** enqueues a task - "run `send_welcome_email` with this argument."
📝 **(2) The broker** (usually Redis or RabbitMQ) is the middleman that *holds the queue* of task
messages until a worker is ready for them.
📝 **(3) The workers** are separate processes that pull tasks off the broker and actually execute the
Python code.
📝 **(4) The result backend** (optional) stores each task's return value and status, so the client can
later ask "is it done? what did it return?"

```mermaid
flowchart LR
  A[Your app<br/>client] -->|enqueue task| B[(Broker<br/>queue)]
  B -->|deliver task| C[Worker<br/>process]
  C -.->|store result| D[(Result<br/>backend)]
```

That broker in the middle is exactly the **message queue** idea from
[Webhooks & Message Queues](/guides/webhooks-and-message-queues): a producer drops a message, the queue
holds it, a consumer picks it up when ready. Celery is a polished, Python-native layer built on top of
that pattern - it gives the message a name (`send_welcome_email`), serializes the arguments, and runs the
matching function on the other side. If the queue concept feels shaky, that guide is worth a detour
first; everything here sits on top of it.

The result backend is the one piece you can often skip. For fire-and-forget jobs like sending an email,
you don't care about a return value - you enqueue it and forget it. You only need one when your app
wants to *check on* a task later (e.g. "is the report ready to download yet?").

## It's separate processes - that's the whole trick

💡 **The realization that makes Celery click:** your web app and your workers are *different programs*.
They don't share memory or call each other's functions. They communicate *only* through the broker, by
passing messages. Your web app says "here's a task" into the broker; a worker, possibly on another
server entirely, picks it up later.

That separation is exactly why Celery is powerful:

- **It scales by adding workers.** Drowning in `generate_report` jobs? Start three more worker processes
  (or spin up another machine running workers). They all pull from the same broker. The web app doesn't
  change at all.
- **It survives restarts.** Deploy new web code, and the tasks already sitting in the broker are still
  there - a worker grabs them whenever it's ready. The queue is a buffer between the two halves.

But that same separation is *why the next two phases exist*. Because the worker is a different process
that didn't run your web code, it can't see your in-memory objects - so arguments have to be
**serialized** (turned into data) to travel through the broker. That's why we passed `user.id` (a plain
integer) earlier and not the `user` object itself. Configuration, serialization, and "how does the
worker even find my tasks" all flow from this one fact: *they're separate processes talking through a
broker.*

## Where it fits - and the alternatives

In a real stack, your web framework does the fast part of a request and hands the slow part to Celery:

- A [FastAPI](/guides/fastapi-from-zero), Django, or Flask view receives the request, does the quick
  database work, enqueues the heavy job (`generate_report.delay(...)`), and returns. Celery workers,
  running as a totally separate deployment, chew through the queue in the background.

💡 **What about built-in background tasks?** FastAPI has `BackgroundTasks`, Flask has extensions, and all
of them let you run a little work *after* sending the response. For genuinely light, best-effort,
fire-and-forget things ("log this, nobody will miss it if it's lost"), that's fine and far simpler - no
broker to run. But the moment you need work that is **reliable** (survives a crash and gets retried),
**scalable** (spread across many workers), or **scheduled** (run nightly, run in an hour), those
in-process helpers run out of road. That's the territory of a real task queue like Celery.

Next phase, we stop talking in diagrams and actually stand up the two pieces that make all this real:
the **broker** and a **worker**.

## Recap

1. Some work (emails, reports, uploads, slow API calls) is too slow to run inside a web request. Doing
   it inline makes the user wait *and* ties up your web worker.
2. The fix is to do that work in the **background** and return to the user immediately - "schedule it
   and answer now" instead of "do it now while they wait."
3. 📝 **Celery** is a distributed task queue for Python: define tasks, enqueue them from your app, and
   separate **worker** processes run them. "Asynchronous" here means *different processes*, not asyncio.
4. The mental model is **four pieces**: client (enqueues) → **broker** (holds the queue) → **worker**
   (runs it) → optional **result backend** (stores return value/status).
5. 💡 Your app and the workers are **separate processes** talking only through the broker. That's why
   Celery scales (add workers) and why serialization and config matter (next phases).
6. Web frameworks hand heavy jobs to Celery. Built-in background tasks suit light fire-and-forget work,
   but reliable, scalable, or scheduled background work calls for a real task queue.

## Quick check

Make sure the core model stuck before we start wiring things up:

```quiz
[
  {
    "q": "Why is sending a welcome email *inline* inside a signup request a problem?",
    "choices": [
      "Email can't be sent from Python at all without a queue",
      "The user waits for the slow email to finish, and the web worker is tied up the whole time",
      "Inline code always crashes the web server",
      "It uses more memory than a background task"
    ],
    "answer": 1,
    "explain": "Running slow work inline holds the response hostage until it finishes, and blocks that web worker from serving anyone else. Moving it to the background lets you answer the user immediately and free the worker."
  },
  {
    "q": "In Celery's mental model, what is the BROKER responsible for?",
    "choices": [
      "Running your task's Python code",
      "Storing each task's return value so the client can check on it",
      "Holding the queue of task messages until a worker is ready to pick them up",
      "Rendering the web response sent back to the user"
    ],
    "answer": 2,
    "explain": "The broker (Redis/RabbitMQ) is the middleman that holds task messages in a queue. Workers run the code, the result backend stores return values, and the broker just buffers and delivers the messages between them."
  },
  {
    "q": "When Celery docs say a task runs 'asynchronously,' what do they mean?",
    "choices": [
      "It runs in a separate worker process, often on a different machine",
      "It runs later on the same asyncio event loop using await",
      "It runs faster than normal code",
      "It runs only when the user refreshes the page"
    ],
    "answer": 0,
    "explain": "In Celery, asynchronous means the work happens in a different process entirely - handed off through the broker - not on the same event loop via async/await. It's about spreading work across processes, not concurrency within one."
  }
]
```


---

# The Broker & Worker

Phase 1 gave you the four-part shape: app → broker → worker → (result). This phase zooms in on the middle two pieces, the ones you actually have to *run*. Here's the mental model to hold onto: **your web app and your workers are different programs that never talk to each other directly.** They only ever talk to a thing in the middle - the broker. Once that clicks, half of Celery's "why isn't this working?" mysteries solve themselves.

## The broker: a mailbox between two processes

📝 **The broker is the message queue that sits between your app and the workers.** When your app wants a job done later, it doesn't call the worker - it drops a small message ("run `send_welcome_email` for user 42") into the broker. The message waits there until a worker is free to pick it up. That's the whole job of the broker: hold pending task messages, hand them out one at a time.

If you've read [Webhooks & Message Queues](/guides/webhooks-and-message-queues), this is exactly the producer/consumer queue from that guide. Your web app is the **producer** dropping messages; your workers are the **consumers** picking them up. Celery is a friendly Python layer on top of that same idea - not magic, a queue with good manners.

Two brokers cover almost everyone:

- **Redis** - an in-memory data store that doubles as a queue. Dead simple to run, one command to start, and the default choice for most people learning Celery or running small-to-medium apps. This guide uses Redis.
- **RabbitMQ** - a dedicated message broker built for exactly this. More moving parts to operate, but stronger delivery guarantees and richer routing (priorities, complex fan-out). Reach for it when you outgrow Redis or need those guarantees.

⚠️ **The broker is a separate service you have to run.** It is not part of your Python code and it doesn't start itself. No broker running means no queue exists, which means your tasks have nowhere to go. With Docker, starting Redis is a one-liner:

```bash
docker run -d -p 6379:6379 redis
```

*What just happened:* we started a Redis container in the background (`-d`) and exposed its default port `6379` to your machine. That URL - `localhost:6379` - is what Celery will connect to in a moment. (No Docker? A native `redis-server` install works the same way.)

## Creating the Celery app

The Celery "app" is a single Python object that holds your configuration - most importantly, *where the broker is*. It's tiny. Put this in a file called `tasks.py`:

```python
from celery import Celery

app = Celery("tasks", broker="redis://localhost:6379/0")
```

*What just happened:* we created one `Celery` instance. The first argument, `"tasks"`, is just a name (conventionally the module it lives in). The important part is `broker="redis://localhost:6379/0"` - the **broker URL**, pointing Celery at the Redis we started (`/0` picks Redis database 0). This `app` object is what you'll attach tasks to in Phase 3.

💡 The broker URL is the single most important line of config you'll write. Get the host, port, or scheme wrong and everything downstream fails quietly - more on that at the end of this phase.

## The worker: the process that does the work

📝 **A worker is a separate process that connects to the broker, pulls task messages off the queue, and runs them.** Your web app does not run tasks - it only enqueues them. Something has to be on the other end of the queue actually executing the code; that something is the worker, and you start it yourself from a terminal:

```bash
celery -A tasks worker --loglevel=info
```

*What just happened:* the `celery` command-line tool started a worker. `-A tasks` tells it which app to use (your `tasks.py` from above, where the `app` object lives); `worker` is the subcommand that says "be a worker"; `--loglevel=info` makes it chatty so you can see what it's doing. This process stays running in the foreground, waiting for jobs.

When it boots, a worker prints a banner that tells you everything about its setup:

```console
 -------------- celery@laptop v5.3.6
--- ***** -----
-- ******* ---- [config]
- *** --- * --- .> app:         tasks:0x7f9a1c
- ** ---------- .> transport:   redis://localhost:6379/0
- ** ---------- .> results:     disabled://
- *** --- * --- .> concurrency: 8 (prefork)
-- ******* ---- 
--- ***** ----- [queues]
 -------------- .> celery           exchange=celery(direct) key=celery

[tasks]
  . tasks.send_welcome_email
  . tasks.generate_report

[2026-06-23 10:14:02,331: INFO/MainProcess] Connected to redis://localhost:6379/0
[2026-06-23 10:14:02,402: INFO/MainProcess] celery@laptop ready.
```

*What just happened:* read this banner top to bottom and it confirms the whole mental model. **transport** is the broker it connected to (your Redis). **results** is disabled - no result backend yet (Phase 4). **concurrency: 8 (prefork)** is how many tasks it can run at once (next section). **[queues]** shows it's listening on the default `celery` queue. **[tasks]** lists every task it knows how to run. The final `ready.` line means it's connected and waiting for work.

💡 Stop and notice: the worker and your web app are **two completely separate processes.** You start each one differently, and the only thing they share is the broker URL. They could run on different machines and it would work identically - slow work runs over *there*, in the worker, while your web request returns instantly over *here*.

## Concurrency: one worker, many tasks at once

📝 **A single worker can run several tasks simultaneously using a pool of sub-workers.** You control how many with `--concurrency`:

```bash
celery -A tasks worker --concurrency=4
```

*What just happened:* we told the worker to run up to 4 tasks at the same time. By default the worker uses the **prefork** pool, forking 4 separate child processes; each one grabs a task and runs it independently. If `--concurrency` is omitted, Celery defaults to one child per CPU core (the `8` you saw in the banner).

📝 Which pool you want depends on what your tasks *do*:

- **prefork** (separate processes, the default) - best for **CPU-bound or blocking work**, like crunching numbers for a `generate_report`. Separate processes sidestep Python's GIL and isolate crashes, so one task blowing up doesn't take its siblings down.
- **gevent / eventlet** (lightweight green threads in one process) - best for **I/O-bound work** that spends most of its time *waiting*: calling an email API for `send_welcome_email`, hitting a database, fetching a URL. You can run hundreds of these cheaply because each one is mostly idle.

A rough rule: if your task burns CPU, use prefork; if it sits around waiting on the network, gevent/eventlet lets one worker handle far more of them. You can tune this later - the default prefork pool is a perfectly good starting point.

## The flow, end to end

💡 Now you can see the full hand-off in motion:

1. Your web app calls a task (Phase 3) → a message lands in the **broker** (Redis).
2. A free child in the **worker** pool pulls that message off the queue.
3. The worker runs your function and (optionally) stores what it returned in a **result backend** (Phase 4).

To make this real you need exactly three things running: a **broker** (Redis), a **worker** (the `celery ... worker` command), and your **app** (which enqueues jobs). With all three up, your web app can finally offload work and return instantly.

⚠️ **The number-one beginner trap: tasks that silently never run.** If you enqueue a job and *nothing happens* - no error, no result, just silence - the cause is almost always the broker. Either no broker is running, or your broker URL is wrong (typo in the host, wrong port, pointing at a Redis that isn't there). Celery happily accepts the task into a queue nobody is draining, and it sits there forever. When a task seems to vanish, **check the broker first**: is Redis up? Does the URL in your `Celery(...)` call match where Redis actually is? Is a worker connected to that same URL?

With the broker and worker understood and running, you're ready to actually write tasks and call them - which is Phase 3.

## Recap

- The **broker** is a message queue (Redis or RabbitMQ) that holds pending task messages between your app and your workers - it's the producer/consumer queue from [Webhooks & Message Queues](/guides/webhooks-and-message-queues).
- The broker is a **separate service you must run**; it's not part of your Python code and won't start itself.
- The **Celery app** is a small Python object (`Celery("tasks", broker=...)`) whose most important config is the **broker URL**.
- A **worker** is a separate process (`celery -A tasks worker`) that connects to the broker, pulls messages, and runs your tasks; its startup banner shows the broker, queues, registered tasks, and concurrency.
- **Concurrency** lets one worker run many tasks at once: `--concurrency=N`, with **prefork** (processes) for CPU/blocking work and **gevent/eventlet** for I/O-bound waiting.
- When a task silently never runs, **check the broker first** - no broker or a wrong broker URL is the classic culprit.

## Quick check

```quiz
[
  {
    "q": "What is the broker's job in Celery?",
    "choices": ["It runs your task functions", "It holds pending task messages between your app and the workers", "It stores the return value of finished tasks"],
    "answer": 1,
    "explain": "The broker is the message queue in the middle: your app drops task messages into it, and workers pull them out. Running the task is the worker's job; storing return values is the result backend's job."
  },
  {
    "q": "Your app enqueues a task but nothing ever happens - no error, no result. What's the most likely cause?",
    "choices": ["A syntax error in your task code", "Too much concurrency", "No broker is running, or the broker URL is wrong"],
    "answer": 2,
    "explain": "Silent no-ops almost always mean the broker. If no broker is up or the URL is wrong, Celery queues the task into a void where no worker can reach it. Check the broker first."
  },
  {
    "q": "Which pool fits a CPU-bound task like generating a heavy report?",
    "choices": ["gevent", "prefork (separate processes, the default)", "eventlet"],
    "answer": 1,
    "explain": "Prefork uses separate processes, which sidestep the GIL and isolate crashes - ideal for CPU-bound or blocking work. gevent/eventlet shine for I/O-bound tasks that mostly wait."
  }
]
```


---

# Defining & Calling Tasks

Here's the mental model to hold onto before we touch any code: **a task is a function you've agreed to run somewhere else.** You define it in your code, you call it from your web process, but the body actually executes inside a worker - a separate program, possibly on a different machine. Everything weird about Celery comes from that one fact: the function and the call live in two different worlds, and a message has to travel between them.

In Phase 2 you stood up a broker and a worker and watched them shake hands. Now we make them earn their keep: defining real work like `send_welcome_email(user_id)` and `generate_report(...)`, then handing it off so your web request can return without waiting.

## Defining a task

📝 You turn an ordinary function into a Celery task by decorating it. That's the whole ceremony.

```python
# tasks.py
from .celery_app import app

@app.task
def send_welcome_email(user_id):
    user = User.objects.get(id=user_id)
    send_email(
        to=user.email,
        subject="Welcome aboard!",
        body=render_welcome(user),
    )
```

*What just happened:* the `@app.task` decorator **registered** this function with your Celery app under a name (by default, the dotted path like `tasks.send_welcome_email`). Registration is the important word - Celery now knows this function exists and can look it up by name later. The body is plain Python, nothing special. The catch is *where* it runs: when a task fires, this code executes **in the worker process**, not your web process. The web side only ever sends a message saying "run the task named `send_welcome_email` with this argument."

💡 If you're in a project with many task modules (or a framework like Django), reach for `@shared_task` instead of `@app.task`. It registers the task without needing to import your concrete `app` object, which keeps your task files from depending on app-creation order. Same idea, fewer import headaches.

```python
from celery import shared_task

@shared_task
def generate_report(account_id, month):
    rows = query_usage(account_id, month)
    pdf = render_pdf(rows)
    store_report(account_id, month, pdf)
```

*What just happened:* `generate_report` is registered the same way, but it didn't have to import `app` - whichever Celery app is active picks it up. For app-style projects this is the friendlier default, and it's why you'll see it everywhere in real codebases.

## Calling it: `.delay()`

Now the part that trips up every newcomer at least once. There are two ways to "call" a task, and they do completely different things.

📝 `send_welcome_email.delay(user_id)` does **not** run the function. It packages up the task name and arguments, drops that message on the broker, and returns **immediately** - handing you back an `AsyncResult` (a receipt you can check later; that's Phase 4). The worker picks the message up and runs the body whenever it gets to it.

⚠️ `send_welcome_email(user_id)` - no `.delay` - runs the function **right now, inline, in your current process.** No broker, no worker, no queue. It's just a normal Python call. This is the classic mistake: someone forgets `.delay`, the email sends synchronously inside the web request, the request blocks, and they wonder why Celery "isn't doing anything." Celery never saw it.

Here's the contrast inside a web view:

```python
# views.py
def signup(request):
    user = create_user(request.POST)

    # WRONG: runs the email send inside the request - user waits for SMTP
    # send_welcome_email(user.id)

    # RIGHT: enqueue it and move on
    send_welcome_email.delay(user.id)

    return redirect("/welcome")  # returns instantly; email goes out later
```

*What just happened:* the `.delay(user.id)` line returns in microseconds because all it did was push a tiny message onto the broker. The HTTP response goes back to the user immediately, and some worker sends the actual email a moment later - outside the request/response cycle entirely. The commented-out direct call would have blocked `signup` until the SMTP handshake finished, exactly the latency you adopted Celery to avoid. **The rule:** if you want it to run in the worker, use `.delay()` (or `.apply_async()`). A bare call always stays inline.

## `.apply_async()` for options

`.delay()` is deliberately minimal - it's shorthand for "send these args, use all the defaults." When you need more control, reach for its longer sibling.

📝 `.apply_async()` takes the arguments as a list (and/or `kwargs` dict) plus a pile of options: `countdown` (wait N seconds before running), `eta` (run at a specific datetime, once), `queue` (route to a named queue), `priority`, retry settings, and more.

```python
from datetime import datetime, timedelta

# Run 60 seconds from now, on the dedicated "reports" queue
generate_report.apply_async(
    args=[account_id, "2026-06"],
    countdown=60,
    queue="reports",
)

# Or schedule it for a specific moment
generate_report.apply_async(
    args=[account_id, "2026-06"],
    eta=datetime(2026, 7, 1, 9, 0),
)
```

*What just happened:* the first call still enqueues immediately and returns an `AsyncResult` just like `.delay()` did - but the worker holds the message for 60 seconds (`countdown`) before executing, and it's routed to the `reports` queue so a report-only worker can pick it up instead of competing with email tasks. The second uses `eta` to pin execution to a wall-clock time. Note the shape: positional task arguments go inside `args=[...]`, and Celery's own options sit alongside them - the whole reason `.apply_async()` exists. For the everyday case with no options, `generate_report.delay(account_id, "2026-06")` is the exact same thing, just terser.

## Arguments must be serializable

This is the gotcha that bites hardest, so let's be blunt about it.

⚠️ Whatever you pass to a task gets **serialized** - by default to JSON - so it can be written into a broker message, travel across the network, and be **deserialized** by a worker in a different process. That means your arguments (and return values) must be simple, JSON-friendly data: numbers, strings, booleans, lists, dicts. You **cannot** pass a live database model object, an open file handle, a database connection, a request object, or anything else that only makes sense inside the process that created it.

💡 The fix is a one-liner habit: **pass an id, not the object.** Hand the task the primitive, and let the task re-load whatever it needs on the worker side.

```python
# WRONG: a model instance can't be JSON-serialized,
# and even if it could, it'd be a stale snapshot by the time the worker runs.
# send_welcome_email.delay(user)

# RIGHT: pass the id; the task fetches a fresh User inside the worker.
send_welcome_email.delay(user.id)
```

*What just happened:* passing `user` asks Celery to cram a whole `User` object into a JSON message, which either errors outright or (with a permissive serializer) ships a frozen copy already drifting out of date. Passing `user.id` ships a single integer. Look back at the task body in the first example - its very first line is `User.objects.get(id=user_id)`, re-loading a **fresh** copy inside the worker, with the worker's own database connection. That's why tasks take ids: the object that exists in your web process does not exist in the worker, so you give the worker the key to go find its own.

## How it travels

Let's zoom out and trace one task end to end, because once you see the pipe, every rule above stops feeling arbitrary.

💡 When you call `send_welcome_email.delay(42)`:

1. Celery looks up the **registered name** of the task (`tasks.send_welcome_email`) and serializes that name plus the args `[42]` into a message.
2. The message lands on the **broker** (the queue).
3. A **worker** pulls the message off the queue, deserializes it, finds the function registered under that name, and calls it with `42`.
4. The body runs in the worker, fetches the user, sends the email.

```mermaid
flowchart LR
  A["your web process<br/>.delay(42)"] -->|"serialize name + args"| B[(broker / queue)]
  B -->|"deserialize"| C["worker<br/>runs the function"]
```

📝 Notice what crosses the wire: a **name** and some **plain data** - never the function itself, never live objects. That's the whole reason arguments must be serializable, and why both sides have to agree on the task's name (the worker must import the same task code your web process does).

This also tells you how to *shape* a good task: keep it a **thin entry point** doing one unit of work. Take an id, load what you need, do the job. Resist stuffing five responsibilities into one task - small tasks are easier to retry, route, and reason about. The next question is "how do I get the result back?" That `AsyncResult` we kept brushing past is exactly what Phase 4 is about.

## Recap

- A function becomes a task by decorating it with `@app.task` (or `@shared_task`); that **registers** it by name, and its body runs **in the worker**, not your web process.
- `.delay(args)` enqueues the task and returns an `AsyncResult` immediately - calling the function **without** `.delay` runs it **inline**, bypassing Celery entirely. That's the most common beginner mistake.
- `.apply_async(args=[...], ...)` is the full-control version: `countdown`, `eta`, `queue`, `priority`, and more. `.delay()` is its shorthand for the common case.
- Arguments and return values are **serialized** (JSON by default) to cross the broker, so pass simple data - **an id, not a model object, file, or request.** Let the task re-load the object on the worker side.
- A task call ships a **name + plain data** to the broker; a worker deserializes it and runs the registered function. Keep each task a thin entry point doing one unit of work.

## Quick check

```quiz
[
  {
    "q": "What does send_welcome_email.delay(user_id) actually do?",
    "choices": [
      "Runs the function immediately in the current process",
      "Sends a message to the broker and returns right away, so a worker runs it later",
      "Blocks until the worker finishes and returns the email"
    ],
    "answer": 1,
    "explain": ".delay() enqueues the task and returns an AsyncResult immediately; the worker runs the body later. A bare call (no .delay) is what runs inline."
  },
  {
    "q": "Why should you pass user.id instead of the user object to a task?",
    "choices": [
      "It's shorter to type",
      "Arguments are serialized to cross the broker, so pass simple data and let the task re-load a fresh object in the worker",
      "Celery automatically converts objects to ids for you"
    ],
    "answer": 1,
    "explain": "Args are serialized (JSON by default) to travel through the broker. A live model object can't be serialized meaningfully, so you pass the id and the worker fetches a fresh copy."
  },
  {
    "q": "You need a task to run 60 seconds from now on a specific queue. Which call do you use?",
    "choices": [
      "task.delay(arg)",
      "task(arg)",
      "task.apply_async(args=[arg], countdown=60, queue=\"reports\")"
    ],
    "answer": 2,
    "explain": ".apply_async() exposes options like countdown, eta, queue, and priority. .delay() is only the shorthand for the no-options common case."
  }
]
```


---

# Results & State

Here's the mental model before anything else: **the broker and the result backend are two different mailboxes.** In Phase 3 we kept saying `.delay()` hands you back an `AsyncResult` - a receipt. That receipt is only useful if there's somewhere for the worker to *write down* what happened. The broker carries the message *to* the worker; it does not carry the answer back. For that you need a second store, the **result backend**, where the worker records "this task finished, here's its return value" or "this task blew up, here's the error."

Once that clicks, the rest of this phase is just: how do I set up that second mailbox, how do I read from it, and - the part most people get wrong - do I even need it?

## The result backend

📝 To capture what a task returned (or even just whether it succeeded), Celery needs a **result backend**: a place to store outcomes, keyed by task id. Configured separately from the broker, it can be Redis, a SQL database, or a few other stores.

```python
# celery_app.py
from celery import Celery

app = Celery(
    "myapp",
    broker="redis://localhost:6379/0",     # where messages go OUT
    backend="redis://localhost:6379/1",    # where results come BACK
)
```

*What just happened:* we pointed `broker` and `backend` at the same Redis server but **different databases** (`/0` vs `/1`) - separate concerns, kept in separate stores even on the same box. The broker holds pending *work*; the backend holds finished *outcomes*. They happen to both be Redis here, which is common, but don't have to be - plenty of setups use Redis as broker and Postgres as backend.

⚠️ Without a backend configured, results are **discarded**. The task still runs perfectly fine - the email still sends - but `result.get()` will hang or error, and `result.status` can never move past `PENDING`, because there's nowhere for the worker to write the outcome. If you can't read a result, first check whether you set `backend=` at all.

## AsyncResult: your handle on a running task

📝 When you call `.delay()` (or `.apply_async()`), you get back an `AsyncResult` - a thin handle to one specific task, identified by its id. It doesn't *contain* the result; it knows how to go *ask the backend* about it. The pieces you'll use:

- `result.id` - the unique task id (a UUID string). Save this if you want to check back later.
- `result.ready()` - `True` if the task has finished (success or failure), `False` if still pending/running.
- `result.status` - the current state, e.g. `"PENDING"`, `"SUCCESS"`, `"FAILURE"`.
- `result.get(timeout=...)` - fetch the actual return value, **blocking** until it's ready (or the timeout expires).

```python
# enqueue and hold onto the receipt
result = generate_report.delay(account_id, "2026-06")

print(result.id)        # "d5b8...e91" - a handle you could store and reload
print(result.ready())   # False - the worker probably hasn't finished yet
print(result.status)    # "PENDING"

# later, when you actually need the answer:
report_path = result.get(timeout=30)   # blocks up to 30s, then returns the value
```

*What just happened:* `.delay()` returned instantly with a receipt. `ready()` and `status` are cheap, non-blocking peeks at the backend - they answer "is it done yet?" without waiting. `get()` is the opposite: it **sits and waits** for the worker to finish, then hands you whatever the task `return`ed. `timeout=30` is a safety valve so you're not blocked forever if the worker is wedged.

⚠️ `get()` **blocks the calling thread until the task completes.** That's fine in a one-off script or the shell. It's a trap inside a web request - you'd hand control to Celery only to immediately sit and wait for it, throwing away the entire point of going async.

⚠️ Worse: **never call `.get()` inside another task.** A worker blocking on another task's result can deadlock the whole pool - if all your workers are busy waiting on results that only a free worker could produce, nothing moves. If you need to chain work, that's what chains and workflows are for, not nested `get()` calls.

## Task states

📝 A task moves through a small set of states, and the backend records the latest one:

- **PENDING** - Celery doesn't know anything yet. Queued but not started, *or* an id it's never heard of (more on that below).
- **STARTED** - a worker has picked it up and begun (only tracked if you opt in via `task_track_started`).
- **SUCCESS** - finished cleanly; the return value is in the backend.
- **FAILURE** - raised an exception; the backend holds the traceback.
- **RETRY** - failed but is scheduled to run again (Phase 5).
- **REVOKED** - cancelled before it could run.

```python
result = send_welcome_email.delay(user.id)

result.status        # "PENDING" - just enqueued
# ...worker picks it up and runs...
result.status        # "SUCCESS"
result.successful()  # True
result.failed()      # False
```

*What just happened:* we watched one task walk from `PENDING` to `SUCCESS`. The convenience methods `successful()` / `failed()` are just readable wrappers over `status`. The flow is always: unknown/queued → (started) → a terminal state (`SUCCESS`, `FAILURE`, or `REVOKED`), possibly looping through `RETRY` on the way.

⚠️ Here's the confusion that bites everyone: **PENDING also means "I have no record of this id."** Celery's backend stores a result *after* a task reaches a terminal state - it does **not** write a row the moment you enqueue. So Celery genuinely cannot tell "queued, waiting for a worker" apart from "this id never existed / was a typo." Both report `PENDING`. If a result is stuck `PENDING` forever, don't assume it's still running - it may have finished and expired, or you may be checking an id the backend never saw.

## When you need results - and when you don't

💡 The practical default for a lot of background work is: **you don't need the result.** `send_welcome_email` is fire-and-forget - nobody is waiting on its return value, and "did the email send?" is answered by your email provider's logs, not by polling Celery. Storing a result for it is pure overhead: every finished task writes a row to the backend nobody will ever read.

📝 You can turn results off per-task or globally:

```python
@app.task(ignore_result=True)
def send_welcome_email(user_id):
    user = User.objects.get(id=user_id)
    send_email(to=user.email, subject="Welcome aboard!", body=render_welcome(user))

# or globally, in config:
# app.conf.task_ignore_result = True
```

*What just happened:* `ignore_result=True` tells Celery not to bother writing this task's outcome to the backend. The task runs identically; we've only stopped recording an answer no one asked for - a real saving in backend writes and storage for a busy email queue.

💡 You **do** need results when something downstream waits on the outcome: a user clicked "Download report" and the file path comes back from `generate_report`; or one step feeds the next in a chain. The rule of thumb: **keep a result only if a real reader exists.** Don't store results you'll never read - they're not free.

## Polling vs pushing

So a user kicks off `generate_report` and wants to know when it's ready. You already know not to block the request with `get()`. What do you do instead?

💡 The simple, robust pattern is **polling by id**:

1. The request enqueues the task and immediately returns `result.id` to the client.
2. The client (or a status endpoint) checks back periodically: "is task `d5b8…` done?"
3. When the status flips to `SUCCESS`, fetch the value and show the download link.

```python
# views.py
def start_report(request):
    result = generate_report.delay(request.user.account_id, "2026-06")
    return json({"task_id": result.id})          # return the receipt, don't wait

def report_status(request, task_id):
    result = generate_report.AsyncResult(task_id)  # rebuild the handle from the id
    if result.ready():
        return json({"state": result.status, "url": result.result})
    return json({"state": result.status})          # still PENDING/STARTED
```

*What just happened:* `start_report` returns instantly with just the id - the request is never blocked. `report_status` reconstructs an `AsyncResult` **from that id alone** (no new task is enqueued; we're just looking one up) and reports whether it's done. The client polls this second endpoint every couple of seconds. Nobody blocks; the worker churns away independently. For richer experiences you can skip polling entirely and **push** completion - a websocket message or a webhook the task fires on success.

⚠️ One last gotcha: **results expire.** Celery deletes them from the backend after `result_expires` (24 hours by default), so a result is a short-lived notification, not durable storage. If a user might come back next week for that report, persist the real outcome (the file, a DB row) yourself. And sometimes the status you poll won't be `SUCCESS` but `FAILURE` or `RETRY` - handling *that* gracefully is exactly where Phase 5 picks up.

## Recap

- The **result backend** is a separate store from the broker - configure it with `backend=`. The broker carries work out; the backend carries outcomes back. Without one, results are discarded and status never leaves `PENDING`.
- `.delay()` returns an **`AsyncResult`**: a handle by id. Use `result.id`, `result.ready()`, `result.status` for cheap non-blocking peeks, and `result.get(timeout=...)` to fetch the return value.
- **`.get()` blocks.** Never call it inside a web request (defeats the purpose) or inside another task (can deadlock the worker pool).
- Task states run **PENDING → STARTED → SUCCESS / FAILURE / RETRY / REVOKED.** `PENDING` is ambiguous: it also means "unknown id," so don't read it as proof a task is still running.
- Skip results for fire-and-forget work (`ignore_result=True`); keep them only when a real reader waits. Expose a **status endpoint and poll by id** rather than blocking - and remember results **expire** (`result_expires`), so persist anything you need long-term yourself.

## Quick check

Check your grip on results and state before we get into failures:

```quiz
[
  {
    "q": "You configured Celery with only broker= and no backend=. A task runs fine, but result.status stays PENDING forever. Why?",
    "choices": [
      "The worker crashed silently",
      "With no result backend there is nowhere to record the outcome, so status can never advance",
      "PENDING means the task is still running"
    ],
    "answer": 1,
    "explain": "The backend is where outcomes get written. With no backend, results are discarded and status is stuck - the task still ran, you just can't observe its outcome."
  },
  {
    "q": "Why is calling result.get() inside a web request a mistake?",
    "choices": [
      "It returns the wrong value",
      "It blocks the request until the task finishes, throwing away the whole point of going async",
      "get() only works in the shell"
    ],
    "answer": 1,
    "explain": ".get() blocks the caller until completion. In a request that means the user waits exactly as long as a synchronous call - return the task id and poll a status endpoint instead."
  },
  {
    "q": "A task is fire-and-forget (send_welcome_email) and nothing ever reads its return value. What's the sensible setup?",
    "choices": [
      "Set ignore_result=True so Celery doesn't store an outcome nobody reads",
      "Always store the result so you have a record",
      "Call .get() after .delay() to confirm it sent"
    ],
    "answer": 0,
    "explain": "Storing results you'll never read is pure overhead that piles up in the backend. ignore_result=True skips the write; the task runs identically."
  }
]
```


---

# Retries & Error Handling

Here's the mental model to carry through this whole phase: **a background task is allowed to fail, and that's exactly the point.** A web request that fails shows the user a 500 and they retry by refreshing. A task has no user staring at it - it's running alone in a worker, minutes after the request that spawned it already returned. So the task has to be its own safety net: when something goes wrong, *it* decides whether to try again, how long to wait, and what to do when it finally gives up. Getting that decision right is most of what separates a toy Celery setup from one you'd trust with real money.

In Phase 4 you learned how to read a task's result and state. Now we deal with the messy reality those states are reporting on: timeouts, flaky third-party APIs, services that are down for ninety seconds, and workers that crash mid-job.

## Why background tasks fail

📝 Most task failures aren't bugs - they're **transient**. A mail server hiccups. A payment gateway times out. An internal service is rebooting. The network drops a packet. None of these mean your code is wrong; they mean the world was briefly uncooperative. The exact same call would succeed if you tried it again ten seconds later.

This is actually one of the core reasons to use a task queue in the first place. Inside a web request, a flaky upstream is a disaster - you can't make the user sit through a thirty-second retry loop, so you give up and show an error. A task has all the time in the world: it can wait, retry, back off, and eventually succeed without anyone noticing the bumps. The user already got their HTTP response; the work happening in the background is free to be patient.

⚠️ The flip side: because nobody is watching, a task that fails *permanently* can vanish silently. We'll deal with that at the end - first, retries.

## Retrying

There are two ways to make a task retry, and they fit different situations. Let's do both.

The first is **imperative**: you catch the error yourself and call `self.retry(...)`. To get access to `self`, you bind the task with `bind=True`.

```python
from celery import shared_task
from smtplib import SMTPException

@shared_task(bind=True, max_retries=5)
def send_welcome_email(self, user_id):
    user = User.objects.get(id=user_id)
    try:
        send_email(
            to=user.email,
            subject="Welcome aboard!",
            body=render_welcome(user),
        )
    except SMTPException as exc:
        # mail server hiccuped - wait 10s and try again
        raise self.retry(exc=exc, countdown=10)
```

*What just happened:* `bind=True` made `self` the first argument - the running task instance, and how you reach `self.retry()`. When the SMTP call throws, we don't let the task die; we call `self.retry(exc=exc, countdown=10)`, telling Celery to **re-enqueue this same task** to run again in 10 seconds. Note the `raise` - `self.retry()` actually raises a special `Retry` exception to abort the current run cleanly, so raising its return value is the idiomatic spelling. We pass `exc=exc` so that if we *do* eventually run out of retries, the original `SMTPException` gets recorded as the failure, not a generic one. `max_retries=5` caps it: after the fifth failed attempt, the task is allowed to fail for real.

📝 The second way is **declarative** - you describe which exceptions should auto-retry and let Celery handle the try/except for you:

```python
from celery import shared_task
from smtplib import SMTPException

@shared_task(
    autoretry_for=(SMTPException,),
    retry_backoff=True,
    retry_backoff_max=600,
    max_retries=5,
)
def send_welcome_email(user_id):
    user = User.objects.get(id=user_id)
    send_email(
        to=user.email,
        subject="Welcome aboard!",
        body=render_welcome(user),
    )
```

*What just happened:* this version has no try/except at all. `autoretry_for=(SMTPException,)` tells Celery: if the body raises that exception, retry automatically - same effect as the hand-written version, far less code. The new piece is `retry_backoff=True`, spacing retries out **exponentially**: roughly 1s, then 2s, 4s, 8s, instead of hammering every 10 seconds. 💡 Backoff matters because the thing you're retrying against is often *already struggling* - an overloaded mail server or API doesn't need your worker pounding it on a fixed interval. `retry_backoff_max=600` caps any single wait at 10 minutes so delays don't grow absurd, and `max_retries=5` still bounds the total attempts.

Reach for the declarative form for the common "retry on these exceptions with backoff" case. Drop to imperative `self.retry()` when you need to decide *at runtime* - for example, reading a `Retry-After` header from a rate-limited API and passing it as the `countdown`.

## Idempotency - the critical idea

This is the most important paragraph in the phase, so let's be blunt. **The moment you add retries, you have accepted that your task might run more than once.**

⚠️ A retried task runs again. A task that the broker redelivers (because the worker crashed, or the acknowledgement got lost) runs again. Even a task you only meant to send once can, under the right network failure, be *delivered* twice. This isn't a Celery quirk - it's the nature of distributed messaging, the same "at-least-once delivery" reality covered in [Webhooks & Message Queues](/guides/webhooks-and-message-queues).

Here's the nightmare made concrete:

```python
@shared_task(autoretry_for=(GatewayTimeout,), max_retries=3)
def charge_payment(order_id):
    order = Order.objects.get(id=order_id)
    # DANGER: the charge can succeed at the gateway,
    # then the *response* times out on the way back.
    gateway.charge(order.amount, order.card_token)  # raises GatewayTimeout
    order.mark_paid()
```

*What just happened:* something genuinely awful, hiding in plain sight. The gateway **successfully charged the card**, but the network dropped the response, so our code saw a `GatewayTimeout` and the task retried - **charging the customer a second time.** Retrying a non-idempotent task doesn't add safety; it adds risk.

💡 The fix is **idempotency**: design the task so that running it twice has the same effect as running it once. The standard tool is an **idempotency key** - a unique token for this logical operation that you check before acting:

```python
@shared_task(autoretry_for=(GatewayTimeout,), max_retries=3)
def charge_payment(order_id):
    order = Order.objects.get(id=order_id)

    if order.is_paid:                 # already done? do nothing.
        return order.payment_id

    # Pass an idempotency key so the gateway itself dedupes
    # a retried charge instead of billing twice.
    result = gateway.charge(
        order.amount,
        order.card_token,
        idempotency_key=f"order-{order.id}",
    )
    order.mark_paid(payment_id=result.id)
    return result.id
```

*What just happened:* two layers of protection. First, the early `if order.is_paid` check means a redelivered task that already finished returns and does nothing - running it again is harmless. Second, we hand the payment gateway an `idempotency_key` tied to the order, so even if our two attempts both reach the gateway, *it* recognizes the second one as a duplicate and returns the original charge instead of billing again. (Every serious payment API supports this mechanism, precisely because retries are universal.) **This is the discipline: before you turn on retries, make the task idempotent.**

## acks_late & worker crashes

There's one more way a task can run twice, and it forces a genuine trade-off you have to choose on purpose.

📝 By default, a worker **acknowledges** a task to the broker the instant it *picks it up* - before running it. The broker hears "got it" and drops the message. Now imagine the worker crashes (out of memory, deploy, machine reboot) one second into a thirty-second job. The message is already gone from the broker - the task is **lost forever**, half-finished, and nothing will ever retry it. This is "at-most-once" delivery: a task runs zero or one times, never more.

📝 Setting `acks_late=True` flips this. The worker acknowledges only **after** the task finishes successfully:

```python
@shared_task(acks_late=True)
def generate_report(account_id, month):
    rows = query_usage(account_id, month)
    pdf = render_pdf(rows)
    store_report(account_id, month, pdf)
```

*What just happened:* the task is no longer acknowledged up front. If this worker dies mid-report, the broker never heard a confirmation, so after a timeout it **re-queues the message** and another worker runs it again. The report survives the crash. ⚠️ But look at the cost: if the crash happens *after* `store_report` runs but *before* the ack is sent, the task gets re-run and the report is generated twice. That's the at-least-once trade - `acks_late=True` guarantees the work isn't lost, at the price of possibly running more than once, which is exactly why it only makes sense for an **idempotent** task. (`generate_report` here re-creates a report keyed by `(account_id, month)`, so running it twice just overwrites the same output - safe.) Rule of thumb: `acks_late` for work you can't afford to lose *and* have made safe to repeat; the default for work where a rare double-run would be worse than a rare miss.

## Failure handling

Retries handle the transient failures. But some failures are permanent - a malformed record, a deleted user, a bug - and no amount of retrying will fix them. You need a plan for when a task truly gives up.

📝 When a task exhausts its retries (or raises an error you didn't auto-retry), Celery marks its state `FAILURE` and, if you've configured a result backend, stores the exception and traceback there (that's what Phase 4's `AsyncResult` reads). You can also hook the moment of final failure with an `on_failure` handler or an error callback:

```python
from celery import Task

class AlertOnFailure(Task):
    def on_failure(self, exc, task_id, args, kwargs, einfo):
        logger.error("Task %s failed permanently: %s", self.name, exc)
        alert_oncall(f"{self.name} failed for args={args}: {exc}")

@shared_task(base=AlertOnFailure, autoretry_for=(SMTPException,), max_retries=5)
def send_welcome_email(user_id):
    ...
```

*What just happened:* by giving the task a custom `base` class, its `on_failure` runs **only after the last retry has failed** - the final-defeat hook. Here we log the error and page on-call. Without something like this, a permanently failing task fails into the void: no exception bubbles up to a user, no stack trace lands in your request logs, nothing. 💡 For high-volume cases, the equivalent pattern at the queue level is a **dead-letter queue** - failed messages get routed to a separate queue you can inspect and replay later.

⚠️ Internalize this: **a silently-failing background task is worse than a failing web request.** A failed request screams at the user, who tells you. A failed task whispers to no one - the welcome email just never arrives, the report never appears, and you find out from an angry customer next week. You only know your tasks are failing if you actively look, which is why monitoring (Phase 7) isn't optional.

💡 So here's the whole discipline of this phase in one breath: **make tasks idempotent, retry transient errors with backoff, and surface permanent failures loudly.** Do those three things and your background jobs become something you can actually trust.

## Recap

- Most task failures are **transient** (flaky APIs, timeouts, brief outages) - a task is the right place to handle them because, unlike a web request, it can afford to wait and retry.
- Retry **imperatively** with `self.retry(exc=..., countdown=..., max_retries=...)` (needs `bind=True`), or **declaratively** with `autoretry_for=(...)`, `retry_backoff=True`, and `max_retries`. Use backoff so you don't hammer a struggling service.
- **Idempotency is non-negotiable once you retry:** a retried or redelivered task may run more than once. Check "already done?" and use an idempotency key so a retried `charge_payment` doesn't double-charge.
- `acks_late=True` acks after completion, so a worker crash re-queues the task instead of losing it - the at-least-once trade. It only makes sense for idempotent tasks.
- On final failure, the state is `FAILURE` (exception stored if you have a backend); use `on_failure`/error callbacks, dead-letter queues, and alerting to surface it. A silently failing task is worse than a failing request - monitor them.

## Quick check

```quiz
[
  {
    "q": "Why must a task be idempotent before you enable retries?",
    "choices": [
      "Retries run faster on idempotent tasks",
      "A retried or redelivered task may run more than once, so running twice must equal running once (e.g. no double charge)",
      "Celery refuses to retry tasks that aren't marked idempotent"
    ],
    "answer": 1,
    "explain": "Retries and broker redelivery mean at-least-once execution. If the task isn't idempotent, a second run causes real damage like a duplicate payment."
  },
  {
    "q": "What does retry_backoff=True do?",
    "choices": [
      "Cancels the task after the first failure",
      "Spaces retries out exponentially (1s, 2s, 4s...) so you don't hammer a struggling service",
      "Retries the task an unlimited number of times"
    ],
    "answer": 1,
    "explain": "Exponential backoff increases the wait between attempts, giving an overloaded upstream room to recover instead of pounding it on a fixed interval."
  },
  {
    "q": "What is the trade-off of setting acks_late=True?",
    "choices": [
      "Tasks run faster but use more memory",
      "If a worker crashes the task is re-queued (not lost), but it may run twice - so the task must be idempotent",
      "Results are stored permanently instead of expiring"
    ],
    "answer": 1,
    "explain": "acks_late acknowledges only after completion, so a crash re-queues the task (at-least-once). The cost is possible double execution, which is only safe for idempotent tasks."
  }
]
```


---

# Scheduled Tasks with Celery Beat

So far every task in this guide has been triggered by *something happening* - a user signs up, a report is requested, a web request comes in and your app fires off a job. But a huge amount of real work isn't request-triggered at all. It's work that just needs to happen *on a schedule*, whether anyone's looking or not.

📝 Think about a typical web app: a digest email that goes out every morning at 7am, a cleanup job that sweeps away old reports every hour, a billing summary that runs on the first of the month. Nobody clicks a button for these - they run because the clock said so. What you want is something cron-like - a way to run your existing Celery tasks on a timetable. That's exactly what Celery Beat gives you.

If you've used Unix cron (covered in [The Terminal & Shell](/guides/the-terminal-and-shell)), the mental model will feel familiar: a clock that fires jobs at set times. Beat is that clock, but for your Celery tasks specifically.

## What Celery Beat is: the clock, not the doer

📝 **Beat is a scheduler process. It does not run your tasks itself.** This is the single most important thing to understand, and it trips people up constantly. When the schedule says "time to run `send_daily_digest`," Beat doesn't execute that function. Instead, it *enqueues a message* into the broker - the same Redis queue from [Phase 2](02-the-broker-and-worker.md) - and a normal worker picks it up and runs it, exactly like any other task.

So the division of labor is clean:

- **Beat = the clock.** It watches the schedule and drops task messages into the broker at the right times.
- **The worker = the doer.** It pulls those messages off the queue and actually runs your code.

You start Beat as its own process, separate from your worker:

```bash
celery -A tasks beat --loglevel=info
```

*What just happened:* we launched the Beat scheduler. `-A tasks` points it at your Celery app (same `tasks.py` as always), and `beat` is the subcommand that says "be the scheduler." This process stays running, watching the clock. When a scheduled time arrives, it enqueues the matching task and goes back to waiting.

⚠️ **Beat only schedules - it never executes.** If you run *only* Beat with no worker, your scheduled tasks will pile up in the broker and never run. You need both processes alive at the same time: Beat to enqueue on schedule, and at least one worker to drain the queue. A common "my nightly job isn't running" bug is forgetting to keep a worker up.

## Defining a schedule

You tell Beat what to run and when through `app.conf.beat_schedule` - a dictionary where each entry names a task, a schedule, and (optionally) arguments. Put this in your `tasks.py`:

```python
from celery import Celery
from celery.schedules import crontab

app = Celery("tasks", broker="redis://localhost:6379/0")

app.conf.beat_schedule = {
    "daily-digest-email": {
        "task": "tasks.send_daily_digest",
        "schedule": crontab(hour=7, minute=0),
    },
    "hourly-report-cleanup": {
        "task": "tasks.cleanup_old_reports",
        "schedule": 3600.0,
    },
}
```

*What just happened:* we defined two scheduled entries. The dictionary *keys* (`"daily-digest-email"`, `"hourly-report-cleanup"`) are just human-readable names for each schedule - pick whatever's clear. Inside each entry, `"task"` is the dotted name of the task to run (the same string you'd see in the worker's `[tasks]` banner), and `"schedule"` says when. The digest uses `crontab(hour=7, minute=0)` - "every day at 7:00am." The cleanup uses a plain number, `3600.0` - "every 3600 seconds," once an hour.

💡 Two flavors of schedule, two use cases. A **plain number** (or a `timedelta`) means "every N seconds" - a fixed *interval*. A **`crontab(...)`** means "at these calendar times" - wall-clock scheduling. The cleaner version of that hourly cleanup uses `timedelta` so the intent reads at a glance:

```python
from datetime import timedelta

app.conf.beat_schedule = {
    "hourly-report-cleanup": {
        "task": "tasks.cleanup_old_reports",
        "schedule": timedelta(hours=1),
    },
}
```

*What just happened:* `timedelta(hours=1)` is exactly equivalent to `3600.0` but says what it means. Reach for intervals when "every so often" is enough and the exact wall-clock time doesn't matter; reach for `crontab` when it has to happen at, say, 7am sharp.

If your task takes arguments, pass them with an `"args"` (tuple) or `"kwargs"` (dict) key:

```python
app.conf.beat_schedule = {
    "weekly-summary": {
        "task": "tasks.send_summary",
        "schedule": crontab(hour=9, minute=0, day_of_week="monday"),
        "args": ("weekly",),
    },
}
```

*What just happened:* each Monday at 9:00am, Beat enqueues `send_summary("weekly")`. The `"args"` tuple passes straight through to your task function, just as if you'd called `send_summary.delay("weekly")` yourself.

## crontab schedules in a bit more depth

📝 The `crontab()` schedule mirrors Unix cron: you specify some combination of `minute`, `hour`, `day_of_week`, `day_of_month`, and `month_of_year`, and Beat fires the task whenever the clock matches. Anything left out defaults to "every" - so `crontab(minute=0)` means "at minute 0 of *every* hour."

A few examples to anchor it:

```python
crontab(minute=0, hour=7)                          # every day at 7:00am
crontab(minute=0, hour=9, day_of_week="monday")    # every Monday at 9:00am
crontab(minute="*/15")                             # every 15 minutes
crontab(minute=0, hour=0, day_of_month=1)          # midnight on the 1st of each month
```

*What just happened:* each line is a wall-clock rule. Note `minute="*/15"` - the `*/N` step syntax is straight from cron and means "every 15th minute." If cron's field syntax is fuzzy, the cron reference in [The Terminal & Shell](/guides/the-terminal-and-shell) covers the same `minute hour day month weekday` grammar Beat borrows from.

⚠️ **Timezones will bite you.** By default Celery interprets `crontab` times in UTC, *not* your local time - so `crontab(hour=7)` might fire at what feels like the middle of the night. Set the timezone explicitly so 7am means 7am where you live:

```python
app.conf.timezone = "America/New_York"
app.conf.enable_utc = True
```

*What just happened:* we told Celery to evaluate schedules against New York time. Now `crontab(hour=7)` fires at 7am Eastern. ⚠️ Double-check this in production: a digest that's supposed to land at breakfast but goes out at 2am is the classic symptom of a forgotten `timezone` setting.

## Production gotchas

⚠️ **Run exactly one Beat process. Ever.** This is the big one. Beat is the clock, and if you run *two* clocks, every scheduled task gets enqueued twice - your daily digest goes out to every user twice, your cleanup runs in duplicate. This "double-scheduled tasks" bug is sneaky because everything looks fine until someone notices the duplicate emails. One worker can (and should) be scaled to many processes, but Beat must be a single instance. Be especially careful deploying multiple app servers - it's dangerously easy to accidentally start Beat on each one.

💡 Because a schedule can fire a task more than once - two Beats, a restart that re-fires a missed slot, a retry from [Phase 5](05-retries-and-error-handling.md) - your **periodic tasks should be idempotent too.** Running `cleanup_old_reports` twice in a row should be harmless; sending the digest twice should not double-charge or double-notify. Same "make running it twice safe" discipline as the retries phase, same underlying reason: at-least-once delivery means *plan for twice*.

For schedules that need to change at runtime - letting admins add or edit scheduled jobs from a UI, or storing schedules in a database instead of hardcoding them in `tasks.py` - reach for a database-backed scheduler:

- **`django-celery-beat`** stores the schedule in your Django database, editable through the admin. Great if you're already on Django.
- **`RedBeat`** stores the schedule in Redis and is designed to coordinate so you can run it more safely in multi-instance setups.

💡 Put it all together and the recipe for reliable scheduled work is short: **Beat (the clock) + workers (the doers) + idempotent tasks (the safety net).** Keep exactly one Beat running, keep workers up to drain what it enqueues, and write your periodic tasks so a double-fire never hurts. That foundation is what Phase 7 builds on when we take all of this - workers, retries, schedules - into production for real.

## Recap

- **Celery Beat is a scheduler process**, started with `celery -A tasks beat`. It runs separately from your worker.
- Beat is **the clock, not the doer**: on schedule it *enqueues* tasks into the broker, and a normal worker runs them. No worker running means scheduled tasks never execute.
- Define schedules in **`app.conf.beat_schedule`** - each entry has a `"task"` name, a `"schedule"`, and optional `"args"`/`"kwargs"`.
- Use a **number/`timedelta`** for fixed intervals ("every N seconds") and **`crontab(...)`** for calendar timing ("every day at 7am," "every Monday 9am"). Set `app.conf.timezone` or schedules run in UTC.
- **Run exactly one Beat process** - two Beats double-schedule every task. For runtime-editable or DB-backed schedules, use `django-celery-beat` or `RedBeat`.
- **Periodic tasks should be idempotent**, since a schedule can fire twice on restarts or duplicate Beats.

## Quick check

```quiz
[
  {
    "q": "When a scheduled time arrives, what does Celery Beat actually do?",
    "choices": ["Runs the task function directly inside the Beat process", "Enqueues a task message into the broker for a worker to run", "Sends the result back to your web app"],
    "answer": 1,
    "explain": "Beat is the clock, not the doer. On schedule it drops a task message into the broker, and a normal worker picks it up and executes it. That's why you still need a worker running alongside Beat."
  },
  {
    "q": "Which schedule means 'every day at 7:00am'?",
    "choices": ["schedule=7.0", "crontab(hour=7, minute=0)", "timedelta(hours=7)"],
    "answer": 1,
    "explain": "crontab(hour=7, minute=0) is calendar/wall-clock timing - 7:00am every day. A plain number or timedelta is a fixed interval ('every N seconds'), not a specific time of day."
  },
  {
    "q": "Why is it critical to run exactly one Beat process?",
    "choices": ["Beat can only connect to one broker at a time", "Two Beat processes double-schedule every task, causing duplicate jobs", "Multiple Beats slow down the workers"],
    "answer": 1,
    "explain": "Each Beat is an independent clock. Two clocks means every scheduled task gets enqueued twice - duplicate emails, double cleanups. Scale workers freely, but keep Beat to a single instance (and make tasks idempotent as a safety net)."
  }
]
```


---

# Production: Scaling, Monitoring & Pitfalls

Here's the mental model for this whole phase: **everything you've learned so far makes Celery *work*; this phase makes it *survive production*.** A toy Celery setup and a rock-solid one run the exact same code. The difference is entirely operational - how you scale it, whether you can see what it's doing, and whether you've sidestepped the handful of traps that bite everyone once. Think of this as a field guide you'll come back to, not a tutorial you read once.

You've got tasks, retries, idempotency, results, and schedules. Now let's keep all of that healthy when real traffic hits it.

## Scaling: more workers, more queues

📝 **The first lever is running more workers.** Every worker process you start connects to the same broker and drains the same queue. Start a second worker - on the same box or a different machine - and the two of them split the pending work between them, no coordination needed. That's the beauty of the broker-in-the-middle shape from [The Broker & Worker](02-the-broker-and-worker.md): scaling out is "add more consumers." You also have the per-worker `--concurrency` knob to run more tasks inside each process. Roughly: `--concurrency` scales *one* machine up; more worker processes scale *out* across machines.

📝 **The second, more important lever is routing different work to different queues.** By default everything lands in one queue called `celery`, and every worker drains it. That's fine until the day a flood of slow `generate_report` jobs piles up - now your quick `send_welcome_email` tasks are stuck behind them, and welcome emails go out an hour late. The fix is to give slow work its own queue with its own dedicated workers, so the two never compete.

```python
# celeryconfig.py - route tasks to named queues by name
task_routes = {
    "tasks.send_welcome_email": {"queue": "fast"},
    "tasks.charge_payment":     {"queue": "fast"},
    "tasks.generate_report":    {"queue": "slow"},
}
```

*What just happened:* we declared a routing table. Now when your app enqueues `generate_report`, Celery drops that message into the `slow` queue instead of the default; the quick tasks go to `fast`. Nothing else in your task code changes - routing is pure configuration. The messages are now *sorted*; the next step is putting workers on each pile.

```bash
# A pool of workers dedicated to quick jobs
celery -A tasks worker --queues=fast --concurrency=8 -n fast@%h

# A separate worker (or pool) for the heavy reports
celery -A tasks worker --queues=slow --concurrency=2 -n slow@%h
```

*What just happened:* we started **two independent worker pools**, each told via `--queues` which queue to drain. The `fast` pool ignores the `slow` queue entirely, so a backlog of reports can never starve your emails. We also gave each a distinct node name with `-n` (`%h` expands to the hostname) so they show up as separate workers in monitoring. 💡 This single pattern - a `fast` queue and a `slow`/reports queue with dedicated workers - solves the most common "Celery feels laggy under load" complaint in production.

## Monitoring: you can't see tasks otherwise

⚠️ Background tasks fail in the dark. There's no user watching a spinner, no 500 page, no red toast. When `send_welcome_email` throws on the tenth retry, *nothing visible happens* - the email never arrives, and you find out when a customer complains next week. This is the exact danger called out in [Retries & Error Handling](05-retries-and-error-handling.md): a silently-failing task is worse than a failing web request, because a failing request at least screams.

📝 **Flower is the first tool to reach for** - a web dashboard built specifically for Celery. You run it as one more process pointed at your broker:

```bash
celery -A tasks flower --port=5555
```

*What just happened:* we started Flower, which connects to the same broker, listens to Celery's event stream, and serves a dashboard at `http://localhost:5555`. Open it and you can see your workers (online/offline, tasks each is running), live task throughput, success and failure rates, and a searchable list of recent tasks with their args, runtime, and tracebacks - the "is anything on fire?" view you were missing.

📝 Flower is the convenient dashboard, but for real production you also want **the three pillars of observability** wired into your existing stack: structured **logs** from each task (use Celery's task logger so every line is tagged with the task name and id), **metrics** (task counts, failure rate, runtime histograms, and crucially **queue depth**) scraped into something like Prometheus, and **traces** that follow a request from your web app into the task it spawned. Full picture in [Observability: Logs, Metrics & Traces](/guides/observability-logs-metrics-traces) - Flower answers "right now," metrics + alerting answer "tell me *before* it's a problem."

💡 The line to internalize: **if you can't see your queue depth and your failure rate, you're flying blind.** A queue that's slowly growing instead of draining is the earliest, clearest signal that you're under-provisioned - but only if something is watching it.

## The pitfall cheat-card

These are the traps that get nearly everyone exactly once. Keep this table handy; when something's mysteriously broken, scan the symptom column first.

| Symptom | Cause | Fix |
|---|---|---|
| Task never runs (silent) | Broker down / wrong broker URL / no worker started / task not imported | Check the broker first (Phase 2): is it up, does the URL match, is a worker connected to that queue? |
| `EncodeError` / pickling error, or task gets stale data | Passing a model object as an argument | Pass an **id**, re-fetch inside the task (Phase 3) |
| Task runs *inline*, blocking the request | Called the function directly instead of `.delay()`/`.apply_async()` | Use `.delay(...)` to enqueue (Phase 3) |
| Double charge / duplicate effect | Non-idempotent task retried or redelivered | Make the task **idempotent** with an idempotency key (Phase 5) |
| Worker hangs / deadlocks | Calling `.get()` on a result *inside* a task or web request | Don't block on results inside tasks; chain or restructure (Phase 4) |
| Scheduled job fires twice | Two Beat processes running | Run **exactly one** Beat (Phase 6) |
| Broker memory bloat / slow enqueues | A giant argument (whole file, huge list) sent through the broker | Pass a reference (id, S3 key); keep messages tiny |

Three of these deserve a closer look because they're the ones people argue with:

- ⚠️ **Task never runs, no error.** This is the number-one Celery mystery, and the instinct is to debug your *code*. Don't - the code never ran. The message went into a queue nobody is draining (broker down, wrong URL, or no worker on that queue). Check the broker and worker before the task body.
- ⚠️ **Passing a model object.** It feels natural to write `send_welcome_email.delay(user)`. But the arguments get **serialized** into a message and may run seconds or minutes later - by then the object is stale, and rich objects often can't be serialized at all. Pass `user.id` and re-fetch fresh inside the task.
- ⚠️ **A giant task argument.** Every argument travels through the broker as part of the message. Shove a 20 MB CSV in there and you bloat broker memory, slow every enqueue, and risk hitting message-size limits. Upload the blob somewhere, pass the key, let the task fetch it.

## Worker reliability

A few knobs keep workers from being their own worst enemy. Brief, but each one has saved a production system.

📝 **Time limits stop runaway tasks.** A task stuck in an infinite loop or a hung network call will occupy a worker slot forever. Two settings cap that:

```python
# celeryconfig.py
task_soft_time_limit = 50   # raise SoftTimeLimitExceeded - task can clean up
task_time_limit      = 60   # hard kill the worker process at 60s, no questions
```

*What just happened:* `task_soft_time_limit` raises a catchable `SoftTimeLimitExceeded` inside the task at 50 seconds, giving it a chance to release a lock or log something; `task_time_limit` is the hard backstop - at 60 seconds Celery kills the child process outright. Together they guarantee no single task can wedge a worker indefinitely.

📝 Two more worth setting:

- `worker_max_tasks_per_child = 1000` - recycle each child process after 1000 tasks. If a task slowly leaks memory (a common C-extension reality), this caps the damage by restarting the process before it bloats.
- `worker_prefetch_multiplier` - how many messages a worker grabs *ahead* of what it's running. The default (4) is great for many short tasks but terrible for long ones: a worker can hoard several slow jobs while siblings sit idle. ⚠️ **For long-running tasks, set `worker_prefetch_multiplier = 1`** so each worker takes one job at a time and work spreads evenly.

## Deployment shape

💡 Step back and look at what a production Celery deployment actually *is* - a small set of long-running processes:

```mermaid
flowchart LR
  Web[Web app] -->|enqueue| Broker[(Broker: Redis/RabbitMQ)]
  Beat[Beat x1] -->|scheduled jobs| Broker
  Broker --> W1[Worker: fast queue]
  Broker --> W2[Worker: slow queue]
  W1 --> Flower[Flower / metrics]
  W2 --> Flower
```

The pieces: **the broker** (Redis or RabbitMQ), **N worker processes** routed by queue and run under a supervisor or in containers so they restart if they die, **exactly one Beat process** if you schedule anything (Phase 6), and **Flower plus metrics** for visibility. That's the whole production footprint.

And the operating discipline, in one breath: **keep tasks small and idempotent, route work by queue, set time limits, run one Beat, and monitor your failure rate and queue depth.** 💡 The difference between "Celery is flaky and we don't trust it" and "Celery is rock-solid, we forget it's even there" is not the library - it's exactly this checklist. Same code, different operations.

## Recap

- **Scale out** by running more worker processes (all draining the same broker) and **scale up** each with `--concurrency`; the bigger win is **routing** slow work to its own queue with dedicated workers so a flood of reports can't starve quick emails.
- Background tasks **fail silently** - use **Flower** for a live dashboard and wire structured logs, metrics (queue depth, failure rate), and traces into your stack ([Observability](/guides/observability-logs-metrics-traces)). If you can't see queue depth and failure rate, you're blind.
- The **pitfall cheat-card**: a silent no-op means check the broker; pass an **id** not a model object; use `.delay()` not a direct call; make retried tasks **idempotent**; don't `.get()` inside a task; run **one** Beat; keep arguments tiny.
- **Reliability knobs**: `task_time_limit`/`task_soft_time_limit` kill runaways, `worker_max_tasks_per_child` recycles memory-leaking workers, and `worker_prefetch_multiplier = 1` spreads long tasks evenly.
- A production deployment is just a **broker + N workers (by queue) + one Beat + Flower/metrics** - and the discipline around it is what makes Celery rock-solid instead of flaky.

## Quick check

```quiz
[
  {
    "q": "A flood of slow report jobs keeps delaying your quick welcome emails. What's the standard fix?",
    "choices": [
      "Increase max_retries on the email task",
      "Route slow and fast work to separate queues, each with its own dedicated workers",
      "Switch the broker from Redis to RabbitMQ"
    ],
    "answer": 1,
    "explain": "Put slow jobs on their own queue with dedicated workers so they run in a separate lane. Quick tasks on the fast queue never wait behind a backlog of reports."
  },
  {
    "q": "Why is monitoring (Flower, metrics, alerts) considered non-optional for production Celery?",
    "choices": [
      "It makes tasks run faster",
      "Background tasks fail silently - with no user watching, you only know about failures and growing queues if something is actively watching",
      "Celery refuses to start workers without a dashboard attached"
    ],
    "answer": 1,
    "explain": "A failed task screams to no one. Without visibility into failure rate and queue depth, problems stay invisible until a customer complains."
  },
  {
    "q": "Your tasks run for several minutes each, but work bunches up on some workers while others sit idle. Which setting helps?",
    "choices": [
      "Set worker_prefetch_multiplier = 1 so each worker takes one long task at a time",
      "Set task_time_limit = 1 to kill tasks faster",
      "Set worker_max_tasks_per_child = 1 to recycle after every task"
    ],
    "answer": 0,
    "explain": "The default prefetch lets a worker hoard several queued jobs ahead of time - fine for short tasks, bad for long ones. Setting it to 1 makes each worker grab one job at a time so long tasks spread evenly."
  }
]
```


---

# Where to Go Next

Take a second and look at what you can actually do now. Define a task, hand it off with `.delay()`, and watch a separate worker pick it up off the broker. Track its state through a result backend with `AsyncResult`, retry it sanely when a flaky API blinks, schedule it to run every night with Beat, and scale a fleet of workers while watching the whole thing breathe in Flower. None of that is a toy - it's the shape of real production background work.

And the part that makes it *yours* is that you understand the model underneath. Every quirk, every config knob, every confusing error traces back to one picture: **your app puts a message on a broker, a worker pulls it off and runs it, and a result backend optionally remembers what happened.** App → broker → worker → result. Hold those four pieces and Celery stops being intimidating.

So this last phase isn't more decorators. It's where Celery meets the web frameworks you already use, a pattern for stitching tasks into workflows, a clear-eyed look at when *not* to reach for Celery, and what to build next.

## How it plugs into your web framework

💡 You'll almost never run Celery alone. It lives next to a web app, and the division of labor is always the same: **the web request enqueues a task and returns immediately; a worker runs it later.** The request stays fast; the slow work happens off to the side.

The wiring differs a little per framework:

- **[Django](/guides/django-from-zero)** - the most paved road. The `django-celery-results` and `django-celery-beat` packages store results and schedules in your database, and you write tasks with `@shared_task` so they don't depend on importing your Celery app directly. A view calls `send_report.delay(user.id)` and returns.
- **[Flask](/guides/flask-from-zero)** - create the Celery app inside your **app factory** and use a custom `ContextTask` base so tasks run inside a Flask application context (config, DB session, extensions). After that it's the same `.delay()` from a route.
- **[FastAPI](/guides/fastapi-from-zero)** - no special integration to learn. Import your task and call `.delay()` straight from an endpoint. FastAPI's built-in `BackgroundTasks` covers light fire-and-forget work; Celery is what you graduate to when the job is heavy, retryable, or scheduled.

The pattern underneath never changes. The framework is the front desk that takes the order; the workers are the kitchen out back.

## Canvas: composing tasks into workflows

📝 So far every task has been a lone job. But real work often comes in steps - fetch the data, *then* transform it, *then* email the result. Celery has a vocabulary for stitching tasks together called **canvas**, and three pieces cover most of it:

- **`chain`** - run tasks in sequence, passing each result into the next. (Fetch → process → notify.)
- **`group`** - run many tasks in parallel and collect their results. (Thumbnail 100 images at once.)
- **`chord`** - a group *plus* a callback that fires once every task in the group finishes. (Process all the rows, then send one summary.)

A tiny chain reads almost like a sentence:

```python
from celery import chain

# Run these in order; each result flows into the next task.
workflow = chain(fetch_data.s(url), clean_data.s(), save_report.s())
workflow.delay()
```

That `.s()` is a **signature** - a task plus its arguments, packaged up so Celery can wire it into the pipeline instead of running it right now. This is only a taste; canvas goes deeper. But the moment you catch yourself chaining tasks by hand with one task calling `.delay()` on the next, that's your cue to reach for it.

## Celery vs the alternatives, plainly

Celery is the standard, and it earned that spot - it's powerful and feature-rich. But "standard" doesn't mean "always right." All that power comes with heavier configuration, and for a modest job it can feel like bringing a forklift to carry a grocery bag. Here's the clear-eyed landscape:

```mermaid
flowchart TD
  Q{What do you need?} --> F[Scheduling, routing,<br/>retries, canvas?]
  Q --> S[Just simple<br/>background jobs?]
  Q --> A[An asyncio app?]
  F --> Celery[Celery]
  S --> RQ[RQ or Dramatiq]
  A --> arq[arq]
```

- **Celery** - the full toolbox: scheduling, routing, retries, canvas, multiple brokers. Reach for it when you actually need those features. The cost is configuration.
- **RQ (Redis Queue)** - simpler and smaller, Redis-only. Genuinely lovely for modest needs where Celery's machinery is more than the job warrants.
- **Dramatiq** - simpler than Celery but with solid, sensible defaults out of the box. A strong middle ground.
- **arq** - async-native, built for `asyncio` apps. If your codebase is already `async def` all the way down, arq fits the grain.
- **Cloud queues & framework built-ins** - managed services like **AWS SQS** handle the broker for you, and tools like FastAPI's `BackgroundTasks` cover the lightest fire-and-forget work with no extra infrastructure at all.

💡 The practical rule: reach for **Celery when you need its features** - scheduling, routing, retries, canvas. Reach for **RQ or Dramatiq when you want simpler**. There's no shame in picking the smaller tool; the best background queue is the one your team can actually operate at 3 a.m.

## What to build - and one last thing

The fastest way to make all of this stick is to wire it into something real. Two projects, both small enough to finish:

1. **Move email off the request.** Take a web app you've got and pull email-sending into a Celery task. Add a `retry` so a flaky mail provider doesn't lose a signup. Then add a **nightly digest** with Beat, and watch all of it run in Flower. That single project exercises tasks, retries, scheduling, and monitoring - most of this guide in one go.
2. **Report generation with a status endpoint.** Kick off a slow report as a task, return its task ID, and add an endpoint that checks `AsyncResult` so the front end can poll for "pending → in progress → done." That's the result-backend story made tangible.

When you want the canonical reference, the **official Celery documentation** is the place to go - start with its *"First Steps with Celery"* tutorial, which walks the broker-and-worker setup from scratch and is maintained by the people who build it.

And remember the through-line: background work was never magic. It's a hand-off through a broker you now understand. The job leaves your request, lands on a queue, gets picked up by a worker, and its outcome gets remembered - every "it just happens in the background" is that one shape, wearing different clothes. You can read what's underneath now, build on top of it, and reason about it when it breaks. Go move something slow off the request path and watch it run.

## Recap

1. **You can do the whole job now** - define, call, retry, schedule, scale, and monitor tasks - and you understand the app → broker → worker → result model that ties it all together.
2. **Celery lives next to a web framework** - Django (`django-celery-*`, `@shared_task`), Flask (app-factory init with a `ContextTask`), FastAPI (`.delay()` straight from an endpoint). The web app enqueues; workers run.
3. **Canvas composes tasks** - `chain` for sequences, `group` for parallel work, `chord` for a callback after a group finishes. Use signatures (`.s()`) to wire them together.
4. **Pick the right tool plainly** - Celery for its full feature set, RQ or Dramatiq when you want simpler, arq for async apps, and cloud queues or built-ins for light work.
5. **Build one thing and finish it** - move email-sending to a task with a retry and a nightly Beat digest, or build a report task with an `AsyncResult` status endpoint. Background work is a hand-off through a broker you now understand.

## Quick check

Test yourself on the decisions that matter most as you leave this guide:

```quiz
[
  {
    "q": "In a web app using Celery, what does the web request itself do with a heavy task?",
    "choices": [
      "It runs the task inline and waits for the result before responding",
      "It enqueues the task on the broker and returns immediately, while a worker runs it later",
      "It blocks until a worker confirms the task finished",
      "It schedules the task with Beat so it runs the next night"
    ],
    "answer": 1,
    "explain": "The pattern is always the same: the web app enqueues the task (for example with .delay()) and returns fast; a separate worker process picks it up off the broker and runs it later."
  },
  {
    "q": "You need to run three tasks in sequence, passing each result into the next. Which canvas primitive fits?",
    "choices": [
      "group, because it runs tasks together",
      "chord, because it adds a callback",
      "chain, because it runs tasks in order and passes results along",
      "AsyncResult, because it tracks status"
    ],
    "answer": 2,
    "explain": "chain runs tasks in sequence and feeds each result into the next. group runs tasks in parallel, and chord is a group plus a callback that fires when the whole group finishes."
  },
  {
    "q": "Your app has modest background needs and already runs on Redis. When is it reasonable NOT to reach for Celery?",
    "choices": [
      "Never - Celery is the standard, so it's always the right choice",
      "When you want something simpler; a lighter tool like RQ or Dramatiq may fit better",
      "Only if you stop using a broker entirely",
      "Only when you have no scheduled jobs at all, ever"
    ],
    "answer": 1,
    "explain": "Celery is powerful but heavier to configure. Reach for it when you need its features (scheduling, routing, retries, canvas); reach for RQ or Dramatiq when you want simpler. Picking the smaller tool is a legitimate choice."
  }
]
```
