# Tokio: The Async Runtime

> Learn the runtime every async Rust web framework sits on: why Rust async needs a runtime at all, futures and how await works, spawning tasks, the multi-threaded work-stealing scheduler and blocking, channels and synchronization, and select with timeouts. The engine under axum, actix, and Rocket - made visible.


---

# Tokio: The Async Runtime

Here's something that surprises people coming to async Rust: the language gives you `async`/`await` and the
`Future` trait, but it does **not** ship anything to actually *run* those futures. A `Future` in Rust is
inert - it does nothing until something polls it to completion. That something is a **runtime**, and in
practice that runtime is **Tokio**. Every async web framework you'll meet - [axum](/guides/axum-from-zero),
[actix-web](/guides/actix-web-from-zero), [Rocket](/guides/rocket-from-zero) - runs on it. This is the
**roots** guide: learn Tokio and the `#[tokio::main]` at the top of every Rust server stops being a magic
incantation.

The mental model is a producer and an engine. Your `async fn`s produce **futures** - lazy descriptions of
work that yields at each `.await`. Tokio is the **engine** that drives them: it has a **scheduler** that
runs many futures concurrently across a small pool of OS threads, **spawns** independent work as tasks,
and wakes a future when the thing it was waiting on (a socket, a timer, a channel) is ready. Hold "futures
are inert plans; the runtime is what executes them," and async Rust clicks.

> 📝 This is a **roots** guide - it assumes you know **Rust**, including the basics of `async`/`await` and
> the `Future` trait ([Rust From Zero](/guides/rust-from-zero)). It pairs with
> [hyper & tower](/guides/hyper-and-tower) (the HTTP layer above Tokio) and underpins every Rust framework
> guide. Examples run as plain Rust programs (`cargo run`), shown with their output.

## How to read this

Short and foundational - read in order. It builds from "why a runtime exists" up to spawning, scheduling,
channels, and `select!`. Phases carry difficulty badges.

## The phases

1. **[What Tokio Is & Why Futures Need a Runtime](01-what-tokio-is.md)** 🟢 - inert futures, the runtime that drives them, and `#[tokio::main]`.
2. **[Async, Await & Futures](02-async-await-futures.md)** 🟡 - the `Future` trait, what `.await` does, and cooperative yielding.
3. **[Tasks & Spawning](03-tasks-and-spawning.md)** 🟡 - `tokio::spawn`, tasks vs threads, and `JoinHandle`.
4. **[The Runtime & Scheduler](04-runtime-and-scheduler.md)** 🔴 - the multi-threaded work-stealing scheduler, blocking the executor, and `spawn_blocking`.
5. **[Channels & Synchronization](05-channels-and-sync.md)** 🔴 - `mpsc`/`oneshot`/`broadcast`, and async-aware `Mutex` vs `std`.
6. **[select! & Timeouts](06-select-and-timeouts.md)** 🔴 - racing futures with `tokio::select!`, timeouts, and cancellation.
7. **[Where to Go Next](07-where-to-go-next.md)** 🟢 - how axum/actix use Tokio, and the async ecosystem.

> The throughline: **`async fn`s make inert futures; Tokio is the engine that polls, schedules, and wakes
> them.** Every Rust web server is a pile of futures running on this runtime.


---

# What Tokio Is & Why Futures Need a Runtime

Here's the thing nobody warns you about when you start writing async Rust: the language hands you `async`,
`await`, and the `Future` trait - and then ships **nothing** to actually run any of it. You can write a
perfect `async fn`, call it, and have *zero lines of its body execute*. The compiler will even shrug and
warn you about an "unused future." Coming from almost any other language, that feels broken. It isn't - once
it clicks, the `#[tokio::main]` sitting at the top of every Rust web server
([axum](/guides/axum-from-zero), actix, Rocket) stops looking like a magic incantation and starts looking
like the plug it actually is.

This is the **roots** guide. The same way [WSGI/ASGI](/guides/wsgi-and-asgi-explained) is the foundation
every Python web app quietly stands on, **Tokio is the foundation every async Rust program stands on**.
Learn this one layer and a whole stack of mysteries above it dissolves.

## The mental model: producers and an engine

Hold this picture and everything in this guide follows from it:

> 📝 **`async fn`s produce inert futures. A runtime (Tokio) is the engine that polls those futures to
> completion.** The function call builds a *plan* for some work; it does not *do* the work. Something else
> has to pick up that plan and drive it. That something is the runtime.

A `Future` in Rust is a value - a little state machine describing "here's the work, and here's how to make
progress on it one step at a time." Calling an `async fn` returns one of these values. That's *all* it does.
Nobody is running it yet. It just sits there, fully formed and completely idle, waiting to be **polled**.

The runtime is the part that polls. It takes your future, asks it "can you make progress?", and keeps asking
- parking the future when it'd block on something slow (a socket, a timer) and waking it back up when that
thing is ready - until the future finally produces its value.

```mermaid
flowchart LR
  A["async fn call"] --> B["Future (inert plan)"]
  B --> C["Runtime polls it"]
  C -->|not ready: park| C
  C -->|ready: resume| C
  C --> D["value produced"]
```

*What just happened:* the diagram is the whole story. An `async fn` call gives you an inert future. Left
alone, it sits at step B forever. The runtime is what carries it from B to D - polling, parking when work
would block, resuming when it can. No runtime, no movement. The future is a plan; Tokio is the worker who
reads the plan and does it.

## Seeing it: `#[tokio::main]`

So how do you actually get a runtime? In a real program you'd build one by hand, but Tokio gives you a
one-line shortcut for the common case. First, pull in the dependency:

```bash
cargo add tokio --features full
```

Then the canonical async program:

```rust
#[tokio::main]
async fn main() {
    println!("hello from an async main");
    say_hi().await;
}

async fn say_hi() {
    println!("hi");
}
```

*What just happened:* `#[tokio::main]` is **sugar**, not magic. At compile time it rewrites your
`async fn main` into an ordinary, non-async `fn main` that does two things: builds a Tokio runtime, then
calls `block_on(...)` on the runtime, handing it the future your async `main` produced. `block_on` is the
bridge from the normal blocking world into async-land - it drives that one future to completion and blocks
the thread until it's done. Inside, `say_hi().await` produces a future *and* tells the runtime "drive this
to completion before moving on," so you see `hello...` then `hi`. Without the macro, `main` can't be `async`
at all (Rust's real entry point is a plain `fn`), which is exactly the gap Tokio is filling.

⚠️ **The gotcha that bites everyone once.** Drop the `.await` and write just `say_hi();`. It compiles. It
runs. And `"hi"` **never prints** - because `say_hi()` only built a future and then dropped it on the floor,
unpolled. The compiler flags this with `warning: unused implementer of Future that must be used` and the
note `futures do nothing unless you .await or poll them`. Read that warning literally: it is not a style
nag, it's the runtime telling you nothing ran. If async code mysteriously "does nothing," a missing
`.await` is the first suspect.

## What a runtime actually does

You'll go deep on the scheduler in Phase 4, but the high-level job of a runtime is worth naming now. A Tokio
runtime is two cooperating parts:

- **The executor** - the part that holds your futures and *polls* them, deciding which one gets to make
  progress next. It runs them concurrently, often across a small pool of OS threads.
- **The reactor** - the part wired into the operating system's I/O and timers. When a future says "I'm
  waiting on this socket / this 5-second timer," the executor **parks** it (stops polling, frees the thread
  for other work) and the reactor remembers to **wake** it the moment that resource is ready.

That park-and-wake dance is the entire point of async: one thread can babysit thousands of in-flight futures
because a future waiting on the network costs nothing while it waits - it's parked, not blocking a thread.

## If you know JavaScript's event loop

📝 If you've written async JavaScript, Tokio plays a role you already recognize: it's the engine that runs
your async code, the rough equivalent of the **event loop** that drives JS promises and `await`. The shape
is familiar - submit async work, let the engine interleave it, get woken when things resolve.

But there's one difference that explains the unused-future warning above. A JavaScript promise is **eager**:
the moment you create it, its work is already scheduled and running. A Rust future is **lazy and polled**:
creating it runs *nothing*, and it only advances when the runtime polls it. That's why a forgotten `.await`
in Rust is silent dead code, where a forgotten `await` in JS still runs the work (you just don't wait for
the result). Same overall job - drive async tasks - opposite default. For the deeper comparison of these
two models, see [/guides/async-await-and-the-event-loop](/guides/async-await-and-the-event-loop).

## Recap

1. Rust gives you `async`/`await` and the `Future` trait but **ships no runtime** - there's nothing built in
   to actually execute futures.
2. A `Future` is **inert**: calling an `async fn` returns a future that does *nothing* until something polls
   it to completion. A missing `.await` runs zero code (and triggers the "unused future" warning).
3. A **runtime** (executor + reactor) is the engine that drives futures: it polls them, parks them when
   they'd block on I/O or timers, and wakes them when ready. **Tokio** is the de-facto one.
4. `#[tokio::main]` is sugar that turns `async fn main` into a normal `fn main` which builds a runtime and
   `block_on`s your async main future.
5. Mental model to keep: **`async fn`s produce inert plans; Tokio is the engine that runs them.** Every Rust
   web server is a pile of futures on this runtime.
6. Like JS's event loop in role, but Rust futures are **lazy and polled**, not eager - which is exactly why
   forgetting `.await` silently runs nothing.

## Quick check

One quick check before we open up the `Future` trait itself in Phase 2:

```quiz
[
  {
    "q": "You call an async fn but forget to .await it (and don't spawn it). What happens?",
    "choices": [
      "None of its body runs - it just builds an inert future that gets dropped, and the compiler warns about an unused future",
      "It runs immediately to completion, you just can't read the return value",
      "It runs on a background thread automatically",
      "The program fails to compile"
    ],
    "answer": 0,
    "explain": "A Rust future is inert: calling the async fn only builds the future. Without .await (or spawning) nothing polls it, so no body runs - and rustc emits 'futures do nothing unless you .await or poll them'. It still compiles; it's just silent dead code."
  },
  {
    "q": "What does #[tokio::main] actually do to your async fn main?",
    "choices": [
      "Rewrites it into a normal non-async fn main that builds a Tokio runtime and block_on()s your async main future",
      "Marks main as a coroutine the OS schedules directly",
      "Spawns every statement in main on its own thread",
      "Replaces println! with an async logging system"
    ],
    "answer": 0,
    "explain": "It's compile-time sugar. Rust's real entry point must be a plain fn, so the macro generates one that constructs a runtime and calls block_on on the future your async main produces - bridging the blocking world into async-land."
  },
  {
    "q": "How does a Rust future differ from a JavaScript promise?",
    "choices": [
      "A Rust future is lazy and only advances when a runtime polls it; a JS promise is eager and starts running the moment it's created",
      "They're identical - both run eagerly on creation",
      "A Rust future runs eagerly, while a JS promise must be polled",
      "Neither does anything until awaited"
    ],
    "answer": 0,
    "explain": "Both fill the 'drive async work' role, but with opposite defaults. JS promises are eager (work is scheduled on creation), while Rust futures are inert until a runtime polls them - which is why a forgotten .await in Rust runs nothing at all."
  }
]
```


---

# Async, Await & Futures

**An `async fn` is a state machine that implements the `Future` trait, and `.await` is the spot where that machine is allowed to pause and hand control back to the runtime.** Phase 1 told you futures are inert plans and Tokio is the engine. Now we open the plan up and look at how it's built - because once you see the state machine, `.await` stops being magic and starts being a place where your function can fall asleep and get woken up later.

You'll rarely touch the machinery directly. The compiler writes it for you. But knowing it's there is the difference between "my async code mysteriously froze the whole server" and "ah, I blocked at a point that has no `.await`, so nothing else could run." That second person is the one you want to be.

## The `Future` trait, briefly

Every async value in Rust is a `Future`. The trait is small:

```rust
trait Future {
    type Output;
    fn poll(self: Pin<&mut Self>, cx: &mut Context) -> Poll<Self::Output>;
}
```

The runtime drives a future by calling `poll`. That call returns one of two things:

- `Poll::Pending` - "not done yet, don't bother me until I tell you I'm ready."
- `Poll::Ready(value)` - "finished, here's the `Output`."

*What just happened:* A future isn't a thread you start and forget. It's a thing the runtime **pokes** by calling `poll`. Each poke either makes progress and finishes (`Ready`) or stalls and says "later" (`Pending`). The runtime's whole job is deciding which future to poke next.

You almost never write `poll` by hand. The `async`/`await` syntax generates the entire `poll` implementation for you - the `Pin`, the `Context`, the state tracking, all of it. Seeing the trait once is enough; you can forget the signature and remember the two outcomes.

## What `.await` actually does

When you write `something.await`, the compiler turns it into roughly this loop, baked into the generated state machine: poll the inner future; if it's `Ready`, take the value and keep going; if it's `Pending`, **return `Pending` from the whole task** so the runtime can go run something else. Later, when the inner thing is ready, the runtime polls your task again, and execution resumes right after the `.await`.

That resume-where-you-left-off is exactly what a state machine gives you. Each `.await` is a numbered state. When the function pauses, the runtime remembers which state it's in and what local variables are still live, so it can pick up later without redoing earlier work.

```rust
async fn fetch_user(id: u64) -> String {
    format!("user-{id}")
}

async fn fetch_orders(user: &str) -> Vec<String> {
    vec![format!("{user}-order-1"), format!("{user}-order-2")]
}

async fn report(id: u64) -> usize {
    let user = fetch_user(id).await;     // state 0: pause here if not ready
    let orders = fetch_orders(&user).await; // state 1: pause here if not ready
    orders.len()
}
```

*What just happened:* `report` is one future, built by composing two smaller ones. At each `.await` it can yield: if `fetch_user` returns `Pending`, the whole `report` task yields and Tokio runs other tasks; when `fetch_user` is ready, `report` resumes and moves to the next `.await`. These two run **in sequence** - `fetch_orders` only starts after `fetch_user` finishes, because the second line needs `user`.

Notice that calling `fetch_user(id)` does **nothing** on its own - it just builds a future. The work happens at `.await`. This is the inertness from Phase 1, up close.

> 💡 Sequential awaits are not concurrency. The example above waits for one thing, then the next. If you wanted both to make progress at the same time, you'd combine them - with something like `tokio::join!` (run several futures concurrently in one task) or `tokio::spawn` (give each its own task). Those are [Phase 3: Tasks & Spawning](03-tasks-and-spawning.md). For now, hold the distinction: `.await` in a straight line is *waiting*, not *parallelism*.

## Composing futures

Because an `async fn` is just a future, futures nest naturally. `report` above is a single future assembled from `fetch_user` and `fetch_orders`. You can also build an anonymous future inline with an `async` block:

```rust
async fn greet(name: String) -> String {
    let fut = async move {
        let upper = name.to_uppercase();
        format!("HELLO, {upper}")
    };
    fut.await
}
```

*What just happened:* `async move { ... }` creates an unnamed future that captures `name` by value (`move`), the same way a closure does. It doesn't run when constructed - only when `.await`ed. This is the building block Tokio's `spawn` and `select!` lean on: chunks of async work you hand to the runtime as values.

## ⚠️ Cooperative scheduling: yield, or starve everyone

This is the part that bites people, so read it twice. Tokio's scheduling is **cooperative**: a task only ever yields control **at an `.await`**. There is no preemption, no timer that interrupts a task mid-computation. Between two `.await` points, your code runs straight through and nothing else on that thread gets a turn.

That's fine for I/O-bound work, where you're constantly awaiting sockets, timers, and channels. But it has a sharp edge:

```rust
async fn handler() {
    // ⚠️ No .await anywhere in this loop.
    let mut total: u64 = 0;
    for i in 0..5_000_000_000u64 {
        total = total.wrapping_add(i);
    }
    println!("{total}");
}
```

*What just happened:* This is a CPU-bound loop with zero `.await` points. Once a worker thread starts running it, that thread is **pinned** until the loop finishes - it never reaches a yield point, so it can't go run other tasks. Every other task waiting on that thread is frozen. You've starved the runtime. The same thing happens if you call a **blocking** API (a synchronous file read, `std::thread::sleep`, a blocking DB driver) inside an async fn: it parks the thread with no way to yield.

The fix is to keep heavy or blocking work off the async worker threads - Tokio gives you `spawn_blocking` for exactly this, which we cover in [Phase 4: The Runtime & Scheduler](04-runtime-and-scheduler.md). The mental model to lock in now: **async Rust is for I/O concurrency - overlapping lots of waiting - not for CPU parallelism.** If a task computes hard without awaiting, it's the wrong tool, and it'll take the whole runtime down with it.

> ⚠️ "It compiled and ran fine in a tiny example" is a trap here. A blocking call only reveals itself under load, when many tasks are competing for the same handful of worker threads and one of them refuses to yield.

## 📝 The Waker: how a future gets polled again

One loose end. When a future returns `Pending`, how does the runtime know *when* to poll it again? It doesn't poll in a busy loop - that would burn a CPU for nothing.

> 📝 Remember the `cx: &mut Context` argument on `poll`? It carries a **Waker**. Before a future returns `Pending`, it stashes that waker with whatever it's waiting on - a socket registered with the OS, a timer, a channel. When that thing becomes ready (the socket has data, the timer fires), it calls `.wake()`, which tells the runtime "this task is worth polling now." The runtime reschedules it, and `poll` runs again, this time likely returning `Ready`.

You don't write waker code when you use `async`/`await` - the leaf futures in Tokio (its socket types, timers, channels) handle registration for you. The intuition is all you need: **`Pending` isn't "try again immediately"; it's "park me, and I'll be woken when there's a reason to look again."** That's what makes a runtime able to juggle thousands of idle connections on a few threads without spinning.

## Recap

- An `async fn` compiles to a **state machine implementing `Future`**; you write `async`/`await`, the compiler writes the `poll`.
- `poll` returns `Poll::Pending` (not ready, wake me later) or `Poll::Ready(value)` (done). The runtime drives futures by polling them.
- `.await` is a **yield point**: on `Pending` the whole task hands control back to the runtime; on `Ready` it resumes right after the `.await`. Sequential awaits are waiting, not concurrency.
- Scheduling is **cooperative** - tasks only yield at `.await`. CPU-bound loops or blocking calls with no `.await` **starve the runtime** (use `spawn_blocking`, Phase 4). Async is for I/O concurrency, not CPU parallelism.
- A **Waker** (carried in `poll`'s `Context`) lets a parked future get rescheduled exactly when the thing it waited on becomes ready - no busy-polling.

## Quick check

```quiz
[
  {
    "q": "What does an async fn compile down to in Rust?",
    "choices": ["A new OS thread", "A state machine that implements the Future trait", "A closure that runs immediately", "A blocking function call"],
    "answer": 1,
    "explain": "The compiler turns an async fn into a state machine implementing Future - it generates the poll method, with each .await as a state where execution can pause and resume."
  },
  {
    "q": "When a future being .awaited returns Poll::Pending, what happens?",
    "choices": ["The thread sleeps for a fixed interval then retries", "The task yields control back to the runtime so other tasks can run", "The program panics", "The runtime polls it again in a tight busy loop"],
    "answer": 1,
    "explain": "Pending means the task yields to the runtime. It registers a Waker so it can be re-polled later when its dependency is ready - not busy-polled and not on a fixed timer."
  },
  {
    "q": "Why can a long CPU-bound loop with no .await inside an async fn break a Tokio app?",
    "choices": ["Tokio preempts it after a timeout, corrupting state", "Scheduling is cooperative, so with no .await the task never yields and starves other tasks on that thread", "async fns can't contain loops", "It uses too much memory for the state machine"],
    "answer": 1,
    "explain": "Tokio scheduling is cooperative: tasks only yield at .await points. A loop with no .await pins its worker thread and starves every other task on it. Offload such work with spawn_blocking (Phase 4)."
  }
]
```


---

# Tasks & Spawning

**A task is a future that the runtime schedules independently.** In [the last phase](02-async-await-futures.md) you saw that a future is an inert plan - it does nothing until something polls it, and `.await` drives one future to completion *from where you are*. A task is the next step up. When you hand a future to `tokio::spawn`, you're telling the runtime: "take this, run it on your own, alongside everything else." You get back a handle, and the work starts running concurrently right away - you don't have to be sitting there awaiting it.

That distinction is the whole phase. A bare future is potential energy. A spawned task is potential energy that the runtime has already plugged in and switched on.

## `tokio::spawn` and the `JoinHandle`

`tokio::spawn(async { ... })` takes a future, registers it as a task with the runtime's scheduler, and immediately returns a **`JoinHandle<T>`**. The task is now alive and being driven concurrently with your current code. The handle is your receipt: hold onto it, and later you can `.await` it to collect whatever the task produced.

```rust
async fn expensive_io() -> u64 {
    // pretend this talks to a database or a network service
    42
}

#[tokio::main]
async fn main() {
    let handle = tokio::spawn(async {
        expensive_io().await
    });

    // ... do other work here while the task runs ...

    let result: u64 = handle.await.unwrap();
    println!("task produced: {result}");
}
```

*What just happened:* `tokio::spawn` kicked off the `async` block as its own task - `expensive_io()` started running without us awaiting it inline. `handle` is a `JoinHandle<u64>`. When we eventually `handle.await`, we block *this* code until the task finishes and we get its output.

One detail that trips people up: awaiting a `JoinHandle` gives you a **`Result<T, JoinError>`**, not a bare `T`. That's why we wrote `.unwrap()`.

```rust
let handle = tokio::spawn(async {
    expensive_io().await
});

let outcome: Result<u64, tokio::task::JoinError> = handle.await;
match outcome {
    Ok(value)  => println!("got {value}"),
    Err(join_err) => println!("the task failed: {join_err}"),
}
```

*What just happened:* the `Result` wrapper exists because a task can fail independently of its return value - most commonly, it **panicked**. A panic inside a spawned task doesn't unwind your `main`; the runtime catches it and surfaces it through the `JoinHandle` as a `JoinError`. So `.await` on a handle is asking two questions at once: "did the task finish?" and "did it finish *cleanly*?"

> 📝 The `T` in `JoinHandle<T>` is whatever the spawned future returns. If the future returns `()`, you get a `JoinHandle<()>` - useful when you only care that the work ran, not what it produced.

## Tasks vs OS threads: why you can have thousands

If you've used `std::thread::spawn`, this all looks familiar - spawn work, get a handle, join it. So why not just use threads?

Because tasks are **green threads**: lightweight, runtime-managed units of work that are *radically* cheaper than OS threads. An OS thread carries its own stack (often a megabyte or more) and is scheduled by the kernel; spawning tens of thousands of them will exhaust memory and bury the scheduler. A Tokio task is, roughly, a heap-allocated future plus a little bookkeeping - small enough that spawning *hundreds of thousands* is routine.

The trick is **M:N scheduling**: Tokio multiplexes many (M) tasks onto a few (N) OS worker threads - typically one worker per CPU core. When a task hits an `.await` that isn't ready, it yields its worker thread back to the runtime, which immediately picks up another ready task on that same thread. The threads never sit idle waiting; they hop between tasks. (Exactly *how* the scheduler does this - work-stealing across workers - is [the next phase](04-runtime-and-scheduler.md).)

> 💡 This is the entire reason async shines for servers. A web server spawns **one task per incoming connection**. Ten thousand connected clients means ten thousand tasks - but they ride on maybe 8 OS threads, because at any instant most of those connections are *waiting* (for the next request byte, for a database reply) and parked at an `.await`, costing nothing. Try that with one OS thread per connection and you fall over at a fraction of the load.

## ⚠️ The `'static` and `Send` rules

Here's where the borrow checker enters and surprises people. A spawned task may be picked up by any worker thread, and it may outlive the function that spawned it. The runtime can't make guarantees about either, so it imposes two requirements on the future you hand to `tokio::spawn`: it must be **`'static`** (it can't borrow anything with a shorter lifetime) and **`Send`** (it can be safely moved between threads).

In practice that means **you cannot borrow local stack data into a spawn.** This will not compile:

```rust
async fn broken() {
    let name = String::from("ada");

    // ❌ does not compile: the task might outlive `name`
    let handle = tokio::spawn(async {
        println!("hello, {name}");
    });

    handle.await.unwrap();
}
```

*What just happened:* the `async` block tries to *borrow* `name`, which lives on `broken`'s stack. The compiler refuses, because once spawned, the task is independent - it could run after `broken` returns and `name` is gone. The lifetime can't be proven `'static`, so it's rejected.

The fix is to give the task **ownership** of what it needs. `move` the data in:

```rust
async fn works() {
    let name = String::from("ada");

    let handle = tokio::spawn(async move {
        // the task now OWNS `name`
        println!("hello, {name}");
    });

    handle.await.unwrap();
}
```

*What just happened:* `async move` moves `name` into the task, so the task owns it outright - no borrow, no lifetime worry, `'static` satisfied. The task can now safely run whenever and wherever the scheduler likes.

When several tasks need to *share* the same data (and you can't hand each one its own copy), reach for an **`Arc`** - an atomically reference-counted pointer. Clone the `Arc` and `move` a clone into each task:

```rust
use std::sync::Arc;

async fn shared() {
    let config = Arc::new(String::from("shared config"));

    let mut handles = Vec::new();
    for i in 0..3 {
        let config = Arc::clone(&config); // cheap: bumps a refcount
        handles.push(tokio::spawn(async move {
            println!("task {i} sees: {config}");
        }));
    }

    for h in handles {
        h.await.unwrap();
    }
}
```

*What just happened:* each task got its own `Arc` clone (a cheap pointer + refcount bump, not a deep copy of the string) and `move`d it in. Every task owns a handle to the *same* underlying data, satisfying both `'static` and `Send`. The original `config` and all clones keep the data alive until the last one drops. (If tasks need to *mutate* shared state, you'll wrap it further - `Arc<Mutex<T>>` - which is the [channels & synchronization](05-channels-and-sync.md) phase.)

## Concurrency without spawning: `join!` and `try_join!`

Spawning is the right tool when you want **independent** tasks - fire-and-forget background work, one task per connection, things that should run on their own schedule. But sometimes you just want to run a handful of futures **together, right here**, and wait for all of them. For that, spawning is overkill. Reach for `tokio::join!`.

```rust
async fn fetch_user() -> String { String::from("ada") }
async fn fetch_orders() -> u32 { 7 }

async fn dashboard() {
    // both futures make progress concurrently, on THIS task
    let (user, orders) = tokio::join!(fetch_user(), fetch_orders());
    println!("{user} has {orders} orders");
}
```

*What just happened:* `join!` polls both futures concurrently on the *current* task and returns a tuple of their results once both finish. There's no `tokio::spawn`, no `JoinHandle`, no new task. While `fetch_user()` is parked at an `.await`, `fetch_orders()` gets a turn - they interleave on one thread. Because nothing crosses a thread boundary, the futures **don't need to be `Send` or `'static`** - you can freely borrow local data into them.

When the futures can fail and you want to bail out the moment any one of them does, use `tokio::try_join!`:

```rust
async fn load_a() -> Result<u32, String> { Ok(1) }
async fn load_b() -> Result<u32, String> { Err("boom".into()) }

async fn load_all() -> Result<(), String> {
    let (a, b) = tokio::try_join!(load_a(), load_b())?;
    println!("{a} {b}");
    Ok(())
}
```

*What just happened:* `try_join!` runs both concurrently but **short-circuits** on the first `Err` - as soon as `load_b()` fails, `try_join!` returns that error and we never reach the `println!`. It's `join!` with early exit baked in, perfect for "do these together, but if any fails the whole batch fails."

So the rule of thumb:

- **`tokio::spawn`** - for *independent* work that should run on its own, possibly on another thread, possibly outliving the current scope. Requires `Send + 'static`.
- **`join!` / `try_join!`** - for "do these few things together, here, and wait for them." No spawn, no `Send` requirement, can borrow locals.

## Recap

- A **task** is a future the runtime schedules **independently**; `tokio::spawn(future)` launches one and returns a **`JoinHandle<T>`**, with the work starting concurrently right away.
- `.await`-ing a `JoinHandle` yields a **`Result<T, JoinError>`** - the `Err` arm fires if the task panicked, since panics are caught and surfaced through the handle rather than crashing your code.
- Tasks are **green threads**: cheap enough to spawn by the thousands, multiplexed M:N onto a few OS worker threads - which is exactly how a server runs one task per connection on a small thread pool.
- Spawned futures must be **`'static` + `Send`**, so you can't borrow local stack data into them - `move` owned data in, or share it via an `Arc`.
- For running a few futures **together on the same task**, use `tokio::join!` (or `try_join!` to short-circuit on the first error) - no spawn and no `Send`/`'static` requirement.

## Quick check

```quiz
[
  {
    "q": "What does tokio::spawn return, and what happens to the future you pass it?",
    "choices": [
      "It returns the future's output directly after running it to completion",
      "It returns a JoinHandle<T>, and the task starts running concurrently right away",
      "It returns a JoinHandle<T>, but the task won't run until you .await the handle",
      "It returns nothing; the future is dropped unless you store it"
    ],
    "answer": 1,
    "explain": "tokio::spawn schedules the future as an independent task that begins running immediately, and hands you a JoinHandle<T> to collect its result later."
  },
  {
    "q": "Why won't this compile: a spawned async block that borrows a local String from the enclosing function?",
    "choices": [
      "Spawned futures must be 'static + Send, so they can't borrow shorter-lived local stack data",
      "tokio::spawn only accepts functions, not async blocks",
      "String isn't allowed inside async blocks",
      "You must always return a value from a spawned task"
    ],
    "answer": 0,
    "explain": "A task may move between threads and outlive the spawning function, so it must be 'static and Send. Borrowing a local fails that; move the data in (async move) or share it via Arc."
  },
  {
    "q": "You want to run two futures concurrently and wait for both, while borrowing local data - no separate task needed. What fits best?",
    "choices": [
      "Two separate tokio::spawn calls, then await both handles",
      "tokio::join!(a, b)",
      "Wrap both in Arc and spawn them",
      "Call .await on each one in sequence"
    ],
    "answer": 1,
    "explain": "tokio::join! drives both futures concurrently on the current task - no spawn, no JoinHandle, and because nothing crosses a thread boundary, no Send/'static requirement, so borrowing locals is fine."
  }
]
```


---

# The Runtime & Scheduler

> 💡 **Tokio runs thousands of tasks on a tiny pool of threads - usually about one thread per CPU core.** Many tasks, few threads. The scheduler keeps those few threads busy by handing each one its own queue of ready tasks, and when a thread runs dry it *steals* work from a neighbor. That's the engine. Once you internalize "few threads, many tasks, balanced by stealing," the one rule that matters falls out on its own: **never let a single task hog a thread.**

You met tasks and `tokio::spawn` in [Tasks & Spawning](03-tasks-and-spawning.md). Now we look under the floorboards at the thing that actually *runs* them - and at the single mistake that brings more async Rust servers to their knees than any other.

## A picture of the engine

When you write the default `#[tokio::main]`, Tokio builds a **multi-threaded runtime**: a pool of **worker threads** (by default, as many as you have CPU cores). Each worker owns a **local queue** of tasks that are ready to make progress. The worker pulls a task off its queue, polls it until it hits an `.await` that isn't ready, sets it aside, and grabs the next one. That's how a handful of threads juggle thousands of tasks - each task only occupies a thread for the brief moment it's actually doing work.

The clever part is **work-stealing**. If one worker empties its queue while another is buried, the idle worker reaches over and steals half the busy worker's tasks. No central dispatcher, no thread sitting idle while another drowns - the load balances itself.

```mermaid
flowchart LR
  T1[task] --> Q1
  T2[task] --> Q1
  T3[task] --> Q2
  T4[task] --> Q2
  Q1[Worker 1 queue] --> W1[Worker 1]
  Q2[Worker 2 queue] --> W2[Worker 2]
  Q2 -. steals .-> W1
```

*What just happened:* Tasks land in per-worker queues; each worker chews through its own queue, and an idle Worker 1 steals from Worker 2's queue to even out the load. The arrows are tasks flowing to threads; the dotted arrow is the steal that keeps everyone busy.

## Runtime flavors

You don't always want a whole pool. Tokio gives you three ways to choose:

```rust
// 1. The default: multi-threaded, work-stealing, ~one worker per core.
#[tokio::main]
async fn main() {
    // ...
}

// 2. Single-threaded: everything runs on one thread. Great for tests,
//    CLIs, and light apps where a thread pool is overkill.
#[tokio::main(flavor = "current_thread")]
async fn main() {
    // ...
}
```

*What just happened:* The first form is the multi-threaded runtime you get for free. The second, `current_thread`, runs every task on the single thread that called `main` - no work-stealing, no cross-thread coordination, and your tasks don't need to be `Send`. It's lighter and easier to reason about; it just can't use more than one core.

When you need to set the knobs yourself, build the runtime explicitly with `Builder`:

```rust
fn main() {
    let runtime = tokio::runtime::Builder::new_multi_thread()
        .worker_threads(4)        // pin the pool to 4 workers
        .enable_all()             // turn on the I/O and timer drivers
        .build()
        .unwrap();

    runtime.block_on(async {
        // your async program runs here
    });
}
```

*What just happened:* `Builder` is what `#[tokio::main]` expands into under the hood. `new_multi_thread()` picks the work-stealing scheduler, `worker_threads(4)` overrides the default (instead of "one per core"), `enable_all()` switches on the I/O and time drivers, and `block_on` hands an async block to the runtime and blocks the calling thread until it finishes. Reach for this when the macro's defaults don't fit - separate runtimes for different subsystems, a fixed thread count, custom thread names.

## ⚠️ The cardinal sin: blocking a worker

Here is the mistake. Remember the model: a worker thread polls one task, and it can only move to the *next* task when the current one yields at an `.await`. So what happens if a task **never yields** - if it sits in a tight CPU loop, or calls a synchronous blocking function that parks the thread?

The worker is stuck. It can't poll any of the other tasks sitting in its queue. They're all frozen behind the one greedy task. On a 4-core machine you've just lost a quarter of your entire capacity to a single task. Do it on enough tasks and the whole server seizes up.

> ⚠️ **Blocking the executor is the one mistake that turns a fast async server into a frozen one.** The symptoms are nasty because they're *spooky*: latency spikes out of nowhere, requests that should take 2ms taking 2000ms, "the server just froze for a few seconds." It rarely looks like the line of code that caused it.

Here's the trap in its most innocent form:

```rust
#[tokio::main]
async fn main() {
    // Spawn 100 tasks. Looks harmless. It is not.
    for id in 0..100 {
        tokio::spawn(async move {
            // ⚠️ THE SIN: std::thread::sleep blocks the OS thread.
            // This task does NOT yield - it parks the whole worker.
            std::thread::sleep(std::time::Duration::from_secs(1));
            println!("task {id} done");
        });
    }
    tokio::time::sleep(std::time::Duration::from_secs(5)).await;
}
```

*What just happened:* `std::thread::sleep` puts the *operating-system thread* to sleep - the worker, not the task. While that worker naps, every other task in its queue is stranded. With only a few workers, those 100 tasks crawl through in batches instead of all finishing in roughly one second. The function name says "sleep," but to the runtime it reads as "freeze a worker for one full second." Anything synchronous and slow does the same: a blocking database driver, `std::fs` file reads, a JSON parse over a 50MB payload, a hash over a huge buffer.

## The fixes

Once you see the disease, the cures are mechanical. There are two questions: *is it a timer, or is it real work?*

### If you're sleeping or timing, use Tokio's async sleep

```rust
#[tokio::main]
async fn main() {
    for id in 0..100 {
        tokio::spawn(async move {
            // ✅ tokio::time::sleep yields the worker back to the runtime.
            tokio::time::sleep(std::time::Duration::from_secs(1)).await;
            println!("task {id} done");
        });
    }
    tokio::time::sleep(std::time::Duration::from_secs(2)).await;
}
```

*What just happened:* `tokio::time::sleep(...).await` doesn't park the thread - it tells the runtime "wake me in a second" and **yields**, freeing the worker to poll the other 99 tasks immediately. Now all 100 genuinely run concurrently and finish in about a second. The difference between this and the broken version is a single `.await` on the right `sleep`, and it's the difference between a healthy server and a frozen one.

### If it's unavoidable blocking or heavy CPU work, use `spawn_blocking`

Sometimes you *can't* make the call async - a legacy synchronous library, a blocking driver, or genuinely CPU-bound work like resizing an image. Don't run it on an async worker. Hand it to `spawn_blocking`, which runs your closure on a **separate pool of blocking threads** that exists precisely so async workers stay free:

```rust
#[tokio::main]
async fn main() {
    let result = tokio::task::spawn_blocking(|| {
        // Heavy, synchronous, no .await anywhere - and that's fine here.
        // This runs on the blocking pool, not on an async worker.
        expensive_hash_computation()
    })
    .await
    .unwrap(); // unwrap the JoinError; result is the closure's return value

    println!("hash = {result}");
}

fn expensive_hash_computation() -> u64 {
    (0..50_000_000u64).fold(0, |acc, n| acc.wrapping_add(n.wrapping_mul(31)))
}
```

*What just happened:* `spawn_blocking` moves the closure onto Tokio's dedicated blocking thread pool, so the synchronous CPU grind never touches an async worker - the workers keep serving other tasks the whole time. It hands back a `JoinHandle` just like `tokio::spawn`, so you `.await` it to get the result; the outer `.unwrap()` handles the `JoinError` (if the closure panicked), and the inner value is whatever the closure returned.

> 📝 For **CPU-bound parallelism** - processing thousands of images, crunching a big dataset across all cores - `spawn_blocking` works, but a dedicated thread pool like [rayon](https://docs.rs/rayon) is often the better tool. Rayon is built for data parallelism (`par_iter` and friends) and won't compete with Tokio's I/O threads. Rule of thumb: **async is for waiting; threads (spawn_blocking or rayon) are for computing.**

### A brief note on `block_in_place`

> 📝 There's a third option, `tokio::task::block_in_place`, for when you must block *inside* a multi-threaded worker without shipping the work elsewhere. It tells the runtime "I'm about to block this worker - go move my sibling tasks to other threads first," so they aren't stranded. It's a sharper, more situational tool than `spawn_blocking` (it only works on the multi-threaded runtime, not `current_thread`), so reach for `spawn_blocking` first and keep `block_in_place` in your back pocket for the cases where moving the work out isn't practical.

## Recap

- The default `#[tokio::main]` builds a **multi-threaded, work-stealing runtime**: a pool of worker threads (≈ one per CPU core), each with a local task queue, with idle workers stealing from busy ones to balance load.
- Flavors: the multi-thread default, `current_thread` (everything on one thread - ideal for tests and light apps), or an explicit `tokio::runtime::Builder` when you need to set `worker_threads` and other knobs.
- **Blocking the executor is the cardinal sin:** a task that loops on CPU or calls a synchronous blocking function never yields, so it parks a whole worker and freezes every other task behind it - the source of mystery latency spikes and "the server froze."
- For sleeping and timers in async code, use `tokio::time::sleep(...).await` (it yields), never `std::thread::sleep` (it parks the thread).
- For unavoidable blocking or heavy CPU work, use `tokio::task::spawn_blocking`, which runs the closure on a **separate blocking thread pool** so async workers stay free; reach for rayon for CPU-bound parallelism.
- `block_in_place` exists for blocking inside a multi-thread worker without moving the work, but it's a niche tool - prefer `spawn_blocking`.

## Quick check

```quiz
[
  {
    "q": "On the default multi-threaded runtime, what does the work-stealing scheduler do when one worker thread empties its task queue?",
    "choices": ["It shuts that worker down to save memory", "It steals ready tasks from a busier worker's queue", "It blocks until the main thread assigns it more work", "It spawns a brand-new OS thread for each idle moment"],
    "answer": 1,
    "explain": "Each worker has a local queue; an idle worker steals tasks from a busy one, so load balances itself without a central dispatcher."
  },
  {
    "q": "Why is calling std::thread::sleep inside a Tokio task a problem?",
    "choices": ["It returns an error at compile time", "It parks the whole worker thread, so other tasks on that worker can't make progress", "It silently does nothing in async code", "It uses more memory than tokio::time::sleep"],
    "answer": 1,
    "explain": "std::thread::sleep parks the OS thread (the worker), not just the task. Every other task queued on that worker is frozen until it returns. Use tokio::time::sleep(...).await, which yields."
  },
  {
    "q": "You must run a synchronous, CPU-heavy library call from async code. What's the right tool?",
    "choices": ["Wrap it in an async block and .await it", "Call it directly inside tokio::spawn", "Run it inside tokio::task::spawn_blocking", "Replace it with tokio::time::sleep"],
    "answer": 2,
    "explain": "spawn_blocking runs the closure on Tokio's separate blocking thread pool, so the async worker threads stay free to keep polling other tasks."
  }
]
```


---

# Channels & Synchronization

You've got tasks now - independent units of work the scheduler runs concurrently ([Tasks & Spawning](03-tasks-and-spawning.md)). But a pile of isolated tasks isn't a program. The moment two tasks need to agree on something - "here's the next job," "I'm done, here's the result," "the config just changed" - they have to **coordinate**. This phase is about how.

There are exactly **two ways** for tasks to coordinate:

- **Pass messages** - one task hands data to another through a **channel**. The data *moves*; only one task owns it at a time.
- **Share state** - multiple tasks touch the same value behind a **lock** (a `Mutex`, `RwLock`, etc.). The data *stays put*; tasks take turns reaching in.

📝 Both are legitimate, but they pull in different directions. Channels keep ownership clean - at any instant exactly one task holds the data, so whole classes of "two tasks stomped on each other" bugs cannot happen. Locks are sometimes the natural fit (a shared counter, a cache), but they invite contention and deadlocks if you're careless. The Rust community's rule of thumb, borrowed from Go: **"share state by communicating, rather than communicate by sharing state."** When in doubt, reach for a channel first.

Let's build up the toolbox.

## mpsc: many senders, one receiver

`mpsc` stands for **multi-producer, single-consumer**. Many tasks can send into it; exactly one task pulls values out. This is the workhorse - a job queue, a request funnel, an event pipeline. It's the channel you'll reach for most.

```rust
use tokio::sync::mpsc;

#[tokio::main]
async fn main() {
    let (tx, mut rx) = mpsc::channel::<i32>(32);

    tokio::spawn(async move {
        for i in 0..5 {
            tx.send(i).await.unwrap();
        }
        // tx is dropped here when the task ends
    });

    while let Some(n) = rx.recv().await {
        println!("got {n}");
    }
    println!("channel closed");
}
```

*What just happened:* `mpsc::channel::<i32>(32)` created a **bounded** channel that holds up to 32 queued values, and handed back a sender `tx` and a receiver `rx` (note `mut rx` - receiving mutates it). We `move`d `tx` into a spawned task that sends `0..5`, each `send(i).await` parking briefly if the buffer were full. The main task loops with `while let Some(n) = rx.recv().await`, printing each value. When the spawned task ends, `tx` drops; with no senders left, `recv()` returns `None`, the loop ends, and we print "channel closed". Output: `got 0` through `got 4`, then `channel closed`.

Two details that trip people up:

**Bounded vs unbounded.** `channel(32)` is bounded - and that bound is a feature. When the buffer fills, `send().await` *suspends the sender* until the consumer drains some. That's **backpressure**: a slow consumer naturally slows fast producers instead of letting an unbounded queue eat all your memory. There's also `mpsc::unbounded_channel()`, whose `send()` never waits and returns immediately - convenient, but it gives you no backpressure, so a producer outrunning a consumer grows the queue without limit. ⚠️ Prefer bounded channels unless you have a specific reason not to; an unbounded channel is an out-of-memory crash waiting for a bad day.

**Multiple producers.** To get the "multi" in multi-producer, **clone the sender**. Each clone feeds the same receiver:

```rust
let (tx, mut rx) = mpsc::channel::<String>(16);

for id in 0..3 {
    let tx = tx.clone();
    tokio::spawn(async move {
        tx.send(format!("hello from worker {id}")).await.unwrap();
    });
}
drop(tx); // drop the original so recv() can eventually see None

while let Some(msg) = rx.recv().await {
    println!("{msg}");
}
```

*What just happened:* We cloned `tx` once per worker task; all three clones point at the same single receiver. Each worker sends one message. The crucial line is `drop(tx)` - `recv()` only returns `None` when **every** sender is gone, so if we kept the original `tx` alive in `main`, the loop would hang forever waiting for a message that never comes. Dropping it explicitly lets the loop end once all workers finish. The three messages print in whatever order the scheduler runs the workers.

💡 That "the loop hangs because a sender is still alive" gotcha is the single most common mpsc mistake. If your `recv()` loop never exits, count your senders.

## oneshot: a single value, once

Sometimes a task doesn't need a stream - it needs to hand back **one** result. "Go do this work and tell me the answer." That's `oneshot`: a channel that carries exactly one value, then closes.

```rust
use tokio::sync::oneshot;

#[tokio::main]
async fn main() {
    let (tx, rx) = oneshot::channel::<u64>();

    tokio::spawn(async move {
        let answer = expensive_computation().await;
        let _ = tx.send(answer); // send takes self - usable once
    });

    match rx.await {
        Ok(value) => println!("the answer is {value}"),
        Err(_) => println!("the worker dropped the sender without sending"),
    }
}

async fn expensive_computation() -> u64 { 42 }
```

*What just happened:* `oneshot::channel()` gave us a sender and receiver for a single `u64`. The spawned task computes a result and calls `tx.send(answer)` - note `oneshot`'s `send` takes `self` and isn't `async`, because it can only ever fire once. The main task does `rx.await` (the receiver itself is a future) and gets `Ok(value)` with the result, or `Err` if the sender was dropped before sending. This is the canonical "task returns a value to its caller" pattern, and it's what request/response plumbing inside servers is built on: send a request *plus a oneshot sender* down an mpsc, and the handler replies on the oneshot.

Two more channel shapes round out the set, each for a different fan-out shape:

- **`broadcast`** - multi-producer, **multi-consumer**, where *every* receiver gets *every* message. Think a chat room fanning one message out to all connected clients, or a shutdown signal every task needs to hear. (Slow receivers can lag and miss messages - `recv()` tells you when that happens.)
- **`watch`** - receivers only care about the **latest** value, not the history. Perfect for config reloads or a shared status: a writer updates the value, and each reader sees the current one whenever it checks. Old values are overwritten.

## Locks: when you do share state

Channels cover most coordination, but sometimes shared mutable state is genuinely the cleaner model - a request counter, an in-memory cache, a connection pool. For that you reach for a lock. And here Tokio gives you a choice that confuses nearly everyone the first time: there are **two** `Mutex` types, and picking the wrong one either won't compile or quietly invites a deadlock.

⚠️ **The rule:** never hold a `std::sync::Mutex` guard across an `.await`.

Here's why. `std::sync::Mutex` is a normal, blocking mutex - fast, and the right default for most code. But its guard (the thing `lock()` returns) is **not `Send`**, and `tokio::spawn` requires futures to be `Send`. So if a guard is still alive at an `.await` point, the borrow checker refuses to compile your spawned task. Even setting `Send` aside, holding a *blocking* lock across an await is dangerous: the task can be suspended while still holding the lock, and another task on the same thread that wants the lock blocks the whole OS thread - a recipe for deadlock.

So the two cases:

```rust
use std::sync::Mutex; // the fast, blocking one

#[tokio::main]
async fn main() {
    let counter = Mutex::new(0u64);

    {
        let mut guard = counter.lock().unwrap();
        *guard += 1;
    } // guard dropped HERE, before any await - perfect

    some_async_thing().await;
    println!("count = {}", *counter.lock().unwrap());
}

async fn some_async_thing() {}
```

*What just happened:* We used `std::sync::Mutex` for a tiny, synchronous critical section: lock, bump the counter, unlock. The guard lives inside its own `{ }` block and is dropped at the closing brace - *before* the `.await` on the next line. No await happens while the lock is held, so `Send` is never an issue and there's no deadlock risk. This is the common case, and `std::sync::Mutex` is faster than the async one, so this is what you should use the vast majority of the time.

The other type, `tokio::sync::Mutex`, is **async-aware**: its `lock()` is itself `.await`-able, and its guard *is* allowed to live across `.await` points. Use it only when you genuinely must hold the lock while awaiting something:

```rust
use tokio::sync::Mutex; // the async-aware one
use std::sync::Arc;

async fn run(shared: Arc<Mutex<Vec<String>>>) {
    let mut guard = shared.lock().await;      // .await to acquire
    guard.push(fetch_line().await);           // holding the lock across an await - allowed
    println!("now {} lines", guard.len());
}

async fn fetch_line() -> String { "data".to_string() }
```

*What just happened:* `shared.lock().await` asynchronously waits its turn for the lock - if another task holds it, this task yields instead of blocking the thread. Because it's `tokio::sync::Mutex`, we can keep `guard` alive across the `await` on `fetch_line()` and the compiler is happy. The cost: it's slower than `std::sync::Mutex`, which is exactly why you reserve it for the case where the standard one literally won't work.

💡 Decision in one line: **`std::sync::Mutex` by default; `tokio::sync::Mutex` only when the lock must survive an `.await`.**

Tokio's `sync` module has a few more tools worth knowing by name:

- **`RwLock`** - many concurrent readers *or* one writer. Use it when reads vastly outnumber writes (a rarely-changing config read on every request).
- **`Semaphore`** - hands out a fixed number of permits to cap concurrency. Want at most 10 simultaneous outbound requests? Acquire a permit before each, with only 10 to go around.
- **`Notify`** - a lightweight "wake me when something happens" primitive, for hand-rolled coordination without passing a value.

## Bringing it back to the mental model

Step back and the whole phase collapses to one choice. Two tasks need to coordinate - do you **move the data between them** (a channel) or **let them share it** (a lock)?

Reach for channels first. They keep ownership unambiguous, they make backpressure explicit, and they sidestep deadlocks by construction. mpsc for streams of work, oneshot for a single reply, broadcast to tell everyone, watch for the latest value. When shared state really is the cleaner design, lock it - and remember the one rule that keeps you out of trouble: `std::sync::Mutex` for quick non-await sections, `tokio::sync::Mutex` only when the guard must cross an `.await`.

Next we'll let a task wait on *several* of these at once and give up after a deadline.

## Recap

- Tasks coordinate two ways: **message passing** (channels) or **shared state** (locks). Prefer channels - "share state by communicating."
- **`mpsc`** is multi-producer, single-consumer: clone `tx` for many senders, `rx.recv().await` yields `Option<T>` and returns `None` once all senders drop. **Bounded** channels apply backpressure; unbounded ones risk unbounded memory.
- **`oneshot`** carries a single value once - the natural way for a task to return a result. **`broadcast`** fans every message out to every receiver; **`watch`** exposes only the latest value.
- ⚠️ Never hold a **`std::sync::Mutex`** guard across an `.await` - its guard isn't `Send` (won't compile in a spawned task) and risks deadlock. Use **`tokio::sync::Mutex`** only when you must lock across an await.
- 💡 `std::sync::Mutex` is the faster default for short, non-await critical sections. `RwLock` caps reads vs writes, `Semaphore` caps concurrency, `Notify` wakes tasks.

## Quick check

```quiz
[
  {
    "q": "Your mpsc `while let Some(n) = rx.recv().await` loop never exits, even after all worker tasks finish. What's the most likely cause?",
    "choices": ["The channel buffer is too small", "A clone of the sender (often the original tx) is still alive, so recv() never sees None", "You used a bounded channel instead of unbounded", "oneshot should have been used instead"],
    "answer": 1,
    "explain": "recv() returns None only when EVERY sender is dropped. A lingering tx clone keeps the loop waiting forever; drop the extra sender."
  },
  {
    "q": "When is it safe to hold a `std::sync::Mutex` guard?",
    "choices": ["Any time, std Mutex is always fine in async code", "Only across an .await point", "For short, synchronous critical sections that do NOT span an .await", "Only inside tokio::spawn"],
    "answer": 2,
    "explain": "std::sync::Mutex is the fast default, but its guard isn't Send and blocks the thread - keep the locked section synchronous and drop the guard before any .await."
  },
  {
    "q": "You need to push config updates so each reader always sees the current settings, never the history. Which channel fits best?",
    "choices": ["mpsc", "oneshot", "broadcast", "watch"],
    "answer": 3,
    "explain": "watch keeps only the latest value; new writes overwrite old ones and each receiver reads the current value when it checks - exactly the config/state-update pattern."
  }
]
```


---

# select! & Timeouts

> 💡 **`tokio::select!` races several futures at once. The *first* one to finish wins - its branch runs.
> The moment it wins, every other future is **dropped right where it stood**. Dropped means cancelled.**

That single fact is the source of both the power and the pain in this phase. The power: you can say "do
this work, but give up if it takes longer than 5 seconds" or "serve requests until someone hits Ctrl-C"
in a few lines. The pain: a future you cancelled might have been *halfway through something* when you
yanked it out from under itself. We'll build up the happy cases first, then spend real time on the part
that bites people - cancellation safety.

## Racing two futures with `select!`

The classic shape is "real work versus a timer." Whichever finishes first decides what happens:

```rust
use tokio::time::{sleep, Duration};

async fn do_work() -> u32 {
    sleep(Duration::from_secs(2)).await; // pretend this is a network call
    42
}

#[tokio::main]
async fn main() {
    tokio::select! {
        res = do_work() => println!("work finished: {res:?}"),
        _ = sleep(Duration::from_secs(5)) => println!("timed out waiting"),
    }
}
```

*What just happened:* `select!` started polling **both** branches concurrently. `do_work()` finishes
after 2 seconds, the 5-second `sleep` is still pending - so the first branch wins, prints
`work finished: 42`, and the sleep future is **dropped without ever completing**. Flip the numbers
(work takes 6 seconds, timer is 5) and the other branch wins instead, printing `timed out waiting` while
`do_work()` gets cancelled mid-flight. Each branch is `pattern = future => body`; the winner's body runs
with its result bound to the pattern, and `select!` evaluates to that body's value.

📝 `select!` polls branches in (effectively) random order when several are ready at once, so you can't
rely on a fixed priority. If you need deterministic priority, add a `biased;` line as the first thing
inside the macro and it'll poll top-to-bottom.

## The careful part: cancellation safety

This is the bit that separates "I used `select!`" from "I understand `select!`." When a losing branch is
dropped, any *partial work that future was holding* goes away with it. If that future had read half a
message off a socket, that half-message is gone - it was living in the future's stack, and the future no
longer exists.

⚠️ **Not every future is safe to cancel mid-poll.** A future is **cancellation-safe** if dropping it
before completion loses *no* committed progress - you can re-create it next loop iteration and nothing
was silently swallowed. Many Tokio ops document this explicitly. Safe examples: `tokio::time::sleep`,
`mpsc::Receiver::recv`, `TcpListener::accept`. Unsafe examples: things that consume input incrementally,
like `AsyncReadExt::read` into a buffer where a partial read can't be un-read.

The danger almost always shows up in a **loop** that re-evaluates the same `select!` over and over:

```rust
use tokio::sync::mpsc;

async fn run(mut rx: mpsc::Receiver<String>, mut shutdown: mpsc::Receiver<()>) {
    loop {
        tokio::select! {
            Some(msg) = rx.recv() => {
                println!("got: {msg}");
            }
            _ = shutdown.recv() => {
                println!("shutting down");
                break;
            }
        }
    }
}
```

*What just happened:* this loop is safe **because both branches are cancellation-safe**. On each
iteration `select!` creates fresh `recv()` futures. Say a message arrives and the first branch wins - the
`shutdown.recv()` future is dropped. That's fine: `recv()` holds no partial state, so dropping it loses
nothing, and next iteration we just call `recv()` again. The whole pattern only works because cancelling
the loser is harmless.

Now picture the unsafe version: a branch that does `socket.read(&mut buf).await` inside this same loop.
If the *other* branch fires while `read` has pulled 300 of an expected 500 bytes into `buf`, that read
future is dropped and those 300 bytes are stranded - buffered in the future's own state, not yet returned
to you. Next iteration starts a brand-new `read` and your protocol is now corrupt. Two standard fixes:

- **Use a cancellation-safe API** instead (e.g. read full framed messages with `tokio_util`'s
  `FramedRead`, whose `next()` *is* cancellation-safe).
- **Move the fragile work out of `select!`** - do the multi-step read in its own task or its own future
  that runs to completion, and only `select!` on a channel/handle that delivers the *finished* result.

The rule of thumb: **before you put a future in a `select!` loop, ask "is it safe to drop this halfway?"**
If you're not sure, check its docs for the words "cancel safety" - Tokio is good about labeling this.

## `timeout`: the common case, done cleanly

Racing your work against a `sleep` works, but for the everyday "bound this one operation" need there's a
purpose-built helper that reads better and can't get the branches wrong:

```rust
use tokio::time::{timeout, sleep, Duration};

async fn fetch() -> String {
    sleep(Duration::from_secs(3)).await;
    "data".to_string()
}

#[tokio::main]
async fn main() {
    match timeout(Duration::from_secs(2), fetch()).await {
        Ok(value) => println!("got {value} in time"),
        Err(_elapsed) => println!("gave up - fetch was too slow"),
    }
}
```

*What just happened:* `timeout(duration, future)` returns a `Result<T, Elapsed>`. Here `fetch()` needs 3
seconds but we allowed 2, so the timer expires first, `fetch()` is **cancelled**, and we get
`Err(Elapsed)`. If `fetch()` had finished in under 2 seconds we'd get `Ok("data")`. Note the cancellation
caveat from above still applies - `timeout` drops the inner future when time runs out, so the same
"is this safe to cancel?" question holds for whatever you wrap.

## `interval`: periodic ticks

For "do something every N seconds," reach for `interval` rather than sleeping in a loop (sleeping drifts;
`interval` accounts for how long the work took and keeps a steady cadence):

```rust
use tokio::time::{interval, Duration};

#[tokio::main]
async fn main() {
    let mut tick = interval(Duration::from_secs(1));
    let mut count = 0;
    loop {
        tick.tick().await; // fires immediately the first time, then every 1s
        count += 1;
        println!("tick {count}");
        if count == 3 { break; }
    }
}
```

*What just happened:* `tick.tick().await` resolves on schedule - note the **first** tick completes
immediately, then subsequent ones land one second apart. This is what you `select!` against when you want
"do periodic work *and* react to other events in the same loop."

## Graceful shutdown: the payoff pattern

Now combine everything. A long-running worker shouldn't loop forever - it should run until told to stop,
then exit *cleanly* (finish the current item, flush, close connections). The shape: `select!` between
"the actual work" and "a shutdown signal." A common signal is Ctrl-C via `tokio::signal::ctrl_c()`:

```rust
use tokio::time::{interval, Duration};

#[tokio::main]
async fn main() {
    let mut work = interval(Duration::from_secs(1));

    loop {
        tokio::select! {
            _ = work.tick() => {
                println!("doing a unit of work");
            }
            _ = tokio::signal::ctrl_c() => {
                println!("\nshutdown signal received - cleaning up");
                break;
            }
        }
    }

    println!("worker stopped cleanly");
}
```

*What just happened:* the loop races a periodic tick against the Ctrl-C future. While no signal arrives,
the tick branch keeps winning and the work runs. The instant you press Ctrl-C, that branch wins, we
`break` out of the loop, and control flows to the cleanup line *after* the loop - instead of the process
being killed mid-task. Both `tick()` and `ctrl_c()` are cancellation-safe, so the loop is sound.

For fanning shutdown out to **many** tasks, a single signal isn't enough - you want one notification that
every worker can listen to. Two good tools:

- A [`watch` channel](05-channels-and-sync.md): the supervisor sets a "should I stop?" value, every
  worker `select!`s on `shutdown.changed()`, and they all wake at once.
- `tokio_util`'s **`CancellationToken`**, which formalizes exactly this: hand each task a clone of the
  token, they `select!` on `token.cancelled()`, and calling `token.cancel()` once trips them all. It also
  supports child tokens for hierarchical shutdown. Reach for it when "stop everything gracefully" is a
  recurring need rather than a one-off.

## Recap

- **`tokio::select!` races futures; the first to finish wins and the rest are dropped (cancelled).** That
  cancellation is the whole point - and the whole hazard.
- **Cancellation safety** is the thing to watch in `select!` loops: a dropped loser can lose partial work.
  Prefer cancellation-safe ops (`recv`, `sleep`, `accept`, `FramedRead::next`) or move fragile multi-step
  work out of the `select!` so it runs to completion.
- **`tokio::time::timeout(dur, fut).await`** returns `Result<T, Elapsed>` - the clean way to bound a single
  operation, no manual sleep-racing required.
- **`tokio::time::interval`** gives you steady periodic ticks via `tick().await` (first tick is immediate);
  ideal to `select!` against alongside other events.
- **Graceful shutdown** is `select!`-ing work against a shutdown signal (`ctrl_c`, a `watch` channel, or
  `tokio_util`'s `CancellationToken`) so the worker exits cleanly instead of being killed mid-task.

## Quick check

```quiz
[
  {
    "q": "In tokio::select!, what happens to the futures in the branches that did NOT finish first?",
    "choices": ["They keep running in the background", "They are dropped (cancelled) where they stood", "They are paused and resumed on the next select!", "They each get their own spawned task"],
    "answer": 1,
    "explain": "select! runs the first branch to complete and drops every other future at that point - dropping a future cancels it."
  },
  {
    "q": "Why can a select! loop containing socket.read(&mut buf).await be unsafe?",
    "choices": ["read() can never be used with select!", "If another branch wins, the read future is dropped and any partially-read bytes are lost", "read() blocks the whole runtime", "select! refuses to compile with read()"],
    "answer": 1,
    "explain": "A partial read holds bytes in the future's own state. If the read loses the race it's dropped, stranding those bytes - that's a cancellation-safety bug."
  },
  {
    "q": "What does tokio::time::timeout(Duration::from_secs(2), some_future()).await return?",
    "choices": ["The future's value, or panics on timeout", "A Result<T, Elapsed> - Err(Elapsed) if it didn't finish in time", "An Option<T> that is None on timeout", "A bool indicating success"],
    "answer": 1,
    "explain": "timeout yields Result<T, Elapsed>: Ok(value) if the future finished in time, Err(Elapsed) if the timer expired first (the inner future is then cancelled)."
  }
]
```


---

# Where to Go Next

Look at how far you've come. Six phases ago, `#[tokio::main]` was a magic incantation you copied off a tutorial and hoped worked. Now you can name every piece under it.

You know **futures** are inert plans and `.await` is a cooperative yield point - the future hands control back to the runtime instead of blocking. You know **tasks** are spawned with `tokio::spawn`, that they're cheap green threads multiplexed onto a small pool of real ones, and that a `JoinHandle` lets you wait on the result. You know the **scheduler** is multi-threaded and work-stealing, that blocking it starves every other task, and that `spawn_blocking` is the escape hatch. You know **channels** - `mpsc`, `oneshot`, `broadcast` - and why an async-aware `Mutex` is held across `.await` where a `std` one must not be. And you know `select!` races futures, drops the losers, and gives you timeouts and cancellation almost for free.

That's the whole engine. The runtime under every async Rust program is no longer magic - it's machinery you can name.

This last phase doesn't add new mechanism. It points your new X-ray vision at the frameworks and libraries you'll actually reach for, and tells you what to build next.

## Your web framework is a pile of Tokio tasks

💡 Here's the realization that pays off the whole guide: **axum, actix-web, and Rocket aren't alternatives to Tokio - they're Tokio applications.**

When you write `#[tokio::main]` at the top of an axum server, that macro builds the runtime you just learned about. The server itself binds a `TcpListener`, then loops accepting connections - and for each connection it calls `tokio::spawn`. Every request your server handles is running inside a task, on the work-stealing scheduler, woken when its socket has bytes ready. The "framework" is a very nice set of conveniences - routing, extractors, middleware - layered over exactly the spawning and polling you already understand.

```mermaid
flowchart TD
  APP["Your handlers / routes"] --> FW["axum / actix / Rocket"]
  FW --> HY["hyper (HTTP) on Tokio"]
  HY --> TASKS["Tokio tasks - one per connection"]
  TASKS --> SCHED["work-stealing scheduler"]
  SCHED --> OS["small pool of OS threads"]
```

Read that top to bottom: your code sits on a framework, the framework sits on **hyper** (the HTTP implementation), hyper sits on Tokio tasks, the scheduler runs those tasks across a handful of real threads, and the OS handles the rest. Nothing in that stack is mysterious to you now.

See [axum from zero](/guides/axum-from-zero) for the framework most new Rust web work starts with, [actix-web from zero](/guides/actix-web-from-zero) for the high-throughput veteran, and [hyper & tower](/guides/hyper-and-tower) for the HTTP layer sitting directly on Tokio - the one place this picture gets concrete.

## The ecosystem worth knowing

Tokio is more than a scheduler. It ships the async building blocks you'll use constantly, and a galaxy of crates is built on top.

From Tokio itself:

- **`tokio::net`** - `TcpListener` and `TcpStream` for TCP, plus UDP sockets. This is where a server actually accepts connections.
- **`tokio::io`** - the `AsyncRead` and `AsyncWrite` traits, the async cousins of `std`'s `Read`/`Write`. Almost everything that moves bytes implements them.
- **`tokio::fs`** - file operations that don't block the scheduler (under the hood it uses `spawn_blocking`, because the OS has no truly async file API on most platforms).
- **`tokio::time`** - `sleep`, `interval`, and the `timeout` you met in [Phase 6](06-select-and-timeouts.md).

And the crates built on Tokio that you'll meet at work:

- **hyper** - the HTTP/1 and HTTP/2 implementation under most Rust web frameworks.
- **tonic** - gRPC, for service-to-service APIs.
- **reqwest** - the ergonomic HTTP *client*, for calling other services.
- **sqlx** - async, compile-time-checked database access (Postgres, MySQL, SQLite).
- **tokio-util** - extras like `CancellationToken` (clean shutdown across many tasks) and codecs for framing byte streams.

📝 You don't need to learn these all at once. The point is that when one shows up in a `Cargo.toml` or a stack trace, you already know the ground it stands on.

## A word on the alternatives

Let me be direct, because a roots guide that pretended Tokio were the only option would be doing you a disservice: it isn't. **`async-std`** and **`smol`** are other async runtimes, and they're real, working projects with thoughtful designs.

But here's the practical reality. Tokio is dominant, and most of the async library ecosystem assumes it. The crates above - hyper, tonic, reqwest, sqlx - are built against Tokio. If you pick a different runtime, you'll find a thinner shelf of compatible libraries and more rough edges.

There's a deeper reason this matters, and it's worth holding onto: **a future generally needs to run on the runtime it was created for.** Tokio's I/O types register with Tokio's reactor; hand one to `async-std`'s executor and it has nothing to wake it. You can't freely mix futures from different runtimes in the same program. So the runtime choice is a foundational one, not a swappable detail - and for almost everyone, the answer that keeps the most doors open is Tokio.

## What to build

💡 Reading got you here. Building is what makes it permanent. You now understand the layer every async Rust program stands on - the fastest way to lock that in is to make the engine do real work.

Pick one of these. Each leans on a different piece you just learned:

- **A chat server.** Accept connections with `tokio::net`, spawn a task per client, and use a `broadcast` channel to fan every message out to all of them. This is the canonical exercise for "many tasks sharing state without locks."
- **A concurrent downloader.** Take a list of URLs, `spawn` a task per download (with `reqwest`), and `join!` them so they all run at once instead of one after another. You'll *feel* the concurrency in the wall-clock time.
- **A worker pool.** Push jobs into an `mpsc` channel from one side; on the other side, a fixed set of worker tasks pull and process them. Backpressure, fan-out, graceful shutdown - all the patterns from [Phase 5](05-channels-and-sync.md) in one small program.

When you've built one and want to see the HTTP layer that sits directly above Tokio - where connections become requests and responses - read [hyper & tower](/guides/hyper-and-tower). That's the natural next rung up the stack.

The line to carry out of this whole guide: **`async fn`s make inert futures; Tokio is the engine that polls, schedules, and wakes them - and now you know the engine.** Go build something that makes it run.

## Recap

1. **Frameworks are Tokio apps.** axum, actix-web, and Rocket build the runtime with `#[tokio::main]` and run as a pile of spawned tasks - one per connection - on the work-stealing scheduler. The framework is conveniences over spawning and polling you already understand.
2. **The stack is legible top to bottom:** your handlers → framework → hyper (HTTP) → Tokio tasks → scheduler → a small pool of OS threads.
3. **Tokio ships the building blocks:** `tokio::net` (TCP/UDP), `tokio::io` (`AsyncRead`/`AsyncWrite`), `tokio::fs`, and `tokio::time`. On top sit hyper, tonic, reqwest, sqlx, and tokio-util.
4. **Alternatives exist but Tokio dominates.** `async-std` and `smol` are real runtimes, but most libraries assume Tokio - and futures generally need the runtime they were made for, so the choice is foundational, not swappable.
5. **Build to cement it:** a chat server (`broadcast`), a concurrent downloader (`spawn` + `join!`), or a worker pool (`mpsc`). Then read hyper & tower for the HTTP layer above Tokio.

## Quick check

One last check - the picture that turns Rust web servers from magic into machinery:

```quiz
[
  {
    "q": "Mechanically, what is an axum (or actix-web, or Rocket) server?",
    "choices": [
      "A Tokio application: #[tokio::main] builds the runtime, and the server spawns a task per connection on the scheduler",
      "A standalone runtime that replaces Tokio entirely",
      "A synchronous, thread-per-request server with no async involved",
      "A browser-side framework that never touches the runtime"
    ],
    "answer": 0,
    "explain": "These frameworks are Tokio apps. The #[tokio::main] macro builds the runtime you learned about, and the server accepts connections in a loop, calling tokio::spawn for each one. Routing and extractors are conveniences over that spawning and polling."
  },
  {
    "q": "Which crate is the HTTP implementation that most Rust web frameworks sit on, directly above Tokio?",
    "choices": [
      "hyper",
      "sqlx",
      "tonic",
      "tokio-util"
    ],
    "answer": 0,
    "explain": "hyper is the HTTP/1 and HTTP/2 implementation under most Rust web frameworks; it runs on Tokio. sqlx is async database access, tonic is gRPC, and tokio-util provides extras like CancellationToken and codecs."
  },
  {
    "q": "Why is choosing a runtime a foundational decision rather than a swappable detail?",
    "choices": [
      "A future generally needs to run on the runtime it was created for, and most async libraries assume Tokio",
      "Runtimes are interchangeable, so any future runs on any executor without issue",
      "async-std is required before Tokio will start",
      "The runtime only matters for file I/O and nothing else"
    ],
    "answer": 0,
    "explain": "Tokio's I/O types register with Tokio's reactor; an executor from a different runtime has no way to wake them. You can't freely mix futures across runtimes, and since most of the ecosystem (hyper, reqwest, sqlx) targets Tokio, it's the choice that keeps the most doors open."
  }
]
```
