# hyper & tower: The HTTP and Middleware Under axum

> Learn the two libraries every Rust web framework is built on: hyper, the low-level HTTP implementation, and tower, the universal async Service abstraction. The Service trait, Layers and middleware composition, the tower-http toolbox, and exactly how axum is a tower Service over hyper. The bottom of the Rust web stack, made visible.


---

# hyper & tower: The HTTP and Middleware Under axum

Underneath [axum](/guides/axum-from-zero) - and a lot of the Rust networking world - sit two libraries
most people never learn directly: **hyper**, a fast, correct, low-level HTTP implementation, and **tower**,
an abstraction for "an async function from a request to a response" that makes middleware composable across
the whole ecosystem. This is the deepest **roots** guide in the Rust set. You'll rarely write hyper or
tower by hand, but understanding them demystifies a stack of things at once: what a `Service` is, why axum
middleware is a `tower::Layer`, how `tower-http` gives you tracing/CORS/timeouts for free, and what
`axum::serve` is really doing.

The mental model is one trait and one wrapper. **`tower::Service`** is the universal shape: *given a
request, asynchronously produce a response* (plus a readiness check for backpressure). Everything - a
whole axum app, a single endpoint, a database client, a rate limiter - can be a `Service`. A **`Layer`**
is a function that wraps one `Service` to make another (that's middleware). hyper, meanwhile, is what
actually speaks HTTP on the socket and calls your top-level `Service` for each request. Hold "a `Service`
turns a request into a response, a `Layer` wraps a `Service`, and hyper drives the whole thing over the
network," and the Rust web stack lays itself out flat.

> 📝 This is the deepest **roots** guide - it assumes **Rust** (traits, generics, `async` - [Rust From
> Zero](/guides/rust-from-zero)) and the runtime beneath it ([Tokio](/guides/tokio-the-async-runtime)).
> It's most rewarding after you've used [axum](/guides/axum-from-zero), so the pieces have somewhere to
> land. Examples run as plain Rust programs.

## How to read this

Read in order - it builds from hyper's HTTP server up through the `Service` trait, `Layer`s, and
`tower-http`, then maps the whole thing back onto axum. Phases carry difficulty badges.

## The phases

1. **[What hyper & tower Are](01-what-hyper-and-tower-are.md)** 🟢 - the HTTP library and the `Service` abstraction, and where they sit in the stack.
2. **[hyper: The HTTP Library](02-hyper-the-http-library.md)** 🟡 - `Request`/`Response`, a bare hyper server, and what it hands your code.
3. **[The Service Trait](03-the-service-trait.md)** 🔴 - `poll_ready` + `call`, the universal "async request → response," and why it's everywhere.
4. **[Layers & Middleware](04-layers-and-middleware.md)** 🔴 - `tower::Layer`, `ServiceBuilder`, and composing middleware as wrapped services.
5. **[The tower-http Toolbox](05-tower-http.md)** 🟡 - tracing, CORS, compression, and timeouts as reusable layers.
6. **[How axum Uses Them](06-how-axum-uses-them.md)** 🟢 - axum's `Router` as a `Service` over hyper, with your layers wrapped around it.
7. **[Where to Go Next](07-where-to-go-next.md)** 🟢 - applying this to gRPC (tonic), clients, and the rest of the tower ecosystem.

> The throughline: a **`Service`** turns a request into a response, a **`Layer`** wraps a `Service`
> (that's middleware), and **hyper** drives it over the socket. axum is a `Service` you assembled.


---

# What hyper & tower Are

When you write an [axum](/guides/axum-from-zero) app, you import a `Router`, hang some handlers off
it, add a `.layer(...)` or two, and call `axum::serve`. It feels like one tidy thing. But underneath it
sit two libraries almost nobody learns directly - **hyper** and **tower** - and nearly the whole Rust
networking world is standing on them too.

This is the deepest **roots** guide in the Rust set. Here's the plain framing up front: you will
rarely write hyper or tower by hand, and this guide is most rewarding *after* you've already built
something with axum, so the pieces have somewhere to land. We're going to assume you've met
[HTTP](/guides/http-explained) (requests, responses, status codes) and the async runtime underneath
it all, [Tokio](/guides/tokio-the-async-runtime). With those in hand, the goal of this phase is small
and sharp: build the mental model, name the two libraries, and see where they sit. The deep code
comes in Phases 2 through 4.

## The mental model: one trait, one wrapper, one driver

Before any names or code, hold this picture. It's the whole guide compressed into one sentence:

> 💡 **A `Service` turns a request into a response. A `Layer` wraps a `Service` (that's middleware).
> hyper drives it over the socket. Tokio runs the whole thing.**

That's it. Everything else is detail hanging off those four clauses. Read them as a stack, bottom to
top - the runtime at the bottom, the network library above it, your wrapped-up app on top:

```mermaid
flowchart TB
  T[Tokio: the async runtime] --> H[hyper: speaks HTTP on the socket]
  H --> L[tower Layers: middleware wrapping your app]
  L --> S[your app: a Service, request to response]
```

*What just happened:* we drew the Rust web stack as four layers. **Tokio** at the bottom is the
engine that actually runs async tasks (it's the *only* one of the four that touches threads and the
OS scheduler). **hyper** sits on top of Tokio and knows how to read and write HTTP over a TCP
connection. Above hyper, **tower Layers** wrap your code with reusable middleware. And at the very
top, **your app is a `Service`** - the thing that takes a request and produces a response. Each layer
only talks to its neighbors. Keep this diagram; the rest of the guide is just zooming into each box.

## What hyper is

📝 **hyper** - a fast, correct, **low-level** HTTP library for Rust. It implements HTTP/1 and HTTP/2,
for both clients and servers. When bytes arrive on a socket, hyper is the thing that parses them into
a proper HTTP request, and when you hand it a response, it serializes that back into bytes on the
wire. It speaks the protocol so you don't have to.

Notice the word **low-level**, because it's doing a lot of work in that sentence. hyper gives you HTTP
*on the socket* and almost nothing above that:

- It does **not** give you routing. There's no "when the path is `/users/:id`, call this function."
- It does **not** give you extractors. Nothing pulls a JSON body or a query parameter out for you.
- It does **not** give you middleware in any built-in sense.

Those conveniences - routing, extractors, middleware - are exactly what a *framework* like axum adds
on top. hyper's job is narrower and deeper: be the correct, performant HTTP implementation that
everything else builds on. Think of it as the part of the stack that handles "this is genuinely valid
HTTP/1.1, and here are the parsed pieces," and then steps out of your way.

> 📝 If you've read [/guides/wsgi-and-asgi-explained](/guides/wsgi-and-asgi-explained), hyper plays a
> role a bit like a WSGI/ASGI *server* (gunicorn, uvicorn): it owns the socket and the HTTP plumbing,
> then calls into your code. The big difference is *how* it calls your code - and that's tower's
> story.

## What tower is

📝 **tower** - a library of reusable, composable components for networking, built around one central
idea: the **`Service`** trait. A `Service` is an abstraction for *"an async function from a request to
a response"* (plus a small readiness check we'll meet in Phase 3). That's the entire concept. Given a
request, asynchronously produce a response.

The power is in how *general* that is. Once "request in, response out" is a named, shared shape,
almost anything fits it:

- A single endpoint is a `Service`.
- A whole axum app is a `Service`.
- A database client can be a `Service` (request: a query; response: rows).
- A rate limiter is a `Service`.

And here's the second half of tower, the part that makes it more than just a trait:

📝 **`Layer`** - a thing that wraps one `Service` to produce another `Service`. That's middleware,
made composable. A logging `Layer` wraps your app and returns a new `Service` that logs, then calls
the inner one. A timeout `Layer` wraps that and returns yet another `Service` that enforces a
deadline. Because every layer takes a `Service` and returns a `Service`, you can stack them like
nesting dolls, in any order, and reuse the same `Layer` across completely different apps.

> 💡 This is the quiet superpower. Because a `Layer` is "Service in, Service out," middleware written
> for one project works in any other - and across whole *libraries*. That's why a crate like
> `tower-http` can ship tracing, CORS, compression, and timeouts as `Layer`s that drop into *any*
> tower-based app, axum included. We're not writing one yet (Phase 4 does); just hold "a `Layer`
> wraps a `Service`."

## Where they sit under axum

Now the payoff - let's map those names back onto the axum you've actually used. The same three lines
you write every day are, underneath, exactly the three concepts above:

```rust
let app = Router::new()
    .route("/", get(handler))   // your app is a Service
    .layer(TraceLayer::new_for_http()); // a tower Layer wraps it

axum::serve(listener, app).await?; // hyper drives it over the socket
```

*What just happened:* three lines, three layers of the stack. `Router::new()...` builds an
`axum::Router`, and a `Router` **is a tower `Service`** - request in, response out, with all the
routing logic living inside its `call`. The `.layer(...)` adds a **tower `Layer`** that wraps that
`Service` with middleware (here, request tracing). And `axum::serve` is **hyper**: it takes your
listening socket, accepts connections, parses HTTP off each one, and calls your top-level `Service`
once per request. Tokio (started by axum's `#[tokio::main]`) is the runtime quietly executing all of
it. Every piece of that line traces to a box in the diagram.

So axum isn't a separate magical universe. It's a set of *ergonomics* - routing, extractors, response
helpers - assembled into a tower `Service`, served by hyper, run on Tokio. The framework is the
convenient face; hyper and tower are the foundation it's standing on.

## Why learn this

Let's look straight at the trade-off, because this project doesn't do hand-waving.

⚠️ **You will rarely write raw hyper or tower at a real job, and that's fine.** Frameworks like axum
exist precisely so you don't have to wire up sockets and `Service` impls by hand. Reaching for bare
hyper when axum would do is usually a mistake, not a badge of honor. This guide is *not* arguing you
should hand-roll HTTP servers.

It's arguing something more useful: knowing this layer is what makes the layers above it stop being
magic. Once you can see the `Service` and the `Layer` underneath, a pile of axum mysteries resolve at
once - 

- **Why axum middleware is a `.layer(...)`** - because middleware *is* a tower `Layer`, the same
  abstraction the whole ecosystem shares.
- **What `tower-http` is** - a box of ready-made `Layer`s (tracing, CORS, compression, timeouts) that
  work in any tower app, not just axum.
- **What `axum::serve` actually does** - it's hyper, accepting connections and calling your `Service`.

Set your expectations accordingly: this is a **conceptual roots** guide, not day-to-day code. The
goal isn't a hyper server you'll deploy. It's an X-ray view of the stack you already use, so the next
time you write `.layer(...)` or read a `tower` error message, you know exactly what's underneath. Next
up, Phase 2 zooms into the bottom of the picture: hyper itself, its `Request`/`Response` types, and a
bare hyper server.

## Recap

1. Underneath axum (and much of Rust networking) sit two libraries most people never learn directly:
   **hyper** (HTTP) and **tower** (the `Service`/middleware abstraction).
2. The mental model: **a `Service` turns a request into a response; a `Layer` wraps a `Service`
   (middleware); hyper drives it over the socket; Tokio runs it all.**
3. **hyper** is a fast, correct, *low-level* HTTP library - HTTP/1 and HTTP/2, client and server. It
   speaks HTTP on the socket but gives you no routing, extractors, or middleware.
4. **tower** centers on the **`Service`** trait ("async request → response") and **`Layer`** (wrap a
   `Service` to get a new one). Anything request-to-response can be a `Service`; middleware is a
   `Layer`.
5. Under axum: a `Router` **is a `Service`**, `.layer(...)` adds **tower `Layer`s**, and `axum::serve`
   **is hyper**. axum is ergonomics assembled into a `Service`, served by hyper, run on Tokio.
6. ⚠️ You'll rarely write hyper/tower by hand - but understanding them demystifies `Service`, why axum
   middleware is a `Layer`, what `tower-http` is, and what `axum::serve` does.

## Quick check

Three questions on the ideas that have to stick before Phase 2:

```quiz
[
  {
    "q": "In one line, what is the difference between hyper and tower?",
    "choices": [
      "hyper is a low-level HTTP library that speaks the protocol on the socket; tower is the Service abstraction (async request-to-response) plus composable Layers for middleware",
      "hyper is the routing framework and tower is the database layer",
      "They are two names for the same library; tower is just the newer one",
      "hyper handles middleware and tower handles parsing HTTP off the wire"
    ],
    "answer": 0,
    "explain": "hyper implements HTTP/1 and HTTP/2 on the socket but gives you no routing, extractors, or middleware. tower provides the Service trait (async request to response) and Layers that wrap Services to make composable middleware."
  },
  {
    "q": "What is a tower `Layer`?",
    "choices": [
      "Something that wraps one Service to produce another Service - that's how middleware is made composable",
      "A TCP connection pool that hyper manages internally",
      "A trait you implement to parse JSON request bodies",
      "The runtime that schedules async tasks onto threads"
    ],
    "answer": 0,
    "explain": "A Layer takes a Service and returns a new Service, so middleware (logging, timeouts, CORS) can be stacked and reused across any tower-based app. The task scheduler is Tokio, not a Layer."
  },
  {
    "q": "When you call `axum::serve(listener, app)`, what is each part underneath?",
    "choices": [
      "axum::serve is hyper driving HTTP over the socket; `app` (the Router) is a tower Service; any `.layer(...)` you added are tower Layers wrapping it",
      "axum::serve is tower and the Router is hyper",
      "Both axum::serve and the Router are pure hyper with no tower involved",
      "axum::serve is Tokio and the Router parses raw bytes itself"
    ],
    "answer": 0,
    "explain": "axum::serve is hyper: it accepts connections, parses HTTP, and calls your top-level Service once per request. The Router is that Service, your .layer(...) calls are tower Layers wrapping it, and Tokio runs the whole thing."
  }
]
```


---

# hyper: The HTTP Library

Here's the mental model to carry through this whole phase, and indeed through the
rest of the guide: **hyper is the thing that speaks HTTP on the socket.** It reads
the raw bytes a client sent, parses them into a `Request`, hands that request to a
function you wrote, waits for you to give back a `Response`, and writes that response
back out as bytes - correctly, including all the fiddly HTTP/1 and HTTP/2 framing you
never want to implement yourself.

That's the job. hyper is the HTTP *engine*. It is not a web framework, and the
difference matters more than you'd expect - we'll get to exactly what it leaves out,
and why leaving it out is a feature, not a gap.

> 💡 If you remember one sentence: hyper takes a `Request` off the wire and calls
> your code, expecting a `Response`. Everything else in this phase is detail about
> what those two types are and how you wire the loop together.

## The types come from the `http` crate

hyper doesn't invent its own request and response types. It uses the ones from the
`http` crate - a tiny, dependency-light crate that the whole Rust HTTP ecosystem
shares. That sharing is the point: hyper, axum, reqwest, and tonic all agree on what
a `Request` is, so a request value can pass between them without conversion.

The two stars are generic over their **body**:

```rust
// from the `http` crate
struct Request<B>  { /* method, uri, headers, ... + a body of type B */ }
struct Response<B> { /* status, headers, ...           + a body of type B */ }
```

*What just happened:* `Request<B>` and `Response<B>` bundle the metadata you'd
expect - method, URI, status code, headers - with a body whose type is a parameter
`B`. The metadata is always the same shape; the body type varies depending on where
the value is in its life. An incoming request's body and an outgoing response's body
are *different* concrete types, and that's normal.

## Bodies are streams of bytes (the `Body` trait)

Why is the body a type parameter instead of, say, a `Vec<u8>`? Because an HTTP body
isn't always a finished buffer sitting in memory. It might be a 4 GB file streaming
in chunk by chunk, or a server-sent-events stream that never really ends. So hyper
abstracts "a body" behind a trait - `http_body::Body` - which describes *a thing you
can pull data frames from over time*, rather than a fixed blob.

You'll meet a few concrete body types constantly:

- **`hyper::body::Incoming`** - the body of a request that just arrived. hyper hands
  you this; you read from it. You don't construct it.
- **`http_body_util::Full<Bytes>`** - a whole buffer you already have in memory,
  presented as a (one-chunk) body. The everyday choice for a simple response.
- **`http_body_util::Empty<Bytes>`** - a body with no data at all (a `204`, say).

```rust
use http_body_util::Full;
use hyper::body::Bytes;

// A response whose body is one complete buffer:
let body = Full::new(Bytes::from("hello from hyper"));
```

*What just happened:* `Bytes` is a cheap-to-clone reference-counted byte buffer (from
the `bytes` crate). Wrapping it in `Full` says "this is the entire body, all at
once." `Full<Bytes>` implements the `Body` trait, so hyper knows how to write it to
the socket. For most handlers returning a string or a JSON blob, `Full<Bytes>` is all
you need.

> ⚠️ Don't reach for the `Body` trait's methods (`poll_frame` and friends) by hand
> unless you're building streaming machinery. Day to day you pick a ready-made body
> type and move on. The trait is there so the *plumbing* is generic, not so you write
> plumbing.

## You hand hyper a service; it calls you

hyper needs something to call for each request. That something is a **service**: in
spirit, an async function from `Request` to `Response`. (Phase 3 makes "service" a
precise trait - for now, "async fn request → response" is the right picture.)

The quickest way to make one is `service_fn`, which turns an async closure or function
into a service hyper accepts:

```rust
use http_body_util::Full;
use hyper::{body::Bytes, Request, Response, service::service_fn};

async fn handle(_req: Request<hyper::body::Incoming>)
    -> Result<Response<Full<Bytes>>, hyper::Error>
{
    Ok(Response::new(Full::new(Bytes::from("hello from hyper"))))
}

let service = service_fn(handle);
```

*What just happened:* `handle` takes a request whose body is `Incoming` (it came off
the wire) and returns a `Result` of a response whose body is `Full<Bytes>` (we built
it in memory). Notice the body types differ on the way in versus the way out - 
exactly as promised. `service_fn(handle)` wraps the function into the shape hyper
wants to call. We ignore the request here (`_req`), because there's no router yet to
look at the path - every request gets the same reply.

## Serving a connection

hyper is **runtime-agnostic** in 1.x: its core doesn't know about Tokio, threads, or
sockets. That keeps the HTTP logic portable, but it means you have to bridge hyper to
whatever I/O you're actually using. With Tokio (the usual choice - see
[/guides/tokio-the-async-runtime](/guides/tokio-the-async-runtime)), that bridge is
**`hyper_util::rt::TokioIo`**, a thin adapter that wraps a Tokio TCP stream so hyper
can read and write through it.

The shape of serving one connection looks like this:

```rust
use hyper::server::conn::http1;
use hyper_util::rt::TokioIo;

// `tcp_stream` is a tokio::net::TcpStream you got from accepting a connection.
let io = TokioIo::new(tcp_stream);

http1::Builder::new()
    .serve_connection(io, service_fn(handle))
    .await?;
```

*What just happened:* `TokioIo::new` wraps the raw Tokio socket into something hyper's
runtime-agnostic core can drive. `http1::Builder::new().serve_connection(...)` then
runs the HTTP/1 protocol over that connection: it parses incoming requests, calls
your service for each one, writes the responses, and keeps the connection alive for
keep-alive requests until the client goes away. `.await` drives that to completion for
this one connection.

> 📝 This is the *shape*, not a paste-and-run server. A real program adds a
> `tokio::net::TcpListener` accept loop, spawns a task per connection, and handles
> errors - and most people reach for `hyper_util::server::conn::auto` instead of
> `http1` so a single setup serves both HTTP/1 and HTTP/2. The takeaway here is the
> three moving parts: **adapt the socket (`TokioIo`), pick a protocol builder, hand it
> a service.**

## What hyper does *not* give you

This is the part that surprises people coming from Express, Flask, or even axum. Look
again at `handle`: it received a `Request` and the only thing it could do was build a
`Response`. There was:

- **No router.** hyper does not look at the path or method for you. `GET /users/42`
  and `POST /login` arrive at the *same* function; matching them is your problem.
- **No path parameters, no query parsing helpers.** You get a `Uri`; pulling `42` out
  of `/users/42` is on you.
- **No middleware.** No built-in logging, auth, CORS, compression, or timeouts.
- **No JSON helpers.** No `req.json()`, no automatic serialization. You read raw
  bytes and call serde yourself.

```rust
// With bare hyper, you'd route by hand - something like:
match (req.method(), req.uri().path()) {
    (&hyper::Method::GET, "/")      => { /* home */ }
    (&hyper::Method::POST, "/login") => { /* login */ }
    _ => { /* 404 */ }
}
```

*What just happened:* that hand-rolled `match` is the kind of thing a framework
generates for you. With bare hyper you'd write it, and it'd get unwieldy fast as
routes multiply - which is precisely *why* frameworks exist. hyper deliberately stops
at "raw request in, raw response out."

⚠️ That bareness is the whole design philosophy, not an oversight. hyper aims to be a
small, fast, correct HTTP implementation that *everything else* can build on. Routing,
middleware, and JSON are opinions; hyper stays out of opinions so axum, tonic, and
your own code can layer theirs on top. When you understand that hyper is "just" the
HTTP engine, the rest of the stack - services, layers, axum - clicks into place,
because you can see exactly what they're adding.

## The 1.0 split: hyper-util and http-body-util

You've now seen three crates working together, and that's intentional. When hyper hit
1.0, it pushed conveniences *out* of the core into companion crates so the core itself
could stay small and promise long-term stability:

- **`hyper`** - the stable core: protocol, `Incoming`, `serve_connection`,
  `service_fn`.
- **`hyper-util`** - runtime glue and helpers that aren't part of the stable
  guarantee: `TokioIo`, the `auto` connection server, client pools.
- **`http-body-util`** - ready-made body types: `Full`, `Empty`, and combinators.

> 📝 The logic: a 1.0 means a strong backward-compatibility promise. The smaller the
> stable surface, the easier that promise is to keep. So hyper kept only the truly
> stable HTTP core under the 1.0 umbrella and parked the more-likely-to-evolve helpers
> in `-util` crates. That's why importing from three crates to write one server is
> normal here - it's a deliberate split, not boilerplate sprawl.

## Recap

- **hyper speaks HTTP on the socket:** it parses bytes into a `Request`, calls your
  service, and writes your `Response` back out. It's the HTTP engine, not a framework.
- **The types come from the `http` crate:** `Request<B>` and `Response<B>`, generic
  over a body type `B`, shared across the whole ecosystem.
- **Bodies are streams behind the `Body` trait.** You usually pick a concrete type:
  `Incoming` for arriving requests, `Full<Bytes>` for an in-memory response.
- **You hand hyper a service** (quickest: `service_fn`), and serve a connection by
  adapting the socket with `TokioIo` and running a protocol builder's
  `serve_connection`. hyper's core is runtime-agnostic; `hyper-util` bridges it to
  Tokio.
- **hyper gives you no router, middleware, or JSON** - on purpose. That bareness is
  what lets frameworks build on it.
- **hyper 1.0 split** conveniences into `hyper-util` and `http-body-util` to keep the
  stable core small.

## Quick check

```quiz
[
  {
    "q": "In hyper 1.x, what does TokioIo do?",
    "choices": [
      "Parses the HTTP request into a Request value",
      "Adapts a Tokio TCP socket so hyper's runtime-agnostic core can drive it",
      "Routes requests to the correct handler by path",
      "Serializes a struct into a JSON response body"
    ],
    "answer": 1,
    "explain": "hyper's core doesn't know about Tokio. TokioIo (from hyper-util) wraps a Tokio stream so hyper can read and write through it."
  },
  {
    "q": "Which body type would you typically use for a response that's a complete in-memory buffer?",
    "choices": [
      "hyper::body::Incoming",
      "http_body_util::Empty<Bytes>",
      "http_body_util::Full<Bytes>",
      "Vec<u8>"
    ],
    "answer": 2,
    "explain": "Full<Bytes> presents a whole buffer as a one-chunk body. Incoming is for arriving requests; Empty is for no body at all."
  },
  {
    "q": "Which of these does bare hyper give you out of the box?",
    "choices": [
      "A router that matches paths to handlers",
      "Middleware for logging and CORS",
      "JSON request/response helpers",
      "The parsed Request and the raw connection, and nothing more"
    ],
    "answer": 3,
    "explain": "hyper hands you the request and the connection and stops there. Routing, middleware, and JSON are deliberately left to frameworks built on top of it."
  }
]
```


---

# The Service Trait

Strip away the type machinery and a `tower::Service` is two things bolted together:

- **`call`** - give it a request, it hands you back a *future* of a response.
- **`poll_ready`** - ask it "are you ready to take work right now?" before you call.

That pair - *do the work* and *am I ready?* - is the universal shape every piece of
the tower ecosystem speaks. An entire axum app, one endpoint, a database client, a
rate limiter, a gRPC server: each is something you hand a request and get a response
back from, with a chance to say "hold on, I'm full." Once that clicks, the rest of the
stack stops looking like magic and starts looking like the same trait wearing different
hats.

> 📝 This is the deepest phase in the guide. It assumes you're comfortable with Rust
> traits, generics, associated types, and `async`/futures. If `Future` and
> `Poll` aren't familiar yet, the [Tokio guide](/guides/tokio-the-async-runtime) is the
> place to shore that up first.

## The trait itself

Here is `tower::Service`, simplified down to the parts that matter (the real one carries
a couple of extra bounds, but the shape is exactly this):

```rust
pub trait Service<Request> {
    type Response;
    type Error;
    type Future: Future<Output = Result<Self::Response, Self::Error>>;

    fn poll_ready(&mut self, cx: &mut Context<'_>) -> Poll<Result<(), Self::Error>>;
    fn call(&mut self, req: Request) -> Self::Future;
}
```

*What just happened:* We declared three associated types - what the service produces
(`Response`), how it fails (`Error`), and the concrete `Future` type `call` returns - 
and two methods. `call` takes a `Request` and returns that future. `poll_ready` returns
`Poll<Result<(), Self::Error>>`: `Poll::Ready(Ok(()))` means "go ahead, call me,"
`Poll::Pending` means "not yet, I'll wake you," and an `Err` means the service has failed.
Notice there's no `async fn` here - `call` returns a future *value* you then await, which
is what lets these compose without boxing everything.

### `call` is the easy half

`call` is the part your intuition already had: a request goes in, a future of
`Result<Response, Error>` comes out. You await that future to get the response. If you've
written an HTTP handler, you've written the body of a `call` - this is just the formal
shape of "handle one request."

### `poll_ready` is the half nobody tells you about

`poll_ready` is the readiness and **backpressure** hook, and it's the genuinely
non-obvious part. The contract is a discipline between caller and service:

> A caller must wait until `poll_ready` returns `Poll::Ready(Ok(()))` *before* it is
> allowed to call `call`.

That single rule is what lets a service push back. Imagine a rate limiter that allows 100
requests per second. When it's used up its budget for this window, its `poll_ready`
returns `Poll::Pending` and registers to wake the caller when the next window opens - so
the caller naturally stalls instead of flooding the service. Or picture a connection pool
that's handed out every connection it has: `poll_ready` stays `Pending` until one frees
up. The service gets to say *"I'm at capacity, hold off"* through the type system, and
well-behaved callers honor it.

> ⚠️ The flip side of that contract: a `Service` may assume `poll_ready` returned `Ready`
> before `call` happens, and is allowed to reserve a resource (a permit, a pool slot)
> when it reports `Ready`. So you don't call `call` twice off one `poll_ready`, and you
> don't skip `poll_ready`. When you compose services through tower, the framework
> upholds this for you - it matters most when you're writing a `Service` by hand.

> 💡 Most simple services are *always* ready - their `poll_ready` is a one-liner that
> returns `Poll::Ready(Ok(()))`. Backpressure is opt-in. You only reach for a meaningful
> `poll_ready` when the service genuinely has a limited resource to protect.

## Why it's generic over the request

Look again at the trait header: `Service<Request>`. The request type is a generic
parameter, not baked in. That one decision is why the whole ecosystem composes.

Because the request is generic, the *same* trait describes:

- an **HTTP** service whose `Request` is `http::Request<Body>`,
- a **gRPC** service whose request is a decoded protobuf message,
- a **database** client whose request is a query,
- any **request → response** thing you can imagine.

The middleware you'll meet next phase - timeouts, retries, concurrency limits - is
written against `Service<Request>` generically, so it works on *all* of them without
caring what flows through. A timeout layer doesn't know or care whether it's wrapping an
HTTP handler or a database call; it just knows it's wrapping a `Service`. That generality
is the entire reason a timeout you learned for axum also applies to a tonic gRPC client.

## The easy way to make one: `service_fn`

Implementing `Service` by hand is genuinely verbose - you have to name an associated
`Future` type, which usually means boxing the future or writing your own. For the common
case where your service is just "an async function," tower gives you a shortcut:
**`tower::service_fn`** turns an async closure into a `Service`.

```rust
use tower::{service_fn, Service, ServiceExt};
use std::convert::Infallible;

// An async closure (request -> response) becomes a full Service.
let mut svc = service_fn(|req: String| async move {
    Ok::<_, Infallible>(format!("handled: {req}"))
});

// Honor the contract: wait until ready, then call.
let ready = svc.ready().await.unwrap();      // drives poll_ready for you
let resp = ready.call("ping".to_string()).await.unwrap();
assert_eq!(resp, "handled: ping");
```

*What just happened:* `service_fn` wrapped our closure into something implementing the
full `Service` trait - `Response`, `Error`, `Future`, `poll_ready`, and `call` are all
synthesized for us, and its `poll_ready` is the always-ready one-liner. The
`ServiceExt::ready` helper drives `poll_ready` to completion and then hands back a
`&mut Service` you can `call` - so it enforces the "ready before call" contract in one
line. We get a real `Service` without naming a single associated type.

When *would* you hand-write a `Service`? When it's **stateful** - a real rate limiter
holding a token bucket, a connection pool, anything whose `poll_ready` needs to inspect
internal state and return `Pending`. A closure can't easily carry and mutate that across
`poll_ready` and `call`, so you implement the trait directly. That's the verbose path,
and it exists for exactly these cases.

## Everything is a Service

Here's the payoff, and it's worth saying plainly:

> 💡 An axum `Router` is a `Service`. A single axum handler is a `Service`. A tonic gRPC
> server is a `Service`. An HTTP *client* is a `Service`. The middleware that wraps any
> of them produces another `Service`.

Learn this one trait and the ecosystem clicks into place. The next phase shows how a
`Layer` *wraps* a `Service` to make a new one - that's all middleware is - and once both
ideas are in hand, "how does axum actually work" answers itself.

## Recap

- A `tower::Service` is **`call`** (request → future of `Result<Response, Error>`) plus
  **`poll_ready`** (am I ready to take work?).
- `poll_ready` is the **backpressure** hook: callers must wait for `Ready` before calling,
  letting a service signal "I'm at capacity" (rate limiters, full connection pools).
- Most simple services are **always ready** - a one-line `poll_ready`. Backpressure is opt-in.
- The trait is **generic over the request type**, which is why one set of middleware
  composes across HTTP, gRPC, database clients, and more.
- **`tower::service_fn`** turns an async closure into a `Service` - the easy path.
  Hand-implement the trait only for **stateful** services that need a real `poll_ready`.
- Everything in the ecosystem - `Router`, handlers, gRPC servers, HTTP clients - is a
  `Service`.

## Quick check

```quiz
[
  {
    "q": "What does a Service's call method return?",
    "choices": ["The response directly", "A future of Result<Response, Error>", "A Poll value", "Nothing; it mutates state"],
    "answer": 1,
    "explain": "call takes a request and returns a future that resolves to Result<Response, Error> - you await it to get the response."
  },
  {
    "q": "What is poll_ready for?",
    "choices": ["Parsing the request body", "Backpressure: signaling whether the service can take work before call", "Returning the response", "Logging each request"],
    "answer": 1,
    "explain": "poll_ready is the readiness/backpressure hook. A caller must see Poll::Ready(Ok(())) before calling call, so a service can say 'I'm at capacity, hold off.'"
  },
  {
    "q": "When would you hand-implement Service instead of using service_fn?",
    "choices": ["Always - service_fn is deprecated", "Never - service_fn covers every case", "For stateful services whose poll_ready must inspect state and return Pending", "Only for HTTP services"],
    "answer": 2,
    "explain": "service_fn is the easy path for stateless 'just an async function' services (always ready). You hand-write the trait for stateful services like a rate limiter or pool whose poll_ready needs real logic."
  }
]
```


---

# Layers & Middleware

In [the last phase](03-the-service-trait.md) you saw the one trait the whole tower world stands on: a
**`Service`** is "an async function from a request to a response," with a readiness check bolted on the
front. Once everything is a `Service`, a beautiful thing falls out of the shape - the whole reason tower
exists.

Here is the mental model, and it's the only thing you need to carry through this phase:

> 💡 **Middleware is a `Service` that wraps another `Service`. A `Layer` is the factory that does the
> wrapping.**

That's it. A piece of middleware (logging, auth, timeouts, compression) is not a special kind of object
with its own special trait. It's a plain `Service` that happens to hold *another* `Service` inside it - 
the `inner` one - and in its `call`, it does a little work *before* handing the request down to `inner`,
and a little more work *after* `inner` hands the response back up. Request goes down through the layers,
response comes back up through them. Like an onion, or nesting dolls, or - the metaphor that'll actually
stick - a stack of `try/finally` blocks wrapped around your real handler.

The `Layer` is the small companion piece: a factory whose only job is "given an inner service, build the
wrapper around it." You need both because the wrapping has to happen *for every service you apply it to*,
and a factory is how you make that reusable.

Let's look at all three pieces: the `Layer` trait, the wrapper service, and `ServiceBuilder` (which makes
stacking them readable).

## The `Layer` trait

The trait is almost insultingly small. It has one associated type and one method:

```rust
pub trait Layer<S> {
    type Service;
    fn layer(&self, inner: S) -> Self::Service;
}
```

*What just happened:* `S` is the type of the service being wrapped (the `inner`). `layer` takes that
inner service and returns a new one - `Self::Service` - which is the wrapper. Read it as a function:
"give me a service, I'll give you back a bigger service with my behavior bolted around it." That's the
entire contract. A `Layer` doesn't handle requests itself; it only *constructs* the thing that does.

> 📝 Notice `layer` takes `&self`, not `self`. The same `Layer` value can wrap many services - which is
> exactly what `ServiceBuilder` relies on below, and what makes a layer reusable across a whole app.

## The wrapper service: before and after `inner.call`

The wrapper is where the actual behavior lives, and it's just a `Service`. The pattern is always the
same: hold the `inner` service in a field, and in `call`, do your work around `inner.call(req)`. Here's
the heart of a logging middleware:

```rust
// shape of a logging middleware service's call:
fn call(&mut self, req: Request) -> Self::Future {
    let start = Instant::now();
    let fut = self.inner.call(req);     // delegate to the wrapped service
    async move {
        let res = fut.await;
        log::info!("took {:?}", start.elapsed());
        res
    }
}
```

*What just happened:* before delegating, we grab a timestamp (the "on the way in" work). Then we call
`self.inner.call(req)` to get the inner service's future - note we don't `.await` it yet, we just hold
the future. Inside the returned `async move` block we `await` it to get the real response, do our "on the
way out" work (logging the elapsed time), and pass the response back up. The middleware sandwiches the
inner service: code before the delegate runs on the way down, code after the `await` runs on the way back
up. Swap the logging for "check an auth header before, add a CORS header after" and you've got auth or
CORS - the *shape* never changes.

> ⚠️ The "before" work runs at `call` time, but the "after" work only runs when the returned future is
> awaited. That's why we capture `start` outside the `async move` block (so it measures the real call
> start) but read `start.elapsed()` inside it (so it measures completion). Mixing those up is the classic
> first-middleware bug - you end up timing how long it took to *create* the future, which is roughly zero.

This is two pieces working together: the **wrapper service** above (does the work) and a small **`Layer`
struct** (a factory whose `layer()` produces that wrapper, plugging in whatever `inner` it's handed).
You'll see those two pieces spelled out at the end of this phase.

## `ServiceBuilder`: composing layers readably

You rarely apply one layer. A real app wants tracing *and* a timeout *and* compression *and* auth - a
stack of them. You *can* wrap by hand (`Timeout::new(Compression::new(Trace::new(app))))`), but that
nests inside-out and reads backwards. tower gives you **`ServiceBuilder`** to write the stack the way you
think about it, top to bottom:

```rust
use tower::ServiceBuilder;

let service = ServiceBuilder::new()
    .layer(trace_layer)        // outermost
    .layer(timeout_layer)
    .layer(compression_layer)  // innermost
    .service(app);             // your real service at the bottom
```

*What just happened:* each `.layer(...)` adds one piece of middleware to the stack, and `.service(app)`
caps it off with the actual service everything wraps. The result is a single `Service` you can hand to
hyper (or that *is* your axum app) - the layers are baked in. It reads as a list instead of a pile of
nested constructors, which matters a lot once you have five of them.

Now the rule everyone gets wrong exactly once:

> ⚠️ **With `ServiceBuilder`, layers wrap top-to-bottom: the FIRST `.layer` is the OUTERMOST.** It runs
> *first* on the way in and *last* on the way out. The last `.layer` is the innermost, closest to your
> real service. In the example above, a request hits `trace_layer` first, then `timeout_layer`, then
> `compression_layer`, then `app` - and the response travels back out in reverse.

This is the *opposite* of what happens when you call `.layer()` repeatedly on a bare service directly: in
that case each new `.layer()` wraps *around* the previous result, so the last one you add ends up
outermost. `ServiceBuilder` deliberately flips this so the reading order (top = first to see the request)
matches the execution order. Pick one mental model - almost everyone uses `ServiceBuilder`, so anchor on
"first `.layer` = outermost = first to touch the request," and you'll be right.

## Why this matters: a Layer works on *any* Service

Here's the payoff, and the reason the next phase exists at all.

> 💡 Because middleware is "a `Service` wrapping a `Service`," a `Layer` doesn't care *what* it's
> wrapping. The same timeout layer works on an axum app, a [tonic](/guides/axum-from-zero) gRPC server,
> or an HTTP *client*. Write the layer once; reuse it everywhere a tower `Service` shows up.

Think about how unusual that is. In most frameworks, "middleware" is a framework-specific concept - 
Express middleware doesn't run on your database client, Django middleware doesn't wrap your gRPC server.
In tower, because the `Service` abstraction is universal, your middleware is universal too. A retry layer
you wrote for outbound HTTP calls can wrap your inbound server. A rate limiter can sit in front of *any*
service. This is the single biggest idea in tower, and it's why a whole crate of ready-made layers can
exist and Just Work across the ecosystem - which is exactly [the tower-http
toolbox](05-tower-http.md), the next phase.

## Do you ever write one by hand?

Mostly, no - and that's the good news. Hand-writing a layer means building **two pieces**: a small
`Layer` struct (the factory) and a wrapper `Service` (the behavior), wired together so the struct's
`layer()` produces the service. It's a fair bit of boilerplate for the type plumbing, and the futures get
fiddly.

The happy path is: **reach for a ready-made layer** from `tower` or `tower-http` (timeouts, tracing,
CORS, compression, concurrency limits - covered next phase), and only hand-write a layer when you need
behavior nobody has packaged for you. When that day comes, you now know the shape: a `Layer` struct whose
`layer()` returns a wrapper `Service` that does work around `inner.call`. Everything else is filling in
types.

## Recap

- **Middleware = a `Service` that wraps another `Service`.** It holds the `inner` service and, in `call`,
  does work before and after delegating to `inner.call(req)`.
- A **`Layer`** is the small factory that does the wrapping: `fn layer(&self, inner: S) -> Self::Service`.
  It builds the wrapper; it doesn't handle requests itself.
- In the wrapper's `call`, "before" work runs immediately; "after" work runs when the returned future is
  awaited. Capture state before, read it after.
- **`ServiceBuilder`** stacks layers readably: `.layer(...).layer(...).service(inner)`. The first
  `.layer` is the **outermost** - first in, last out. (This is the reverse of calling `.layer()` on a
  bare service repeatedly.)
- Because a `Service` is universal, a `Layer` works on **any** service - app, server, or client. Write
  once, reuse everywhere. That's why `tower-http` exists.
- Hand-writing a layer means two pieces (a `Layer` struct + a wrapper `Service`); most of the time you
  use ready-made ones.

## Quick check

```quiz
[
  {
    "q": "What is a piece of tower middleware, structurally?",
    "choices": ["A special trait separate from Service", "A Service that wraps another Service and delegates to inner.call", "A function hyper calls before any Service", "A configuration struct read at startup"],
    "answer": 1,
    "explain": "Middleware is just a Service that holds an inner Service and does work before/after calling inner.call(req). A Layer is the factory that builds it."
  },
  {
    "q": "Given ServiceBuilder::new().layer(a).layer(b).service(app), which layer touches an incoming request first?",
    "choices": ["app", "b, because it was added last", "a, because the first .layer is the outermost", "They run concurrently"],
    "answer": 2,
    "explain": "With ServiceBuilder, the first .layer is the outermost: it runs first on the way in and last on the way out. So 'a' sees the request first, then 'b', then 'app'."
  },
  {
    "q": "Why can the same tower Layer wrap an axum app, a tonic server, and an HTTP client?",
    "choices": ["Because tower copies the layer's code into each framework", "Because every Layer is generated at compile time per framework", "Because all of them are tower Services, and a Layer just wraps a Service", "Because hyper rewrites the layer for each target"],
    "answer": 2,
    "explain": "A Layer only knows it's wrapping some Service. Since app, server, and client are all Services, the same layer composes with any of them - write once, reuse everywhere. That's the basis for tower-http."
  }
]
```


---

# The tower-http Toolbox

Here's the mental model for the whole phase: **`tower-http` is a box of ready-made `Layer`s.** In Phases 3 and 4 you learned the abstract machinery - a `Service` turns a request into a response, a `Layer` wraps a `Service` to make a new one. That was the theory. `tower-http` is where you cash it in. It's a crate full of pre-built `Layer`s that do the things every real HTTP service needs - log requests, handle CORS, compress responses, enforce timeouts - and because they're plain tower `Layer`s, they snap onto *any* HTTP tower service. Your axum app, a tonic gRPC server, a bare hyper service, even an HTTP *client*: same layers, same `.layer(...)` move.

That's the dividend the `Service`/`Layer` abstraction was paying toward the whole time. You don't write a tracing middleware or a CORS handler - you add a crate and hang a layer.

> 📝 This phase is the *applied* counterpart to Phase 4. Phase 4 showed you how `Layer` and `ServiceBuilder` compose middleware in the abstract; here you get the concrete, production-grade layers you'll actually reach for. The exact same crate, and the exact same layers, are what axum users add with `.layer` - see [axum's middleware phase](/guides/axum-from-zero). We're looking at it from underneath.

## Why tower-http works everywhere

`tower-http` is built on the `http` and `http-body` crates - the shared vocabulary types (`Request`, `Response`, body streams) that the whole Rust HTTP ecosystem agrees on. It is *not* built on axum, or hyper, or tonic specifically. It targets the lowest common denominator: "a tower `Service` whose request and response are `http::Request` / `http::Response`."

That's why one `TimeoutLayer` can wrap an axum router today and a hyper-based HTTP client tomorrow. The layer doesn't know or care what's inside - it only speaks `http`.

You add layers à la carte, enabling a Cargo feature per family of layers you want:

```bash
cargo add tower-http --features trace,cors,compression,timeout,limit,fs
```

*What just happened:* each feature flag (`trace`, `cors`, `compression`, …) pulls in one group of layers and nothing else. `tower-http` is heavily feature-gated so you only compile the middleware you actually use. Forgetting the feature is the usual "why won't `tower_http::trace` resolve?" - the module is gated off until you enable its flag.

## The key layers

Here's the toolbox, layer by layer. Each one is a `Layer` you construct and wrap around a service.

- **`TraceLayer::new_for_http()`** - emits request and response events (method, path, status, latency) through the `tracing` crate. The single most useful layer in the box.
- **`CorsLayer`** - adds CORS headers and handles preflight `OPTIONS` requests. `CorsLayer::permissive()` allows everything (fine for local dev); the builder (`CorsLayer::new().allow_origin(...)`) locks it down for production.
- **`CompressionLayer`** - compresses response bodies (gzip, brotli, deflate, zstd) based on the client's `Accept-Encoding`. Its mirror, **`DecompressionLayer`**, transparently decompresses *request* bodies.
- **`TimeoutLayer`** - aborts a request that runs longer than a set `Duration`, returning a `408`-style response instead of hanging forever. (`tower` itself also ships a generic `timeout`; the `tower-http` one is HTTP-aware.)
- **`RequestBodyLimitLayer`** - rejects requests whose body exceeds a byte limit, so a client can't exhaust your memory by streaming a giant upload.
- **`SetResponseHeaderLayer`** - sets (or overrides) a response header on every response, e.g. a `cache-control` or a custom `x-powered-by`.
- **`ServeDir`** - not a layer but a ready-made `Service` that serves static files from a directory. You mount it as a route's service rather than wrapping something.

```rust
use std::time::Duration;
use tower_http::{
    compression::CompressionLayer,
    cors::CorsLayer,
    limit::RequestBodyLimitLayer,
    timeout::TimeoutLayer,
    trace::TraceLayer,
};

let trace = TraceLayer::new_for_http();
let cors = CorsLayer::permissive();
let compress = CompressionLayer::new();
let timeout = TimeoutLayer::new(Duration::from_secs(10));
let body_limit = RequestBodyLimitLayer::new(2 * 1024 * 1024); // 2 MiB
```

*What just happened:* every line constructs a `Layer` value - and nothing more. Constructing a layer does no work; it's just a recipe for how to wrap a service. None of these are attached to anything yet. That's the next step.

## Composing the stack with ServiceBuilder

A real service wants several of these at once, in a deliberate order. As you saw in Phase 4, **`ServiceBuilder`** stacks layers so they read top-to-bottom - the first `.layer(...)` becomes the **outermost** wrapper, the one a request hits first on the way in.

```rust
use tower::ServiceBuilder;
use tower_http::{trace::TraceLayer, compression::CompressionLayer, cors::CorsLayer};

let service = ServiceBuilder::new()
    .layer(TraceLayer::new_for_http())   // outermost: logs everything inside it
    .layer(CompressionLayer::new())
    .layer(CorsLayer::permissive())
    .service(inner);                     // your app / hyper service
```

*What just happened:* `ServiceBuilder::new()` starts an empty stack; each `.layer(...)` wraps another ring around whatever comes after it; `.service(inner)` plugs your actual service in at the center. The result is a *new* `Service` - `inner` wrapped in CORS, wrapped in compression, wrapped in tracing. A request flows in `trace → compress → cors → inner`, and the response unwinds back out in reverse. The ordering matters for real reasons: tracing is outermost so it times and logs *everything*, including the work compression does, and it sees responses *before* they're compressed into an opaque blob. (This is exactly the Phase 4 ordering rule - `ServiceBuilder` reads in execution order; bare chained `.layer()` calls read bottom-up.)

> 💡 Step back and feel the payoff. You wrote zero middleware. You got request tracing, response compression, and CORS - three things every production HTTP service needs - by adding a crate and stacking three values. *That* is what the `Service`/`Layer` abstraction from Phases 3 and 4 buys you: an ecosystem of drop-in, composable middleware that works on anything shaped like an HTTP service.

## ⚠️ TraceLayer logs nothing on its own

The most common surprise with `TraceLayer` - and with any `tracing`-based layer - is that you add it, run your server, hit it with requests, and see **no logs at all**. Nothing is broken.

> ⚠️ `TraceLayer` *emits* `tracing` events; it does not *print* them. The `tracing` crate splits those two jobs deliberately. Until you install a **subscriber** to receive and render events, they go nowhere. The fix is one line at startup:
>
> ```rust
> tracing_subscriber::fmt::init();
> ```
>
> Put that at the top of `main`, before you start serving, and your `TraceLayer` events appear on stderr. No subscriber, no output - every Rust web developer hits this once.

*What just happened:* `tracing_subscriber::fmt::init()` registers a global subscriber that formats events and writes them out. `TraceLayer` was doing its job all along - emitting structured events into the `tracing` system - but with no subscriber listening, there was nobody on the other end of the line. This separation is a feature: the same layer can feed a pretty dev console, structured JSON in production, or an OpenTelemetry pipeline, just by swapping the subscriber.

## The same layers axum hands you

If you've used axum's middleware, none of this is new - and that's the point.

> 💡 When an axum guide tells you to `cargo add tower-http --features trace` and write `.layer(TraceLayer::new_for_http())`, *this is the crate it means*. axum has no middleware of its own. `TraceLayer`, `CorsLayer`, `CompressionLayer`, `TimeoutLayer` - they're `tower-http` layers, and they work on an axum `Router` for one reason only: a `Router` is a tower `Service`, so a tower `Layer` wraps it like any other. The skill transfers in both directions. Anything you learned reaching for layers in axum applies to a bare hyper service here; anything here applies straight back to [axum](/guides/axum-from-zero). One abstraction, one toolbox, everywhere.

## Recap

- **`tower-http` is a box of ready-made `Layer`s** - the concrete payoff of the `Service`/`Layer` abstraction. You add a crate and hang a layer instead of writing middleware.
- It's built on the shared `http`/`http-body` types, not on any one framework, so its layers work on **any HTTP tower service**: axum, tonic, a raw hyper service, even a client.
- The staples: **`TraceLayer`** (request/response tracing), **`CorsLayer`** (CORS + preflight), **`CompressionLayer`**/`DecompressionLayer`, **`TimeoutLayer`**, **`RequestBodyLimitLayer`**, `SetResponseHeaderLayer`, and `ServeDir` for static files. Enable each with a `cargo add tower-http --features ...` flag.
- Compose them with **`ServiceBuilder`** (top-to-bottom = outermost-first, per Phase 4's ordering rule) or, on axum, with `.layer(...)`.
- ⚠️ **`TraceLayer` prints nothing without a `tracing_subscriber`** - install one (`tracing_subscriber::fmt::init()`) at startup, or you'll see no logs.
- 💡 These are the exact layers axum users add with `.layer` - same crate, because a `Router` is just another tower `Service` (see [axum](/guides/axum-from-zero)).

## Quick check

```quiz
[
  {
    "q": "Why can a tower-http layer like TimeoutLayer wrap an axum app, a tonic gRPC server, AND an HTTP client?",
    "choices": ["It has special-cased code for each framework", "It's built on the shared http/http-body types, so it works on any HTTP tower Service", "axum, tonic, and the client all re-export it", "It only works on axum; the others reimplement it"],
    "answer": 1,
    "explain": "tower-http targets the lowest common denominator: a tower Service over http::Request/http::Response. It doesn't know what's inside, so it wraps anything shaped that way."
  },
  {
    "q": "You add TraceLayer::new_for_http(), send requests, and see no logs. What's wrong?",
    "choices": ["TraceLayer is broken; use a different layer", "You forgot to call .layer()", "No tracing_subscriber is installed - TraceLayer emits events but a subscriber must render them", "Tracing only works in release builds"],
    "answer": 2,
    "explain": "TraceLayer emits tracing events; it doesn't print them. Without a subscriber (e.g. tracing_subscriber::fmt::init()) listening, the events go nowhere."
  },
  {
    "q": "In a ServiceBuilder stack, which layer is the outermost - the one a request hits first on the way in?",
    "choices": ["The last .layer() added", "The first .layer() added", "Whichever calls .service()", "Ordering is undefined in ServiceBuilder"],
    "answer": 1,
    "explain": "ServiceBuilder reads top-to-bottom: the first .layer() is the outermost wrapper and runs first inbound. (Bare chained .layer() calls are the reverse - bottom-up.)"
  }
]
```


---

# How axum Uses Them

This is the phase where the magic disappears. You've spent five phases on the bare metal - hyper speaking HTTP on the socket, the `Service` trait, `Layer`s wrapping services, the `tower-http` toolbox. Hold one idea and watch [axum](/guides/axum-from-zero) dissolve into things you already understand.

📝 **An axum app is a `tower::Service` you assemble. Everything else - extractors, `IntoResponse`, the router DSL - is ergonomics layered on top of the raw `Request`/`Response` that hyper deals in.** The friendly façade you learned in the axum guide was never a separate world. It was Service + hyper + Tokio, wearing a nicer coat.

> 💡 If that lands, the rest of this phase is just confirmation. We'll take each axum concept you know and point at the root it's standing on. Router is a Service. Middleware is a Layer. `serve` is hyper. `async` is Tokio. Four mappings, and the framework stops being a black box.

## `axum::Router` IS a `Service`

When you wrote `Router::new().route("/books", get(list_books))` over in [axum from zero](/guides/axum-from-zero), it felt like a special framework object - a routing table with its own rules. Underneath, it's the exact trait from [Phase 3](03-the-service-trait.md):

```rust
// axum's Router implements the tower Service trait:
//   impl Service<Request> for Router {
//       type Response = Response;
//       type Error = Infallible;
//       ...
//   }

let app: Router = Router::new()
    .route("/books", get(list_books))
    .route("/books/{id}", get(get_book));

// Because `app` is a Service, you can call it directly - no HTTP needed:
let response = app.oneshot(request).await.unwrap();
```

*What just happened:* `app` looks like a router, but its type implements `Service<Request, Response = Response>` - the same `poll_ready` + `call` shape every other Service has. Routing, the `get(...)` wrappers, and your handlers are a *convenience layer* that, once compiled, produce a plain Service. axum does the trait gymnastics so your `async fn` handlers slot in, but the result is nothing exotic: a value that turns a `Request` into a `Response`. `oneshot` (from `tower::ServiceExt`) proves it - you can drive the whole app with one request and no socket in sight, which is exactly how axum's own tests work.

## `axum::serve` is hyper

Remember the `TokioIo` + `serve_connection` dance from [Phase 2](02-hyper-the-http-library.md) - accept a TCP connection, wrap the stream, hand it to hyper's HTTP engine along with your service? You don't write that in an axum app. You write one line:

```rust
let listener = tokio::net::TcpListener::bind("0.0.0.0:3000").await.unwrap();
axum::serve(listener, app).await.unwrap();
```

*What just happened:* `axum::serve` is the convenience wrapper over the exact Phase-2 ceremony. It runs Tokio's accept loop on the `listener`, and for each incoming connection it drives your `Router` Service with hyper's HTTP engine - parsing requests off the socket, calling `app`, writing responses back. The two arguments say it plainly: a Tokio listener (connections) and a Service (your app). `serve` is the glue between them, and that glue *is* hyper. Nothing in this line is axum-specific machinery; it's the bare server from Phase 2 with the boilerplate folded away.

## `.layer` is a tower `Layer`

In the axum middleware phase you wrapped your router with `.layer(TraceLayer::new_for_http())` and stacked layers with `ServiceBuilder`. Those are not "axum middleware." They're the [Phase 4](04-layers-and-middleware.md) `Layer` trait and the [Phase 5](05-tower-http.md) `tower-http` crate, unchanged:

```rust
use tower_http::trace::TraceLayer;
use tower_http::timeout::TimeoutLayer;
use std::time::Duration;

let app = Router::new()
    .route("/books", get(list_books))
    .layer(TraceLayer::new_for_http())      // a tower Layer (Phase 5)
    .layer(TimeoutLayer::new(Duration::from_secs(10)));

// And custom middleware via from_fn is just sugar that produces a Layer:
// axum::middleware::from_fn(require_auth)  ->  impl tower::Layer
```

*What just happened:* `.layer(...)` on a Router applies a `tower::Layer` - it wraps your Service in another Service, the onion you built by hand in Phase 4. `TraceLayer` and `TimeoutLayer` are the *same* layers from `tower-http` in Phase 5; they don't know or care that axum exists, because they operate on the `Service` trait, not on axum. Even `axum::middleware::from_fn` is sugar: it takes your `async fn(Request, Next)` and produces a Layer. "axum middleware" is tower middleware with a friendlier on-ramp - which is exactly why the same layers also wrap HTTP clients and gRPC services.

## Extractors and `IntoResponse` are typed sugar over hyper

Here's the one piece that feels most like magic, and it's the smallest trick of all. hyper deals in a raw `Request` (headers and a stream of body bytes) and a raw `Response`. You almost never touch those in axum. Instead:

```rust
use axum::{extract::{Path, Json, State}, response::IntoResponse, http::StatusCode};

async fn get_book(
    State(db): State<Db>,   // pulled from app state
    Path(id): Path<u32>,    // parsed from the URL path
) -> impl IntoResponse {
    match db.find(id) {
        Some(book) => (StatusCode::OK, Json(book)),         // -> Response
        None => (StatusCode::NOT_FOUND, "no such book").into_response(),
    }
}
```

*What just happened:* the arguments are **extractors**. Before your function runs, axum reads the raw hyper `Request` and turns pieces of it into typed values - `Path<u32>` parses `id` out of the URL, `Json<T>` would deserialize the body, `State` hands you shared app state. On the way out, your return type implements **`IntoResponse`**, so axum turns `(StatusCode, Json(book))` back into a real hyper `Response` - status line, `content-type: application/json` header, serialized body. You wrote typed Rust; axum did the byte-level translation in both directions. Extractors are the request side of the convenience, `IntoResponse` is the response side, and together they're why you never hand-parse headers or hand-build a `Response` the way the bare hyper handler in Phase 2 had to.

## The whole stack, top to bottom

Every layer you've met in this guide stacks into one picture. From your code down to the runtime:

```mermaid
flowchart TD
  H[Your handlers<br/>async fn + extractors] --> R[axum Router<br/>a tower Service]
  R --> L[tower Layers<br/>TraceLayer, Timeout, auth]
  L --> HY[hyper<br/>HTTP over the socket]
  HY --> TK[Tokio<br/>the async runtime]
```

*What just happened:* a request enters at the bottom - Tokio accepts the connection, hyper parses the HTTP, the Layers wrap inward, the Router routes to your handler, and the response travels back out the same path. Read top-down, it's also your authoring experience: you write handlers, axum assembles them into a Service, you wrap Layers around it, and `axum::serve` plugs that into hyper on Tokio. Same stack, two directions.

💡 So "learning axum" was really learning a friendly façade over `Service` + hyper + Tokio. Every axum concept you know maps to something in this guide: **Router = Service. Middleware = Layer. `serve` = hyper. `async` = Tokio.** The framework didn't invent a new universe - it gave ergonomic names to the roots you've now seen directly. That's the whole payoff: nothing in axum is magic, and you can read its source the same way you'd read your own.

## Recap

- **`axum::Router` is a `tower::Service`** (`Service<Request, Response = Response>`). The routing DSL, `get(...)` wrappers, and `async fn` handlers compile down to one plain Service - provable with `oneshot`, no socket required.
- **`axum::serve(listener, app)` is hyper** - the convenience wrapper over the Phase-2 `TokioIo` + `serve_connection` dance: Tokio accepts connections, hyper drives your Router Service for each one.
- **`.layer(...)` applies a `tower::Layer`** - the same `Layer` trait from Phase 4 and the same `tower-http` layers from Phase 5. `axum::middleware::from_fn` is sugar that produces a Layer.
- **Extractors and `IntoResponse` are typed sugar over hyper's raw `Request`/`Response`** - axum parses the request into typed arguments and turns your return value into a `Response`, so you never touch the bytes.
- **The full stack:** your handlers → axum Router (a Service) → tower Layers → hyper (HTTP on the socket) → Tokio (the runtime). Router = Service, middleware = Layer, `serve` = hyper, `async` = Tokio.

## Quick check

```quiz
[
  {
    "q": "What is an axum::Router, underneath the routing DSL?",
    "choices": ["A tower::Service that turns a Request into a Response", "A hyper connection handler with its own trait", "A macro that generates route-matching code at compile time", "A Tokio task that loops over incoming requests"],
    "answer": 0,
    "explain": "Router implements Service<Request, Response = Response>. The route(...) and get(...) calls are ergonomics that compile down to a plain Service - you can even drive it with oneshot and no socket."
  },
  {
    "q": "What is `axum::serve(listener, app)` actually doing?",
    "choices": ["Compiling the router into a standalone binary", "Wrapping the Phase-2 hyper dance: Tokio accepts connections and hyper drives the Router Service for each", "Registering routes in a global table that hyper reads later", "Starting a separate process per request"],
    "answer": 1,
    "explain": "axum::serve is the convenience wrapper over the TokioIo + serve_connection ceremony from Phase 2 - Tokio runs the accept loop, and hyper's HTTP engine drives your Router Service connection by connection."
  },
  {
    "q": "When you call `.layer(TraceLayer::new_for_http())` on a Router, what kind of thing is TraceLayer?",
    "choices": ["An axum-specific middleware type that only works with Router", "A Tokio runtime hook", "The same tower Layer from tower-http used everywhere else - it wraps the Service", "A hyper request parser"],
    "answer": 2,
    "explain": "TraceLayer is a plain tower::Layer from tower-http (Phase 5). .layer applies it to wrap your Service, and because it operates on the Service trait - not on axum - the same layer also works with HTTP clients and gRPC."
  }
]
```


---

# Where to Go Next

Take a second to notice how far you've come. When you started, the bottom of the Rust web stack was a wall of names you used but couldn't picture: `axum::serve`, "tower middleware," `tower-http`, hyper somewhere underneath. Now you can name every one of them. **hyper** speaks HTTP on the socket. A **`Service`** is the universal shape - request in, response out, plus a readiness check. A **`Layer`** wraps a `Service` to make a new one, and that's all middleware ever was. `tower-http` is a box of ready-made layers. And axum's `Router` is itself a `Service`, with your layers wrapped around it, handed to hyper to drive.

That's not a small thing. This was the deepest **roots** guide in the Rust set, and the wall is gone. What's left in this last phase isn't new mechanism - it's showing you how far that one model reaches. The `Service`/`Layer` abstraction wasn't built for axum, or even for servers. It works across the whole ecosystem, and once you see that, a lot of other libraries stop being separate things to learn.

## The same layers, on gRPC

You don't only build REST APIs in Rust. When you reach for **gRPC**, the library is **tonic** - and tonic is built on the exact same foundation you just learned: **tower + hyper**.

That has a concrete, lovely consequence. The layers you wrote for an axum server - tracing, timeouts, authentication - are tower `Layer`s, and a tonic gRPC server is also a tower `Service`. So the same layers compose onto your gRPC server too. You don't learn a second middleware system for gRPC. You learn the protocol differences (which the [gRPC guide](/guides/grpc-explained) covers), and the middleware story is one you already know.

> 💡 This is the quiet reward of learning roots instead of one framework: tonic looked like a whole new world from the outside. From here, it's "the same `Service` and `Layer`, speaking gRPC instead of REST."

## The same model, pointed outward

Here's the idea that tends to surprise people, and it's where the `poll_ready` backpressure check from [Phase 3](03-the-service-trait.md) finally earns its keep.

So far we've talked about a `Service` as the thing that *receives* a request. But a `Service` is just "async request → response" - and an outbound HTTP or gRPC **client** fits that shape exactly. The request is the one you're *sending*; the response is the one you get back. So a client is a `Service` too.

The moment a client is a `Service`, every tower `Layer` composes onto the requests you send out, the same way they wrap requests coming in. tower ships a toolbox of client-side layers:

- **`tower::retry`** - retry a failed request, with a policy you control.
- **`tower::timeout`** - give up on a slow backend instead of hanging.
- **`tower::limit`** - cap concurrency or rate so you don't overwhelm a downstream service.
- **`tower::load` / `tower::balance`** - spread requests across several backend instances (load balancing).
- **`tower::buffer`** - queue requests and hand the client out to many callers.

This is where `poll_ready` stops being abstract. A load balancer needs to ask each backend "are you ready?" before sending. A concurrency limiter needs to say "not right now" when it's full. That readiness check - the half of the `Service` trait that felt like ceremony in Phase 3 - is exactly the hook that makes retry, rate limiting, and load balancing possible. Backpressure was always the point.

On the HTTP side, **hyper** has a client of its own (in `hyper-util`), and the friendly `reqwest` library sits on top of it - so when you drop a level under `reqwest`, you land back on hyper, in territory you now recognize.

## One model, the whole stack

Step back and look at the shape. The same `Service`-wrapped-by-`Layer`s picture describes both ends of the wire:

```mermaid
flowchart LR
  C[Incoming request] --> SL[Server Layers<br/>tracing · auth · timeout]
  SL --> SVC[Your Service<br/>axum / tonic]
  SVC --> CL[Client Layers<br/>retry · timeout · balance]
  CL --> CS[Client Service]
  CS --> B[Backend / DB / API]
```

Requests arrive, pass through server layers into your `Service`; when that `Service` needs to call out, the request leaves through client layers into a client `Service`. Same trait, same wrapping, both directions.

> 💡 Hold onto this one sentence and it covers an enormous amount of Rust: **a `Service` turns a request into a response, a `Layer` wraps a `Service`, hyper drives it, and Tokio runs it.** Servers, clients, gRPC, proxies - they're all that same model wearing different clothes.

## What to actually do next

You don't lock this in by reading more. You lock it in by going back and *recognizing*.

- **Revisit [axum](/guides/axum-from-zero) with new eyes.** Open a project you've built and find the pieces: the `Router` that's a `Service`, the `.layer(...)` calls that are tower `Layer`s, the `tower-http` middleware, `axum::serve` calling into hyper. Nothing should be opaque now. That feeling - "oh, I know what every line is" - is the whole guide paying off at once.
- **Add one client-side layer.** Take an HTTP call your app makes and wrap it with a `tower::timeout` or a small `tower::retry`. Watch a `Service` you *send* requests through behave like the ones you *receive* them through. It's the fastest way to feel that the model really is symmetric.
- **Read tonic if gRPC is in your future.** When you do, you'll find the layers already familiar - only the protocol is new.

And here's the thing to carry out of all of it. You started this set wanting to understand the runtime ([Tokio](/guides/tokio-the-async-runtime)), then the HTTP and middleware (hyper and tower), then the framework on top. You have all three now. No part of a Rust web stack has to be a black box to you again - and that's a rare, durable kind of confidence. The next time something deep in the stack shows up in a stack trace, you won't flinch. You'll know exactly what it is and where to look.

## Recap

1. **The wall is gone.** hyper speaks HTTP, a `Service` is "async request → response," a `Layer` wraps a `Service` (that's middleware), `tower-http` is a box of ready-made layers, and axum is a `Service` you assembled - driven by hyper, run by Tokio.
2. **tonic = the same model, on gRPC.** tonic is built on tower + hyper, so a gRPC server is a tower `Service` and your existing tracing/timeout/auth `Layer`s compose onto it. No second middleware system to learn.
3. **A client is a `Service` too.** Point the model outward and tower's client-side layers - `retry`, `timeout`, `limit`, `load`/`balance`, `buffer` - wrap the requests you *send*. hyper has its own client (under `reqwest`).
4. **`poll_ready` was the point.** The readiness check from Phase 3 is exactly what makes load balancing, rate limiting, and retries possible. Backpressure isn't ceremony.
5. **One sentence covers the stack.** A `Service` turns a request into a response, a `Layer` wraps a `Service`, hyper drives it, Tokio runs it - servers, clients, gRPC, and proxies alike.

## Quick check

A last look at how far one model stretches:

```quiz
[
  {
    "q": "tonic (gRPC) is built on tower + hyper. What does that mean for your middleware?",
    "choices": [
      "The same tower Layers (tracing, timeout, auth) compose onto a gRPC server, because it's a tower Service",
      "gRPC needs a completely separate middleware system you have to learn from scratch",
      "Middleware doesn't work with gRPC at all",
      "You must rewrite every layer in a gRPC-specific dialect"
    ],
    "answer": 0,
    "explain": "A tonic gRPC server is a tower Service, so the Layers you already wrote for an axum server wrap it the same way. The protocol differs; the middleware story is one you already know."
  },
  {
    "q": "Why can tower layers like retry, timeout, and load-balancing apply to an outbound client?",
    "choices": [
      "Because a client is a Service too - async request out, response back - so Layers wrap the requests you send",
      "Because clients secretly run their own hidden server",
      "Because tower copies the server's layers onto the client automatically at startup",
      "They can't - client-side layers don't exist"
    ],
    "answer": 0,
    "explain": "An outbound HTTP/gRPC client fits the Service shape exactly: the request is the one you're sending. Once it's a Service, every tower Layer composes onto it, including retry, timeout, limit, and load/balance."
  },
  {
    "q": "Which single sentence best captures the whole Rust web stack you've now learned?",
    "choices": [
      "A Service turns a request into a response, a Layer wraps a Service, hyper drives it, and Tokio runs it",
      "axum is the only abstraction that matters and everything else is internal detail",
      "hyper handles middleware while tower opens the network socket",
      "gRPC and HTTP each require their own unrelated runtime and middleware model"
    ],
    "answer": 0,
    "explain": "That one sentence covers servers, clients, gRPC, and proxies: the Service/Layer model is the request-to-response shape, hyper drives it over the socket, and Tokio is the runtime underneath."
  }
]
```
