# WSGI & ASGI Explained

> Learn the protocol every Python web framework sits on: what WSGI is and the problem it solved, a bare WSGI app with no framework, WSGI servers and middleware, why ASGI exists for async, a bare ASGI app and its servers, and how Flask, Django, and FastAPI are all built on these contracts. The bottom layer, made visible.


---

# WSGI & ASGI Explained

Underneath Flask, Django, FastAPI - every Python web framework - there's one small contract that lets a
web server and your Python code talk to each other. For synchronous frameworks it's **WSGI**; for async
ones it's **ASGI**. Almost nobody learns these directly, which is exactly why "how does `flask run`
actually serve a request?" feels like magic. It isn't: a WSGI app is a plain Python *callable* the server
invokes, and an ASGI app is an async callable. Once you've written each by hand - in a dozen lines, with
no framework - every framework reads as conveniences over that callable.

This is a **roots guide**. You'll rarely write raw WSGI/ASGI apps in a real job (frameworks exist for
good reasons), but understanding them demystifies a pile of things at once: why you run `gunicorn` or
`uvicorn` in production, what "middleware" really is, why FastAPI is async and Flask isn't, and how the
whole Python web stack fits together. We build it bare-metal first, then point at the frameworks and watch
the shapes line up.

> 📝 This assumes **Python** (functions, async/await - [Python From Zero](/guides/python-from-zero)) and a
> grasp of **HTTP** ([HTTP, Explained](/guides/http-explained)). It's the Python parallel to
> [The Servlet API](/guides/the-servlet-api) (Java's equivalent foundation) and is most useful after
> you've used [Flask](/guides/flask-from-zero) or [FastAPI](/guides/fastapi-from-zero) and want to see
> beneath them. Examples are shown with their output rather than run on the page.

## How to read this

Short and foundational - read in order. It builds a bare WSGI app, then a bare ASGI app, then maps both
onto the frameworks you know. Phases carry difficulty badges.

## The phases

1. **[What WSGI Is](01-what-wsgi-is.md)** 🟢 - the contract between a web server and a Python app, and the problem it solved.
2. **[A WSGI App From Scratch](02-a-wsgi-app-from-scratch.md)** 🟡 - write one in a dozen lines with no framework, and see what Flask *is* underneath.
3. **[The WSGI Server & Middleware](03-the-wsgi-server-and-middleware.md)** 🟡 - gunicorn/uWSGI run your app; middleware wraps it - the root of framework "middleware."
4. **[Why ASGI Exists](04-why-asgi-exists.md)** 🔴 - WSGI is synchronous; async (websockets, high concurrency) needs a new contract.
5. **[An ASGI App & the Servers](05-an-asgi-app-and-the-servers.md)** 🟡 - a bare ASGI app, uvicorn, and how FastAPI/Starlette sit on top.
6. **[From Protocol to Framework](06-from-protocol-to-framework.md)** 🟢 - Flask = a WSGI app, FastAPI = an ASGI app; the whole stack, seen.

> Once you've seen the callable, "a Python web framework" reads as "conveniences over a WSGI or ASGI
> app." The magic was always this contract.


---

# What WSGI Is

Every Python web app you'll ever run in production - a Flask side project, a sprawling Django monolith,
a tiny internal tool - is standing on the same foundation. Strip away the routing decorators and the ORM
and the templating, and underneath there's a single, almost embarrassingly small agreement: a way for a
web server to hand your code a request and get a response back. That agreement is **WSGI**, and this guide
is about the bedrock it sits on.

This is the roots guide. You'll rarely write raw WSGI by hand at work - frameworks exist precisely so you
don't have to - but every Python web framework is a convenience layer over the thing you're about to meet.

## The problem WSGI solved

Let's start with the world *before* WSGI, because the contract makes no sense until you've felt the pain it
removed.

📝 **The pre-WSGI mess** - in the early 2000s, every Python web framework talked to web servers its own
way. A framework written for one server (say, mod_python on Apache) wouldn't run on another without
glue code. Servers and frameworks weren't interchangeable: pick a framework and you were often locked into
a particular way of deploying it. There was no shared plug.

The fix was to agree on one standard plug shape, written down so everyone could build to it.

📝 **WSGI** (Web Server Gateway Interface, defined by **PEP 3333**) - the standard contract between a
Python **web server** (gunicorn, uWSGI, mod_wsgi) and a Python **web application** (Flask, Django, or
something you write by hand). Build your app to the WSGI interface once, and it runs on *any* WSGI-compliant
server. Pick a different server next year - your app doesn't change.

> 💡 **Key point.** WSGI is not a server, not a framework, not a library you install. It's an *interface* -
> an agreement about function shapes. Its whole job is to make Python web apps and Python web servers
> mix-and-match, the way a standard wall socket lets any appliance run off any outlet.

## The contract: a callable

So what does this "standard plug" actually look like? It's smaller than you'd guess.

📝 **WSGI application** - a **callable** (a function, or any object with a `__call__` method) with this exact
signature:

```python
def app(environ, start_response):
    ...
```

That's the entire interface. Three moving parts:

- **`environ`** - a plain dict describing the incoming request: the HTTP method, the path, the headers, the
  query string, the request body stream, and more. The server fills this in and hands it to you.
- **`start_response`** - a function the server gives you. You *call* it with the status line and the response
  headers when you're ready to begin replying.
- **the return value** - you return the response **body** as an iterable of `bytes` (commonly a list with
  one bytestring in it).

```python
def app(environ, start_response):
    # 1. read the request from environ (method, path, headers...)
    # 2. call start_response(status, headers) to set the reply line
    # 3. return an iterable of bytes as the body
    ...
```

*What just happened:* this isn't a working app yet - it's the *shape* every WSGI app shares, and the whole
contract fits in three comments. The server promises to call you with `environ` and `start_response`; you
promise to call `start_response` once and return some bytes. Phase 2 fills in the body and runs it for real.
For now, sit with how little there is: no base class to inherit, no framework to import. A WSGI app is just
a callable that follows the rules.

## Server vs app: who does what

Here's the part that trips people up first. You write the app callable - but you never *call* it yourself.
Something else owns that job.

📝 **WSGI server** (gunicorn, uWSGI, Waitress) - a running program that accepts incoming TCP connections,
parses the raw HTTP off the wire, packs the request into an `environ` dict, and then **calls your app
callable** - once per request. It takes whatever you return and serializes it back into a proper HTTP
response over the socket. The server is the thing that's actually running; your app is code it invokes when
a request arrives.

Picture the division of labor:

```mermaid
flowchart LR
  B[Browser] -->|raw HTTP over TCP| S[WSGI server]
  S -->|builds environ, calls callable| A["your WSGI app(environ, start_response)"]
  A -->|status, headers, body bytes| S
  S -->|HTTP response| B
```

*What just happened:* the browser opens a connection and sends bytes. The **server** does every piece of
unglamorous plumbing - accepting the socket, reading and parsing the HTTP, building the tidy `environ` dict,
calling your callable, and afterward turning your status/headers/body back into HTTP and shipping it. Your
app sits in the middle and does the one interesting part: looks at the request, decides what to send back.

💡 **The server does the grunt work; you write the handler.** You don't manage sockets, parse headers, or
speak HTTP/1.1 by hand - the server hands you a parsed request and waits for your response. In exchange, you
write your logic to *its* shape: a callable it knows how to invoke. The server owns the loop and calls your
code, not the other way around - that's the inversion of control at the heart of Python web.

## Where your frameworks fit

Here's the reveal that justifies the whole guide. That `app(environ, start_response)` callable? It's not a
relic you'll skip past on the way to "real" frameworks. It *is* what the real frameworks are made of.

When you write a Flask app and create `app = Flask(__name__)`, that `app` object **is a WSGI application** -
a callable with exactly the `(environ, start_response)` signature, with all the routing and request parsing
hidden inside its `__call__`. When you run `flask run` in development or `gunicorn myapp:app` in production,
the WSGI *server* is calling that *app* - the same handshake from the diagram above.

- **Flask** - `app` is the WSGI callable; the `@app.route(...)` decorators are a routing layer it consults
  inside `__call__` to decide which view function handles the request. See
  [/guides/flask-from-zero](/guides/flask-from-zero).
- **Django** - `get_wsgi_application()` (the thing in your project's `wsgi.py`) returns the WSGI callable
  gunicorn invokes. Django's middleware, URL resolver, and views all live behind it.
- **Frameworks are conveniences over the callable.** Routing, request objects, templating, sessions - the
  framework adds ergonomics, but the request still enters through the server and lands on one WSGI callable.

💡 If this rings a bell, it should: it's the same idea as Java's Servlet API
([/guides/the-servlet-api](/guides/the-servlet-api)), where a servlet container (Tomcat) calls a servlet
object that handles one request. Same architecture - a standard interface between server and app - in a
different ecosystem. Python calls its version WSGI; Java calls its version the Servlet API. The shape is the
same: the server runs, your handler gets called. (If HTTP status codes and headers are fuzzy,
[/guides/http-explained](/guides/http-explained) is the companion read.)

## Why learn it

Let's look straight at the trade-off, because this project's whole voice is anti-hand-waving.

⚠️ **You will rarely write a raw WSGI app at a real job - and that's fine.** Frameworks exist for good
reasons: they spare you boilerplate and give you routing, sessions, and validation for free. Reaching for
raw WSGI when Flask would do is usually a mistake, not a badge of honor. So this guide is *not* arguing you
should hand-write WSGI apps in production.

It's arguing something more useful: knowing this layer is what makes the layers above it stop being magic.
Once you've seen the callable underneath, a pile of mysteries resolve at once -

- **Why you run gunicorn in prod** - because Flask's dev server is a WSGI server too, but a slow,
  single-purpose one; gunicorn is a production-grade server that calls the very same app callable.
- **What middleware actually is** - a WSGI app that wraps another WSGI app, intercepting `environ` on the way
  in and the response on the way out. (We'll build one.)
- **Why ASGI exists** - WSGI's one-call-per-request shape is synchronous to the bone, which is exactly why
  async Python needed a different contract. That's the rest of this guide.

The fastest way to make all of that concrete is to write a WSGI app by hand - no framework, just the
callable - and run it. That's Phase 2.

## Recap

1. **The problem:** before WSGI, every Python framework talked to servers differently, so frameworks and
   servers weren't interchangeable. WSGI (**PEP 3333**) is the standard contract that fixed it.
2. A **WSGI application** is a **callable** with the signature `app(environ, start_response)`. That's the
   whole interface - no base class, no framework required.
3. `environ` is a dict describing the request; `start_response` is a function you call with the status and
   headers; you **return the body as an iterable of bytes**.
4. The **WSGI server** (gunicorn, uWSGI) handles sockets and HTTP parsing and **calls your app** once per
   request. The server does the grunt work; you write the handler.
5. A **Flask or Django app *is* a WSGI application** - `flask run` and `gunicorn` are the WSGI servers
   calling it. Frameworks are conveniences over this callable, mirroring Java's Servlet API.
6. ⚠️ You'll rarely write raw WSGI at work - but it explains why you run gunicorn, what middleware is, and
   (next) why ASGI exists.

## Quick check

Three questions on the ideas that have to stick before Phase 2:

```quiz
[
  {
    "q": "What is WSGI, in one sentence?",
    "choices": [
      "The standard contract (PEP 3333) between a Python web server and a Python web app, so any compliant app runs on any compliant server",
      "A Python web server you install and run, like gunicorn",
      "A web framework that competes with Flask and Django",
      "The HTTP protocol, reimplemented in Python"
    ],
    "answer": 0,
    "explain": "WSGI is an interface - an agreement about function shapes - not a server, framework, or protocol. Its job is to make Python apps and servers mix-and-match. gunicorn is a server that speaks WSGI; Flask is a framework whose app object is a WSGI callable."
  },
  {
    "q": "What is the signature of a WSGI application callable?",
    "choices": [
      "app(environ, start_response) - environ is the request dict, start_response sets status/headers, and you return the body as bytes",
      "app(request, response) - two framework objects you mutate in place",
      "app() with no arguments, returning an HTML string",
      "handle(socket) - you read and write the raw TCP socket yourself"
    ],
    "answer": 0,
    "explain": "A WSGI app is a callable taking environ (a dict describing the request) and start_response (a function you call with the status line and headers), and it returns the response body as an iterable of bytes. That's the entire contract."
  },
  {
    "q": "When you run `gunicorn myapp:app`, who calls whom?",
    "choices": [
      "The WSGI server (gunicorn) calls your app callable once per request, after parsing the HTTP and building environ",
      "Your app starts gunicorn and hands it requests to serve",
      "The browser calls your app directly, with no server in between",
      "gunicorn and your app each parse half of the HTTP request"
    ],
    "answer": 0,
    "explain": "The server owns the running process and all the HTTP plumbing: it accepts the socket, parses the request into environ, and calls your app callable, then ships whatever you return. The server does the grunt work; your app just handles the request."
  }
]
```


---

# A WSGI App From Scratch

In Phase 1 you learned the contract: a WSGI app is just a Python *callable* the server invokes with two
arguments, and it hands back a response. This phase makes it real. We're going to write a complete,
working web app - one a browser can hit - with no Flask, no Django, no framework of any kind. Just a
function.

The mental model to hold the whole way through: **a WSGI app is one function the server calls per request.**
The server hands it everything it knows about the incoming request (in a dict called `environ`), gives it a
callback to set the status and headers (`start_response`), and expects back the body as bytes. Match that
shape and you have a web app. Everything a framework does - routing, request objects, response helpers - is
*convenience layered over this one function*. Once you've written the function yourself, the frameworks stop
being magic and start being ergonomics.

## The simplest WSGI app

Here's the whole thing - a web app you can run right now:

```python
from wsgiref.simple_server import make_server

def app(environ, start_response):
    start_response("200 OK", [("Content-Type", "text/plain")])
    return [b"Hello"]

if __name__ == "__main__":
    server = make_server("localhost", 8000, app)
    print("Serving on http://localhost:8000")
    server.serve_forever()
```

Save it as `bare.py` and run it:

```bash
python bare.py
```

Then in another terminal (or your browser at `http://localhost:8000`):

```console
$ curl http://localhost:8000
Hello
```

*What just happened:* `app` is the WSGI callable - a plain function taking `environ` and `start_response`.
When a request comes in, the server calls `app(environ, start_response)`. Inside, we call `start_response`
with the status line (`"200 OK"`) and a list of header tuples, then return the body as a list containing one
bytestring, `b"Hello"`. The `make_server(...)` line uses **`wsgiref.simple_server`** - a tiny WSGI server
that ships *with Python itself*, no install needed - to wire that function to a real TCP socket on port 8000.

💡 Read that again, because it's the whole point of this guide: **this is a complete web app.** It accepts
HTTP requests and returns HTTP responses. Flask, Django, FastAPI - every one of them, stripped to the studs,
is a function shaped exactly like `app` above. Everything else those frameworks ship is convenience built on
top of this. You are looking at the foundation.

## Reading the request from `environ`

So far `app` ignores its input entirely - it says "Hello" no matter what you ask for. To actually *respond*
to the request, you read it from `environ`.

📝 **`environ` is a plain dict** the server fills in with everything about the incoming request. The keys
you'll reach for constantly:

| Key | Holds | Example value |
|-----|-------|---------------|
| `environ["REQUEST_METHOD"]` | the HTTP verb | `"GET"`, `"POST"` |
| `environ["PATH_INFO"]` | the URL path | `"/notes"` |
| `environ["QUERY_STRING"]` | everything after `?` | `"q=milk&sort=date"` |
| `environ["HTTP_USER_AGENT"]` | a request header | `"curl/8.4.0"` |
| `environ["wsgi.input"]` | the request body, as a stream | a file-like object |

Two things worth pinning down. **Headers arrive as `HTTP_*` keys**: the server takes each incoming header,
uppercases it, swaps dashes for underscores, and prefixes `HTTP_`. So `User-Agent` becomes
`environ["HTTP_USER_AGENT"]`, `Accept-Language` becomes `environ["HTTP_ACCEPT_LANGUAGE"]`. And **the body is
a stream, not a string** - you read it from `environ["wsgi.input"]` (using the byte length in
`environ["CONTENT_LENGTH"]`), and what you get back is bytes.

Here's `app` actually looking at the request:

```python
def app(environ, start_response):
    method = environ["REQUEST_METHOD"]
    path = environ["PATH_INFO"]
    start_response("200 OK", [("Content-Type", "text/plain")])
    return [f"You sent a {method} to {path}".encode("utf-8")]
```

*What just happened:* we pulled the verb and path straight out of `environ`, built a response string from
them, and `.encode("utf-8")`'d it into bytes before returning. Hit `http://localhost:8000/notes` and you get
back `You sent a GET to /notes`. That's it - that's "reading the request." There's no `request` object with
friendly attributes here; there's a dict, and you fish what you need out of it by key.

## Routing by hand

A real app does different things at different URLs. With no framework, you do that the most direct way
imaginable: look at `PATH_INFO` and branch.

```python
def app(environ, start_response):
    path = environ["PATH_INFO"]

    if path == "/":
        status, body = "200 OK", b"Welcome home"
    elif path == "/notes":
        status, body = "200 OK", b"Here are your notes"
    else:
        status, body = "404 Not Found", b"No such page"

    start_response(status, [("Content-Type", "text/plain")])
    return [body]
```

*What just happened:* we read the path once, then an `if`/`elif`/`else` decides the status and body. A request
to `/` returns "Welcome home"; `/notes` returns the notes line; anything else falls through to a real
`404 Not Found`. This little dispatcher - one place that reads the path and routes to the right response - is
called a **front controller**, and you just wrote one by hand.

💡 Look hard at that `if`/`elif` chain, because it's the thing a framework's URL router *automates*. When you
write `@app.route("/notes")` in Flask, Flask is maintaining this exact branching for you under the hood -
matching `PATH_INFO` against registered routes and calling the right function. You're doing manually,
explicitly, what a router does generically. Same idea; the framework just hides the `if` ladder behind a
decorator.

## Returning the response correctly

The WSGI return contract is strict, and getting it slightly wrong is the #1 way a from-scratch app blows up.
⚠️ Three rules, no exceptions:

1. **Status is a string** - `"200 OK"`, including the number *and* the reason phrase. Not `200`, not `"200"`.
2. **Headers are a list of `(str, str)` tuples** - `[("Content-Type", "text/plain")]`. Both sides strings.
3. **The body is an iterable of *bytes*** - `[b"Hello"]`, not `"Hello"`. Strings must be `.encode()`'d first.

Here's a correct response next to the two mistakes everyone makes:

```python
# CORRECT
start_response("200 OK", [("Content-Type", "text/plain")])
return [b"Hello"]

# BROKEN #1 - body is a str, not bytes
start_response("200 OK", [("Content-Type", "text/plain")])
return ["Hello"]          # TypeError: a bytes-like object is required

# BROKEN #2 - forgot to encode a built string
start_response("200 OK", [("Content-Type", "text/plain")])
name = "Ada"
return [f"Hello {name}"]  # same TypeError - it's still a str
```

*What just happened:* the correct version returns a list holding one bytestring. Both broken versions return
*strings*, and the server raises a `TypeError` because WSGI demands bytes on the wire - HTTP bodies are bytes,
and WSGI refuses to guess an encoding for you. ⚠️ The fix is always the same: `.encode("utf-8")` any string
before returning it, and write byte literals (`b"..."`) for constants. A second classic trip-up: returning
the bytestring *bare* instead of wrapped in a list - `return b"Hello"` technically works because a bytestring
is iterable, but the server then iterates it *one byte at a time*, which is almost never what you want. Return
an iterable *of* byte chunks: `[b"Hello"]`.

## What the framework adds

You now have, by hand, every moving part of a web app: a callable the server invokes, request data read from
`environ`, routing by branching on the path, and a correctly-shaped bytes response. So what does a framework
like [Flask](/guides/flask-from-zero) actually *give* you? Map it piece for piece against what you just wrote:

| You wrote, by hand | Flask gives you | What it's doing underneath |
|--------------------|-----------------|----------------------------|
| `if path == "/notes": ...` | `@app.route("/notes")` | the same `PATH_INFO` branching, registered for you |
| `environ["REQUEST_METHOD"]`, `environ["HTTP_*"]` | `request.method`, `request.headers` | a friendly object wrapping `environ` |
| `start_response(...)` + `return [body.encode()]` | `return "some html"` | builds the status, headers, and bytes body for you |
| your `def app(environ, start_response)` | Flask's `app` object | Flask's `app` **is** a WSGI callable too |

💡 That last row is the kicker. The Flask `app` you create with `Flask(__name__)` is itself a WSGI callable -
it has the same `(environ, start_response)` shape your function has, which is *exactly* why a WSGI server can
run it. When you saw `request.method` and `request.form` over in
[Routing & Views](/guides/flask-from-zero), you were looking at a polished wrapper over the same `environ`
dict you just read by hand. `@app.route` is the `if` ladder. `return "html"` is the `start_response` plus the
`.encode()`. None of it is new machinery - it's ergonomics layered over the function you already understand.

You wrote the core. A framework is the comfortable seat bolted on top. Next phase we look at the *other* half
of running a real app: the **WSGI server** that actually invokes your callable in production (you won't ship
`wsgiref` - you'll reach for gunicorn or uWSGI), and **middleware**, the trick of wrapping one WSGI app in
another - which turns out to be the root of every framework's "middleware" feature.

## Recap

1. **A WSGI app is one function** the server calls per request: `def app(environ, start_response)`. With
   `wsgiref.simple_server` (built into Python) that function is already a complete, runnable web app - no
   framework required.
2. **`environ` is a plain dict** holding the request: `REQUEST_METHOD`, `PATH_INFO`, `QUERY_STRING`, headers
   as `HTTP_*` keys, and the body as a stream at `environ["wsgi.input"]`.
3. **Routing by hand** is branching on `PATH_INFO` - a front controller. 💡 A framework's URL router automates
   exactly this `if`/`elif` ladder behind `@app.route`.
4. ⚠️ **The response contract is strict:** status is a string (`"200 OK"`), headers are a list of `(str, str)`
   tuples, and the body is an iterable of *bytes* (`[b"..."]`). Returning a `str` is the most common error -
   `.encode("utf-8")` first.
5. **A framework is ergonomics over this callable:** `@app.route` = manual path branching, `request` = a
   wrapper over `environ`, `return "html"` = building the bytes body and headers. Flask's `app` object *is* a
   WSGI callable - same shape as the one you wrote.

You can now read a WSGI app and see straight through to the function underneath. Next: the server that runs it
in the real world, and middleware.

## Quick check

Make sure the bare callable stuck:

```quiz
[
  {
    "q": "Your WSGI app does `return \"Hello\"` (a plain string). What happens?",
    "choices": [
      "It errors - the WSGI body must be an iterable of bytes, so you need `return [b\"Hello\"]`",
      "It works fine; WSGI encodes strings to bytes automatically",
      "It returns an empty response because strings aren't allowed",
      "It works but sends the wrong Content-Type header"
    ],
    "answer": 0,
    "explain": "WSGI requires the body to be an iterable of bytes - HTTP bodies are bytes and WSGI won't guess an encoding. Return `[b\"Hello\"]`, or `.encode(\"utf-8\")` a built string before returning it."
  },
  {
    "q": "A request comes in with a `User-Agent: curl/8.4.0` header. Where do you read it inside the app?",
    "choices": [
      "`environ[\"HTTP_USER_AGENT\"]` - headers arrive uppercased, dashes-to-underscores, with an `HTTP_` prefix",
      "`environ[\"User-Agent\"]` - headers keep their original name",
      "`environ[\"wsgi.input\"]` - all headers live in the body stream",
      "`start_response[\"User-Agent\"]` - headers come through the callback"
    ],
    "answer": 0,
    "explain": "The server maps each incoming header into `environ` by uppercasing it, replacing dashes with underscores, and prefixing `HTTP_`. So `User-Agent` becomes `environ[\"HTTP_USER_AGENT\"]`."
  },
  {
    "q": "In Flask, what is `@app.route(\"/notes\")` actually automating compared to the bare WSGI app?",
    "choices": [
      "The manual `if path == \"/notes\"` branching on `PATH_INFO` - the router registers routes and dispatches for you",
      "The TCP socket setup that `make_server` does",
      "The `.encode(\"utf-8\")` call on the response body",
      "Nothing - `@app.route` is unrelated to WSGI"
    ],
    "answer": 0,
    "explain": "Routing by hand means branching on `environ[\"PATH_INFO\"]`. `@app.route` is sugar over exactly that: Flask keeps a table of paths and runs the matching view, so you don't write the `if`/`elif` ladder yourself."
  }
]
```


---

# The WSGI Server & Middleware

In [Phase 2](02-a-wsgi-app-from-scratch.md) you wrote a WSGI app - a plain callable that takes `(environ, start_response)` and returns bytes - and you ran it with `wsgiref.simple_server` to see it answer a real HTTP request. That was the *app* half of the contract working. This phase is about the *other* half: the **server** that actually calls your app in production, and a single trick - wrapping one WSGI app in another - that turns out to be the root of every "middleware" you've ever met.

## The dev server isn't enough

Here's the thing nobody warns you about until it bites: ⚠️ **the server you ran in Phase 2 - `wsgiref.simple_server` - is a *toy*.** So is `flask run`. So is Django's `runserver`. They exist to let you develop on your laptop with one command and zero config. They are not built to face the internet.

What's wrong with them? They're typically single-threaded - one request at a time, so a second visitor waits behind the first. They have no process management, no graceful restarts, no protection against slow clients, and the Python docs and framework docs say so out loud. Flask literally prints a warning: *"This is a development server. Do not use it in a production deployment."*

The mental model: 📝 **a dev server is a demo car with no seatbelts - fine for the parking lot, lethal on the highway.** It runs your WSGI app, but it can't run it *for real users at real volume*. For that you need a production WSGI server.

## Production WSGI servers

📝 **A production WSGI server is a program whose entire job is to accept HTTP connections, build the `environ` dict for each request, call your app callable, and ship the bytes back out.** It's the same job `wsgiref` did in Phase 2 - just hardened, concurrent, and built to stay up for months.

The popular ones in Python:

- **gunicorn** ("Green Unicorn") - the default choice, simple and battle-tested.
- **uWSGI** - older, extremely configurable, more knobs than most people need.
- **waitress** - pure-Python, works on Windows, nice when you can't compile C extensions.

They all speak the same WSGI contract, so they're interchangeable from your app's point of view. To run your Phase 2 app - say it lives in `myapp.py` and the callable is named `app` - you point gunicorn at it:

```bash
gunicorn myapp:app
```

```console
[2026-06-23 10:14:02 +0000] [4821] [INFO] Starting gunicorn 22.0.0
[2026-06-23 10:14:02 +0000] [4821] [INFO] Listening at: http://127.0.0.1:8000 (4821)
[2026-06-23 10:14:02 +0000] [4821] [INFO] Using worker: sync
[2026-06-23 10:14:02 +0000] [4824] [INFO] Booting worker with pid: 4824
[2026-06-23 10:14:02 +0000] [4825] [INFO] Booting worker with pid: 4825
[2026-06-23 10:14:02 +0000] [4826] [INFO] Booting worker with pid: 4826
```

*What just happened:* The argument `myapp:app` is `module:callable` - gunicorn **imports** the `myapp` module and grabs the `app` object out of it, exactly the callable you wrote by hand. That's the whole handshake. gunicorn now owns the socket on port 8000; when a request arrives, it builds `environ`, calls `app(environ, start_response)`, and writes the returned bytes back to the client. Notice it booted *three* workers (pids 4824–4826) - hold that thought, it's the next section. 💡 This is the answer to a question you've probably had: *why does every Python deploy guide say "run it with gunicorn"?* Because gunicorn **is** the WSGI server - the thing that calls your app. Your framework provides the app; gunicorn provides the runtime.

## Workers & concurrency

Look again at those three "Booting worker" lines. 📝 **A worker is a separate OS process that runs a full copy of your app and handles requests on its own.** gunicorn itself is a *master* process that doesn't touch requests - it just spawns workers, watches them, and restarts any that die. The workers do the actual serving.

Why more than one? Because WSGI is **synchronous**. A sync worker handles exactly one request at a time, start to finish - it reads the request, calls your app, waits for your code (including any database query or API call) to finish, sends the response, and only *then* is free for the next request. So your concurrency is, roughly, your worker count. Three workers means three requests being served at once; a fourth visitor waits for a worker to free up. Need more parallelism, add more workers (a common starting rule of thumb is `2 × CPU cores + 1`).

In a real deployment, gunicorn rarely faces the internet alone. It sits **behind a reverse proxy** - usually **nginx** - which handles TLS, serves static files, and buffers slow clients so a visitor on bad wifi can't tie up a worker just by sending bytes slowly.

```mermaid
flowchart LR
  Client[Browser] --> N[nginx<br/>reverse proxy]
  N --> M[gunicorn master]
  M --> W1[worker 1 → your app]
  M --> W2[worker 2 → your app]
  M --> W3[worker 3 → your app]
```

*What just happened:* nginx terminates the connection from the browser and forwards the request to gunicorn's master, which has already handed the socket to its pool of workers; whichever worker is free picks up the request and runs your app. The reverse proxy is the bouncer at the door; the workers are the staff actually doing the work inside.

⚠️ Here's the limitation that matters for the rest of this guide: because a sync worker is **busy for the entire duration of a request**, one slow request - a 10-second external API call, a heavy report - ties up a whole worker for those 10 seconds, doing nothing but waiting. With three workers, three slow requests and your whole site stalls. This is the wall WSGI's synchronous model hits, and it's exactly the problem [Phase 4](04-why-asgi-exists.md) introduces ASGI to solve. File it away.

## WSGI middleware

Now the second idea - and it's a beautiful one, because it reuses everything you already know.

Suppose you want to time every request, or log it, or check an auth token before any request reaches your app. You *could* paste that code into the top of your app callable. But there's a cleaner move that the WSGI contract makes almost free.

📝 **A middleware is itself a WSGI app that wraps another WSGI app.** It's a callable with the exact same `(environ, start_response)` signature as your app - so the *server* can't tell the difference, it just calls the outermost one. Inside, the middleware does some work, then calls the **inner** app it's wrapping, then optionally does more work on the way back. The inner app has no idea it's been wrapped.

That's the entire trick. Same signature, one app holding another. Here's a timing-and-logging middleware around the app:

```python
import time

# Your actual app - the inner WSGI app from Phase 2.
def app(environ, start_response):
    start_response("200 OK", [("Content-Type", "text/plain")])
    return [b"Hello from the app"]

# The middleware: a WSGI app that wraps another WSGI app.
def timing_middleware(inner_app):
    def wrapped(environ, start_response):
        # --- BEFORE: runs on the way in ---
        start = time.perf_counter()
        path = environ.get("PATH_INFO", "/")

        result = inner_app(environ, start_response)   # call the wrapped app

        # --- AFTER: runs on the way back out ---
        elapsed = (time.perf_counter() - start) * 1000
        print(f"{path} took {elapsed:.1f}ms")
        return result
    return wrapped

# Wrap it. THIS is what the server now calls.
app = timing_middleware(app)
```

*What just happened:* `timing_middleware` takes the inner app and returns a *new* callable, `wrapped`, that has the same `(environ, start_response)` shape - so it **is** a WSGI app, indistinguishable to the server. When a request comes in, `wrapped` records a timestamp, calls `inner_app(environ, start_response)` to do the real work, then measures elapsed time and logs it before returning the inner app's result. The line `app = timing_middleware(app)` is the load-bearing one: the name `app` now points at the wrapper, so when gunicorn does `myapp:app` it calls the *middleware*, which calls your original app. Your handler stayed completely untouched - the cross-cutting concern lives entirely outside it.

💡 If that "before / call inner / after" shape feels familiar, it should. It is the *exact same sandwich* as a Java servlet filter's `doFilter` from [The Servlet API](/guides/the-servlet-api): code before, a single call that passes control inward, code after on the way back. Different language, identical idea. And it's the universal "middleware" pattern from [What a Framework Even Is](/guides/what-a-framework-even-is) - Express's `app.use(...)`, Django's middleware classes, ASP.NET's pipeline. Every one of them is *this*: something in the request's path, wrapping what comes next. You're looking at the root.

## The chain

You rarely have just one middleware. You have several - logging, then auth, then maybe compression - and you stack them by wrapping each around the previous result:

```python
app = logging_middleware(auth_middleware(timing_middleware(app)))
```

*What just happened:* Each call wraps the one inside it, building layers. The request enters `logging_middleware` first (outermost), which calls `auth_middleware`, which calls `timing_middleware`, which finally calls your real `app` at the core - and then the response unwinds back out through each layer in reverse. 💡 **It's an onion** - the same onion you saw with servlet filters. The request travels inward through every layer's "before" half to your app, and the response travels back outward through every "after" half. The first middleware to see the request is the last to see the response.

Frameworks dress this up with nicer APIs - `app.add_middleware(...)`, decorators, config lists - so you rarely hand-wrap callables like this in a real job. But underneath the syntax, it is *always* a WSGI app wrapping a WSGI app. There is no other mechanism. When a framework says "register this middleware," it is building this exact chain for you.

And now you can see the full production picture: a real WSGI server (gunicorn) running worker processes behind a reverse proxy, calling a stack of middleware that wraps your app. That's how Python web apps actually run. The one crack in it - the synchronous worker stuck waiting on slow work - is what [Phase 4](04-why-asgi-exists.md) cracks open.

## Recap

1. ⚠️ Dev servers (`wsgiref.simple_server`, `flask run`, Django's `runserver`) are single-threaded and unhardened - fine for development, never for production. You need a real WSGI server.
2. 📝 A production WSGI server (**gunicorn**, uWSGI, waitress) imports your app callable via `module:callable` and serves it. `gunicorn myapp:app` is why every deploy guide says "run gunicorn" - it's the runtime that calls your app.
3. 📝 gunicorn runs **worker processes** for concurrency; each sync worker handles one request at a time, usually behind nginx as a reverse proxy. ⚠️ A slow request ties up a whole worker - the wall that motivates ASGI.
4. 📝 **Middleware is a WSGI app that wraps another WSGI app** - same `(environ, start_response)` signature, doing work before/after it calls the inner app. It's the root of framework "middleware" and the twin of Java's servlet filter.
5. 💡 Middleware **stacks into a chain** - an onion the request passes through before reaching your app, and back out in reverse. Frameworks give nicer APIs, but underneath it's always one WSGI app wrapping another.

## Quick check

See whether the two big ideas - the server that calls your app, and the app that wraps your app - actually landed:

```quiz
[
  {
    "q": "What does the command `gunicorn myapp:app` actually do?",
    "choices": [
      "Imports the myapp module, grabs the `app` callable, and serves it by calling it for each request",
      "Compiles myapp.py into a standalone executable web server",
      "Starts Flask's built-in development server with extra logging",
      "Sends an HTTP request to the app and prints the response"
    ],
    "answer": 0,
    "explain": "`myapp:app` is module:callable. gunicorn imports myapp, takes the `app` object (your WSGI callable), and for each incoming request builds environ and calls app(environ, start_response). gunicorn is the WSGI server - the thing that calls your app."
  },
  {
    "q": "Why does a sync WSGI server like gunicorn run multiple worker processes?",
    "choices": [
      "Because each sync worker handles only one request at a time, so more workers means more concurrent requests",
      "Because each worker handles a different URL path",
      "Because Python code cannot be imported more than once per process",
      "Because one worker is for HTTP and the others are for HTTPS"
    ],
    "answer": 0,
    "explain": "WSGI is synchronous: a worker is busy for the whole duration of a request. So concurrency is roughly the worker count - three workers serve three requests at once. A slow request ties up its whole worker, which is exactly the limitation ASGI later addresses."
  },
  {
    "q": "What IS a piece of WSGI middleware, mechanically?",
    "choices": [
      "A WSGI app that wraps another WSGI app - same (environ, start_response) signature, doing work before and after it calls the inner app",
      "A special configuration file the server reads at startup",
      "A separate process that runs alongside gunicorn",
      "A subclass your app must inherit from to be servable"
    ],
    "answer": 0,
    "explain": "Middleware is just a WSGI app with the same signature as your app, holding the inner app and calling it in the middle. The server can't tell the difference. Stacked, they form a chain - an onion - exactly like servlet filters and every framework's 'middleware'."
  }
]
```


---

# Why ASGI Exists

You've seen the whole WSGI machine: an app callable, a production server like gunicorn running worker processes, middleware wrapping middleware. It's a clean, proven design - Python ran the web on it for over a decade. So why does a *second* contract exist at all?

Because WSGI has a ceiling, and you hit it the moment your app spends its time *waiting*. This phase is about that ceiling, the idea that breaks through it, and the new contract - ASGI - that bakes that idea in. We won't write a full ASGI app yet (that's [Phase 5](05-an-asgi-app-and-the-servers.md)); the goal here is the mental model, so the code in Phase 5 reads as inevitable rather than mysterious.

## WSGI's ceiling

📝 Recall the core fact from [Phase 3](03-the-wsgi-server-and-middleware.md): **WSGI is synchronous.** A sync worker is busy for the *entire* duration of a request - it reads the request, calls your app, and then sits there until your code returns, including every second your code spends waiting on a database query, a third-party API, or a slow client dribbling bytes over bad wifi.

That word - *waiting* - is the whole problem. Most web work is **I/O-bound**: your code isn't computing, it's blocked on something external coming back. And while a sync worker waits, it does *nothing else*. It can't pick up another request. It's a cashier standing frozen at the register because one customer is on the phone with their bank.

So your concurrency is capped at your worker count, and each worker is a full OS process holding a full copy of your app in memory. Want to serve 500 slow requests at once? You'd need something like 500 workers, and the RAM bill for that is brutal. The synchronous model makes high I/O concurrency expensive by construction.

⚠️ And there's a second limit that's even harder: **WSGI can't do long-lived connections at all.** Websockets, server-sent events, streaming responses - these stay open and exchange data over time. But the WSGI contract is shaped as *one request in, one response body out, done.* There is no place in `(environ, start_response) → bytes` for "keep this connection open and send more later." It's not that WSGI does websockets slowly; it can't express them. The shape doesn't fit.

## What async buys you

Here's the move that changes everything. 📝 **An async server runs many requests on a single thread using an event loop: while one request is parked waiting on I/O, the loop runs another request that's ready to make progress.** No request is ever frozen holding the thread hostage - the instant it says "I'm waiting on the database," the loop sets it aside and serves someone else, then comes back when the data arrives.

If that sounds familiar, it's the exact idea from [Async/Await & the Event Loop](/guides/async-await-and-the-event-loop): waiting is wasteful, so a single worker juggles many in-flight jobs instead of standing idle on any one of them. Same engine, now pointed at HTTP connections.

The payoff: one process can hold *thousands* of concurrent connections, because an idle connection costs almost nothing - it's just a paused task waiting for its turn, not a whole process consuming a worker slot. For I/O-bound work, that's a different order of magnitude.

```mermaid
flowchart TB
  subgraph WSGI["WSGI · sync"]
    R1[req 1] --> W1[worker 1<br/>BLOCKED waiting]
    R2[req 2] --> W2[worker 2<br/>BLOCKED waiting]
    R3[req 3 waits...] -.-> Q[no free worker]
  end
  subgraph ASGI["ASGI · async"]
    L[one event loop] --> A1[req 1 awaiting DB]
    L --> A2[req 2 awaiting API]
    L --> A3[req 3 running now]
    L --> A4[...thousands more]
  end
```

*What just happened:* On the WSGI side, each request owns a worker for its whole life, so a blocked request wastes a whole worker and a third request has to wait for one to free up. On the ASGI side, a single event loop holds every connection at once; the ones awaiting I/O cost nothing while the loop drives whichever request is actually ready to run. Same hardware, vastly more concurrent connections - *for I/O-bound work*.

## ASGI: the async contract

ASGI (Asynchronous Server Gateway Interface) is WSGI's async successor - the same *idea* (a standard contract between a server and your app, so they stay swappable) rebuilt around `async`. 📝 **The app is an async callable with three arguments instead of two:**

```python
async def app(scope, receive, send):
    # scope: dict describing the connection (like WSGI's environ)
    # receive: await this to get incoming events (request body, ws messages)
    # send: await this to push outgoing events (response, ws messages)
    ...
```

*What just happened:* Compare this to WSGI's `def app(environ, start_response)`. Three things changed, each on purpose. First, `async def` - the app runs *on* the event loop, so it can `await` without freezing the loop. Second, `scope` replaces `environ`: a dict with the connection's details (type, path, headers), handed to you once when the connection opens. Third - and this is the big one - instead of returning a body, you're given two async callables: `receive`, which you `await` to pull the next incoming event, and `send`, which you `await` to push events back out. The full body of this app comes in [Phase 5](05-an-asgi-app-and-the-servers.md); right now just sit with the *shape*.

📝 That `receive`/`send` pair is the heart of it. ASGI is an **event-based** interface, not a single-return interface. A connection isn't "one request, one response" - it's a stream of messages flowing both ways over time. That's precisely what lets ASGI handle websockets and streaming, which WSGI structurally cannot.

## Why three args, and not a return value

💡 This is the question worth pausing on, because it's the *whole reason* ASGI exists - not cosmetics, not "async is trendy." It's structural.

In WSGI, your app **returns** a body. Returning is a one-time act: you hand back the bytes, and you're done. That's a perfect fit for request → response, and a dead end for anything else. Once you've returned, the conversation is over.

In ASGI, `receive` and `send` are async callables you can `await` **as many times as you like, over the life of the connection.** A websocket can `await send(...)` to push a message, `await receive()` to read the client's reply, and loop like that for minutes. A streaming response can `await send(...)` chunk after chunk as data becomes available. The connection is an ongoing exchange, not a single transaction - and you can only model an ongoing exchange with repeatable, awaitable calls, never with one return statement.

So the three-argument, event-driven signature isn't a restyle of WSGI. It's the minimum shape that can express "this connection stays open and trades many messages." That capability is the reason ASGI was created.

## WSGI vs ASGI - when to reach for each

Now the plain-spoken part, because "newer" does not mean "use it for everything."

💡 **WSGI is completely fine - and simpler - for ordinary synchronous apps.** A classic Flask app, a traditional Django project, a CRUD service whose database is fast and local: run it on gunicorn with sync workers and move on. You get a smaller mental model, no async footguns, and battle-tested tooling. Reaching for ASGI here buys you nothing and costs you complexity.

**Reach for ASGI when you have:** high-concurrency I/O-bound traffic (lots of requests, each mostly waiting on slow databases or external APIs), websockets, server-sent events, streaming responses, or a codebase that's already `async` end to end. This is the home turf of frameworks like FastAPI and Starlette - see [FastAPI From Zero](/guides/fastapi-from-zero) - which are ASGI-native and lean on exactly these strengths.

⚠️ One myth to kill before it bites you: **async is not automatically "faster."** It shines at *I/O-bound concurrency* - many connections that mostly wait. It does **nothing** for CPU-bound work; in fact, a heavy computation in an async handler *blocks the event loop* and freezes every other connection on that worker, which is worse than the sync model would have been. Async trades "many idle workers" for "one busy loop juggling many waits." If your bottleneck is the CPU crunching numbers, neither the loop nor `await` saves you - that's a job for more processes or a background queue, not ASGI.

With the *why* firmly in hand, [Phase 5](05-an-asgi-app-and-the-servers.md) fills in that three-argument app for real and introduces the servers (uvicorn, hypercorn) that run it.

## Recap

1. ⚠️ WSGI is **synchronous**: a worker is busy for the entire request, so it sits idle while waiting on I/O. High I/O concurrency means many workers, which means lots of memory.
2. ⚠️ WSGI **can't do long-lived connections at all** - websockets, SSE, streaming don't fit the "one request, one response body" shape.
3. 📝 An **async server** runs many requests on one event loop: while one awaits I/O, the loop serves others. Far more concurrent connections per process for I/O-bound work - the same idea as [the event loop guide](/guides/async-await-and-the-event-loop).
4. 📝 **ASGI** is the async contract: `async def app(scope, receive, send)`. `scope` is the connection info (like `environ`); `receive`/`send` are awaitable event channels, not a single return value.
5. 💡 You `await` `receive`/`send` repeatedly over a connection's life - that's the structural reason ASGI exists (it can model websockets and streaming), not a cosmetic change. WSGI stays fine and simpler for ordinary sync apps; ASGI wins for high-concurrency I/O, websockets, and streaming - but async ≠ faster for CPU work.

## Quick check

See whether the *why* behind ASGI actually landed - not the syntax, the reasoning:

```quiz
[
  {
    "q": "Why does a synchronous WSGI worker struggle under high I/O-bound concurrency?",
    "choices": [
      "It is busy for the entire request, sitting idle while waiting on I/O, so concurrency is capped at the worker count",
      "It can only handle GET requests, not POST",
      "It recompiles your app on every request, which is slow",
      "It uses too little memory to hold many connections"
    ],
    "answer": 0,
    "explain": "A sync worker is occupied for the whole request, including time spent waiting on a database or API. While it waits it can't serve anyone else, so concurrency equals worker count - and each worker is a full process, so scaling that up is expensive."
  },
  {
    "q": "What is the structural reason ASGI uses `receive`/`send` instead of returning a response body like WSGI?",
    "choices": [
      "Returning a body is a one-time act; awaitable `receive`/`send` can be called repeatedly, so a connection can exchange many messages over time (e.g. a websocket)",
      "It makes the code shorter to type",
      "Async functions are not allowed to use return statements",
      "It lets ASGI apps skip building the scope dict"
    ],
    "answer": 0,
    "explain": "Returning ends the conversation after one body. `receive`/`send` are awaitable callables you can use many times across a connection's life, which is the only way to model an ongoing, two-way exchange like a websocket or a streamed response. That capability is why ASGI exists."
  },
  {
    "q": "Your service does heavy CPU-bound number crunching per request. Will switching from WSGI to ASGI speed it up?",
    "choices": [
      "No - async helps I/O-bound concurrency, not CPU work; a heavy computation even blocks the event loop and freezes other connections",
      "Yes - async makes all Python code run faster",
      "Yes - the event loop runs your computation on multiple cores automatically",
      "No - but only because ASGI servers are slower than gunicorn"
    ],
    "answer": 0,
    "explain": "Async is for connections that mostly wait. CPU-bound work doesn't wait - and worse, a long computation in an async handler blocks the single event loop, stalling every other connection on that worker. The fix for CPU work is more processes or a background queue, not ASGI."
  }
]
```


---

# An ASGI App & the Servers

Back in [Phase 2](02-a-wsgi-app-from-scratch.md) you wrote a complete WSGI app - one function the
server calls per request, handed an `environ` dict, returning bytes. [Phase 4](04-why-asgi-exists.md)
explained *why* that shape couldn't go async, and what ASGI replaced it with. Now we make ASGI real
the same way: by writing the whole thing by hand, with no framework in sight.

The mental model to carry through, and it's a direct echo of the WSGI one: **an ASGI app is one
`async` function the server calls per connection.** Instead of `(environ, start_response)`, it takes
three arguments - `scope` (what the connection is), `receive` (an awaitable you call to *get* events
coming in), and `send` (an awaitable you call to *push* events out). Match that shape and you have a
web app that can `await`. Everything FastAPI does is ergonomics layered over this one function - exactly
as Flask was ergonomics over the WSGI callable. Once you've written it yourself, FastAPI stops being
magic.

## A bare ASGI app

Here's the whole thing - a working ASGI app, no framework:

```python
async def app(scope, receive, send):
    assert scope["type"] == "http"

    await send({
        "type": "http.response.start",
        "status": 200,
        "headers": [(b"content-type", b"text/plain")],
    })
    await send({
        "type": "http.response.body",
        "body": b"Hello",
    })
```

*What just happened:* `app` is the ASGI callable - an `async def` function taking `scope`, `receive`,
and `send`. The server calls `await app(scope, receive, send)` per connection. First we check
`scope["type"]` is `"http"` (ASGI also delivers websocket and lifespan connections through this same
function - more below). Then we **send the response as two events**: an `http.response.start` carrying
the status and headers, followed by an `http.response.body` carrying the bytes. Notice the headers are
`(bytes, bytes)` tuples here, not `(str, str)` like WSGI - ASGI works in raw bytes on both sides.

💡 Look at what's *missing*: there is no `return`. In WSGI you returned the body; here the response is
**sent** as events through `send`, not returned. That's the ASGI shape from [Phase 4](04-why-asgi-exists.md)
in the flesh - a response isn't a value you hand back, it's a stream of messages you push out, each one
an `await` point where the event loop can go do other work. The two-event split (`start` then `body`)
is why you can begin a response, then stream the body in chunks later, all without blocking a worker.

## Reading the request

`scope` is to ASGI what `environ` was to WSGI: a dict the server fills in with everything *static* about
the connection - the method, the path, the headers. The difference is the *body*. In WSGI the body was a
single stream you read from `environ["wsgi.input"]`. In ASGI the body arrives as **`http.request` events
you pull in by `await receive()`** - and it can come in several chunks.

```python
async def app(scope, receive, send):
    assert scope["type"] == "http"
    method = scope["method"]          # "GET"
    path = scope["path"]              # "/notes"

    # Pull the request body in, chunk by chunk.
    body = b""
    more = True
    while more:
        event = await receive()       # an {"type": "http.request", ...} message
        body += event.get("body", b"")
        more = event.get("more_body", False)

    reply = f"{method} {path} ({len(body)} bytes of body)".encode("utf-8")
    await send({"type": "http.response.start", "status": 200,
                "headers": [(b"content-type", b"text/plain")]})
    await send({"type": "http.response.body", "body": reply})
```

*What just happened:* the method and path come straight off `scope` - no fishing through `HTTP_*` keys
like WSGI, they're plain `scope["method"]` and `scope["path"]`. The body is different: each
`await receive()` hands back one `http.request` event with a `body` chunk and a `more_body` flag. We
loop, appending chunks, until `more_body` is `False` and we've got the whole body. Then we send the
response back out as the same two events.

⚠️ This is the trap if you're coming from WSGI: **the body is not one read.** WSGI gave you a single
input stream; ASGI dribbles the body in as a series of `receive` events, and you must loop until
`more_body` is false or you'll silently process a half-empty request. The upside is exactly the point of
async - each `await receive()` is a yield point, so a worker waiting on a slow upload isn't blocked, it's
free to serve other connections.

## Running it: uvicorn

A WSGI app needs a WSGI server (gunicorn, uWSGI) to invoke it. 📝 **An ASGI app needs an ASGI server** -
the three common ones are **uvicorn**, **hypercorn**, and **daphne**. They speak HTTP on the socket and
translate it into the `scope` / `receive` / `send` calls your app expects. Save the app above as
`myapp.py` and point uvicorn at it:

```bash
uvicorn myapp:app
```

```console
$ uvicorn myapp:app
INFO:     Started server process [12345]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://127.0.0.1:8000 (Press CTRL+C to quit)
```

*What just happened:* `uvicorn myapp:app` means "import `myapp.py`, find the `app` object, and run it as
an ASGI app." Uvicorn binds a socket on port 8000, accepts HTTP connections, and for each one calls
`await app(scope, receive, send)` - the exact function you wrote. Curl `http://127.0.0.1:8000/notes` and
you get your method/path line back.

💡 Hold the contrast next to [Phase 3](03-the-wsgi-server-and-middleware.md): **gunicorn alone is a WSGI
server; uvicorn is an ASGI server.** They are not interchangeable - a WSGI server can't drive an
`async def app(scope, ...)`, and an ASGI server can't drive a `def app(environ, start_response)`. In
production you usually combine them: run **gunicorn with uvicorn worker processes**
(`gunicorn -k uvicorn.workers.UvicornWorker myapp:app`), behind nginx. Gunicorn gives you the robust
process manager (multiple workers, restarts); the uvicorn workers give you the async event loop. nginx
out front terminates TLS and serves static files. Same layered shape as the WSGI stack - just async-aware
workers in the middle.

## A glimpse: lifespan and websockets

📝 Remember the `assert scope["type"] == "http"` at the top? That guard exists because **HTTP is not the
only thing that comes through your app.** ASGI delivers other connection types through the *same*
`scope` / `receive` / `send` function:

- **`scope["type"] == "lifespan"`** - sent once at startup and once at shutdown. The server calls your
  app with a lifespan scope, you `await receive()` a `lifespan.startup` event (open your DB pool here),
  and later a `lifespan.shutdown` event (close it). It's how an ASGI app runs setup/teardown code.
- **`scope["type"] == "websocket"`** - a long-lived two-way connection. You `await receive()` incoming
  messages and `await send()` outgoing ones, for as long as the socket stays open.

💡 That's the whole reason ASGI exists in one sentence: **one protocol shape - `scope` / `receive` /
`send` - handles HTTP requests, websockets, and the app lifecycle alike.** WSGI could only ever do one
request-response over HTTP; ASGI's three-argument async contract is general enough to carry all three.
You don't need the details today - just register that your `app` function is a single door that
*everything* comes through, and the `scope["type"]` tells you what kind of connection you're holding.

## What FastAPI and Starlette add

Now the reveal. You just wrote a bare ASGI app - `scope` / `receive` / `send`, headers as byte tuples,
the body looped in chunk by chunk, the response pushed out as two events. That is tedious, and nobody
ships it by hand. So what does [FastAPI](/guides/fastapi-from-zero) do?

💡 **FastAPI is built on Starlette, and Starlette is an ASGI framework** - meaning Starlette's
application object is itself an ASGI callable, the same `async def app(scope, receive, send)` shape you
wrote, with routing, request parsing, and response serialization built on top. When you write this in
FastAPI:

```python
from fastapi import FastAPI

app = FastAPI()

@app.get("/notes")
async def list_notes():
    return {"notes": ["milk", "bread"]}
```

*What just happened:* underneath, `app` is an ASGI application - uvicorn calls it with exactly
`scope` / `receive` / `send`. The `@app.get("/notes")` decorator does the `scope["path"]` matching you'd
otherwise hand-write. FastAPI reads the body via `await receive()` for you, parses it, and turns the
`dict` you return into the two `send` events - `http.response.start` with `content-type: application/json`,
then `http.response.body` with the JSON bytes. Your tidy `async def` endpoint *is* the bare ASGI machinery
from the top of this page, with all the ceremony automated away.

Map it the same way [Phase 2](02-a-wsgi-app-from-scratch.md) mapped Flask onto bare WSGI:

| You wrote, by hand | FastAPI / Starlette gives you | What it's doing underneath |
|--------------------|-------------------------------|----------------------------|
| `assert scope["type"] == "http"`, `scope["path"]` branching | `@app.get("/notes")` | the same scope inspection + path matching, registered |
| the `await receive()` body loop | typed request models, parsed for you | reads and assembles the body events |
| two `send(...)` events with byte headers | `return {...}` | builds the `start` + `body` events and serializes JSON |
| your `async def app(scope, receive, send)` | the `FastAPI()` object | **is** an ASGI callable too - same shape |

That last row is the kicker, and it's the same kicker as WSGI: the `app` you create with `FastAPI()` is
itself an ASGI callable, which is *exactly* why uvicorn can run it. You wrote the bare ASGI app. FastAPI
is the comfortable seat bolted on top. The final phase ties the two halves of this guide together -
WSGI and ASGI, frameworks and servers - into one picture.

## Recap

1. **An ASGI app is one `async` function** the server calls per connection:
   `async def app(scope, receive, send)`. It's the async sibling of the WSGI callable from
   [Phase 2](02-a-wsgi-app-from-scratch.md).
2. **The response is *sent*, not returned** - you `await send(...)` an `http.response.start` event
   (status + byte headers) then an `http.response.body` event (the bytes). No `return` of a body.
3. **`scope` holds the static request** (method, path, headers, plus `type`), while **the body arrives
   as `http.request` events you pull in with `await receive()`** - looping on `more_body`, ⚠️ unlike
   WSGI's single input stream.
4. **ASGI apps need an ASGI server** - uvicorn, hypercorn, or daphne. gunicorn alone is a *WSGI* server;
   in production, run gunicorn with uvicorn workers behind nginx.
5. **The same `scope`/`receive`/`send` shape carries HTTP, `websocket`, and `lifespan`** connections -
   one protocol, many connection types. That generality is the whole reason ASGI exists.
6. 💡 **FastAPI is built on Starlette, an ASGI framework** - `@app.get` async endpoints are this exact
   `scope`/`receive`/`send` machinery with routing, parsing, and serialization on top. You wrote the
   bare app; FastAPI is the ergonomics.

## Quick check

Make sure the ASGI shape stuck:

```quiz
[
  {
    "q": "How does a bare ASGI app return its response body to the client?",
    "choices": [
      "It `await send(...)`s an `http.response.start` event then an `http.response.body` event - the response is sent, not returned",
      "It `return`s a list of bytes, exactly like a WSGI app",
      "It assigns the body to `scope[\"body\"]` before the function ends",
      "It calls `start_response(...)` then returns the bytes"
    ],
    "answer": 0,
    "explain": "ASGI apps push the response out as events: an `http.response.start` (status + headers) followed by one or more `http.response.body` events. There is no `return` of a body - that streaming-of-events shape is what lets the app stay async."
  },
  {
    "q": "In an ASGI app, where does the request body come from?",
    "choices": [
      "From `await receive()` events - `http.request` messages whose `body` chunks you loop over until `more_body` is false",
      "From `scope[\"body\"]`, fully assembled before the app runs",
      "From `environ[\"wsgi.input\"]`, read as one stream",
      "From the `send` callable, which both sends and receives"
    ],
    "answer": 0,
    "explain": "Unlike WSGI's single input stream, ASGI delivers the body as `http.request` events. You `await receive()` repeatedly, appending each chunk's `body`, until `more_body` is False."
  },
  {
    "q": "What is the relationship between FastAPI and the bare `scope`/`receive`/`send` app?",
    "choices": [
      "FastAPI is built on Starlette (an ASGI framework); a FastAPI app IS an ASGI callable, with routing, parsing, and serialization layered over the same scope/receive/send machinery",
      "FastAPI replaces ASGI with its own faster protocol that uvicorn translates",
      "FastAPI is a WSGI framework, so it uses environ/start_response instead",
      "FastAPI has nothing to do with ASGI - it runs directly on raw sockets"
    ],
    "answer": 0,
    "explain": "FastAPI sits on Starlette, an ASGI framework, so its app object is itself an ASGI callable - which is why uvicorn can run it. Your `@app.get` async endpoint is the bare scope/receive/send machinery with the ceremony automated."
  }
]
```


---

# From Protocol to Framework

Go back to the start of this guide for a second. A Python web framework was a black box: you wrote `@app.route` or `@app.get`, requests showed up as a tidy `request` object, you returned something, and a response went out. A lot happened in the middle that you couldn't name.

Look at what you can see now. You know a request arrives at a **server** - gunicorn or uvicorn - which calls your app as a **callable**. For synchronous frameworks that callable takes `(environ, start_response)`; for async ones it takes `(scope, receive, send)`. You know **middleware** is an app that wraps another app. You know why `flask run` is fine for your laptop and wrong for production. That's not trivia - that's the whole shape of Python web, and you've now seen all of it bare.

This last phase is the payoff. We're not adding new mechanism. We're pointing the X-ray vision you just built at the frameworks you'll actually use at work, and watching the magic turn into machinery you can already name.

## Mapping the magic to the mechanism

💡 Here's the thing worth reading slowly: every "feature" in a Python web framework is a convenience over something in this guide. Once you've seen the bare version, the framework version stops being a spell and becomes a name for a thing you understand.

```mermaid
flowchart LR
  FL["Flask / Django (classic)"] --> WS["WSGI app (Phases 1-2)"]
  FA["FastAPI / Starlette"] --> AS["ASGI app (Phase 5)"]
  RT["@app.route / @app.get"] --> PA["Routing over the raw path"]
  RQ["request / Request"] --> EN["Wrapper over environ / scope"]
  MW["Framework 'middleware'"] --> WR["An app wrapping an app (Phase 3)"]
  RUN["flask run vs production"] --> SRV["dev server vs gunicorn/uvicorn"]
```

Reading that left to right, in plain words:

- **Flask and classic Django are WSGI apps.** Under the decorators and the `request` object, each one *is* a WSGI callable - the same `(environ, start_response)` contract you wrote by hand in [Phase 2](02-a-wsgi-app-from-scratch.md). You serve them in production with **gunicorn** (or uWSGI). See [Flask From Zero](/guides/flask-from-zero) and [Django From Zero](/guides/django-from-zero).
- **FastAPI and Starlette are ASGI apps.** Underneath, each is an async callable taking `(scope, receive, send)` - exactly what you built in [Phase 5](05-an-asgi-app-and-the-servers.md). You serve them with **uvicorn**. See [FastAPI From Zero](/guides/fastapi-from-zero).
- **`@app.route("/users")` / `@app.get("/users")`** is routing over the raw path. In the bare apps you read `environ["PATH_INFO"]` (or `scope["path"]`) and branched by hand. The decorator is the framework filling in that same routing table for you.
- **`request` / `Request`** is a friendly wrapper over `environ` or `scope`. Method, path, headers, query string, body - the framework parses the raw dict you saw and hands you attributes instead of keys.
- **Framework "middleware"** is a WSGI/ASGI app wrapping an app - [Phase 3](03-the-wsgi-server-and-middleware.md), generalized. Code that runs before and after your handler, able to read or rewrite the request and response, or short-circuit it entirely. You built one.
- **`flask run` vs production** is the dev server vs gunicorn/uvicorn. The built-in server is a convenience for your laptop; in production a real WSGI/ASGI server runs your same app under load.

That covers the whole Python web world. The frameworks differ in ergonomics and features - but underneath, every one of them is a WSGI or ASGI callable that a server invokes.

## Why you still reach for a framework

Here's the plain truth, because a roots guide that pretends raw WSGI/ASGI is enough would be doing you a disservice: you don't want to build a real app out of bare callables. It's genuinely tedious.

You'd hand-write the routing and keep it in sync. You'd parse query strings and JSON bodies yourself, then serialize responses by hand. There's no dependency injection, no validation, no friendly request object - you'd reach into `environ` and `scope` for every value and write the same boilerplate a thousand times. The frameworks exist because smart people got tired of doing exactly that, and the conveniences they add are real and worth having.

💡 So here's the point of having learned this: you almost certainly won't write raw WSGI or ASGI at work - you'll write Flask, Django, or FastAPI. What changed is that you now understand *what those frameworks are conveniences over*. That's the entire purpose of a roots guide. When uvicorn shows up in a traceback, you don't flinch - you know it's calling your ASGI callable. When middleware swallows a request, you know it's an app wrapping an app and you know where to look. The framework didn't get simpler; you got the map.

## The deployment payoff

💡 There's a bonus here that pays off across every framework guide on this site: you now understand the deploy story, full stop.

It's the same picture under Flask, Django, and FastAPI, because they're all WSGI or ASGI apps. For a **WSGI** app (Flask, classic Django) you run **gunicorn**. For an **ASGI** app (FastAPI, Starlette) you run **uvicorn** - or, for the best of both, **gunicorn with uvicorn workers** (process management from gunicorn, async handling from uvicorn). In front of either you put **nginx** for TLS and static files, and you set `DEBUG=False` so you're not leaking stack traces in production.

When a deploy guide tells you to "run gunicorn behind nginx," that's no longer a recipe you follow on faith. It's gunicorn calling your WSGI callable, with nginx in front - the exact shape this guide drew.

## What to build - and a last word

📝 Reading got you here. One small build will lock it in for good. Here's the exercise that cements everything:

Write a **bare WSGI app** (about a dozen lines, `(environ, start_response)`) and a **bare ASGI app** (about a dozen lines, `async def app(scope, receive, send)`). Run the first with `gunicorn` and the second with `uvicorn`, and hit each with a browser or `curl`. Feel how little is actually there.

Then do the magic trick: open the source of a framework you use. Find Flask's `wsgi_app` / `__call__`, or Starlette's `async def __call__`, and locate the `(environ, start_response)` or `(scope, receive, send)` entry point. Seeing the exact signature you just wrote, sitting inside real framework code that powers production apps, is the moment it all fuses together. You'll never read a framework the same way again.

When you want the authoritative reference, go to **PEP 3333** (the WSGI spec) and the **ASGI specification**. They're precise, they're the source of truth, and now that you have the concepts, they'll read as confirmation rather than fog.

The line to carry out of this whole guide: **every Python web framework is conveniences over one small callable - and now you can see it.**

## Recap

1. **The X-ray vision is the whole point.** You can now see, under any Python web framework, a WSGI or ASGI callable, a server that calls it, and middleware that wraps it - the bare mechanism this guide built.
2. **Each framework maps to something you know:** Flask and classic Django are WSGI apps (served by gunicorn); FastAPI and Starlette are ASGI apps (served by uvicorn); `@app.route`/`@app.get` is routing over the raw path; `request`/`Request` wraps `environ`/`scope`; "middleware" is an app wrapping an app; `flask run` is the dev server, not production.
3. **Frameworks earn their keep.** Raw WSGI/ASGI is tedious - manual routing, manual parsing and serialization, no DI, endless boilerplate. Frameworks add the conveniences; this guide showed you *what they're conveniences over*.
4. **The deploy story is one picture.** gunicorn (WSGI) or uvicorn / gunicorn-with-uvicorn-workers (ASGI), behind nginx, with `DEBUG=False` - the same under Flask, Django, and FastAPI, because they're all WSGI/ASGI apps.
5. **Build to cement it:** a bare WSGI app and a bare ASGI app, run under gunicorn and uvicorn - then open a real framework's source and find the callable's entry point. PEP 3333 and the ASGI spec are your authoritative sources.

## Quick check

One last check - the mappings that turn frameworks from magic into mechanism:

```quiz
[
  {
    "q": "Mechanically, what are Flask and classic Django?",
    "choices": [
      "WSGI apps - callables taking (environ, start_response), served in production by gunicorn",
      "ASGI apps that only run under uvicorn",
      "Standalone web servers that replace gunicorn",
      "Browser-side JavaScript frameworks"
    ],
    "answer": 0,
    "explain": "Flask and classic Django are WSGI applications underneath the decorators and the request object - the same (environ, start_response) callable you wrote by hand in Phase 2. In production you run them with gunicorn (or uWSGI)."
  },
  {
    "q": "When a framework talks about 'middleware', what is it built on?",
    "choices": [
      "A WSGI or ASGI app that wraps another app",
      "A second web server running alongside the first",
      "A database connection pool",
      "The browser's fetch API"
    ],
    "answer": 0,
    "explain": "Framework 'middleware' is an app wrapping an app - Phase 3, generalized. It runs before and after your handler, can read or rewrite the request and response, or short-circuit the chain entirely."
  },
  {
    "q": "Which deployment statement is accurate across these frameworks?",
    "choices": [
      "WSGI apps run under gunicorn and ASGI apps under uvicorn (or gunicorn with uvicorn workers), typically behind nginx with DEBUG=False",
      "flask run is the recommended way to serve production traffic",
      "FastAPI must be served by gunicorn with no async support",
      "Django can only be deployed as static files"
    ],
    "answer": 0,
    "explain": "It's the same picture under Flask, Django, and FastAPI because they're all WSGI/ASGI apps: gunicorn for WSGI, uvicorn (or gunicorn with uvicorn workers) for ASGI, behind nginx, with DEBUG=False. The built-in flask run dev server is for your laptop, not production."
  }
]
```
