# API Gateway, Explained

> A single front door in front of many backend services - what an API gateway actually does, and when it earns its keep versus when it's overkill.


---

# API Gateway, Explained

Once a system grows past one service, clients face an awkward question: which address do I call for what? The orders service, the payments service, the search service - each with its own host, its own auth, its own quirks. An API gateway exists to make that question disappear: one address, one door, everything behind it becomes the gateway's problem instead of the client's - this guide covers why that door exists, what it actually does once it's there, and the tradeoffs nobody mentions in the pitch.

## How to read this

Read it in order. Phase 1 is the problem the gateway solves - what life looks like without one; Phase 2 is what a gateway actually does day to day: routing, auth, rate limiting, and more. Phase 3 is the clear-eyed tradeoffs, real products you'll encounter, and when a gateway is more machinery than a small system needs.

## The phases

1. [The single front door](01-the-single-front-door.md) - the problem: clients otherwise juggle N services, N addresses, N auth schemes.
2. [What a gateway actually does](02-what-a-gateway-actually-does.md) - routing, authentication, rate limiting, transformation, and aggregation.
3. [Tradeoffs and real examples](03-tradeoffs-and-real-examples.md) - the new failure point, the latency cost, real products, and when to skip it.


---

# The single front door

Imagine a mobile app for an online store. It needs product data, order history, inventory status, and to process a payment. If the backend is split into an orders service, a products service, an inventory service, and a payments service, the app now needs to know all four addresses, speak whatever protocol each one prefers, and handle four different ways of proving who the user is.

That's the situation an API gateway exists to erase.

## What the client has to deal with, without a gateway

Picture the app calling each service directly:

```text
GET  https://orders-svc.internal.example.com/v2/orders/482
GET  https://products-svc.internal.example.com/api/products?ids=9,14,22
GET  https://inventory.internal.example.com/stock/check
POST https://payments-svc.internal.example.com/v1/charge
```

*What just happened:* the client needs to know four hostnames, four different URL shapes, and - this is the painful part - probably four different ways of authenticating, because four teams built these services at different times with different conventions. One might expect an API key in a header, another a bearer token, another a signed request. The client also now has a hard dependency on internal details: if the payments team renames their service or moves it to a new host, every client needs to know.

```text
Client's job without a gateway:
  - know every service's address
  - know every service's auth scheme
  - handle failures from each one separately
  - update itself whenever a service moves or changes shape
```

*What just happened:* the client just wants "give me this order" and "charge this card" - not a directory of internal infrastructure. None of those bullet points touch the app's actual job, and the list grows with every service you add.

## What the client deals with, with a gateway

Now put one thing in front of all four services - a gateway that the client is the only thing it talks to:

```text
GET  https://api.example.com/orders/482
GET  https://api.example.com/products?ids=9,14,22
GET  https://api.example.com/inventory/check
POST https://api.example.com/charge
```

*What just happened:* one hostname. One auth scheme - the client authenticates once, to the gateway, and the gateway is trusted to sort out the rest. The client has no idea there are four services back there, doesn't know their addresses, and doesn't care if the payments team rewrites their entire service next month, as long as the gateway's contract with the client stays the same.

> The gateway's whole job is to be the one thing a client has to know about. Everything behind it can change shape without the client noticing.

## Why this matters more as a system grows

With one or two backend services, this is barely a problem - hardcode two addresses and move on. The pain scales worse than you'd guess, because it isn't just "one more address": every new service brings its own auth quirks, failure modes, versioning scheme, and error conventions. A system with fifteen services and no gateway means every client re-solves the same integration problem fifteen times over.

```text
2 services, no gateway   -> mildly annoying
5 services, no gateway   -> every client hardcodes 5 addresses and 5 auth flows
15 services, no gateway  -> nobody actually knows the full list anymore
15 services, one gateway -> clients still see exactly one address
```

*What just happened:* the gateway doesn't remove the complexity of having many services - that complexity is real and it's still there. What it does is move that complexity to one place, behind one door, instead of scattering it across every client that ever needs to talk to the backend. Phase 2 covers what actually happens at that door once traffic arrives there.


---

# What a gateway actually does

"One front door" describes the shape, not the job. Once a request walks through that door, the gateway does real work on it before deciding where it goes. Five things show up in almost every gateway you'll encounter.

## Path-based routing

The most basic job: look at the incoming request and decide which backend service actually handles it. The client only sees one host; the gateway is the thing that knows the real map.

```text
Request path              -> routed to
/orders/*                 -> orders-service
/products/*                -> products-service
/inventory/*               -> inventory-service
/charge                    -> payments-service
```

*What just happened:* the gateway inspects the path (sometimes the hostname, sometimes a header) and forwards the request to whichever service owns that piece of the API. This is the piece that makes "one address, many services" actually work - the client's illusion of a single API is a routing table on the gateway's side.

## Authentication and authorization at the edge

Instead of every backend service independently verifying who's calling and what they're allowed to do, the gateway does it once, at the door, before the request goes anywhere.

```text
1. Request arrives with a token
2. Gateway validates the token (is it real? is it expired?)
3. Gateway checks: is this user allowed to hit this route?
4. Only if both pass -> forward to the backend service
```

*What just happened:* the backend services can now largely trust that anything reaching them has already been vetted, instead of every one of them reimplementing token validation and permission checks. This isn't just convenience - it's a real security posture. One well-tested piece of auth code protecting every service beats fifteen homegrown copies of the same logic, some of which will inevitably be wrong.

## Rate limiting

The gateway sees every request to every service, which makes it the natural place to enforce "this client can make at most N requests per minute."

```text
client_id: acme-corp   requests this minute: 118   limit: 120   -> allow
client_id: acme-corp   requests this minute: 121   limit: 120   -> reject (429)
```

*What just happened:* one client hammering the API gets stopped at the door, before its flood of requests ever reaches - and potentially overwhelms - a backend service. Without a gateway, you'd need to build rate limiting into every service separately, or hope none of them ever need it.

## Request and response transformation

Clients and backend services don't always want to speak in exactly the same shape. The gateway can reshape a request or response in flight.

```text
Client sends:     { "user_id": "482" }
Gateway forwards: { "userId": "482", "apiVersion": "2" }   <- old service still expects this shape

Service returns:  { "userId": "482", "internal_flags": {...}, "balance_cents": 4200 }
Gateway returns:  { "userId": "482", "balance": 42.00 }    <- strips internals, client-friendly units
```

*What just happened:* the gateway acts as a translator in both directions. This is genuinely useful when a backend service is old, written by a different team with different conventions, or exposes internal fields that have no business reaching a client. The client gets a clean, stable contract; the backend keeps whatever shape is convenient for it.

## Aggregating multiple calls into one

Sometimes a single client request logically needs data from several backend services. Instead of making the client fire off three requests and stitch the results together itself, the gateway can do the fan-out internally and hand back one combined response.

```text
Client requests:  GET /order-summary/482

Gateway internally calls:
  orders-service     -> order details
  products-service   -> product names for the items in the order
  inventory-service  -> current stock status

Gateway returns one combined JSON response to the client.
```

*What just happened:* the client made one request and got one response, even though three services were involved. This pattern is sometimes called "backend for frontend" when it's tailored to a specific client type (mobile vs. web), but the underlying move is the same: push the fan-out and stitching work to the gateway, which is already positioned to talk to everything, instead of making every client re-implement that orchestration.

## Why all five live in one place

None of these five jobs strictly requires a gateway - you could build auth checks and rate limiting into every service yourself. The reason they cluster into one component is that they're all *cross-cutting*: every service needs some version of routing awareness, auth, and rate limiting, and duplicating that logic N times means N chances to get it wrong and N places to update when the policy changes. Centralizing it in the gateway means one implementation, one place to patch a bug, one dashboard to watch.

```quiz
[
  {
    "q": "Why do authentication and rate limiting typically live in the gateway rather than in each backend service?",
    "choices": [
      "Backend services are incapable of running auth code",
      "They're cross-cutting concerns - duplicating them in every service means more chances to get it wrong",
      "It's required by the HTTP specification",
      "Gateways are the only place that can see a bearer token"
    ],
    "answer": 1,
    "explain": "Auth and rate limiting apply to nearly every request. Centralizing them means one correct implementation instead of N slightly-different copies scattered across services."
  },
  {
    "q": "A client requests order-summary data, and the gateway internally calls three separate backend services, then returns one combined response. What is this pattern called?",
    "choices": [
      "Path-based routing",
      "Rate limiting",
      "Aggregation (sometimes backend-for-frontend)",
      "Request transformation"
    ],
    "answer": 2,
    "explain": "Aggregation is the gateway fanning a single client request out to multiple backend services and stitching the results into one response."
  },
  {
    "q": "What does path-based routing let a gateway do?",
    "choices": [
      "Encrypt traffic between the client and the gateway",
      "Decide which backend service should handle a request based on its path or host",
      "Automatically retry failed requests forever",
      "Convert JSON responses into XML"
    ],
    "answer": 1,
    "explain": "Path-based routing is the gateway's map from incoming request paths (or hosts) to the backend service that actually owns that piece of the API."
  }
]
```

Watch it animated: [an API gateway](/explainers/APIGateway.dc.html)


---

# Tradeoffs and real examples

A gateway solves a real problem, but it isn't free. It's a piece of infrastructure sitting on the path of every single request, which means its costs are just as centralized as its benefits.

## Cost 1: a new single point of failure

Before the gateway existed, if one backend service went down, only the features depending on that service broke. The gateway sits in front of *everything*, which means if the gateway itself goes down, every service behind it becomes unreachable, even the ones that are perfectly healthy.

```text
Without a gateway:
  payments-service down -> only checkout breaks

With a gateway, if the gateway itself is down:
  gateway down -> orders, products, inventory, payments -- ALL unreachable
```

*What just happened:* you traded "one service can fail independently" for "one component's failure takes down the whole API." This is manageable - gateways are typically run in a redundant, load-balanced cluster specifically because of this risk - but it's not automatic. A gateway you didn't design for high availability is a bigger risk than not having one.

## Cost 2: an extra latency hop

Every request now passes through one more piece of infrastructure before it reaches the service that actually does the work, and the response passes back through it again on the way out.

```text
Without gateway: client -> service                    (1 hop each way)
With gateway:    client -> gateway -> service          (2 hops each way)
```

*What just happened:* the gateway adds processing time - routing logic, auth checks, possibly transformation - on top of a genuine network hop. For most systems this is small, often single-digit milliseconds, and it's a reasonable price for what you get back. But if you're building something latency-critical, it's a real cost to measure, not assume away.

## Cost 3: configuration complexity

The gateway's routing rules, auth policies, and rate limits all live in one place - which is exactly the benefit from Phase 2, and also a new liability. Get a routing rule wrong and you can misroute traffic for every service at once, not just one, because that configuration is now critical infrastructure in its own right. Someone has to own it, version it, and test changes with the same care you'd give to code.

```text
One misconfigured route in the gateway
  -> can break access to a service for every client, all at once,
     even though the service itself is fine
```

*What just happened:* centralizing cross-cutting concerns concentrates risk along with the benefit. A bug in one backend service's auth code affects that one service. A bug in the gateway's auth config can affect all of them simultaneously.

## Real products you'll run into

You don't build a gateway from scratch in most cases - you configure one. A few you'll see repeatedly:

```text
Kong              -> open-source, plugin-based, popular for self-hosted setups
AWS API Gateway    -> managed, deeply integrated with Lambda and other AWS services
nginx              -> not a dedicated gateway product, but its reverse-proxy and routing
                      features cover a genuinely lightweight version of the same idea
```

*What just happened:* these differ mostly in how much you manage yourself versus how much a cloud provider manages for you, and how much built-in tooling (plugins, dashboards, auth integrations) comes out of the box. nginx is worth calling out specifically: many small systems run something that is, functionally, a stripped-down API gateway - a reverse proxy doing path-based routing and maybe basic auth - without ever calling it a "gateway" or reaching for a dedicated product.

## When a gateway is overkill

A gateway earns its cost when there are genuinely multiple backend services, or when the cross-cutting concerns from Phase 2 - auth, rate limiting, routing - are complex enough to be worth centralizing. It earns its cost less when:

```text
One backend service, no plans to split it soon
  -> a gateway adds a hop and a new failure point, for nothing it's solving

Small internal tool, trusted network, a handful of users
  -> auth/rate-limiting needs are minimal; a gateway is machinery you'll maintain
     for a problem you don't have yet

A load balancer already does everything you actually need
  -> if all you need is "spread traffic across instances of one service,"
     that's a load balancer's job, not a gateway's
```

*What just happened:* the pattern to watch for is reaching for a gateway because it's the "correct" architecture for a system with many services, when your system doesn't actually have many services yet. The same instinct that says "don't split into microservices before you feel real pain" applies here - a gateway in front of one service is complexity paid for in advance, on the hope that you'll need it later.

> The gateway is a tool for a specific shape of problem: many services, one client-facing contract. If you don't have the "many services" part yet, you're paying the tradeoffs from this phase for a benefit you can't cash in yet.

The plain read: gateways are close to mandatory once a system has real service sprawl, and unnecessary weight before that point. The decision isn't about which is more modern - it's about whether the specific problems in Phase 1 are ones you actually have.
