# OAuth2 and OpenID Connect

> The standard behind 'Log in with Google': OAuth2 grants delegated access, OIDC adds identity on top, and the authorization-code-with-PKCE flow ties it together.


---

# OAuth2 and OpenID Connect

You clicked "Log in with Google" a thousand times and never thought about it. Now you have to build it, and the spec reads like a tax form: grants, scopes, redirect URIs, three kinds of token that all look like the same blob of base64. The fear underneath is real, because the cost of getting auth wrong is an account takeover, not a styling bug.

Here is the relief. OAuth2 is one idea wearing a lot of jargon: let an app act on your behalf without handing it your password. OIDC is one small addition on top: prove who you are while you're at it. Once you see the four roles and the one flow that matters in 2026, the spec stops being a wall and becomes a checklist.

## How to read this

Go in order. Phase 1 builds the mental model: the problem these protocols solve and the four roles passing tokens around. Phase 2 walks the authorization-code-with-PKCE flow step by step and pulls apart the three token types. Phase 3 is production reality: token storage, the gotchas that leak accounts, and why "roll your own" is the line you do not cross.

This is a concepts-and-protocol guide. There is no single CLI to install; the examples are HTTP requests and token payloads, which is what you will actually be staring at in your network tab.

## The phases

1. [Phase 1: The Problem and the Four Roles](01-the-problem-and-the-roles.md) - why delegated access exists, and who plays what part.
2. [Phase 2: The Authorization Code Flow and the Three Tokens](02-the-flow-and-the-tokens.md) - the dance, PKCE, scopes, and access vs refresh vs ID tokens.
3. [Phase 3: Production Reality and the Gotchas](03-production-reality-and-gotchas.md) - storing tokens, the mistakes that leak accounts, and why you never roll your own.


---

# The Problem and the Four Roles

Picture the moment before any of this existed. A printing service wants to pull the photos out of your Google account so it can mail you a calendar. The only way it knows to ask is the old way: "give me your Google username and password, I'll log in and grab them."

Stop and feel how bad that is. You're handing a third party the keys to your entire account - email, contacts, everything - so they can read one folder of photos. They store your password somewhere. They can do anything you can do, forever, until you change it. And the only way to revoke them is to change your password, which logs out everything else too. This pattern even had a name: the *password anti-pattern*.

OAuth2 exists to kill that anti-pattern. The whole point is **delegated access**: let an app do one specific thing on your behalf, without ever seeing your password, with access you can revoke independently.

## Authorization is not authentication

Before the roles, nail down the one distinction that confuses everyone. There are two different questions hiding inside "log in":

- **Authentication** - *who are you?* Proving identity.
- **Authorization** - *what are you allowed to do?* Granting permission.

OAuth2 was designed for the second question. It is an **authorization** framework. It answers "is this app allowed to read your photos?" - not "who is the person sitting at the keyboard?"

That gap is exactly why OpenID Connect (OIDC) was created. OIDC is a thin layer on top of OAuth2 that adds the missing piece: a trustworthy answer to *who are you?*. So the slogan to carry through this whole guide is:

> **OAuth2 = authorization (delegated access). OIDC = authentication (identity), built on top of OAuth2.**

When you click "Log in with Google," you are using both at once: OIDC to learn who you are, OAuth2 to (optionally) grant the app some access. If you want the deeper split between these two words, see [/guides/auth-vs-authz](/guides/auth-vs-authz).

## The four roles

Every OAuth2 interaction is four parties passing messages. Learn the roles and the rest of the protocol is bookkeeping. Using the photo-printing example:

- **Resource Owner** - *you*. The human who owns the data and can grant access to it.
- **Client** - *the printing app*. The thing that wants access. It is never trusted with your password; it gets tokens instead.
- **Authorization Server** - *Google's login + consent screens*. It authenticates you, asks "do you allow this app to read your photos?", and issues tokens. This is the party that holds your real credentials.
- **Resource Server** - *the Google Photos API*. It holds the actual data and accepts a token as proof the client is allowed in.

```text
   Resource Owner (you)
        │  approve
        ▼
  Authorization Server  ──issues token──►  Client (the app)
  (Google login/consent)                        │
                                                │ token
                                                ▼
                                      Resource Server (Photos API)
```

*What just happened:* You approve once at the Authorization Server. It hands the Client a token. The Client shows that token to the Resource Server to get the data - and at no point did the password anti-pattern happen. The Client never touched your password.

A subtle but important point: the **Authorization Server and Resource Server are often run by the same company** (Google runs both), but they are *different roles*. One issues tokens, the other accepts them. Keeping them separate in your head explains why a token is a thing that gets handed from one place to another, rather than a magic password.

## Why tokens instead of passwords

The token is the whole innovation. Instead of a password - which is one secret that unlocks everything, forever - the Authorization Server issues a token that is:

- **Scoped** - it grants only specific permissions (read photos), not full account access.
- **Expiring** - it stops working after a while, so a leaked token has a short blast radius.
- **Revocable independently** - you can kill *this app's* access without changing your password or affecting any other app.

That trio - scoped, expiring, revocable - is the entire reason OAuth2 is worth its complexity. Hold onto it; everything in Phase 2 is the machinery that produces such a token safely.

> A common myth: "OAuth logs me in." Plain OAuth2 does *not* log you in - it gets the client permission to call an API. The "log in" feeling comes from OIDC's ID token, which you'll meet next phase. An app that uses raw OAuth2 access tokens as a login mechanism is making a classic security mistake, because an access token says *what you can do*, not *who you are*.

## In the wild

Look at any "Connect your GitHub account" or "Allow this app to post to your calendar" button. That consent screen listing exactly what the app may do - "See your email address," "Manage your repositories" - is the Authorization Server showing you the scopes before it issues a token. The fact that you can later visit a "Connected apps" page and revoke one app without touching the others is the *revocable independently* property, made visible.

```quiz
[
  {
    "q": "What problem was OAuth2 primarily designed to solve?",
    "choices": [
      "Encrypting data in transit between servers",
      "Letting an app act on your behalf without handing it your password",
      "Making login pages load faster",
      "Storing passwords securely in a database"
    ],
    "answer": 1,
    "explain": "OAuth2 kills the 'password anti-pattern' by issuing scoped, revocable tokens for delegated access instead of sharing credentials."
  },
  {
    "q": "Which statement correctly separates the two protocols?",
    "choices": [
      "OAuth2 handles authentication; OIDC handles authorization",
      "OAuth2 handles authorization; OIDC adds authentication on top",
      "They are two names for exactly the same thing",
      "OIDC replaces OAuth2 entirely"
    ],
    "answer": 1,
    "explain": "OAuth2 is an authorization framework (delegated access). OIDC is a thin identity layer (authentication) built on top of it."
  },
  {
    "q": "In the photo-printing example, who is the Resource Server?",
    "choices": [
      "You, the human owner of the photos",
      "The printing app that wants the photos",
      "The Google Photos API that holds the photos and accepts tokens",
      "The Google login and consent screen"
    ],
    "answer": 2,
    "explain": "The Resource Server holds the data and accepts a token as proof of access. The Authorization Server (login/consent) is a separate role even when the same company runs both."
  }
]
```


---

# The Authorization Code Flow and the Three Tokens

You know the four roles. Now watch them actually move. OAuth2 defines several "grant types" (ways to get a token), but in 2026 there is essentially one you should use for apps where a user is present: **Authorization Code flow with PKCE**. Older flows like the *implicit grant* exist in the spec but are deprecated for being leaky - skip them. Learn this one well and you've learned the flow that powers virtually every "Log in with…" button you'll build.

## The flow, step by step

Here is the dance. A user wants to log into your app ("the Client") using their Google account ("the Authorization Server"). Follow the redirects.

```text
1. User clicks "Log in with Google"
2. Client → browser redirect → Authorization Server:
     GET /authorize?response_type=code
                    &client_id=abc123
                    &redirect_uri=https://yourapp.com/callback
                    &scope=openid email profile
                    &state=xyz789
                    &code_challenge=BASE64URL(SHA256(verifier))
                    &code_challenge_method=S256
3. User authenticates at Google + approves the consent screen
4. Authorization Server → browser redirect → Client:
     GET https://yourapp.com/callback?code=AUTH_CODE&state=xyz789
5. Client (back end) → Authorization Server, direct POST:
     POST /token
       grant_type=authorization_code
       code=AUTH_CODE
       redirect_uri=https://yourapp.com/callback
       client_id=abc123
       code_verifier=ORIGINAL_RANDOM_STRING
6. Authorization Server returns tokens (JSON):
     { "access_token": "...", "id_token": "...", "refresh_token": "...", "expires_in": 3600 }
```

*What just happened:* The browser only ever carries a short-lived **authorization code** (steps 2–4), never the real tokens. The actual tokens come back over a direct server-to-server POST (step 5) that the browser never sees. That two-step shuffle - code in the front channel, tokens in the back channel - is the entire security design.

Why bother with the intermediate code at all? Because the redirect in step 4 travels through the user's browser, where it can land in logs, history, or a malicious extension. A code is useless on its own - exchanging it requires the second request. So even if someone steals the code, they're missing a piece.

## PKCE: the piece that stops code theft

That missing piece is **PKCE** (Proof Key for Code Exchange, pronounced "pixy"). It closes the hole where an attacker intercepts the authorization code and tries to redeem it themselves.

It works with a one-time secret the client invents at the start:

```text
At step 2 (start):
  code_verifier  = a random high-entropy string the client generates and keeps
  code_challenge = BASE64URL( SHA256(code_verifier) )   ← sent in the /authorize request

At step 5 (token exchange):
  client sends the original code_verifier
  Authorization Server checks:  SHA256(code_verifier) == the code_challenge it stored?
```

*What just happened:* The client commits to a secret up front by sending only its *hash* (the challenge). To redeem the code later, it must reveal the original (the verifier). An attacker who stole the code from the browser never saw the verifier, can't reverse the SHA-256 hash, and so can't complete the exchange. The stolen code is dead in their hands.

PKCE started life as protection for mobile and single-page apps that can't keep a client secret, but current guidance is to use it for **every** authorization-code flow, server-side ones included. Treat it as mandatory.

> **What about `state`?** Different job. `code_challenge`/PKCE stops *code interception*. The `state` parameter (a random value you send and check came back unchanged) stops **CSRF** - an attacker tricking your app into completing *their* login. Send both, always. They guard against different attacks.

## Scopes: asking for exactly what you need

In step 2 you saw `scope=openid email profile`. **Scopes** are the specific permissions the client requests. The Authorization Server shows them on the consent screen and bakes the granted ones into the access token.

```text
scope=openid               → "I want an ID token" (the OIDC opt-in)
scope=email                → access to the user's email address
scope=profile              → access to basic profile (name, picture)
scope=https://www.googleapis.com/auth/calendar.readonly
                           → read-only access to their calendar
```

*What just happened:* Each scope is one slice of permission. The golden rule is **least privilege**: ask only for what your feature actually needs. Requesting `calendar` (read-write) when you only display events trains users to rubber-stamp scary permissions and widens your blast radius if a token leaks. Note the literal scope `openid` - that single word is what turns a plain OAuth2 request into an OIDC request and makes the server return an ID token.

## The three tokens - the part everyone confuses

Step 6 returned three different tokens. They look identical (often base64-ish blobs) but have completely different jobs. Mixing them up is the single most common OAuth mistake.

| Token | Answers | Who reads it | Lifetime |
|-------|---------|--------------|----------|
| **Access token** | "What may this client do?" | The **Resource Server** (the API) | Short (minutes to an hour) |
| **ID token** | "Who is the user?" | The **Client** (your app) | Short, single-use at login |
| **Refresh token** | "Get me a new access token" | The **Authorization Server** only | Long (days to months) |

**Access token** - your proof to the *API*. You attach it to API calls and the Resource Server checks it:

```text
GET /v1/photos HTTP/1.1
Host: photos.googleapis.com
Authorization: Bearer <access_token>
```

*What just happened:* The API trusts the bearer token and returns the data. Critically, the access token is **opaque to the client** - your app should not crack it open to learn who the user is. It is addressed to the API, not to you. Using it for login is the classic mistake from Phase 1.

**ID token** - this is the OIDC addition and the thing that actually logs the user in. It is always a **JWT** (a signed JSON Web Token) with claims about the user:

```text
{
  "iss": "https://accounts.google.com",   ← issuer: who minted this
  "aud": "abc123",                         ← audience: YOUR client_id
  "sub": "10769150350006150715",           ← subject: stable unique user id
  "email": "ada@example.com",
  "name": "Ada Lovelace",
  "iat": 1709400000,                       ← issued-at
  "exp": 1709403600                         ← expiry
}
```

*What just happened:* Your app reads these claims to know who logged in. The `sub` ("subject") is the stable user identifier - use that as your primary key, never the email, because emails change. But you must **verify the signature and check the claims** before trusting any of it: confirm `iss` is the expected issuer, `aud` equals your own `client_id`, and `exp` is in the future. An ID token you didn't validate is just a base64 string anyone could forge.

**Refresh token** - the long-lived ticket for getting fresh access tokens without dragging the user back through login. When the access token expires:

```text
POST /token
  grant_type=refresh_token
  refresh_token=<refresh_token>
  client_id=abc123
```

*What just happened:* You trade the refresh token for a brand-new access token (and sometimes a new refresh token). This is why you stay logged into apps for weeks despite access tokens expiring in an hour. Because it's long-lived and powerful, the refresh token is the **most sensitive** of the three - it goes only to the Authorization Server, lives only on a trusted back end, and never near a browser if you can help it. More on guarding it in Phase 3.

## For builders

A clean mental shorthand: the **access token is for machines** (one API checking another caller), the **ID token is for you** (your app learning the user's identity), and the **refresh token is for time** (surviving past the access token's short life). When you wire up a login, your back end validates the ID token to create a session, stashes the refresh token securely, and uses access tokens to call downstream APIs. Three tokens, three jobs, no overlap.

```quiz
[
  {
    "q": "Why does the Authorization Code flow return a short-lived code through the browser instead of the tokens directly?",
    "choices": [
      "Codes are smaller and load faster",
      "The browser can't store tokens at all",
      "The code travels the risky front channel, but is useless without a second back-channel exchange",
      "It lets the user copy the code manually"
    ],
    "answer": 2,
    "explain": "A stolen code is worthless alone - redeeming it needs the direct server-to-server token request, keeping real tokens off the front channel."
  },
  {
    "q": "What specific attack does PKCE defend against?",
    "choices": [
      "CSRF, where an attacker forces your app to complete their login",
      "Authorization code interception, where a stolen code gets redeemed by an attacker",
      "Brute-forcing the user's password",
      "Replaying an expired access token"
    ],
    "answer": 1,
    "explain": "PKCE ties the code to a secret verifier the attacker never saw; the stolen code can't be redeemed. CSRF is the job of the separate 'state' parameter."
  },
  {
    "q": "Your app needs to know which user just logged in. Which token do you read, and how?",
    "choices": [
      "The access token - decode it to read the user's name",
      "The refresh token - it contains the user id",
      "The ID token - verify its signature and claims, then read 'sub'",
      "Any of them - they all carry identity"
    ],
    "answer": 2,
    "explain": "Identity lives in the ID token (a JWT). Verify signature, iss, aud, and exp, then use the stable 'sub' as the user key. Access tokens are opaque and addressed to the API, not to you."
  }
]
```

Watch it animated: [the OAuth authorization code flow](/explainers/OAuthFlow.dc.html)


---

# Production Reality and the Gotchas

The flow works in the demo. Now it has to survive real users, real attackers, and the 3am page. This phase is the stuff the spec mentions in passing and the tutorials skip - the parts that decide whether your auth is solid or a breach waiting to happen.

## Where do you put the tokens?

The first real decision is storage, and it trips up a lot of single-page apps. The question is: after login, where do the tokens live?

```text
localStorage          → readable by ANY JavaScript on the page.
                        One XSS bug = every token stolen. Avoid for tokens.

JS-readable memory    → gone on refresh; XSS during the session can still read it.

httpOnly cookie       → set by the server, invisible to JavaScript.
                        XSS can't read it. Needs CSRF protection (SameSite).

Back-end session      → tokens never reach the browser at all.
                        Browser holds only a session cookie. Safest.
```

*What just happened:* The safer you go down that list, the less an XSS bug can steal. The pattern that ages best is the **Backend-for-Frontend (BFF)**: your server completes the OAuth flow, keeps the access and refresh tokens server-side, and gives the browser only an `httpOnly`, `Secure`, `SameSite` session cookie. The tokens never touch JavaScript, so a script-injection bug can't exfiltrate them.

If tokens *must* live in the browser (a pure SPA with no back end), keep them in memory, never in `localStorage`, and lean on short access-token lifetimes plus PKCE. But given the choice, push token handling to a back end.

> **Cookies still need TLS.** Everything here assumes the whole flow runs over HTTPS. Authorization codes, tokens, and session cookies are bearer secrets - anyone who reads them in transit owns the session. Mark cookies `Secure` and serve over TLS end to end. If TLS itself is fuzzy, read [/guides/https-and-tls](/guides/https-and-tls).

## The gotchas that actually leak accounts

These are the recurring mistakes. Each one has caused real breaches.

**1. Not validating the ID token.** Reading the JWT claims without checking the signature, `iss`, `aud`, and `exp` means accepting forged identities. A signed token you don't verify is decoration.

**2. Open redirect via `redirect_uri`.** The Authorization Server must match `redirect_uri` against a **pre-registered exact value**, not a prefix or wildcard. A loose match lets an attacker redirect the code to their own server.

```text
Registered:  https://yourapp.com/callback
Attacker tries:
  https://yourapp.com.evil.com/callback     ← different host, must be rejected
  https://yourapp.com/callback/../steal     ← path trick, must be rejected
```

*What just happened:* Exact-match registration is what keeps the authorization code flowing only to *you*. Most providers enforce this; the gotcha is registering a sloppy URI (or a wildcard) on your side and opening the hole yourself.

**3. Forgetting `state`.** Skip the CSRF check from Phase 2 and an attacker can stitch their account onto your user's session. Generate `state` randomly, store it, and reject any callback where it doesn't match.

**4. Using the access token to identify the user.** Said in Phase 2, repeated because it keeps happening: the access token is for the *API*, the ID token is for *identity*. Treating an access token as proof of who someone is - especially one minted for a different app - is a known account-takeover vector.

**5. Mismatched `aud`.** A token issued for *another* client is still a valid, correctly-signed token. If you don't check that `aud` equals *your* `client_id`, you'll accept tokens minted for someone else. This is the "confused deputy" trap.

## Expiry, revocation, and logout

Tokens expire, and "logout" is messier than it looks.

- **Access tokens expire fast** (often around an hour). That's a feature - a leaked one dies quickly. Your back end refreshes silently using the refresh token.
- **Refresh tokens can be revoked.** When a user clicks "remove this app" or you detect compromise, revoke the refresh token at the Authorization Server. Better providers also do **refresh-token rotation**: each use issues a new refresh token and invalidates the old one, so a stolen-and-reused token gets caught.
- **Logout has layers.** Clearing your own session cookie logs the user out of *your* app. It does **not** log them out of Google. OIDC defines separate end-session / front-channel logout mechanisms for that - know that "log out" usually means only your session unless you wire up the rest.

```text
User clicks "Log out":
  1. Destroy your back-end session  ← logs out of YOUR app  (the common case)
  2. (optional) Revoke refresh token at the Authorization Server
  3. (optional) Redirect to provider end-session endpoint  ← logs out of the IdP
```

*What just happened:* Step 1 is what most apps mean by logout. Steps 2 and 3 are deliberate extra work; if you assume a single "logout" nukes everything, you'll be surprised when the user is still signed into the provider.

## Why you never roll your own

You might be tempted to hand-build the token endpoints, the JWT verification, the PKCE math. Don't. Not because you can't - because the failure mode is silent.

A bug in your business logic throws an error someone notices. A bug in your auth - a skipped `aud` check, a timing leak in signature comparison, a wildcard redirect - produces software that *works perfectly in every demo* and quietly grants account takeover. There's no failing test, no error in the logs, until it's a disclosure email.

The lazy move here is also the correct one:

- **Use a battle-tested library** for the client side (the well-known certified OIDC/OAuth client for your language and framework). It handles PKCE, state, and JWT validation correctly so you don't reinvent the subtle parts.
- **Use an established Authorization Server / IdP** rather than writing one - a hosted provider or a mature self-hostable identity server. Issuing tokens correctly (key rotation, consent, revocation, rate limits) is a product, not a weekend.

You still need to understand the flow - that's why this guide exists - so you can configure it correctly, read the network tab when it breaks, and catch a misconfigured `redirect_uri`. But the cryptographic and protocol plumbing is solved. Reach for the proven implementation.

## In the wild

When you wire up "Log in with Google" through a mature library, almost everything in this guide happens for you: the library builds the `/authorize` URL with PKCE and `state`, handles the callback, exchanges the code, and validates the ID token's signature and claims. Your job shrinks to three things - register an exact `redirect_uri`, request least-privilege scopes, and keep the refresh token on a trusted back end. Get those three right and you've used the standard the way it was meant to be used.

```quiz
[
  {
    "q": "What is the safest place to keep OAuth tokens for a web app?",
    "choices": [
      "In localStorage so they survive page refreshes",
      "On a trusted back end, with the browser holding only an httpOnly session cookie",
      "In a JavaScript variable shared across all scripts",
      "In a regular cookie readable by JavaScript"
    ],
    "answer": 1,
    "explain": "The Backend-for-Frontend pattern keeps tokens server-side so an XSS bug can't read them; the browser only gets an httpOnly, Secure, SameSite session cookie."
  },
  {
    "q": "Why must the Authorization Server match redirect_uri against a pre-registered exact value?",
    "choices": [
      "To make the URL shorter",
      "To let the user bookmark the callback",
      "A loose or wildcard match lets an attacker redirect the authorization code to their own server",
      "Exact matching makes the token larger and harder to forge"
    ],
    "answer": 2,
    "explain": "Exact-match registration ensures the code only ever lands at your real callback. A wildcard or prefix match opens an account-takeover hole."
  },
  {
    "q": "What's the strongest argument against hand-rolling your own OAuth/OIDC implementation?",
    "choices": [
      "Libraries are always faster to run",
      "Auth bugs fail silently - they work in every demo while granting account takeover",
      "It's against the OAuth2 specification to write your own",
      "Modern languages can't do the required cryptography"
    ],
    "answer": 1,
    "explain": "A skipped aud check or wildcard redirect produces software that passes every demo and quietly enables takeover, with no error until a breach. Use battle-tested libraries and IdPs."
  }
]
```
