# Realtime APIs: WebSockets vs SSE vs Polling

> How to push live updates: the real tradeoffs between polling, Server-Sent Events, and WebSockets, and which to reach for when.


---

# Realtime APIs: WebSockets vs SSE vs Polling

You've built the feature, and now someone wants it *live*. The dashboard should tick up without a
refresh. The chat should land instantly. The "3 people are typing" dot should actually move. And you're
staring at a plain HTTP API thinking: it answers when I ask, but it can't tap me on the shoulder. That
gap - between request/response and "tell me the moment something changes" - is exactly the one this
guide closes.

The relief: there are only three patterns worth knowing, and they line up from simplest to most powerful.
**Polling** fakes realtime by asking on a loop. **Server-Sent Events** opens a one-way pipe so the server
can stream updates to the browser over plain HTTP. **WebSockets** open a two-way pipe for when both sides
need to talk. The whole skill is picking the *simplest one that fits* - and knowing the scaling traps
before they page you at 3am.

## How to read this

- **Need to ship something live this week?** Phase 1 names the three patterns and gives you a decision
  rule you can apply in a minute. Phase 2 is the hands-on core: real SSE and WebSocket code you can
  copy.
- **Want it to actually make sense?** Read in order. Phase 1 builds the mental model, Phase 2 shows how
  each one really works, and Phase 3 covers the parts that bite at scale - sticky sessions, fan-out, and
  knowing when *not* to reach for WebSockets.

## The phases

1. **[Why HTTP Can't Push](01-why-http-cant-push.md)** - the core mental model: request/response is a
   pull, realtime needs a push, and the three patterns (polling, SSE, WebSockets) are three different
   ways to fake or build that push.
2. **[The Three Patterns in Practice](02-the-three-patterns.md)** - the everyday core: long-polling,
   a real SSE stream with auto-reconnect, and a full-duplex WebSocket, with the comparison table that
   tells you which to grab.
3. **[When It Breaks at Scale](03-when-it-breaks-at-scale.md)** - the deeper payoff: why one server is
   easy and many servers are hard, sticky sessions, fan-out with a pub/sub backplane, and the clear
   case for reaching for the *simplest* pattern.


---

# Why HTTP Can't Push

Here's the thing nobody says out loud when you first learn web development: a normal HTTP request is a
phone call where *only the client is allowed to dial.* You (the browser) call the server, say what you
want, the server answers, and then the line hangs up. The server never calls you back. It *can't* - it
doesn't have your number, and even if it did, the connection is already closed.

That's completely fine for "show me my profile" or "save this form." It falls apart the instant you want
something *live*. A new message, a price tick, a build that turned green - those happen on the
server's side, on the server's schedule, and the server has no way to lean over and tell you. This phase
is about why that limitation exists, and the three practical ways around it.

## The shape of a normal request

HTTP is request/response. One request, one response, then done. The client always speaks first; the
server only ever *replies*. There is no slot in the protocol for the server to say something unprompted.

```console
$ curl https://api.example/v1/notifications
[{"id":1,"text":"Welcome!"}]
$
```
You asked once, got the current list, and the connection closed. If a new notification arrives one
second later, this terminal has no idea - the server can't reach back through a hung-up line, you'd have
to run `curl` again to find out.

📝 **Terminology.** "Realtime" on the web rarely means microsecond-precise. It means *the user sees the
change within a second or so of it happening, without doing anything.* That's the bar these patterns
clear - not hard-realtime like a flight controller.

## The pull-vs-push problem, in one picture

```mermaid
sequenceDiagram
  participant C as Client
  participant S as Server
  Note over C,S: Normal HTTP - client always dials
  C->>S: GET /notifications
  S-->>C: here's the list (then hang up)
  Note over C,S: Something new happens on the server...
  Note over S: ...but the server can't call back
```

The whole problem fits in that gap after "something new happens." Every realtime pattern is an answer to
*one* question: how does the server get a message to a client who isn't currently asking?

## The three answers, from simplest to strongest

There are exactly three patterns you need, and they sit on a ladder. Climb only as high as you must.

**1. Polling - fake the push by asking on a loop.**
The client keeps asking, over and over: "anything new? anything new?" It's not really a push at
all - it's a pull on a timer. Dead simple, works with the HTTP you already have, and wasteful: most
answers are "nope." There's a smarter cousin, **long-polling**, where the server *holds the request
open* until it has something to say - we'll build it in Phase 2.

**2. Server-Sent Events (SSE) - a one-way pipe, server to client.**
The client opens *one* HTTP request and the server, instead of answering and hanging up, keeps the line
open and *streams* events down it as they happen. Traffic flows one direction: server → client. It runs
over plain HTTP, the browser has it built in (`EventSource`), and it reconnects automatically if the
line drops. Perfect for feeds, notifications, live dashboards, progress bars - anything where the client
mostly listens.

**3. WebSockets - a two-way pipe, both directions.**
Both sides can send at any time, independently, over a single long-lived connection. This is *full-
duplex*: the server can push to you and you can push to the server in the same breath. It's the right
tool when both ends genuinely talk - chat, multiplayer games, collaborative editing, a shared cursor.
It's also the most machinery, so don't reach for it out of habit.

💡 **Key point.** These aren't competitors to grade. They're a ladder of capability and cost. The
question is never "which is best" - it's "what's the *least* I need?"

## The one-line decision rule

When someone asks for realtime, run this in your head before writing any code:

- **Does the client only need to *receive* updates?** → SSE. (Or polling, if updates are rare.)
- **Do *both* sides need to send, constantly and independently?** → WebSockets.
- **Are updates rare and latency-tolerant (every 30s is fine)?** → Polling. Don't open a persistent
  connection for a number that changes twice an hour.

```text
receive only ............. SSE   (or polling if rare)
two-way, constant ........ WebSockets
rare + lazy .............. Polling
```
You turned a vague "make it realtime" into a concrete pick in three questions. The default answer for
"show me live stuff" is SSE, not WebSockets - most teams reach for the heavy tool out of reflex and
regret the scaling bill later.

⚠️ **WebSockets are not the default.** They feel like the "real" realtime tech, so people grab them for
a one-way notification feed and inherit a pile of complexity (no automatic reconnect, harder to scale,
trickier through proxies) they didn't need. If data flows one way, SSE is almost always the calmer
choice.

🪖 **War story.** A team built a stock-ticker dashboard - pure server-to-client price updates - on
WebSockets because "realtime means WebSockets." Six months in, they were hand-rolling reconnect logic and
heartbeats, and fighting a corporate proxy that kept killing the socket. They rewrote it on SSE in an
afternoon: the browser handled reconnect for free, and the proxy stopped complaining because it was plain
HTTP the whole time.

## For builders

Before you add *any* realtime layer, ask whether the feature needs it at all. A "live" view that updates
on the user's next click or navigation often feels fine and costs nothing. Realtime is a recurring
operational cost - open connections, more servers, more failure modes. Spend it where the user actually
notices the delay (chat, collaboration, fast-moving numbers), not everywhere you *could*.

If you're coming from the request/response world, these patterns sit on top of the same foundation - 
[REST APIs, Explained](/guides/rest-apis-explained) is the model they extend, and for "the server tells
me later" *between systems* rather than to a browser, that's [Webhooks & Message
Queues](/guides/webhooks-and-message-queues), a different tool for a different shoulder-tap.

## Recap

1. **Plain HTTP can't push** - it's request/response, the client always dials, the server only replies
   and then hangs up.
2. Every realtime pattern answers one question: **how does the server reach a client who isn't asking
   right now?**
3. **Three answers, on a ladder:** polling (ask on a loop - simple, wasteful), SSE (one-way stream over
   HTTP - great for feeds), WebSockets (two-way pipe - for chat and collaboration).
4. **Pick the simplest that fits.** Receive-only → SSE. Both sides talk → WebSockets. Rare and lazy →
   polling. WebSockets are not the default.

Next, we stop talking about them and build all three - including the long-polling trick and a real SSE
stream that reconnects itself.

```quiz
[
  {
    "q": "Why can't a plain HTTP server push an update to a client on its own?",
    "choices": [
      "HTTP is encrypted, so the server can't read the client's address",
      "HTTP is request/response - the client always initiates, and the connection closes after the reply",
      "Servers are too slow to send unprompted messages",
      "Browsers block all incoming server messages for security"
    ],
    "answer": 1,
    "explain": "HTTP is a pull model: the client dials, the server replies, the line closes. There's no slot for the server to speak first."
  },
  {
    "q": "A feature needs to stream live notifications from the server to the browser, one direction only. What's the calmest fit?",
    "choices": [
      "WebSockets, because realtime always means WebSockets",
      "Polling every 100ms",
      "Server-Sent Events (SSE) - a one-way stream over plain HTTP with built-in reconnect",
      "A new HTTP request for every notification"
    ],
    "answer": 2,
    "explain": "One-way server-to-client is exactly SSE's job: plain HTTP, built-in EventSource, automatic reconnect. WebSockets would be extra machinery you don't need."
  },
  {
    "q": "What distinguishes WebSockets from SSE?",
    "choices": [
      "WebSockets are full-duplex - both sides can send at any time; SSE is one-way, server to client",
      "WebSockets don't need a connection",
      "SSE is faster than WebSockets in every case",
      "WebSockets only work for chat apps"
    ],
    "answer": 0,
    "explain": "WebSockets are two-way (full-duplex); SSE streams in one direction only. Reach for WebSockets when both ends genuinely talk."
  }
]
```


---

# The Three Patterns in Practice

You've got the mental model: HTTP can't push, and there are three ways around it. Now let's make them
real. We'll go up the ladder - polling, then SSE, then WebSockets - and for each one you'll see the
actual code on both ends and exactly what travels over the wire. By the end you'll be able to copy the
right pattern, not only name it.

A small grounding note before we start: all three of these run *over the same TCP connections your
browser already uses.* Nothing magic, no new network. They're different conventions for keeping a
conversation going past the usual one-request-one-reply.

## Polling and its smarter cousin, long-polling

**Plain polling - ask on a timer.**
```javascript
// Client: ask every 5 seconds, forever.
setInterval(async () => {
  const res = await fetch("/api/messages?since=" + lastId);
  const msgs = await res.json();
  if (msgs.length) render(msgs);
}, 5000);
```
Every 5 seconds you fire a request asking "anything after `lastId`?" Most of the time the answer is an
empty array - wasted round-trips - and a new message can sit unseen for up to 5 seconds. Cheap to build,
but it's a tax you pay forever.

**Long-polling - let the server hold the line.** The trick: the server *doesn't answer immediately.* It
holds your request open until it actually has something, or until a timeout, then replies. The client
gets the answer the moment it exists, then immediately reconnects.

```javascript
// Client: reconnect the instant the server answers.
async function longPoll() {
  while (true) {
    const res = await fetch("/api/messages?since=" + lastId);
    const msgs = await res.json();   // resolves only when there's news (or a timeout)
    if (msgs.length) { render(msgs); lastId = msgs.at(-1).id; }
  }
}
longPoll();
```
The `await` parks here until the server sends something back, instead of you hammering on a timer. You
get near-instant delivery with no persistent connection - but you're still opening a fresh HTTP request
per message, and a held-open request ties up a server slot. Long-polling is the plain fallback when SSE
and WebSockets aren't available; otherwise, climb the ladder.

📝 **Terminology.** "Long-polling" sounds like a kind of polling, but it behaves like a push: the latency
is "as soon as there's news," not "next time the timer fires." The cost is connection churn - a new
request after every single delivery.

## Server-Sent Events: one stream, server to client

This is the one most people *should* be reaching for and don't. SSE keeps a single HTTP response open
and dribbles events down it. The browser's `EventSource` handles the connection - including reconnecting
if it drops - so you write almost nothing.

**The client - the whole thing.**
```javascript
const stream = new EventSource("/api/stream");

stream.onmessage = (event) => {
  const data = JSON.parse(event.data);
  render(data);
};

stream.onerror = () => {
  // EventSource is ALREADY retrying. You don't reconnect by hand.
  console.log("connection dropped - browser is reconnecting...");
};
```
Three lines and you have a live feed. `onmessage` fires every time the server sends an event. When the
line drops, `EventSource` reconnects on its own and even tells the server where you left off - that
auto-reconnect is the single biggest reason to prefer SSE over a hand-rolled WebSocket for one-way data.

**The server - what it actually sends.** SSE isn't JSON-over-HTTP; it's a tiny text format. The
content type is `text/event-stream`, and each event is `data:` lines ending in a blank line.

```text
HTTP/1.1 200 OK
Content-Type: text/event-stream
Cache-Control: no-cache
Connection: keep-alive

data: {"price": 142.10}

data: {"price": 142.35}

id: 48
data: {"price": 142.80}
```
One response that never ends. Each `data:` block is one event the browser hands to `onmessage`. The `id:`
line is the magic behind reconnect - the browser remembers the last id and sends it back as a
`Last-Event-ID` header when it reconnects, so the server can resume from where you dropped instead of
replaying everything.

**The server side, in code.**
```javascript
// Express-style handler.
app.get("/api/stream", (req, res) => {
  res.writeHead(200, {
    "Content-Type": "text/event-stream",
    "Cache-Control": "no-cache",
    "Connection": "keep-alive",
  });

  const tick = setInterval(() => {
    res.write(`data: ${JSON.stringify({ price: nextPrice() })}\n\n`);
  }, 1000);

  req.on("close", () => clearInterval(tick));  // stop when the client leaves
});
```
You set the streaming headers, then `res.write()` an event whenever you have news - note the `\n\n` that
ends each event. Crucially, you clean up on `close`, or you'll leak a timer (and eventually a connection)
for every client who ever wandered off. That cleanup is the most-forgotten line in SSE code.

⚠️ **The six-connection cap (HTTP/1.1).** Over HTTP/1.1 a browser allows only ~6 connections per domain,
and an open SSE stream eats one of them *per tab.* Open the app in a few tabs and you can starve your own
page of connections. The fix is HTTP/2, which multiplexes many streams over one connection - so serve
SSE over HTTP/2 in production and the cap effectively disappears.

## WebSockets: a two-way pipe

When both sides need to talk, you want a WebSocket. It starts life as a normal HTTP request that asks to
*upgrade* the connection, and once that handshake succeeds, the same TCP connection becomes a two-way
channel where either side sends whenever it likes.

**The handshake - how a WebSocket is born.**
```console
GET /chat HTTP/1.1
Host: yourapp.example
Upgrade: websocket
Connection: Upgrade
Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==

HTTP/1.1 101 Switching Protocols
Upgrade: websocket
Connection: Upgrade
```
The client asked HTTP to "upgrade" to the WebSocket protocol; the server agreed with a
`101 Switching Protocols`. From this point the connection is no longer request/response - it's an open,
two-way pipe. That `101` is the moment a WebSocket exists.

**The client.**
```javascript
const ws = new WebSocket("wss://yourapp.example/chat");

ws.onopen  = () => ws.send(JSON.stringify({ type: "join", room: "general" }));
ws.onmessage = (event) => render(JSON.parse(event.data));

ws.onclose = () => {
  // Unlike SSE, NOTHING reconnects for you. You must do it yourself.
  setTimeout(connect, 1000);
};
```
You can `send()` to the server *and* receive via `onmessage`, both directions, any time - that's the
full-duplex superpower SSE doesn't have. But notice `onclose`: WebSockets give you *no* automatic
reconnect. If you want it (and you do), you write the retry loop, the backoff, and the heartbeats
yourself. That extra labor is the real price of WebSockets.

💡 **Key point.** Use `wss://` (TLS), never `ws://`, in production - same reasoning as `https`. Plenty of
corporate proxies also drop plain `ws://` while letting `wss://` through, so TLS isn't only about
secrecy here, it's about the connection surviving at all.

## The comparison table

Pin this. It's the whole guide compressed:

| | Polling | Long-polling | SSE | WebSockets |
|---|---|---|---|---|
| **Direction** | client pulls | client pulls | server → client | both ways |
| **Connection** | new each time | held, then re-opened | one, long-lived | one, long-lived |
| **Protocol** | plain HTTP | plain HTTP | plain HTTP | upgraded from HTTP |
| **Auto-reconnect** | n/a | manual | **built in** | **manual** |
| **Browser API** | `fetch` | `fetch` | `EventSource` | `WebSocket` |
| **Latency** | up to interval | near-instant | near-instant | near-instant |
| **Best for** | rare updates | fallback | feeds, notifications, dashboards | chat, games, collaboration |
| **Main cost** | wasted requests | connection churn | one-way only | most complexity |

⚠️ **Don't open a connection per number.** A persistent SSE or WebSocket connection has a real cost - it
holds a server resource for as long as it's open. For a value that changes every few minutes, polling is
genuinely the *better* engineering, not the lazy one. Reserve persistent connections for genuinely live
data.

## For builders

Start at the bottom of the table and stop the moment a row fits. Most "live" features are feeds or
dashboards - that's SSE, and the browser does the hard part (reconnect, resume) for free. Only climb to
WebSockets when you can point at a client→server message that has to fly *upward* constantly. And before
either, ask whether plain polling at a sane interval is genuinely good enough; very often it is, and it's
the only one of the four with no persistent-connection bill.

## Recap

1. **Polling** asks on a timer (wasteful); **long-polling** holds the request open so the answer arrives
   the instant there's news - the plain fallback when SSE/WebSockets aren't options.
2. **SSE** is a one-way `text/event-stream` the server keeps open; `EventSource` reconnects and resumes
   for free. Clean up on client `close`, and serve over HTTP/2 to dodge the 6-connection cap.
3. **WebSockets** upgrade an HTTP request (the `101`) into a two-way pipe - full-duplex, but you write
   your own reconnect, backoff, and heartbeats. Use `wss://`.
4. **The table picks for you:** receive-only feed → SSE; both sides talk constantly → WebSockets; rare
   updates → polling.

You can now build any of the three. The last thing that separates a demo from production is what happens
when one server becomes ten - and that's where realtime gets genuinely hard.

```quiz
[
  {
    "q": "What makes long-polling feel like a push even though it uses ordinary HTTP requests?",
    "choices": [
      "It sends many requests per second",
      "The server holds the request open until it has news, so the answer arrives the moment it exists",
      "It uses a special long-polling protocol",
      "It keeps one connection open forever like SSE"
    ],
    "answer": 1,
    "explain": "Long-polling parks the request on the server until there's something to say (or a timeout), so latency is 'as soon as there's news' rather than 'next timer tick'."
  },
  {
    "q": "Why is SSE often preferred over a hand-rolled WebSocket for a one-way feed?",
    "choices": [
      "SSE is binary and therefore faster",
      "SSE supports two-way messaging",
      "The browser's EventSource reconnects automatically and can resume via Last-Event-ID, so you write almost nothing",
      "WebSockets can't carry JSON"
    ],
    "answer": 2,
    "explain": "For one-way data, EventSource handles reconnect and resume for free. WebSockets give you no automatic reconnect - you'd build it yourself."
  },
  {
    "q": "What does the HTTP 101 Switching Protocols response signify in a WebSocket setup?",
    "choices": [
      "An error opening the stream",
      "The server agreed to upgrade the HTTP connection into a two-way WebSocket pipe",
      "The client must poll instead",
      "The TLS handshake completed"
    ],
    "answer": 1,
    "explain": "A WebSocket starts as an HTTP request asking to upgrade; the server's 101 Switching Protocols confirms it, and the connection becomes full-duplex from then on."
  }
]
```


---

# When It Breaks at Scale

Everything in Phase 2 works beautifully - on one server. You demo it, it's instant, everyone's happy.
Then traffic grows, you add a second server behind a load balancer, and realtime quietly breaks in a way
that's maddening to debug: messages arrive for *some* users and vanish for others, seemingly at random.
Nothing is wrong with your code. The problem is that a persistent connection and a stateless load
balancer fundamentally disagree, and this phase is about that fight.

This is the payoff phase. Once you see *why* one server is easy and many is hard, the whole landscape of
"realtime at scale" - sticky sessions, backplanes, fan-out - stops being jargon and becomes one
problem with a couple of standard solutions.

## Why one server is easy and many is hard

On a single server, every connected client is *right there* - they're entries in one in-memory list. To
broadcast a message, you loop over the list and write to each. Done.

Connections are long-lived and *pinned to whichever server accepted them.* Alice's WebSocket lives on
Server A; Bob's lives on Server B. Now Alice sends a chat message. It arrives at Server A. Server A loops
over *its* list of connections - which doesn't include Bob. Bob never hears it. The message is stranded
on the wrong machine.

```mermaid
flowchart TB
  A[Alice] -->|connected to| S1[Server A]
  B[Bob] -->|connected to| S2[Server B]
  A -->|sends message| S1
  S1 -.->|knows only A| A
  S1 -.-x|can't reach Bob| B
```
Alice's message reached Server A, but Bob's connection lives on Server B. Server A has no way to push to
a client it isn't holding. With persistent connections, your clients are scattered across machines, and
no single machine can see them all.

📝 **Terminology.** This is the **fan-out** problem: one incoming event has to reach *many* connected
clients, but those clients are spread across servers that don't share memory. Solving fan-out is the
core of scaling any realtime system.

## Problem one: the load balancer keeps cutting the line

Before fan-out even bites, there's a more basic break. A normal load balancer spreads each *request*
across servers - that's its whole job. But a WebSocket or SSE stream is *one* long-lived connection that
must stay on the server that's holding it. Worse, the handshake and the upgrade can land on different
servers, and the connection dies before it starts.

**The fix: sticky sessions.** You configure the load balancer so that once a client lands on a server,
it *stays* on that server for the life of the connection (often keyed by a cookie or source IP).

```text
Without stickiness:  Alice's upgrade → Server A
                     Alice's frames  → Server B   ✗ (B never saw the handshake - dies)

With stickiness:     Alice's upgrade → Server A
                     Alice's frames  → Server A   ✓ (same server, connection survives)
```
Sticky sessions pin a client to one server so the long-lived connection isn't ripped apart by ordinary
load balancing. It's the first thing to check when WebSockets "randomly" disconnect behind a load
balancer - and a config most teams forget until it bites.

⚠️ **Stickiness is necessary but not sufficient.** Pinning Alice to Server A keeps *her* connection
alive, but it does nothing to get her message to Bob on Server B. Sticky sessions fix the connection
problem; they don't fix fan-out. People frequently turn on stickiness, see connections stabilize, and
are baffled that cross-server messages still vanish.

## Problem two: getting the message to the other servers

To deliver Alice's message to Bob, Server A needs a way to shout to *every* server "here's a message for
room `general`," and each server then pushes it to its own local connections in that room. That shared
shout-channel is a **backplane**, and it's almost always a pub/sub system.

**How it works.**
1. Every server **subscribes** to the backplane (commonly Redis pub/sub, or a message broker).
2. Alice's message hits Server A. Server A **publishes** it to the backplane instead of only looping its
   own list.
3. The backplane **fans it out** to every subscribed server.
4. Each server pushes the message to *its* local connections for that room. Bob, on Server B, finally
   gets it.

```mermaid
flowchart LR
  A[Alice] --> S1[Server A]
  S1 -->|publish| PS[(Pub/Sub backplane)]
  PS -->|fan-out| S1
  PS -->|fan-out| S2[Server B]
  S2 --> B[Bob]
```
No single server has to know every client anymore. Servers only know their *own* connections; the
backplane is the shared nervous system that carries every message to every server. This one pattern - 
publish to a backplane, fan out to local connections - is how essentially all large realtime systems
scale.

🪖 **War story.** A chat app worked flawlessly in staging (one server) and broke the day it scaled to
three in production. Messages reached maybe a third of users - exactly the fraction who happened to share
a server with the sender. The team chased "dropped packets" for a week. The actual fix was one
component they'd never added because they'd never needed it on one box: a Redis pub/sub backplane. The
lesson - realtime bugs that only appear with more than one server are almost always fan-out.

💡 **Key point.** SSE has the *exact same* fan-out problem as WebSockets. It's one-directional, but the
server-to-many-clients broadcast still has to reach clients scattered across machines. SSE being simpler
on the connection side does not make it simpler to scale the *broadcast* - you still need a backplane.

## The real tradeoff: cost grows with connections

Persistent connections don't scale like stateless requests. A stateless API server can handle a request
and forget it; a realtime server holds *every* connection open simultaneously, each consuming memory and
a file descriptor whether or not it's doing anything. Ten thousand idle WebSockets still cost ten
thousand connections.

That reframes the whole "which pattern" question one last time:

- **Polling** spreads load across short requests your existing infrastructure already handles - no sticky
  sessions, no backplane, no held connections. For rare updates it's not the lazy choice, it's the
  *operationally cheapest* one.
- **SSE** needs a backplane for fan-out but inherits the browser's reconnect and rides plain HTTP/2 - the
  middle of the road.
- **WebSockets** need sticky sessions, a backplane, *and* your own reconnect/heartbeat logic. Powerful,
  and the most to operate.

⚠️ **Don't pay for realtime you don't use.** Every persistent connection is a standing cost and a thing
that can break at scale. If a feature is fine updating on the user's next action, give it nothing. If it
needs live data one way, SSE. Only the genuinely two-way, high-traffic features have earned WebSockets
and the backplane that comes with them.

## For builders

When you design a realtime feature, design the *scaled* version on paper from day one, even if you ship
single-server first. Ask: when there are N servers, how does a message reach a client on a different one?
If the answer isn't "via a backplane," you have a bug that's invisible until your second server. And keep
climbing *down* the ladder when you can - the cheapest realtime system is the one that's actually plain
polling because the data didn't move fast enough to justify more.

For the cross-system cousin of this problem - services handing events to each other rather than to
browsers - the durable, queue-based tools live in [Webhooks & Message
Queues](/guides/webhooks-and-message-queues); a pub/sub backplane is the realtime, in-memory relative of
those same ideas.

## Recap

1. **One server is easy** (one in-memory list of clients); **many servers is hard** because connections
   are pinned to whichever machine accepted them.
2. **Sticky sessions** keep a long-lived connection on one server so the load balancer doesn't tear it
   apart - necessary, but it does *not* solve fan-out.
3. **Fan-out** (one event → many clients across servers) is solved with a **pub/sub backplane**: servers
   publish to it and each pushes to its own local connections. SSE needs this too.
4. **Connection cost is the real tradeoff.** Polling rides existing infra; SSE adds a backplane;
   WebSockets add stickiness, a backplane, and DIY reconnect. Pick the simplest that fits, and design the
   N-server version up front.

That's the whole arc: HTTP can't push, three patterns fake or build the push, and at scale they all bow
to the same fan-out problem. Reach for the lightest tool that does the job, and you'll ship realtime that
stays calm at 3am.

```quiz
[
  {
    "q": "Why do realtime messages reach only some users once you scale from one server to several?",
    "choices": [
      "The database can't keep up",
      "Each persistent connection is pinned to one server, and a server can only push to clients it's holding",
      "The browser limits messages per second",
      "TLS drops messages across servers"
    ],
    "answer": 1,
    "explain": "Connections live on whichever server accepted them. A message arriving at Server A can't reach a client whose connection lives on Server B - that's the fan-out problem."
  },
  {
    "q": "What do sticky sessions fix, and what do they NOT fix?",
    "choices": [
      "They fix fan-out but not reconnection",
      "They keep a connection pinned to one server, but they do not deliver a message to clients on other servers",
      "They fix everything about scaling realtime",
      "They replace the need for a backplane"
    ],
    "answer": 1,
    "explain": "Stickiness keeps a long-lived connection alive on one server. Getting a message to clients on other servers still requires a pub/sub backplane."
  },
  {
    "q": "How does a pub/sub backplane solve cross-server fan-out?",
    "choices": [
      "It moves all clients onto one server",
      "Each server publishes incoming messages to the backplane, which fans them out to all servers so each pushes to its own local connections",
      "It disables sticky sessions",
      "It converts WebSockets to polling"
    ],
    "answer": 1,
    "explain": "No server knows every client. Servers publish to the backplane, it relays to all subscribers, and each server delivers to the connections it holds - that's the standard scaling pattern."
  }
]
```
