# NATS and Amazon SQS

> Two lighter messaging options: NATS for fast, simple pub/sub and request-reply, and SQS for a fully managed, zero-ops cloud queue.


---

# NATS and Amazon SQS

You read about Kafka, sketched a topology with partitions and consumer groups and a ZooKeeper-shaped hole in your heart, and then looked at what you actually need: a service that fires an event and another that reacts. That's it. The gap between what Kafka costs to run and what your problem needs is enormous, and you can feel it.

This guide is about the two brokers you reach for when Kafka is overkill. NATS is a tiny, blisteringly fast broker you can run yourself in one binary. SQS is a queue AWS runs for you so you never think about a broker again. Both are boring in the best way, and that's the relief.

## How to read this

Read the phases in order the first time. Phase 1 builds the mental model so the words "subject," "queue group," and "visibility timeout" stop being noise. Phase 2 is the everyday work: publishing, consuming, request-reply, and the SQS receive loop you'll actually write. Phase 3 is where things go wrong in production and how to not get paged for it. If you've already used one of the two, you can skim to the half you don't know.

## The phases

1. [Phase 1: Not Everything Needs Kafka](01-not-everything-needs-kafka.md) - the mental model: what NATS and SQS are, and how to pick a broker.
2. [Phase 2: Sending and Receiving for Real](02-sending-and-receiving-for-real.md) - pub/sub, request-reply, JetStream, and the SQS poll loop.
3. [Phase 3: Production Reality](03-production-reality.md) - visibility timeouts, dead-letter queues, duplicates, and the failure modes.


---

# Phase 1: Not Everything Needs Kafka

Here's the trap. You have two services that need to talk without being wired directly together. You search "message broker," and the internet hands you Kafka: partitions, replication factors, consumer group rebalances, a cluster to babysit. You start to feel that your modest need now requires an operations team.

It doesn't. Most messaging needs are small. A service emits "an order was placed," and one or two other services want to know. Or a request comes in and you want a reply, but over a network instead of a function call. For those, you want the smallest broker that fits. NATS and SQS are two ends of "small": one you run yourself and one AWS runs for you.

## What a broker actually does

Strip away the marketing and a broker does one thing: it sits between a sender and a receiver so they don't have to know about each other. The sender drops a message in. The receiver picks it up later. Neither has to be online at the same moment, neither has to know the other's address, and you can add more receivers without touching the sender.

That decoupling is the whole point. If you've read [the message queues guide](/guides/webhooks-and-message-queues), this is the same mental model - we're now looking at two specific tools and when to pick each.

## NATS: a broker that fits in one binary

NATS is a single small server. You download one file, run it, and you have a broker. No JVM, no ZooKeeper, no cluster required to get started. It speaks a simple text protocol and connects clients in milliseconds.

The core idea in NATS is the **subject** - a dotted string that names what a message is about, like `orders.placed` or `sensors.kitchen.temp`. Publishers send to a subject. Subscribers listen on a subject, and they can use wildcards: `orders.*` matches one token, `sensors.>` matches everything below `sensors`.

```text
publisher  ──▶  subject: orders.placed  ──▶  subscriber A (orders.*)
                                          └─▶  subscriber B (orders.placed)
```

*What just happened:* the publisher named the message by subject and walked away. NATS fanned it out to every subscriber whose pattern matched. The publisher has no idea how many subscribers exist - zero, one, or a thousand - and doesn't care.

By default, core NATS is **fire-and-forget**. If no one is listening when you publish, the message is gone. That sounds scary until you realize how often it's exactly right: live metrics, presence pings, cache-invalidation signals. You don't want yesterday's "user is typing" event. When you *do* need messages to survive, NATS has **JetStream**, a persistence layer that stores messages in a stream and lets consumers replay them. We'll use it in Phase 2.

> NATS also gives you **request-reply** for free: a client publishes a request and waits for one response on a private reply subject. It's RPC over a broker, which means load balancing and failover come along without extra plumbing.

## SQS: a queue you never operate

Amazon SQS is the opposite philosophy. You don't run anything. There's no server, no version to upgrade, no disk to fill. You create a queue through the AWS console or API, and from then on you only send and receive messages. AWS handles durability, scaling, and availability. (For the broader picture of managed services, see [cloud platforms explained](/guides/cloud-platforms-explained).)

SQS is a **queue**, not a pub/sub bus. A message goes in once and is delivered to one consumer, who then deletes it. There's no built-in fan-out to multiple independent readers - if you need that on AWS, you put SNS in front of several SQS queues. Keep SQS in the "work queue" mental slot: a list of tasks that workers pull from.

SQS comes in two flavors, and the difference matters:

```text
Standard queue   →  near-unlimited throughput
                 →  at-least-once delivery  (you may see a message twice)
                 →  best-effort ordering    (not guaranteed)

FIFO queue       →  strict ordering within a message group
                 →  exactly-once processing (dedup within a window)
                 →  lower throughput than Standard
```

*What just happened:* you traded one property for another. Standard gives you huge throughput but can deliver a message more than once and out of order. FIFO guarantees order and de-duplicates, at the cost of throughput. Most workloads want Standard plus idempotent consumers - we'll get to why in Phase 3.

## How to choose

The straightforward decision tree is short:

- **You're already on AWS and want zero ops** → SQS. It's the path of least resistance for background jobs and decoupling services.
- **You want very low latency, request-reply, or true pub/sub fan-out, and you can run a binary** → NATS.
- **You need durable, replayable event streams with many independent consumers and high throughput** → this is where Kafka earns its weight. Don't force NATS or SQS into that shape.

```mermaid
flowchart TD
  A[Need messaging] --> B{On AWS, want zero ops?}
  B -->|Yes| C[SQS]
  B -->|No| D{Need fan-out or request-reply?}
  D -->|Yes| E[NATS]
  D -->|No, simple queue| E
  C --> F{Need event replay + many readers?}
  E --> F
  F -->|Yes, at scale| G[Reconsider Kafka]
```

The thing to internalize: brokers are not a ladder where bigger is better. They're tools with different shapes. Picking the smallest one that fits is a sign you understand the problem, not that you cut a corner.

**In the wild:** plenty of production systems run NATS for internal service-to-service chatter and SQS for cloud background jobs *in the same company*. They're not competitors so much as different drawers in the same toolbox.

```quiz
[
  {
    "q": "By default, what happens to a core NATS message published to a subject with no subscribers?",
    "choices": ["It's stored until a subscriber connects", "It's discarded - core NATS is fire-and-forget", "It's sent to a dead-letter subject", "The publish call blocks until someone subscribes"],
    "answer": 1,
    "explain": "Core NATS is fire-and-forget; no subscriber means the message is gone. JetStream is what you add when you need persistence."
  },
  {
    "q": "What's the key delivery difference between an SQS Standard queue and a FIFO queue?",
    "choices": ["Standard is cheaper; FIFO is free", "FIFO has higher throughput than Standard", "Standard is at-least-once and best-effort ordered; FIFO is ordered and de-duplicated", "Standard supports pub/sub; FIFO does not"],
    "answer": 2,
    "explain": "Standard trades exactness for throughput (at-least-once, best-effort order). FIFO guarantees order and dedup at lower throughput."
  },
  {
    "q": "Which need is the best fit for plain SQS rather than NATS?",
    "choices": ["Sub-millisecond request-reply between services", "Fanning one event out to five independent subscribers", "A zero-ops background work queue on AWS", "Replaying a week of events to a new consumer"],
    "answer": 2,
    "explain": "SQS shines as a managed work queue. Fan-out and request-reply are NATS strengths; replay at scale points toward Kafka."
  }
]
```


---

# Phase 2: Sending and Receiving for Real

Mental models are nice, but you came here to send a message. This phase is the everyday work: standing up NATS, publishing and subscribing, doing request-reply, adding persistence with JetStream, and writing the SQS loop that pulls work and deletes it. The commands here are the ones you'll actually type.

## Run NATS in one command

You need the server (`nats-server`) and, to poke at it by hand, the CLI (`nats`). Once you have the binary on your path:

```bash
# Start the server. Defaults to listening on port 4222.
nats-server

# In another terminal, subscribe to a subject and wait.
nats sub "orders.>"

# In a third, publish a message to a matching subject.
nats pub orders.placed '{"id": 42, "total": 19.99}'
```

*What just happened:* the subscriber, listening on the wildcard `orders.>`, received the message published to `orders.placed`. The server matched the subject pattern and delivered it. No topic to pre-create, no schema to register - the subject sprang into existence the moment you used it.

That last point surprises people coming from Kafka. In NATS, subjects aren't declared up front. They're addresses, and any publish or subscribe can name a new one.

## Pub/sub with a client library

Here's the same flow in code. The shape is identical across the official client libraries: connect, then subscribe with a handler or publish a payload.

```js
import { connect } from "nats";

const nc = await connect({ servers: "localhost:4222" });

// Subscriber: handle every message on orders.placed
const sub = nc.subscribe("orders.placed");
(async () => {
  for await (const msg of sub) {
    console.log("got order:", msg.string());
  }
})();

// Publisher: fire an event and move on
nc.publish("orders.placed", JSON.stringify({ id: 42, total: 19.99 }));
```

*What just happened:* `publish` returned immediately - it didn't wait for the subscriber to process anything. That's the fire-and-forget default. The subscriber's `for await` loop runs independently as messages arrive.

## Queue groups: load balancing built in

What if you have three workers and you want each message handled by exactly *one* of them, not all three? That's a **queue group**. Subscribers that share a queue group name split the messages between them; the broker picks one member per message.

```js
// Start this on three separate workers - same queue name.
const sub = nc.subscribe("jobs.resize", { queue: "resizers" });
for await (const msg of sub) {
  await resizeImage(msg.data);
}
```

*What just happened:* by passing `{ queue: "resizers" }`, the three subscribers became a single competing-consumer pool. A message to `jobs.resize` goes to one resizer, not all three. This is how you scale a worker horizontally in NATS - no partition math, no rebalancing config.

## Request-reply: RPC over the broker

Sometimes you don't want fire-and-forget. You want an answer. NATS request-reply gives you that: the requester sends and waits, a responder replies on a private inbox subject the requester created.

```js
// Responder: listen, compute, reply.
const sub = nc.subscribe("price.lookup");
for await (const msg of sub) {
  msg.respond(JSON.stringify({ sku: "ABC", price: 9.99 }));
}

// Requester: send and await one reply, with a timeout.
const reply = await nc.request(
  "price.lookup",
  JSON.stringify({ sku: "ABC" }),
  { timeout: 1000 }
);
console.log(reply.string());
```

*What just happened:* `nc.request` published the message and blocked until a reply landed on its temporary inbox subject, or until the 1-second timeout fired. If you ran several responders in a queue group, NATS would load-balance the requests across them - you get failover and scaling for free.

> Always set a timeout on `request`. Without one, a missing responder means your caller waits forever. A timeout turns a silent hang into a clean, catchable error.

## JetStream: when messages must survive

Core NATS forgets. When you need messages to persist - to survive a restart, or to be replayed by a consumer that wasn't online yet - you turn on **JetStream**. You define a **stream** that captures subjects, and **consumers** that read from it with delivery tracking and acknowledgments.

```bash
# Enable JetStream on the server.
nats-server -js

# Create a stream that captures everything under orders.
nats stream add ORDERS --subjects "orders.>" --storage file

# Publish - now it's stored, not just broadcast.
nats pub orders.placed '{"id": 99}'

# Create a durable consumer and pull messages with acks.
nats consumer add ORDERS workers --pull --ack explicit
nats consumer next ORDERS workers --ack
```

*What just happened:* the message to `orders.placed` was written to disk inside the `ORDERS` stream. The `workers` consumer pulled it and acknowledged it. If that consumer had been offline at publish time, the message would still be waiting when it came back - the opposite of core NATS. With `--ack explicit`, an unacknowledged message is redelivered, so a crashed worker doesn't drop the job.

The trade is real: JetStream costs disk and adds the acknowledgment dance. Reach for it when loss is unacceptable; stay on core NATS when it isn't.

## SQS: send, receive, delete

SQS has no server for you to start - you create the queue once, then it's all API calls. The receive loop has a shape you must internalize, because it's where every SQS bug lives. Receiving a message does **not** remove it. You have to delete it yourself after you've processed it.

```bash
# Create a standard queue (one-time setup).
aws sqs create-queue --queue-name jobs

# Send a message.
aws sqs send-message \
  --queue-url "$Q" \
  --message-body '{"task": "resize", "id": 7}'

# Receive up to 10 messages, waiting up to 20s for them (long polling).
aws sqs receive-message \
  --queue-url "$Q" \
  --max-number-of-messages 10 \
  --wait-time-seconds 20
```

*What just happened:* `receive-message` returned messages *and a `ReceiptHandle` for each one*. The message is now invisible to other consumers for the visibility timeout - but it still exists in the queue. If you walk away now, it reappears later and gets processed again. The `--wait-time-seconds 20` is **long polling**: instead of returning instantly empty, SQS waits up to 20 seconds for a message to show up, which cuts your empty-receive count and your bill.

The third step is the one beginners forget - deleting the message once you're done:

```bash
# After successfully processing, delete it using the receipt handle.
aws sqs delete-message \
  --queue-url "$Q" \
  --receipt-handle "$RECEIPT_HANDLE"
```

*What just happened:* the message is now gone for good. The receipt handle is a one-time token tied to *this* receive, not a permanent message ID - receive the same message again and you get a fresh handle. The contract is: **receive → process → delete.** Skip the delete and you'll process the same work twice; we cover why that's survivable in Phase 3.

**For builders:** in real code you use the AWS SDK, not the CLI, and you loop forever: long-poll for a batch, process each message, delete it, repeat. The CLI commands above map one-to-one onto SDK calls (`SendMessage`, `ReceiveMessage`, `DeleteMessage`), so what you learned at the terminal is exactly what you'll write.

```quiz
[
  {
    "q": "In SQS, what does calling ReceiveMessage do to the message?",
    "choices": ["Permanently removes it from the queue", "Makes it invisible for the visibility timeout but leaves it in the queue until you delete it", "Marks it processed automatically", "Copies it to a dead-letter queue"],
    "answer": 1,
    "explain": "Receiving hides the message temporarily and returns a receipt handle. It stays in the queue until you explicitly delete it - receive, process, delete."
  },
  {
    "q": "You have three NATS subscribers and want each message handled by exactly one of them. What do you use?",
    "choices": ["A subject wildcard", "Three separate subjects", "A shared queue group name on the subscriptions", "JetStream with explicit acks"],
    "answer": 2,
    "explain": "Subscribers sharing a queue group name form a competing-consumer pool - the broker delivers each message to just one member."
  },
  {
    "q": "When does JetStream earn its extra disk and acknowledgment overhead over core NATS?",
    "choices": ["When you want the lowest possible latency", "When messages must survive restarts or be replayed by a later consumer", "When you have no subscribers", "When you need request-reply"],
    "answer": 1,
    "explain": "JetStream adds persistence and redelivery so messages survive and can be replayed. Core NATS is fire-and-forget and forgets on restart."
  }
]
```


---

# Phase 3: Production Reality

Everything in Phase 2 works on your laptop. Production is where the broker meets a slow database, a worker that crashes mid-task, a message that can never succeed, and a duplicate you didn't expect. None of these are exotic - they're the default behavior you have to design around. This phase is the set of gotchas that, once you've felt them, you never forget.

## The visibility timeout is a deadline, not a suggestion

This is the single most common SQS bug. When you receive a message, SQS hides it for the **visibility timeout** (default 30 seconds). The unspoken contract: process and delete it before that window closes. If your processing takes longer than the timeout, SQS assumes you died, makes the message visible again, and hands it to another worker - while you're still working on it.

```text
t=0    worker A receives msg  (hidden for 30s)
t=30   timeout expires, msg visible again
t=31   worker B receives the SAME msg  ← now two workers process it
t=45   worker A finishes, deletes msg
t=60   worker B finishes, tries to delete - receipt handle stale
```

*What just happened:* a job that took 45 seconds under a 30-second timeout got processed twice. The fix is to set the visibility timeout comfortably above your worst-case processing time, or to call `change-message-visibility` to extend it (a "heartbeat") while a long job runs.

```bash
# Buy more time on an in-flight message before the timeout expires.
aws sqs change-message-visibility \
  --queue-url "$Q" \
  --receipt-handle "$RECEIPT_HANDLE" \
  --visibility-timeout 120
```

*What just happened:* you pushed this message's deadline out to 120 seconds, so a slow job won't be redelivered out from under you. Don't just set a giant static timeout, though - if a worker truly crashes, a huge timeout means the message sits stuck for that long before anyone retries it.

## Design for at-least-once: make consumers idempotent

SQS Standard is **at-least-once**. Even with a sane visibility timeout, a network hiccup between your `delete` call and SQS can leave the message in the queue, and it'll be redelivered. So you cannot assume a message arrives exactly once. The same is true of JetStream redelivery after a missed ack.

The cure isn't to fight duplicates - it's to make processing the same message twice harmless. That property is **idempotency**.

```text
NOT idempotent:  balance = balance + amount        (runs twice → double charge)
idempotent:      if not already_applied(msg_id):
                     balance = balance + amount
                     mark_applied(msg_id)
```

*What just happened:* by recording which message IDs you've already applied and skipping repeats, a second delivery becomes a no-op. Now "at-least-once" stops being scary. This is why most teams pick Standard over FIFO: idempotent consumers are good engineering anyway, and they unlock Standard's throughput without the FIFO penalty.

> FIFO's "exactly-once" only de-duplicates within a roughly five-minute window, and only for messages you mark as duplicates. It is not a license to skip idempotency. Build idempotent consumers regardless of queue type.

## Dead-letter queues: where poison messages go to rest

Some messages will never succeed. A malformed payload, a referenced record that was deleted, a bug in your handler - process it, fail, it reappears, fail again, forever. That's a **poison message**, and it can wedge a queue.

A **dead-letter queue (DLQ)** is the release valve. You configure a redrive policy: after a message has been received `maxReceiveCount` times without being deleted, SQS moves it to a separate queue instead of redelivering it.

```text
main queue ──(received 5×, never deleted)──▶ DLQ
                                              │
                                    you inspect, fix, or replay
```

*What just happened:* the failing message stops poisoning your main queue after five attempts and lands in the DLQ, where it can't block healthy traffic. The DLQ becomes your "needs a human" inbox. Always set up a DLQ for any real queue, and *alarm on its depth* - a DLQ filling up is one of the clearest "something is broken" signals you'll get.

JetStream has the same concept in different clothes: you cap redelivery with `max_deliver` on a consumer and route terminal failures to an advisory subject or another stream.

## NATS gotchas: slow consumers and the persistence cliff

NATS is fast, which creates its own failure mode. If a subscriber can't keep up with the rate of incoming messages, NATS won't block the publisher to wait for it. Instead it protects the system by dropping that subscriber's backlog - a **slow consumer**. You'll see a warning, and that subscriber will have holes in its message history.

The mental fix: in core NATS, the producer's speed is not your consumer's problem to throttle. If you can't afford to drop messages, you are on JetStream, full stop. Pull-based JetStream consumers fetch at their own pace and won't be dropped for being slow - that's the entire point of the persistence layer.

The other cliff is forgetting which mode you're in. Core NATS and JetStream look similar in code but have opposite durability guarantees. A message published to a subject that *isn't* captured by any stream is fire-and-forget, even on a server with JetStream enabled. Confirm your stream's `--subjects` actually covers the subjects you publish to, or you'll think you have persistence you don't.

## A short operational checklist

```text
SQS:
  [ ] visibility timeout > worst-case processing time (or heartbeat it)
  [ ] consumers are idempotent (assume at-least-once)
  [ ] DLQ configured with a sane maxReceiveCount
  [ ] CloudWatch alarm on DLQ depth and queue age
  [ ] long polling on (wait-time-seconds) to cut cost

NATS:
  [ ] core NATS only where message loss is acceptable
  [ ] JetStream stream subjects actually cover your publish subjects
  [ ] consumers ack explicitly; max_deliver caps redelivery
  [ ] request() calls always set a timeout
```

*What just happened:* you turned four sharp failure modes into routine config. None of this is heroic - it's the boring discipline that separates a queue that pages you at 3am from one you forget exists.

**In the wild:** the teams who sleep well aren't the ones with the fanciest broker. They're the ones who assumed duplicates, set a DLQ, alarmed on it, and made their consumers idempotent before they ever needed to. The broker is small; the discipline is the product.

```quiz
[
  {
    "q": "Your SQS worker takes 45 seconds but the visibility timeout is 30 seconds. What goes wrong?",
    "choices": ["The message is deleted before processing finishes", "The message becomes visible again at 30s and a second worker processes it too", "SQS rejects the message", "The queue switches to FIFO mode"],
    "answer": 1,
    "explain": "Once the timeout expires, SQS assumes the worker died and redelivers the message - so it gets processed twice. Raise the timeout or heartbeat with change-message-visibility."
  },
  {
    "q": "What's the right response to SQS Standard's at-least-once delivery?",
    "choices": ["Switch everything to FIFO to get exactly-once", "Make consumers idempotent so processing twice is harmless", "Disable the visibility timeout", "Delete messages before processing them"],
    "answer": 1,
    "explain": "Idempotent consumers turn duplicate delivery into a no-op, letting you keep Standard's throughput without fearing repeats."
  },
  {
    "q": "What is a dead-letter queue for?",
    "choices": ["Storing every message for replay", "Holding messages that repeatedly fail so they stop poisoning the main queue", "Speeding up delivery", "Encrypting messages at rest"],
    "answer": 1,
    "explain": "After maxReceiveCount failed attempts, SQS moves the message to the DLQ so a poison message can't wedge the main queue. Alarm on its depth."
  }
]
```
