# Load Testing: k6 and JMeter

> Find the breaking point before your users do: k6's scriptable load tests and JMeter's mature GUI approach, plus how to read the results.


---

# Load Testing: k6 and JMeter

You shipped a feature. It works on your laptop, it works in staging, it works for the ten people in the demo. Then a launch, a sale, or a link goes wide, and the thing that ran in 80 milliseconds is now timing out, the database is on fire, and you are reading logs at midnight trying to figure out what the limit actually was. The whole point of load testing is to learn that limit on a Tuesday afternoon instead of during the incident.

This guide gives you two tools that do the same job from opposite ends. **k6** is load testing as code: you write a small JavaScript file, run it from the terminal, and drop it into CI. **JMeter** is the older, GUI-driven heavyweight that speaks more protocols than you will ever need. By the end you will know which to reach for, how to design a test that mimics real traffic instead of a meaningless flood, and how to read the numbers so you trust them.

## How to read this

Read it in order the first time. Phase 1 builds the mental model: what load, stress, and soak tests actually measure, and why the average response time lies to you. Phase 2 is the everyday core, the same realistic scenario written in both k6 and JMeter so you can compare them on equal footing. Phase 3 is the part people skip and regret: the gotchas that quietly invalidate a test and what production reality does to your tidy numbers.

If you already run load tests and only want the comparison, skim Phase 1 and live in Phase 2. If you have never written one, start at the top.

## The phases

1. [What load testing actually measures](01-what-load-testing-measures.md) - load vs stress vs soak, virtual users, and why p95 beats the average.
2. [Writing the same test in k6 and JMeter](02-k6-and-jmeter-in-practice.md) - a realistic ramping scenario, side by side, plus thresholds and CI.
3. [When the numbers lie and the system breaks](03-gotchas-and-production-reality.md) - the traps that fake a passing test and what breaks under real load.

For a wider view of where load testing fits, see [/guides/load-and-performance-testing](/guides/load-and-performance-testing) and [/guides/what-performance-means](/guides/what-performance-means).


---

# What load testing actually measures

Here is the reality you are starting from. Your service feels fast. You click around, it responds instantly, the unit tests are green. That feeling is built on a sample size of one: you, alone, with a warm cache and an idle database. Load testing exists to answer a different question, the one you cannot feel from your chair: *what happens when a thousand of you show up at the same second?*

The trap is thinking load testing is about making traffic. Anyone can flood a server. The skill is making traffic that resembles your real users, then reading the response so you learn the truth instead of a number that sounds impressive in a meeting.

## Three tests, three questions

People say "load test" to mean three different experiments. They overlap in tooling but answer different questions, and mixing them up is how you end up confidently wrong.

- **Load test** - "Does the system meet its target under expected traffic?" You drive the load you actually expect (say, your busy-hour traffic plus some headroom) and check that response times and error rates stay inside your goals. This is the everyday test, the one you run in CI.
- **Stress test** - "Where does it break, and how?" You push past expected traffic, deliberately, until something gives. The goal is not to pass; the goal is to find the cliff and watch how the system falls off it. Does it slow down gracefully, or does it fall over and refuse to recover?
- **Soak test** - "Does it survive over time?" You hold a moderate, realistic load for hours. This is how you catch the slow leaks: memory creeping up, connection pools that never release, a disk filling with logs. A system can pass a ten-minute load test and die at hour six.

```text
load     ████████░░░░  expected traffic, short, "do we meet target?"
stress   ████████████  push past the limit, "where's the cliff?"
soak     ██████░░░░░░  moderate, for hours, "does it leak?"
```

*What just happened:* the same tool can run all three. The difference is the shape and duration of the load you apply and the question you brought with you. Decide the question first, then design the test.

## What a virtual user actually is

Both k6 and JMeter work in terms of **virtual users** (VUs in k6, threads in JMeter). A virtual user is a loop: do a request, maybe wait a moment like a human reading the page, do the next request, repeat. Concurrency comes from running many of these loops at once.

This matters because *VU count is not request count.* Fifty virtual users does not mean fifty requests per second. If each user pauses one second between requests and each request takes 200 ms, each user completes roughly one iteration every 1.2 seconds, so fifty users produce around 40 requests per second. The thing your server actually feels is **throughput** - requests per second - and it falls out of VUs, think time, and how fast the server responds. Slow the server down and throughput drops even though your VU count never moved.

> Two numbers describe load from opposite sides. **Concurrency** (virtual users) is how many clients are in flight. **Throughput** (requests/sec) is how much work lands on the server. Report both - one without the other is half a story.

## Why the average lies to you

This is the single most important idea in the guide, so sit with it. Suppose you make 100 requests. Ninety-five come back in 50 ms. Five take 4 seconds because they hit a cold cache, a lock, or a slow query. The **average** is about 250 ms, which sounds fine. But five percent of your users waited four full seconds. The average smeared their pain across everyone and made it disappear.

**Percentiles** put the pain back. The **p95** (95th percentile) is the value that 95% of requests came in *under*. In our example, p95 is around 4 seconds - and that number screams what the average whispered.

```text
sorted response times (ms): 48 49 50 ... 51 | 3900 3950 4000 4100 4200
                            └── 95 fast ──┘   └──── 5 slow ────┘
average = ~250 ms   (looks fine)
p95     = ~4000 ms  (tells the truth)
p99     = ~4200 ms  (the worst real users see)
```

*What just happened:* the average got dragged toward the middle of two clusters and described nobody. p95 and p99 describe the slow tail - the users who notice, complain, and leave. Always read percentiles. Treat the average as decoration.

The numbers you should report from any run are: **p95 latency** (and p99 if your goals are strict), **throughput** (requests/sec), and **error rate** (percent of requests that failed). Those three, together, tell you whether the system met its target, how much work it was doing when it did, and whether it was quietly dropping requests to get there.

## For builders

Before you write a single line of test script, write down your **goals as numbers**: "p95 under 300 ms at 200 requests/sec with error rate under 1%." That sentence is the difference between a test that passes or fails and a test that only produces a graph. A graph invites debate; a threshold gives you a verdict. Phase 2 turns exactly this kind of sentence into a runnable check in both tools.

```quiz
[
  {
    "q": "You hold a moderate, realistic load steady for six hours to catch a slow memory leak. Which test is this?",
    "choices": ["Load test", "Stress test", "Soak test", "Smoke test"],
    "answer": 2,
    "explain": "A soak test runs a moderate load over a long duration specifically to surface slow problems like memory leaks and pool exhaustion."
  },
  {
    "q": "95 requests return in 50 ms and 5 return in 4000 ms. Why is the average a misleading metric here?",
    "choices": ["It is calculated wrong by most tools", "It blends the slow tail into the fast majority, hiding the 5% who waited 4 seconds", "Averages only work for error rates", "It overstates how slow the system is"],
    "answer": 1,
    "explain": "The average smears the painful tail across all requests. p95/p99 expose the slow requests real users actually feel."
  },
  {
    "q": "You run 50 virtual users, each pausing ~1s between requests that take ~200ms. Roughly what does the server feel?",
    "choices": ["50 requests per second", "Around 40 requests per second", "200 requests per second", "Exactly 50 concurrent requests at all times"],
    "answer": 1,
    "explain": "VU count is not throughput. Each user does ~1 iteration per 1.2s, so 50 users produce roughly 40 req/s. Report concurrency and throughput separately."
  }
]
```


---

# Writing the same test in k6 and JMeter

Now we build something real. The scenario: your API has a login endpoint and a "list my orders" endpoint, and you want to know whether it holds up when traffic ramps from nothing to 100 concurrent users over a couple of minutes, then comes back down. We will build this twice - once in k6, once in JMeter - so you can feel the difference in your hands instead of reading a feature matrix.

The shape of the load is the same in both: **ramp up, hold, ramp down**. Ramping matters. Slamming a server from zero to full load tests a cold, panicked system; ramping up the way real traffic arrives gives caches and pools a chance to warm and shows you a curve, not a single point.

## The scenario in k6

k6 is a single binary you install once. Your test is a JavaScript file with a `default` function (what each virtual user does on each loop) and an exported `options` object (how many users, for how long, and what counts as passing).

```javascript
import http from 'k6/http';
import { check, sleep } from 'k6';

export const options = {
  stages: [
    { duration: '30s', target: 100 },  // ramp 0 -> 100 VUs
    { duration: '1m',  target: 100 },  // hold at 100
    { duration: '30s', target: 0 },    // ramp back down
  ],
  thresholds: {
    http_req_duration: ['p(95)<300'],   // p95 latency under 300ms
    http_req_failed:   ['rate<0.01'],   // error rate under 1%
  },
};

export default function () {
  const login = http.post('https://api.example.com/login',
    JSON.stringify({ user: 'demo', pass: 'demo' }),
    { headers: { 'Content-Type': 'application/json' } });

  check(login, { 'login is 200': (r) => r.status === 200 });

  const token = login.json('token');
  http.get('https://api.example.com/orders',
    { headers: { Authorization: `Bearer ${token}` } });

  sleep(1);  // think time: pause like a real user
}
```

*What just happened:* `stages` drew the ramp-hold-ramp curve. `thresholds` encoded your goals from Phase 1 as machine-checkable rules. Each VU logs in, reads the token, lists orders, then pauses one second - a small but realistic user journey, not a blind request flood.

You run it from the terminal. This is the whole workflow:

```console
$ k6 run orders-test.js

     ✓ login is 200

     http_req_duration..............: avg=88ms  p(95)=214ms
     http_req_failed................: 0.42%  ✓ 23  ✗ 5478
     http_reqs......................: 5501   42.3/s
     vus............................: 100    max=100

   ✓ THRESHOLDS PASSED
```

*What just happened:* the run printed exactly the three numbers that matter - p95 (214 ms), error rate (0.42%), and throughput (42.3 req/s) - and checked them against your thresholds. `THRESHOLDS PASSED` means k6 will exit with code 0; if a threshold fails, it exits non-zero, which is the hook that makes CI work.

## The same scenario in JMeter

JMeter comes at this from the GUI. You open the desktop app and build the test as a **test plan** tree: a **Thread Group** holds the load shape, **Samplers** are the requests, and **Listeners** show the results. You are clicking and filling in fields, not writing code.

```text
Test Plan
└── Thread Group         (100 threads, 30s ramp-up, loop for duration)
    ├── HTTP Request: POST /login
    │   └── JSON Extractor: pull "token" into a variable
    ├── HTTP Header Manager: Authorization: Bearer ${token}
    ├── HTTP Request: GET /orders
    ├── Constant Timer: 1000 ms        (think time)
    └── Listener: Summary Report / Aggregate Report
```

*What just happened:* every concept from the k6 script has a one-to-one twin here. Threads are VUs, ramp-up time is the first stage, the JSON Extractor replaces `login.json('token')`, the Header Manager carries the token, and the Constant Timer is `sleep(1)`. Same scenario, assembled with a mouse.

Reading results lives in the **Aggregate Report** listener, which gives you a table with a column literally labeled **95% Line** alongside **Throughput** and **Error %** - the same three numbers, named slightly differently.

> JMeter's built-in ramp shape is one number: ramp-up time, then hold. k6's `stages` describe an arbitrary curve (up, hold, spike, down) directly. For multi-step ramps in JMeter you reach for the Ultimate Thread Group plugin. If your load shape is complex, that difference alone may decide the tool.

## Running it without the GUI

Here is the rule that surprises newcomers: **never run a real JMeter load test from the GUI.** The graphical interface is for *building* and *debugging* the plan. The GUI itself consumes memory and CPU drawing live graphs, which steals resources from load generation and skews your results. For the real run, you go headless:

```console
$ jmeter -n -t orders-test.jmx -l results.jtl
```

*What just happened:* `-n` is non-GUI mode, `-t` points at the test plan you built in the GUI, and `-l` writes raw results to a `.jtl` file you analyze afterward. This is the mode you use for any run whose numbers you intend to trust, and the only mode that belongs anywhere near CI.

k6 has no such split - it is headless by nature. That is most of why it slots into pipelines so cleanly:

```yaml
# a CI step, conceptually
- run: k6 run orders-test.js
  # threshold failure -> non-zero exit -> red build
```

*What just happened:* because k6 returns a failing exit code when a threshold is breached, your performance goal becomes a build gate with no extra glue. JMeter can do this too, but you assert on the `.jtl` afterward rather than getting it from the run itself.

## Which one, and when

Neither tool is "better." They fit different teams.

- **Reach for k6** when your team lives in code and Git, when you want the test reviewed in a pull request, and when CI gating is the goal. Tests are diffable text; the learning curve is "do you know a little JavaScript."
- **Reach for JMeter** when you need protocols beyond HTTP (JDBC, JMS, FTP, LDAP and more), when the people writing tests prefer a GUI over code, or when you are inheriting an organization that already has a wall of `.jmx` files. Its maturity and protocol breadth are real and hard to match.

```quiz
[
  {
    "q": "In a k6 script, what does the `thresholds` block do?",
    "choices": ["Sets how many virtual users to run", "Defines pass/fail rules so k6 exits non-zero when goals are missed", "Controls the ramp-up duration", "Adds think time between requests"],
    "answer": 1,
    "explain": "Thresholds encode your performance goals (e.g. p(95)<300). A breach makes k6 exit non-zero, which is what gates a CI build."
  },
  {
    "q": "Why should you run a real JMeter load test with `-n` (non-GUI) instead of from the GUI?",
    "choices": ["The GUI cannot generate enough load", "The GUI consumes CPU/memory drawing live graphs, skewing results", "Non-GUI mode is the only one that supports HTTP", "Thresholds only work in non-GUI mode"],
    "answer": 1,
    "explain": "The GUI is for building and debugging. Its live rendering steals resources from load generation, so real runs go headless with -n -t -l."
  },
  {
    "q": "Your team wants load tests reviewed in pull requests and gating CI, written by developers comfortable with JavaScript. Which tool fits best?",
    "choices": ["JMeter, for its GUI", "k6, because tests are diffable code and it exits non-zero on threshold failure", "Neither can run in CI", "JMeter, because it supports more protocols"],
    "answer": 1,
    "explain": "k6's scripts are plain text (great for code review) and its threshold-driven exit code makes CI gating trivial. JMeter shines elsewhere: protocol breadth and GUI authoring."
  }
]
```


---

# When the numbers lie and the system breaks

You ran the test. It passed. The graph is green, p95 is comfortable, and you are ready to call it done. This phase is the friend who pulls you aside and asks whether the test measured what you think it did - because a load test that passes for the wrong reason is more dangerous than no test at all. It hands you confidence you did not earn, and you spend it in production.

Let us walk through the ways a green run lies, and then what production does that no test fully captures.

## The load generator is the bottleneck

The first thing to suspect when your numbers look strange is the machine *running the test*, not the machine under test. Generating thousands of concurrent requests is itself hard work. If your laptop (or the tiny CI runner) maxes out its CPU or runs out of open file descriptors, it cannot send requests fast enough, and you measure *its* limit while believing you measured the server's.

```console
$ k6 run --vus 2000 orders-test.js
WARN[0012] Request Failed   error="dial: i/o timeout"
WARN[0014] Request Failed   error="socket: too many open files"
```

*What just happened:* those errors are not your server failing - they are the *generator* failing. The server may be perfectly healthy and starved of traffic. Always watch the generator's own CPU and network during a run. When one box cannot push enough load, you move to distributed generation (k6 supports running across multiple machines; JMeter has a controller/worker setup) rather than trusting numbers from a saturated client.

## Your test data is too clean

This one is subtle and bites almost everyone. If every virtual user logs in as the same account and requests the same record, your database serves that row from cache after the first hit. Your test then measures cache performance, which is dazzlingly fast and completely unlike production, where users hit *different* rows and the cache misses constantly.

```javascript
// trap: every VU reads the same id -> warm cache, fake-fast results
http.get('https://api.example.com/orders/42');

// better: spread reads across many ids the way real users do
const id = Math.floor(Math.random() * 100000) + 1;
http.get(`https://api.example.com/orders/${id}`);
```

*What just happened:* the first line lets one cached row carry the whole test and reports a latency you will never see in production. The second spreads reads across the dataset, forcing real cache misses and real database work. Realistic, varied test data is often the difference between a useful test and a comforting fiction. The same applies to JMeter via a CSV Data Set Config that feeds different values per thread.

## You forgot to check whether requests succeeded

A fast response is worthless if it is an error. Under load, a server protecting itself often returns `429 Too Many Requests` or `503 Service Unavailable` *instantly* - those failures are blazing fast, and they will drag your average latency *down* while your error rate quietly climbs. If you only watch latency, the test looks like it got faster under load. It got faster because it stopped working.

```console
http_req_duration..........: avg=31ms  p(95)=44ms     ← looks great!
http_req_failed............: 73.0%   ✗ 14209         ← it's on fire
```

*What just happened:* p95 dropped to 44 ms not because the system sped up but because three-quarters of requests were fast rejections. This is why error rate sits next to latency in every report from Phase 1 - latency without error rate is a number that can lie to your face. In k6, an `http_req_failed` threshold catches it; in JMeter, watch the **Error %** column.

## Find the bottleneck, do not only name a number

A load test tells you *that* the system slowed at 200 req/s. It does not tell you *why*. The "why" lives in metrics from the system under test, observed during the run: CPU, memory, database connection pool usage, disk and network I/O. The number from the load tool and the resource graphs from the server are two halves of one diagnosis.

```text
req/s climbs ──► p95 climbs ──► look at the server, NOT the tool:
   CPU pinned at 100%?        → compute-bound, scale or optimize code
   DB connections maxed?      → pool too small / queries too slow
   memory climbing, no plateau?→ probable leak (run a soak test)
   CPU idle but slow anyway?  → waiting on a downstream / lock contention
```

*What just happened:* the load tool found the symptom; the server's own metrics name the cause. A passing or failing number with no resource graphs behind it is a verdict with no evidence. Watch both, always.

## Production reality the test never showed you

Even a careful test is a model, and the territory has features the map omits. Keep these in mind before you trust a green run too far:

- **The network is real.** Testing against `localhost` removes latency, TLS handshakes, and bandwidth limits that real users carry on every request. Run the generator from somewhere with a realistic network path to the server.
- **Caches were warm (or cold) differently.** A short test may ride a warm cache that a real cold start would not have. A real deploy restarts everything; consider whether your test reflects that.
- **Real traffic is spiky and mixed.** Users do not arrive in a tidy ramp. They spike, they hit a mix of cheap and expensive endpoints, and they retry on failure (which *amplifies* load right when you can least afford it). A single-endpoint test misses this entirely.
- **Downstream dependencies have their own limits.** Your service may scale fine, but the third-party payment API, the shared database, or the rate-limited search backend may not. The bottleneck is often something you do not own.

> The plain summary of any load test: it tells you a *lower bound* on your problems under one specific, simplified scenario. It cannot prove the system is fine. It can only prove it is not *clearly* broken in the way you tested. Treat a green run as one piece of evidence, not a guarantee.

## In the wild

Mature teams do not run a load test once before launch and forget it. They keep a small, fast k6 test in CI as a regression gate (catching the day a code change quietly doubles a query count), and they run a larger, more realistic stress and soak test against a production-like environment before big events. The cheap test guards every commit; the expensive test guards the launch. For where this fits in the broader practice, see [/guides/load-and-performance-testing](/guides/load-and-performance-testing), and for what "fast enough" even means, [/guides/what-performance-means](/guides/what-performance-means).

```quiz
[
  {
    "q": "Under heavy load your p95 latency suddenly drops to 40ms but error rate jumps to 70%. What is the most likely explanation?",
    "choices": ["The server got faster under load", "The server is returning instant 429/503 rejections, which are fast and pull latency down", "The load generator sped up", "Caching finally kicked in"],
    "answer": 1,
    "explain": "Fast failures (429/503) lower average latency while the system stops working. Always read error rate alongside latency."
  },
  {
    "q": "Every virtual user requests `/orders/42`. Why does this produce misleading results?",
    "choices": ["It overloads the network", "The same row stays cached, so you measure cache speed, not real database load", "It causes too many open files", "k6 cannot handle a static URL"],
    "answer": 1,
    "explain": "Identical requests ride a warm cache and hide real database work. Spread reads across varied data (random ids, CSV data sets) to force realistic cache misses."
  },
  {
    "q": "Your load tool reports p95 climbing past your goal at 200 req/s. What is the single best next step to find the cause?",
    "choices": ["Add more virtual users", "Lower the threshold so the test passes", "Look at the server's own resource metrics (CPU, memory, DB pool) during the run", "Rerun against localhost"],
    "answer": 2,
    "explain": "The load tool shows the symptom; the system's resource graphs name the cause. Diagnosis needs both halves together."
  }
]
```
