# Evaluating LLM Output

> Evals, not vibes: how to measure whether an LLM feature actually works, catch prompt regressions, and ship changes with confidence.


---

# Evaluating LLM Output

You changed the prompt, ran it on three examples, and it looked better - so you shipped. A week later support tells you the summaries got worse for a whole category of inputs you never tried. The plain, uncomfortable truth is that you didn't *know* the change was an improvement. You felt it. And feelings don't catch the case you forgot to look at.

This guide is about replacing the feeling with a measurement. Not a research lab, not a leaderboard - a small, boring set of real inputs with a way to score the output, so that "is this better?" becomes a number you can compare instead of a vibe you can argue about. Once you have that, prompt changes and model upgrades stop being scary leaps in the dark.

## How to read this

- **Want the core idea fast?** Read [Phase 1: Why Vibes Don't Scale](01-why-vibes-dont-scale.md) - what an eval actually is and why three eyeballed examples lie to you.
- **Already convinced, need the how?** Go to [Phase 2: How to Actually Grade Output](02-how-to-grade-output.md) - exact checks, reference-based scoring, and LLM-as-judge with its caveats.
- **Shipping changes and afraid of regressions?** [Phase 3: Evals as a Habit](03-evals-as-a-habit.md) covers regression testing, model upgrades, tracking quality over time, and where offline evals end and production monitoring begins.

## The phases

1. **[Why Vibes Don't Scale](01-why-vibes-dont-scale.md)** - the mental model: you can't improve what you don't measure, eyeballing doesn't scale, and an eval is a real input set plus an expected behavior you can score.
2. **[How to Actually Grade Output](02-how-to-grade-output.md)** - the three grading methods: exact and rule-based checks, reference-based metrics, and LLM-as-judge - when each fits and where each lies to you.
3. **[Evals as a Habit](03-evals-as-a-habit.md)** - regression testing prompts and model upgrades, tracking quality over time, and the line between offline evals, production monitoring, and real user feedback.

> This guide assumes you've already wired a model into your app. If "call the model" is still fuzzy, read [Using an LLM API](/guides/using-an-llm-api) first; the craft of writing the instructions you'll be evaluating lives in [Prompt Engineering, Plainly](/guides/prompt-engineering-plainly).


---

# Why Vibes Don't Scale

Here's the loop almost everyone runs at first. Write a prompt. Paste in an example. Read the output. "Yeah, that's good." Ship it. Then tweak the prompt to fix one annoying thing, paste in the *same* example, read it again, "better," ship. That loop feels like progress because each step ends with you nodding at a screen.

The problem is what the loop can't see. You tested the input you happened to pick. You judged it with a brain that already knew what you wanted and read it in charitably. And you have no record of yesterday's output to compare against, so you can't actually tell whether today's change helped, hurt, or moved the problem somewhere you didn't look.

## You can't improve what you don't measure

This isn't an AI-specific law; it's an engineering one. If "better" is a feeling in your head, then two people can look at the same output and disagree, and *you* can disagree with yourself next week. There's nothing to point at.

A model feature has a sharper version of this problem than normal code, for three reasons:

- **The output space is huge.** A function returns `true` or `false`; a model returns one of effectively infinite strings. "Correct" is often a range, not a single value.
- **It's non-deterministic.** The same input can give different output on different runs, so a single good result doesn't prove the next one will be good.
- **Changes have spooky reach.** Editing one line of a prompt can improve the case you were looking at and silently break a category you weren't. There's no compiler to catch it.

Put together: the thing most likely to mislead you is exactly the thing you're using to judge it - your eyes, on a handful of cases, once.

## What "eyeballing doesn't scale" really means

It's not that looking at output is wrong. Looking is essential. The failure is *only* looking, *a few times*, *with no fixed set*.

```text
The vibe loop                         What it misses
─────────────                         ──────────────
1 input you picked       ──────────▶  the 20 inputs you didn't
read it once, charitably ──────────▶  the reader who isn't you
no saved baseline        ──────────▶  "is this better than before?"
ship on a nod            ──────────▶  the regression three categories over
```

*What just happened:* Each thing the vibe loop skips is a place a real bug hides. Eyeballing scales fine for one look at one case - it falls apart the moment you need to compare versions or cover the inputs you didn't think of.

The fix isn't to look harder. It's to **write down what you're looking for, once, against a set of inputs you keep** - so the judgment is fixed, repeatable, and the same every time you run it. That fixed set is an eval.

## What an eval actually is

Strip away the jargon and an eval is two things:

1. **A set of real inputs** - the kinds of things your feature actually receives. Not toy examples; the messy, boring, edge-y cases from real usage.
2. **A way to decide if the output was acceptable** for each input - an *expected behavior* you can check, whether that's an exact value, a rule, or a judgment.

If you've written a test suite, this will feel familiar, because it *is* a test suite - for a component whose output you can't pin to a single exact string. Each row is "given this input, the output should behave like *this*," and running the eval scores the current model-plus-prompt against the whole set at once.

A single row might look like this:

```json
{
  "input": "Reset my password, I'm locked out and the reset email never arrives.",
  "expected_intent": "account_access",
  "must_mention": ["spam folder", "support"],
  "must_not": ["refund", "make up a ticket number"]
}
```

*What just happened:* This row turns a fuzzy hope ("the bot should handle locked-out users well") into checkable facts: it should classify the intent as `account_access`, it should mention the spam folder and a way to reach support, and it must not wander into refunds or invent a ticket number. None of that needs a human re-reading it each run.

The magic isn't any one row. It's that you have *thirty* of them, drawn from real inputs, and you can run all thirty in seconds - so "did my change help?" becomes "27/30 passed, up from 24/30" instead of "felt good to me."

## Where the inputs come from

The most common eval mistake is making up clean, easy inputs that the model was always going to nail. An eval full of softballs tells you nothing. Get real ones:

- **Production logs.** The actual inputs your feature received. Gold.
- **The failures you remember.** Every time the model embarrassed you, that case belongs in the set - permanently. An eval is also a graveyard of past bugs that must never come back.
- **Edge cases on purpose.** Empty input, very long input, a different language, an adversarial "ignore your instructions" attempt, the input that's *almost* in scope but isn't.

You don't need hundreds to start. Twenty to fifty real, varied cases beat a thousand synthetic clean ones, because they exercise the places the model actually fails.

> 🪖 **War story.** A team "improved" their summarizer prompt and shipped on a glowing demo. The new prompt was tuned, unknowingly, for short articles - the demo input. On long articles it now truncated halfway and dropped the conclusion. A ten-row eval with two long articles in it would have turned the celebration into a red row before anyone shipped. They built the eval *after* the incident, which is the most common time teams build their first one.

## For builders

Start smaller than feels respectable. A `evals.jsonl` file with fifteen real inputs and a five-line script that runs them and prints `passed/total` is already better than every team still shipping on vibes - including, until recently, you. The discipline that matters is *keeping* the file and *adding to it* every time something breaks. Phase 2 is about the scoring; this phase is about accepting that without a fixed set to score against, you're guessing with extra steps.

```quiz
[
  {
    "q": "Why is eyeballing a few outputs especially misleading for an LLM feature compared to ordinary code?",
    "choices": [
      "Models are always wrong, so any output looks bad",
      "The output space is huge and non-deterministic, and prompt changes can silently break cases you didn't look at",
      "You can't read model output without special tools",
      "Ordinary code never has bugs"
    ],
    "answer": 1,
    "explain": "Huge output space, run-to-run variation, and far-reaching prompt changes mean a few charitable looks miss the cases that actually break."
  },
  {
    "q": "What are the two core ingredients of an eval?",
    "choices": [
      "A bigger model and a faster GPU",
      "A leaderboard score and a benchmark name",
      "A set of real inputs plus a way to decide if each output is acceptable",
      "A long prompt and a low temperature"
    ],
    "answer": 2,
    "explain": "An eval is real inputs paired with an expected behavior you can score - essentially a test suite for fuzzy output."
  },
  {
    "q": "Where do the best eval inputs come from?",
    "choices": [
      "Clean made-up examples the model is sure to pass",
      "Production logs, remembered failures, and deliberate edge cases",
      "The model's own training data",
      "Whatever input is shortest to type"
    ],
    "answer": 1,
    "explain": "Real, messy, and previously-failing inputs exercise where the model actually breaks; softball synthetic inputs tell you nothing."
  }
]
```


---

# How to Actually Grade Output

You've got your input set from Phase 1. Now the real question: for each output, how does a *machine* decide pass or fail, so you can run the whole set without re-reading every line by hand?

There isn't one answer, because "correct" means different things for different tasks. Extracting a date has one right answer; summarizing an article has a thousand acceptable ones. So you reach for one of three grading methods, from cheapest-and-strictest to most-flexible-and-fuzziest. The skill is matching the method to the task - and knowing how each one can fool you.

## Method 1: Exact and rule-based checks

**What it is.** Code that inspects the output and returns true or false. Exact match (`output == expected`), or looser rules: does it contain this substring, parse as valid JSON, fall in this set of allowed values, match this regex, stay under this length?

**When it fits.** Any task with a constrained, checkable answer:

- **Classification** - the intent should be exactly `account_access`. Exact match.
- **Extraction** - the pulled-out date should equal `2026-03-14`. Exact match.
- **Format / structure** - it must be valid JSON with these fields, or it must not contain a phone number. Rule check.

```python
def grade(output: str, row: dict) -> bool:
    # must be one of the allowed intents, and exactly the expected one
    if output.strip() not in {"account_access", "billing", "other"}:
        return False
    return output.strip() == row["expected_intent"]

print(grade("account_access", {"expected_intent": "account_access"}))  # True
print(grade("Account Access!", {"expected_intent": "account_access"})) # False
```

*What just happened:* The grader returned a hard pass/fail with zero ambiguity and zero cost - perfect for classification. Note the second case fails on a trailing `!` and capitalization: that's the method working as designed, not a bug. Exact checks are unforgiving, which is their strength and their trap.

**Where it lies to you.** It can't see meaning. "Yes." and "Absolutely, that's correct." are the same answer to a human and a fail/pass split to exact match. The moment the acceptable output is a *range* of phrasings - any summary, explanation, or chat reply - exact match either rejects good answers or you loosen it into uselessness.

💡 **Key point.** Always reach for rule-based first. It's free, instant, deterministic, and never argues. Use it for everything it *can* cover, and only escalate to the fuzzier methods for the parts it genuinely can't.

## Method 2: Reference-based metrics

**What it is.** You write down one or more *reference* (gold) answers and score how close the model's output is to them, using a similarity measure rather than exact equality. The crude classic is word overlap; the modern version compares meaning by turning both texts into vectors and measuring how close they point - semantic similarity.

**When it fits.** Tasks where there's a known good answer but the exact words can vary - translation, short factual answers, "did the summary capture the same key points as the reference summary?"

```text
output:    "The deploy failed because the database migration timed out."
reference: "Deployment broke due to a database migration that timed out."

word-overlap score:    moderate  (different wording trips it up)
semantic similarity:   high      (same meaning, scored as close)
```

*What just happened:* Two sentences that mean the same thing get a mediocre word-overlap score but a high semantic-similarity score. That gap is the whole reason semantic measures exist: they grade what was *said*, not which exact words were used.

**Where it lies to you.** Three ways, and they bite:

- **A score isn't a verdict.** Reference methods give you a number like `0.82`, not a pass/fail - you pick a threshold, and that threshold is a judgment call you can get wrong.
- **High similarity can still be wrong.** An output can be semantically close to the reference and yet have flipped a critical fact - "the migration *succeeded*" is very similar to "the migration *failed*" by overlap, and disastrously different in truth.
- **It's only as good as your reference.** A mediocre gold answer rewards mediocre outputs and punishes ones that are actually *better* than your reference. Garbage reference in, garbage score out.

⚠️ **Gotcha.** Treat reference scores as a signal, not a referee. They're great for catching big drops ("similarity fell off a cliff after this prompt change") and weak at certifying any single output as correct. Don't let a green number lull you past a flipped negation.

## Method 3: LLM-as-judge

**What it is.** You use a *second* model call to grade the first. You give a judge model the input, the output, and a rubric ("Score 1–5 on whether the answer is helpful, grounded in the provided context, and free of invented facts"), and it returns a score and a reason.

**When it fits.** The fuzzy, open-ended tasks the other two methods can't touch - "is this summary good?", "is this reply polite and on-brand?", "does this answer actually address the question?" When acceptable output is a wide range and you can articulate *what good looks like* in words, a judge can apply that rubric across hundreds of outputs far faster than you can.

```text
JUDGE PROMPT (sketch)
─────────────────────
You are grading a support reply. Given the user message and the reply,
score 1–5 on each: groundedness, helpfulness, tone.
A reply that invents facts scores 1 on groundedness regardless of tone.
Return JSON: {"groundedness": n, "helpfulness": n, "tone": n, "why": "..."}
```

*What just happened:* You encoded your definition of "good" into a rubric a model can apply at scale. The judge isn't smarter than you - it's *you, written down and run a thousand times*, which is exactly the leverage Phase 1 promised.

**Where it lies to you - and this is the big one.** The judge is itself an LLM, so it inherits every flaw you're trying to measure:

- **It can be confidently wrong** about the grade, the same way the thing it's grading can be.
- **It has biases.** Many judges favor longer, more confident-sounding answers, or favor an answer that's positioned first, or rate their own model's style highly. A flattering output can score well for being flattering.
- **A vague rubric gets vague grades.** "Rate the quality 1–10" gives you noise. Specific, behavior-anchored criteria ("invents facts → groundedness = 1") give you something repeatable.

🪖 **War story.** A team automated grading with a judge and watched scores climb release after release - then a customer flagged answers that were polished, friendly, and *wrong*. The judge had a length-and-confidence bias: prompt changes had made outputs longer and more assertive, not more correct, and the judge happily rewarded that. The fix: tighten the rubric to score groundedness explicitly, and validate the judge against a small human-graded set.

💡 **Key point.** Before you trust an LLM judge, grade a few dozen outputs *by hand*, then have the judge grade the same ones, and check that they agree. If the judge disagrees with humans, fix the rubric - don't ship a grader you haven't graded.

## Choosing a method

You'll usually use more than one. A single eval row can be checked by rules *and* a judge.

```text
Task shape                          Reach for
──────────                          ─────────
one exact right answer        ──▶   rule-based / exact match
known answer, wording varies  ──▶   reference-based (as a signal)
open-ended, "is it good?"     ──▶   LLM-as-judge (validated)
```

*What just happened:* The grading method follows the *shape* of the task, not your preference. Default to the cheapest method the task allows - rules over references over judges - and only climb to fuzzier, costlier, more-fallible grading when the task genuinely demands it.

## For builders

Layer them. Run the rule checks first as a hard gate (valid JSON? required field present? not over length?) - free, and they catch the dumb breakages instantly. Only send outputs that *pass* the gate to a judge, since judge calls cost tokens and time. Keep a small human-graded sample as ground truth to periodically check your judge still agrees with actual humans. The grader is code; like any code, it can rot, and an eval you trust blindly is vibes with extra latency.

```quiz
[
  {
    "q": "For grading a classification task where the intent must be exactly 'billing', which method fits best?",
    "choices": [
      "LLM-as-judge with a detailed rubric",
      "Reference-based semantic similarity",
      "Rule-based / exact match",
      "Reading every output by hand"
    ],
    "answer": 2,
    "explain": "A constrained, single-correct-answer task is exactly where cheap, deterministic exact-match checks shine."
  },
  {
    "q": "What's the most important precaution before trusting an LLM-as-judge?",
    "choices": [
      "Use the largest available model as the judge",
      "Validate it against a small human-graded set and confirm it agrees with people",
      "Set the judge's temperature to zero",
      "Ask the judge to also fix the output"
    ],
    "answer": 1,
    "explain": "A judge is itself an LLM with biases; you must check it agrees with humans before trusting its scores at scale."
  },
  {
    "q": "Why are reference-based metric scores a signal rather than a verdict?",
    "choices": [
      "They are always exactly right",
      "They produce a similarity number that needs a threshold and can rate a fact-flipped answer as 'close'",
      "They only work on numbers, never text",
      "They require a human to read every output"
    ],
    "answer": 1,
    "explain": "Similarity gives a number, not pass/fail, and a flipped negation can score 'close' while being completely wrong - so treat it as a signal."
  }
]
```


---

# Evals as a Habit

A one-time eval is a snapshot - useful, but it ages the moment you change anything. The payoff comes when running the eval is a *reflex*: something that happens before every prompt edit ships, every time the provider releases a new model, every time you're tempted to nod at a demo and call it done. That's when the eval stops being homework and becomes the thing that lets you move fast without breaking cases you already fixed.

This phase is about the habit: catching regressions, surviving model upgrades, watching quality drift over time, and - crucially - knowing where the eval's reach ends and the real world begins.

## Regression testing: the eval's main job

A regression is when a change makes something that *used to work* stop working. In normal code, your test suite catches it. With a model feature, the eval is that test suite.

The workflow is mechanical and that's the point:

```text
1. Establish a baseline    →  run the eval on what you ship today: 27/30 pass
2. Make your change        →  edit the prompt (or swap the model)
3. Re-run the eval         →  same 30 inputs, same graders
4. Compare                 →  28/30? ship it.  24/30? do NOT ship.
5. Inspect what moved      →  which rows flipped, in BOTH directions
```

*What just happened:* You turned "I think this is better" into a before-and-after diff on a fixed set. The number going up isn't even the most important part - step 5 is, because a change can lift the total while *breaking* specific rows you care about.

⚠️ **Gotcha.** Don't read only the headline score. Going from 27 to 28 can hide a disaster: you fixed *four* easy rows and broke *three* important ones. Always look at which rows flipped from pass to fail, not only the net. A green aggregate over a broken critical case is how regressions ship.

💡 **Key point.** Every bug you fix becomes a permanent eval row. The case the model got wrong in production goes into the set with its expected behavior, *forever*. That's how the eval grows teeth over time - it accumulates every mistake you've ever made so none of them can come back unnoticed.

## Surviving a model upgrade

The provider releases a newer, "better," cheaper model and you want to switch. Here's the trap: "better on the provider's benchmarks" is not "better on *your* task with *your* prompts." A new model can be smarter in general and still regress *your* specific use case - it may follow your formatting instructions differently, be more verbose, or interpret an edge case the old one handled.

Your eval is exactly the tool for this. Don't swap and pray:

- **Run your eval on the new model before switching.** Same inputs, same graders, new model. Compare to your current baseline.
- **Expect to re-tune the prompt.** A prompt squeezed to work around the *old* model's quirks may be fighting the new one. The eval tells you whether a tweak helped.
- **Watch cost and latency, not only quality.** A model that scores the same but is slower or pricier is a regression in a dimension the pass-count doesn't show. Track those alongside accuracy.

🪖 **War story.** A team upgraded to a newer model the day it launched, on the strength of the announcement post. Quality on their extraction task quietly dropped - the new model wrapped its JSON in a markdown code fence the old one never used, and their parser choked. Their eval would have caught it in one run, because "valid parseable JSON" was a rule check. They didn't run it. Lesson: never let a model swap skip the eval, no matter how shiny the launch.

## Tracking quality over time

A single before/after is good. A *trend* is better. If you save each eval run - the date, the model, the prompt version, the scores - you can see quality as a line over weeks, not a feeling per release.

```text
date        prompt  model       pass   notes
────────    ──────  ─────────   ────   ─────────────────────────
2026-05-02  v3      model-A     24/30  baseline
2026-05-14  v4      model-A     27/30  added few-shot examples
2026-06-01  v4      model-B     25/30  upgraded model - REGRESSED
2026-06-03  v5      model-B     28/30  re-tuned prompt for model-B
```

*What just happened:* The history makes the model-B regression and the recovery visible at a glance, tying each score to the exact prompt and model that produced it. Without this record, the June 1st dip would have been invisible until a user complained.

This doesn't need a fancy dashboard to start. A committed results file, or appended rows in a spreadsheet, gives you the trend. The discipline is recording *what changed* next to *what the score did*, so a drop always has a suspect.

## Offline evals are not the whole story

Here's the real limit, and it's an important one: your eval set is a *sample of the past*. It tells you how your system does on inputs you've already collected. It cannot tell you about:

- **Inputs you've never seen.** Real users will send things your set doesn't contain. Always.
- **Real-world drift.** User behavior, slang, the topics they ask about, even the model behind the API can shift under you over time.
- **What users actually feel.** Your grader's definition of "good" is a proxy. Sometimes the thing that passes every check still isn't what the user wanted.

So evals are necessary but not sufficient. They're the *offline* half. The other half is watching production:

- **Production monitoring.** Log real inputs and outputs (privacy permitting), track error rates, latency, cost, refusals, and parse failures. Sample real traffic and run your graders - even your judge - on *live* outputs, not only the frozen set.
- **User feedback.** A thumbs up/down, a "regenerate," a copy-to-clipboard, an edit-before-sending, a support ticket - these are real signals of quality your offline set can't fake. Feed them back into the loop.
- **Close the loop.** When monitoring or feedback surfaces a failure, that case becomes a new eval row. Production *finds* the gaps; the eval set *remembers* them. That circle - offline catches regressions, production catches the unknown, and the unknown becomes part of offline - is the whole quality system.

```text
   OFFLINE EVALS                    PRODUCTION
   ─────────────                    ──────────
   fixed input set                  real, novel inputs
   catches regressions      ◀────▶  catches the unknown unknowns
   run before you ship              watched after you ship
        ▲                                │
        └──── new failure becomes ───────┘
                a new eval row
```

*What just happened:* The two halves feed each other. Offline evals stop you from re-breaking known cases; production tells you about cases you never imagined; and every new production failure graduates into the offline set so it's covered from then on.

## For builders

The minimum viable habit: a committed eval file, a script that runs it and prints the score, and a rule that *no prompt or model change merges without a before/after run pasted into the change*. That alone puts you ahead of most teams shipping AI features. Layer monitoring on once you're live - even a thumbs-up button and logged inputs - and pipe every real failure back into the eval set. You don't need a platform; you need the loop to close.

## Recap

1. **Run the eval on every change** - establish a baseline, re-run after editing, and compare. The diff replaces the vibe.
2. **Read which rows flipped, not only the total** - a higher aggregate can hide a broken critical case.
3. **Never swap models on the announcement alone** - run your eval on the new model first, expect to re-tune the prompt, and track cost and latency too.
4. **Save every run** so quality is a trend tied to specific prompt and model versions, and a drop always has a suspect.
5. **Offline evals catch regressions; production monitoring and user feedback catch the unknown** - and every real failure becomes a new permanent eval row.

You now have the full loop: a real input set, a way to grade it, and the habit of running it before every change while watching production after. That's the difference between hoping your AI feature works and knowing it does.

```quiz
[
  {
    "q": "After a prompt change, your eval goes from 27/30 to 28/30. What should you do before shipping?",
    "choices": [
      "Ship - the number went up",
      "Inspect which rows flipped in both directions, in case important cases broke",
      "Revert - any change is risky",
      "Delete the failing rows so it reads 30/30"
    ],
    "answer": 1,
    "explain": "A higher total can hide regressions: you may have fixed easy rows and broken important ones. Always check what moved, not just the net."
  },
  {
    "q": "A provider ships a newer, cheaper model. What's the safe way to adopt it?",
    "choices": [
      "Swap it in immediately because newer is better",
      "Trust the provider's benchmark scores",
      "Run your own eval on it first, expect to re-tune the prompt, and track cost and latency",
      "Wait a year for reviews"
    ],
    "answer": 2,
    "explain": "Better on general benchmarks isn't better on your task; run your eval, re-tune, and watch cost and latency before switching."
  },
  {
    "q": "What's the key limit of offline evals, and what complements them?",
    "choices": [
      "They're too slow; nothing complements them",
      "They only sample past inputs, so production monitoring and user feedback catch the unknowns - which then become new eval rows",
      "They're perfect and need nothing else",
      "They replace the need to ever look at production"
    ],
    "answer": 1,
    "explain": "Offline evals cover inputs you've collected; production monitoring and feedback catch novel cases, and those failures feed back into the eval set."
  }
]
```
