# RAG in Plain English

> Retrieval-augmented generation is how you get an AI to answer from your documents instead of its memory. The idea, the moving parts, and why it goes wrong.


---

# RAG in Plain English

A regular AI chatbot answers from what it absorbed during training. That training stopped on a certain date, and it never saw your company wiki, your contracts, your support tickets, or last week's policy change. So when you ask it about any of that, it does one of two things: it tells you it doesn't know, or worse, it makes something up that sounds right. Neither is useful when you need a real answer about your own stuff.

RAG - retrieval-augmented generation - is the standard fix. The shape of it is straightforward: before the AI answers, a search step pulls the most relevant passages out of your documents and pastes them into the question. The model then answers using those passages, the same way you'd answer a question with the right page open in front of you. The "memory" doing the work is your documents, not the model's training.

This guide is for anyone deciding whether to build, buy, or trust a RAG system - founders, ops people, support leads, writers wiring up an "ask our docs" feature. You don't need to be an engineer and there's no math here. Phase 1 covers the actual problem RAG solves and when you need it. Phase 2 walks through the moving parts: chunking, turning text into searchable vectors, retrieving the right bits, and handing them to the model. Phase 3 is the clear-eyed part - the four ways these systems quietly produce wrong answers, what each one looks like, and what to do about it. By the end you'll be able to look at a RAG product and reason about whether to believe what it tells you.


---

# Give the AI Your Documents

Ask a plain AI chatbot "what's our refund window for enterprise customers?" and watch what happens. It has never seen your refund policy. So it answers from the average of everything it read during training - "typically 30 days" - in a confident tone that gives you no signal that it's guessing. The answer might be right. It might be six months out of date. You have no way to tell from the reply.

This is the core problem, and it has two halves.

## The model's memory is frozen and generic

An AI model learns from a giant pile of text up to a cutoff date, and then it stops. It's a snapshot. After that, it knows nothing new unless you tell it. Last week's price change, the policy you shipped yesterday, the contract you signed this morning - none of it is in there.

And even for things that *were* in the training data, the model absorbed the public internet, not your private drive. Your internal runbooks, your customer records, your unreleased product specs were never part of it. The model is a brilliant generalist who has read a library but never set foot in your office.

So when you ask it about your world, it has two failure modes. It admits it doesn't know - truthful but useless. Or it pattern-matches to something plausible and states it as fact. That second one is the dangerous default, because the confidence never drops to match the ignorance.

## Two ways to fix it (and why one wins)

There are really only two ways to get private or current knowledge into an AI's answers.

**Fine-tuning** retrains the model on your data so the knowledge gets baked in. It sounds like the obvious move, but it's the wrong tool for facts. It's slow and expensive to do, you have to redo it every time a document changes, and - the killer - the model still can't tell you *where* an answer came from. Your refund policy changes Tuesday, and your fine-tuned model happily quotes the old one until you retrain. Fine-tuning is good at teaching a model a *style* or a *format*. It's bad at keeping it current on *facts*.

**Retrieval** leaves the model alone and changes what you hand it at question time. Right before the model answers, a search step finds the relevant passages in your documents and pastes them into the prompt. The model reads them on the spot and answers from them. This is RAG.

The difference matters in practice:

| | Fine-tuning | Retrieval (RAG) |
|---|---|---|
| Update a fact | Retrain the model | Edit the document |
| Cost to change | High, every time | Roughly free |
| Can cite the source | No | Yes |
| Good for | Tone, format, behavior | Current, private facts |

For "answer from our documents," retrieval wins on every line that matters. That's why almost every "chat with your docs" product you've seen is RAG under the hood, not a custom-trained model.

## The open-book exam analogy

Here's the mental model to carry through the rest of this guide.

A plain chatbot is a student taking an exam from memory. Smart, well-read, but stuck with whatever they happened to study, and prone to bluffing on the questions they didn't.

RAG is the same student taking an *open-book* exam. Before each question, someone hands them the two or three pages most likely to contain the answer. The student is as smart, but now they're reading the actual source instead of straining to recall it. The answers get more accurate, more current, and - because they're working from specific pages - they can point at where the answer came from.

That last part is the quiet superpower. A good RAG system can show you the passages it used. "Here's the answer, and here are the three paragraphs from the employee handbook I based it on." Now you can check its work. A plain chatbot gives you a verdict with no evidence; RAG gives you a verdict with a paper trail.

## When you actually need it

You don't need RAG for everything. If you're asking the AI to brainstorm names, rewrite an email, or explain a general concept, its built-in memory is plenty.

You need RAG when the answer has to come from a specific, trusted body of text:

- A support bot that answers from *your* product docs, not generic advice.
- An internal tool that answers questions from policies, contracts, or wikis.
- Anything where being current matters - prices, inventory, this quarter's numbers.
- Anything where a wrong answer is expensive and you need to show the source.

The pattern in all of these is the same: the truth lives in documents you control, and you want the AI to read from them instead of from a hazy memory of the public internet. The next phase opens up the box and shows how that retrieval step actually finds the right pages.


---

# How It Works

RAG is a pipeline with four steps. Two of them happen once, ahead of time, when you load your documents in. The other two happen every time someone asks a question. None of it requires math to understand - the whole thing is "make the documents searchable, then search them before answering."

Here's the shape:

```mermaid
graph LR
  A[Your documents] --> B[Split into chunks]
  B --> C[Turn into vectors]
  C --> D[(Vector store)]
  E[Question] --> F[Retrieve top chunks]
  D --> F
  F --> G[Model answers]
```

Steps B and C are the one-time setup. Steps F and G run on every question. Let's walk each one.

## Step 1: Chunk the documents

You can't hand the model a 200-page manual every time someone asks a question - it's too much to read, and most of it is irrelevant to any single question. So the first job is to chop your documents into bite-sized pieces. A chunk might be a paragraph, a section, or a few hundred words. These chunks are what the system searches over and retrieves.

Chunking sounds boring. It is the single most important setup decision you'll make, and Phase 3 explains why. For now, hold this thought: a chunk should be big enough to contain a complete idea, but small enough that it's mostly *about* one thing. A chunk that's half about refunds and half about shipping will haunt you later.

## Step 2: Turn each chunk into a vector

This is the part that sounds like magic, so let's strip the magic off.

A computer can't search text by *meaning* the way you do. If your document says "reimbursement" and someone searches "getting my money back," a plain keyword search misses it - no shared words. RAG fixes this by converting each chunk into a list of numbers called an **embedding**, or a **vector**. The trick of these numbers is that chunks with similar *meaning* get similar numbers. "Reimbursement policy" and "how to get money back" land close together, even with no words in common.

The way to picture it: imagine every chunk gets dropped onto a giant map, and things that mean similar things land near each other. All the refund-related chunks cluster in one neighborhood, the shipping chunks in another, the password-reset chunks somewhere else entirely. The vector is only the chunk's address on that map. (The real "map" has hundreds of dimensions, not two - but you never have to think about that. "Similar meaning lands nearby" is the whole idea.)

A separate AI model, called an embedding model, does this conversion. You run all your chunks through it once and store the resulting vectors in a **vector store** (or vector database) - a system built to answer one question fast: "which stored vectors are closest to this one?"

## Step 3: Retrieve the relevant chunks

Now someone asks a question. The system runs the *question* through the same embedding model, turning it into a vector too - an address on the same map. Then it asks the vector store: which chunks live closest to this question?

The store hands back the top few - often the closest three to ten chunks. Those are your open-book pages: the passages most likely to contain the answer. This is the "retrieval" in retrieval-augmented generation, and it usually takes a fraction of a second even across millions of chunks.

Note what this step does *not* do: it doesn't understand the answer, it doesn't reason, it doesn't check anything. It's a similarity match. It returns the chunks that look most related to the question. Whether they actually answer it is a gamble that Phase 3 will make you respect.

## Step 4: Hand them to the model and answer

Finally, the system builds a prompt that stitches the retrieved chunks together with the original question. Roughly:

```text
Use the following passages to answer the question.
If the answer isn't in them, say you don't know.

[chunk 1: ...]
[chunk 2: ...]
[chunk 3: ...]

Question: What's our refund window for enterprise customers?
```

The model reads the passages and writes an answer grounded in them. Because the relevant text is right there in front of it, it doesn't have to recall anything from training - it reads off the page. A well-built system also returns *which* chunks it used, so the answer comes with citations you can click and verify.

That instruction - "if the answer isn't in the passages, say you don't know" - is doing heavy lifting. It's the system's attempt to stop the model from filling gaps with invention. It helps. It does not fully work, which is the heart of the next phase.

## The whole thing in one breath

Set-up, once: split your documents into chunks, convert each chunk to a vector, store the vectors. Per question: convert the question to a vector, find the nearest chunks, paste them in front of the model, let it answer from them.

That's RAG. Every "chat with your PDF," every support bot trained on a help center, every internal "ask the wiki" tool is some version of these four steps. The concept is clean. The reason real systems still give wrong answers isn't the concept - it's that each of these four steps has a way to quietly fail. That's where we go next.


---

# Why It Goes Wrong

RAG is a clean idea, and the demo always works. You point it at a tidy PDF, ask an obvious question, and it nails the answer with a citation. Then you ship it on real documents with real users, and it starts handing out confident, wrong answers. Every time, the failure traces back to one of four steps in the pipeline. Learn these four and you can diagnose almost any RAG problem - and decide how much to trust one you didn't build.

## Failure 1: Bad chunking

Remember chunking - splitting documents into searchable pieces. Get it wrong and everything downstream inherits the mistake, because you can't retrieve a clean answer that was never stored as a clean chunk.

The classic break is **splitting in the middle of an idea**. Your policy says "Refunds are available within 30 days" and then, in the next paragraph, "...except for enterprise contracts, which are non-refundable." If the chunk boundary falls between those two paragraphs, the system can retrieve the first half and never see the exception. The answer it gives is exactly half right, which is worse than no answer.

The other break is **chunks that are too big or too small**. Too big, and a single chunk covers five topics, so it gets retrieved for everything and dilutes the relevant part. Too small, and a chunk lacks the surrounding context needed to make sense - a line that says "this is not permitted" with no nearby clue what "this" is.

What it looks like in the wild: answers that are partially correct, that miss exceptions and edge cases, or that confidently state a rule without its qualifier. The fix is unglamorous - chunk along the document's natural structure (headings, sections), keep related material together, and test with real questions. There's no universal right chunk size; it depends on your documents, which is exactly why teams skip the work and then wonder why the bot is flaky.

## Failure 2: Stale data

RAG's whole pitch is staying current - but only if someone keeps the index current. The retrieval step searches the vectors you stored, not the documents as they exist *right now*. Edit the source document and the stored vectors don't change until something re-processes it.

So the policy team updates the refund window in the wiki on Tuesday. The RAG index was built last month. Until someone re-runs the chunk-and-embed step, the bot keeps quoting the old number - with full confidence and a citation pointing at a document that no longer says that. The citation makes it *more* believable, not less.

This one is sneaky because nothing errors out. The system works perfectly; it's only answering from a stale snapshot. What it looks like: answers that were right last quarter, prices that don't match the current ones, references to a process you've since changed. The fix is operational, not technical - re-index on a schedule or whenever source documents change, and show the document's last-updated date so a human can smell when something's off.

## Failure 3: Retrieving the wrong passages

This is the heart of it, because retrieval is a similarity match, not an understanding. The vector store returns the chunks whose *wording and topic* sit closest to the question - which is usually, but not always, the same as the chunks that *answer* it.

A few ways it whiffs:

- **The right answer exists but loses the race.** It's chunk number eleven, the system only grabs the top five, and the answer never reaches the model. The model then answers from five chunks that don't contain it.
- **It grabs a close cousin.** Ask about the *enterprise* refund policy and it returns the *consumer* refund policy - same topic, same words, wrong audience. They live in the same neighborhood on the map.
- **The question is vague.** "How does this work?" matches a hundred chunks equally well, so the system returns a grab-bag and the answer is mush.

What it looks like: answers that are about the right *subject* but wrong in the *specifics*, or answers that miss information you know is sitting in the documents. The fixes range from cheap to involved - retrieve more chunks and let a second pass re-rank them, combine vector search with old-fashioned keyword search so exact terms aren't lost, and write clearer source documents. But you can never assume retrieval is perfect. It is the step most likely to silently hand the model the wrong page.

## Failure 4: Trusting an answer the sources don't support

Suppose retrieval did its job and handed the model the right passages. The model can *still* produce an answer the passages don't actually back up. It might blend the retrieved text with something half-remembered from training. It might stretch a passage to cover a question it doesn't quite address. It might be handed nothing useful and, instead of saying "I don't know," fill the silence with a plausible guess.

That instruction from Phase 2 - "if the answer isn't in the passages, say you don't know" - reduces this. It does not eliminate it. Models are built to be helpful, and "helpful" and "upfront about not knowing" are in tension. Under pressure, helpfulness often wins.

This is the failure that erodes trust fastest, because the answer *sounds* grounded - it's fluent, it's specific, it may even carry a citation that, on a click, doesn't actually say what the answer claims. What it looks like: confident claims that aren't in the cited source, details that go beyond what the passages contain, an answer where there should have been a plain "I couldn't find this."

The real defense is to **check the citations**, not the prose. A good RAG system shows you the exact passages it used. The discipline - for you and for anyone relying on the tool - is to read those passages, not the AI's summary, whenever the answer matters. If the system can't show its sources, treat its answers like a confident stranger's: possibly right, not to be trusted on anything that counts.

## The diagnostic, in one table

When a RAG answer is wrong, it's almost always one of these. Walk them in order:

| Symptom | Likely culprit |
|---|---|
| Half-right, misses the exception | Bad chunking |
| Was right months ago, wrong now | Stale data |
| Right topic, wrong specifics | Wrong passages retrieved |
| Confident claim not in the source | Ungrounded answer |

RAG is the right tool for "answer from our documents," and when it's built and maintained with care it's genuinely useful. But it is a pipeline of four fallible steps, not a black box that knows things. Understanding where it breaks is what lets you use it without getting burned - and what lets you ask the right question of any RAG product before you trust it with anything that matters.
