# Elasticsearch and OpenSearch

> Full-text search at scale: the inverted index, analyzers and relevance scoring, and how a search engine differs from a database WHERE clause.


---

# Elasticsearch and OpenSearch

You added a search box. You wired it to `WHERE title LIKE '%term%'` and it worked on your laptop with fifty rows. Then real data arrived, the queries crawled, typos returned nothing, and "the most relevant result" turned out to mean "whatever row the database happened to scan first." A search engine exists because that problem is genuinely different from a database lookup, and this guide shows you why and how to reach for the right tool.

## How to read this

Read it in order the first time. Phase 1 builds the one mental model that makes everything else click: search engines flip the data structure inside out. Phase 2 is the everyday work, indexing documents, mappings, analyzers, and getting back ranked results. Phase 3 is the reality nobody warns you about, near-real-time refresh, eventual consistency, and the real question of whether you needed search at all. Skim later for reference, but the first pass should be front to back.

## The phases

1. [Phase 1: The inverted index, or why search is a different problem](01-the-inverted-index.md)
2. [Phase 2: Indexing, mappings, analyzers, and getting ranked results](02-indexing-and-relevance.md)
3. [Phase 3: Near-real-time, consistency, and when to add search at all](03-production-reality.md)


---

# Phase 1: The inverted index, or why search is a different problem

Here is the moment that sends most people toward a search engine. You have a `products` table, a search box, and a query that started life looking reasonable:

```sql
SELECT * FROM products WHERE description LIKE '%wireless%';
```

On your laptop, with a few hundred rows, this is instant. In production, with millions of rows, it is a slow, sweaty full-table scan every single time. And a search engine does not make this query faster. It makes the query unnecessary, by storing the data in a completely different shape. Understanding that shape is the whole game.

## Why LIKE cannot scale, in one picture

A normal database index, the kind you create on a column, is sorted by the *value*. That lets the database jump straight to a row when you ask for an exact value or a prefix:

```sql
-- A normal B-tree index helps with this. It can binary-search to "wireless".
WHERE name = 'wireless mouse'
WHERE name LIKE 'wireless%'   -- prefix: still usable

-- It is useless for this. The wildcard is on the LEFT.
WHERE description LIKE '%wireless%'
```

*What just happened:* `LIKE '%wireless%'` puts a wildcard at the front, so the database cannot use its sorted index, there is no "first letter" to seek to. It has to read every row and check each one. That is O(n), and n only grows.

The deeper problem is that a `LIKE` scan does not understand *words*. `'%cat%'` matches "category" and "scatter". It cannot find "running" when you search "run". It has no idea that "the" appears in every document and carries no signal. It returns rows in whatever order they were scanned, with no notion of which result is *better*. A database is built to answer "give me the rows where this is exactly true." Search is a different question: "give me the best documents for these words, even with typos, ranked by relevance." Different question, different data structure.

## The inverted index: flip the table on its side

A normal table maps a row to its words. You start with a document and read its contents:

```text
doc 1 -> "wireless mouse with usb receiver"
doc 2 -> "wireless keyboard and mouse combo"
doc 3 -> "wired gaming mouse"
```

An *inverted* index flips that. It maps each word to the list of documents that contain it. You start with a word and instantly get every document:

```text
wireless -> [1, 2]
mouse    -> [1, 2, 3]
keyboard -> [2]
wired    -> [3]
usb      -> [1]
gaming   -> [3]
```

*What just happened:* searching for "mouse" is no longer a scan. The engine looks up the single word "mouse" and reads back `[1, 2, 3]` directly, the answer is precomputed and sorted. Searching "wireless mouse" intersects `[1, 2]` with `[1, 2, 3]` to get `[1, 2]`. This is the same trick a book's index uses: you do not read all 400 pages to find every mention of "recursion," you flip to the back, find the word, and it lists the pages. The index was built once; every lookup after that is cheap.

That inversion is the one idea everything else hangs off. The cost moved: building the index is work done *at write time*, so reads, the thing your users wait on, become a lookup instead of a scan.

```mermaid
flowchart LR
  A["Document:<br/>wireless mouse"] --> B["Analyzer:<br/>tokenize + normalize"]
  B --> C["Terms:<br/>wireless, mouse"]
  C --> D["Inverted index<br/>term -> doc ids"]
  E["Query:<br/>Wireless Mice"] --> B
  D --> F["Matching docs,<br/>ranked"]
```

*What just happened:* notice the query goes through the *same* analyzer as the document. That is the secret to matching "Wireless Mice" against a document that said "wireless mouse," covered in Phase 2. Both sides get reduced to the same terms before they ever meet.

## Elasticsearch and OpenSearch: same engine, forked history

Both products are HTTP servers wrapped around Apache Lucene, the Java library that actually implements the inverted index. You talk to them with JSON over REST. They speak almost the same API, store data the same way, and solve the same problem.

The split is a licensing story, not a technical one. Elasticsearch is made by Elastic. In 2021 Elastic changed Elasticsearch's license away from the permissive Apache 2.0 terms, and Amazon (along with the community) forked the last Apache-licensed version into **OpenSearch**, which stays under Apache 2.0. Since then the two have drifted apart in some advanced features, but the core, indices, mappings, analyzers, the query DSL, BM25 scoring, is the same mental model in both.

> For this guide, when we say "the engine," it applies to both. Where a name or detail differs, that is usually a feature or paid-tier difference, not a difference in how search itself works. Learn one, you can read the other's docs.

## For builders

Reach for this when your search must understand language: relevance ranking, typo tolerance, matching "run" against "running," searching across many fields at once, faceted filters, autocomplete. Stay with your database when you need exact lookups, joins, transactions, or your "search" is really just a filter on a few well-indexed columns. The two are not rivals; most real systems keep the database as the source of truth and feed a search engine alongside it. Phase 3 makes that call concrete.

```quiz
[
  {
    "q": "Why can a normal B-tree database index not speed up WHERE description LIKE '%wireless%'?",
    "choices": [
      "B-tree indexes only work on integer columns",
      "The leading wildcard means there is no sorted prefix to seek to, forcing a full scan",
      "LIKE is disabled on indexed columns by default",
      "The index is rebuilt on every query, which is slow"
    ],
    "answer": 1,
    "explain": "A B-tree is sorted by value, so it can seek a prefix. A leading % gives nothing to seek to, so every row must be read."
  },
  {
    "q": "What does an inverted index map?",
    "choices": [
      "Each document to the words it contains",
      "Each table to its foreign keys",
      "Each term to the list of documents that contain it",
      "Each row to its primary key"
    ],
    "answer": 2,
    "explain": "It is inverted from the normal document-to-words layout: term to document ids, so a word lookup returns matches directly."
  },
  {
    "q": "What is the main difference between Elasticsearch and OpenSearch?",
    "choices": [
      "OpenSearch uses a different index structure, not an inverted index",
      "They are unrelated products that happen to share a name",
      "OpenSearch is a fork created over a licensing change; the core search model is the same",
      "Elasticsearch cannot do full-text search"
    ],
    "answer": 2,
    "explain": "OpenSearch forked the last Apache-licensed Elasticsearch after a license change. Both wrap Lucene and share the core search model."
  }
]
```


---

# Phase 2: Indexing, mappings, analyzers, and getting ranked results

Phase 1 was the model. Now you actually use it. You will put documents in, control how the words get processed, and pull results back out *ranked by relevance* instead of by accident. Everything here is JSON over HTTP, so the examples are `curl`-shaped and read the same against Elasticsearch or OpenSearch.

## Indexing a document

In search-engine vocabulary an **index** is roughly a table, and a **document** is roughly a row, stored as JSON. You add a document with a PUT:

```bash
PUT /products/_doc/1
{
  "name": "Wireless Optical Mouse",
  "description": "Ergonomic wireless mouse with USB receiver",
  "price": 24.99,
  "in_stock": true
}
```

*What just happened:* you created (or replaced) the document with id `1` in the `products` index. The engine ran each text field through an analyzer, broke it into terms, and wrote those terms into the inverted index. The document is now findable, almost, there is a short delay before it shows up, which is the whole story of Phase 3.

## Analyzers: the step that makes text searchable

An **analyzer** is the pipeline that turns a blob of text into the clean terms that go in the index. It runs at write time on your documents and again at query time on the search string, so both sides end up speaking the same reduced language. A typical analyzer does three things:

```text
input:      "The Wireless Mice, Running!"
1 tokenize: ["The", "Wireless", "Mice", "Running"]   (split on spaces/punctuation)
2 lowercase:["the", "wireless", "mice", "running"]
3 stem/stop:["wireless", "mouse", "run"]             (drop "the", stem to roots)
```

*What just happened:* this is why a search for "Wireless Mice" finds a product described as "wireless mouse." Both strings get lowercased, stripped of the stopword "the," and stemmed to root forms, so "Mice" and "mouse" both collapse toward the same term. Without the analyzer, you would be doing exact string matching and the search would feel broken. The default analyzer handles most English text well; you swap in language-specific or custom analyzers when you need them.

## Mappings: telling the engine what each field is

A **mapping** is the schema for an index, it says each field's type and how it should be analyzed. The engine can guess (dynamic mapping), but guessing causes pain later, so for anything real you define it:

```bash
PUT /products
{
  "mappings": {
    "properties": {
      "name":        { "type": "text" },
      "price":       { "type": "float" },
      "in_stock":    { "type": "boolean" },
      "category":    { "type": "keyword" }
    }
  }
}
```

*What just happened:* the critical distinction is `text` versus `keyword`. A **`text`** field is run through the analyzer and broken into terms, that is what you do full-text search against. A **`keyword`** field is stored whole, exactly as given, not analyzed, so it is right for things you filter, sort, or group by exactly: status codes, tags, category slugs, IDs. Mark a category as `text` and you cannot reliably filter on the exact value "kitchen-tools"; mark a description as `keyword` and full-text search against it stops working. Choosing the type per field is the bulk of mapping work, and it is the same instinct as designing a clean schema in [/guides/designing-apis-that-last](/guides/designing-apis-that-last): name the shape of your data deliberately, up front.

## Querying, and the word that matters most: score

Here is a real search across two fields:

```bash
POST /products/_search
{
  "query": {
    "multi_match": {
      "query": "wireless mouse",
      "fields": ["name", "description"]
    }
  }
}
```

The response comes back sorted by `_score`, highest first:

```json
{
  "hits": {
    "max_score": 2.41,
    "hits": [
      { "_id": "1", "_score": 2.41, "_source": { "name": "Wireless Optical Mouse" } },
      { "_id": "2", "_score": 0.93, "_source": { "name": "Wireless Keyboard and Mouse Combo" } }
    ]
  }
}
```

*What just happened:* this is the thing a database `WHERE` clause cannot give you. Document 1 scored higher than document 2 because the query terms are denser and more prominent in it. The engine did not return "the rows that match," it returned "the best matches, ranked." That ranking number is `_score`.

## BM25: where the score comes from

The default scoring algorithm is **BM25**. You do not need its formula, you need its intuition, which is three forces:

- **Term frequency** - a document that uses your search word more often is more relevant, but with *diminishing returns*. The tenth mention of "mouse" barely moves the needle past the third.
- **Inverse document frequency** - a word that appears in *few* documents is more informative. Matching "ergonomic" tells you far more than matching "the," so rare terms are weighted up and common ones down.
- **Field length** - a match in a short field counts for more. Your term appearing in a five-word title is a stronger signal than the same term buried in a thousand-word description.

```text
search "wireless mouse":
  doc A title "Wireless Mouse"            -> short field, both terms        -> high score
  doc B desc "...a wireless mouse pad..." -> long field, terms less central -> lower score
```

*What just happened:* BM25 combined those three forces to decide A beats B. This is also why returning the literal `LIKE` scan order felt random, there was no scoring at all. When relevance looks wrong, you are usually fighting one of these three knobs (often field length or a too-common term), and the fix is in your mapping and analyzer choices, not in the algorithm.

## Filter vs. query: not everything needs a score

Often you want to *narrow* results without affecting ranking, "in stock only," "under $50." That is a filter, and it is both faster and cacheable because the engine skips scoring entirely:

```bash
POST /products/_search
{
  "query": {
    "bool": {
      "must":   { "multi_match": { "query": "wireless mouse", "fields": ["name", "description"] } },
      "filter": [
        { "term":  { "in_stock": true } },
        { "range": { "price": { "lte": 50 } } }
      ]
    }
  }
}
```

*What just happened:* the `must` clause scores documents for relevance; the `filter` clause just includes or excludes them with no effect on `_score`. Rule of thumb, full-text matching where ranking matters goes in `must`; exact, yes/no constraints go in `filter`. This split is most of practical query writing.

> Keep your relevance-affecting clauses (`must`) separate from your hard constraints (`filter`). Mixing a `term` filter into the scoring path wastes scoring work and can subtly distort ranking. Filters are the lazy win: faster and cached.

```quiz
[
  {
    "q": "You have a field of category slugs like 'kitchen-tools' that you want to filter on exactly. Which type should it be?",
    "choices": [
      "text, so the analyzer breaks it into searchable terms",
      "keyword, so it is stored whole and unanalyzed for exact filtering",
      "float, so it sorts correctly",
      "boolean, since a category is either matched or not"
    ],
    "answer": 1,
    "explain": "keyword stores the value exactly and is not analyzed, which is what exact filtering, sorting, and grouping need. text would tokenize it."
  },
  {
    "q": "Which is NOT one of the three forces in BM25 scoring?",
    "choices": [
      "Term frequency, with diminishing returns",
      "Inverse document frequency, favoring rare terms",
      "Field length, favoring matches in shorter fields",
      "Document creation date, favoring newer documents"
    ],
    "answer": 3,
    "explain": "BM25 uses term frequency, inverse document frequency, and field length. Recency is not part of it; you would add that explicitly if you wanted it."
  },
  {
    "q": "You want 'in stock only' to narrow results without changing the relevance ranking. Where does it go?",
    "choices": [
      "In the must clause, so it contributes to the score",
      "In the filter clause, which includes/excludes with no scoring",
      "In the analyzer, as a stopword",
      "In the mapping, as a text field"
    ],
    "answer": 1,
    "explain": "filter clauses include or exclude without affecting _score, and they are cacheable, so hard yes/no constraints belong there."
  }
]
```


---

# Phase 3: Near-real-time, consistency, and when to add search at all

The model is clear and the queries work. Now meet the parts that bite in production, the behaviors that feel like bugs until you understand they are deliberate trade-offs. Most search outages and "why is this data wrong" tickets come from expecting a search engine to behave like the database it sits next to. It does not, on purpose.

## Near-real-time: your write is not instantly searchable

You index a document, immediately search for it, and it is not there. Nothing is broken. Search engines are **near-real-time**, not real-time. There is a small delay, often around a second, between writing a document and it appearing in results.

```text
t=0.0s  PUT /products/_doc/99   -> 201 Created   (document is saved)
t=0.2s  search for it           -> 0 hits        (not visible yet)
t=1.1s  search for it           -> 1 hit         (now visible)
```

*What just happened:* the engine batches new documents in memory and periodically does a **refresh** that makes them searchable. Refreshing on every single write would destroy throughput, building the inverted-index segments is expensive, so the engine trades a sliver of latency for the ability to ingest fast. The fix is almost never "force a refresh after every write" (that throws away the whole performance win); it is to design your app to tolerate the delay, or to refresh explicitly only in a test where you need determinism.

## Eventual consistency: the source of truth is not here

This is the rule that saves you from data-loss incidents: **a search engine is not your system of record.** It is built for fast, distributed search, not for transactions. It has no joins, no foreign keys, and no all-or-nothing commits across documents. If your indexing job dies halfway, you can end up with a search index that disagrees with your database.

The pattern that survives contact with production is one-directional:

```text
[ your database ]  <- source of truth, transactions live here
        |
        |  on change: enqueue a job
        v
[ indexing worker ] -> writes/updates the document in the search engine
        |
        v
[ search engine ]  <- a fast, rebuildable copy, NOT authoritative
```

*What just happened:* the database stays authoritative; the search engine is a derived, eventually-consistent copy you can rebuild from scratch at any time. That last property is the safety net, if the index gets corrupted, falls behind, or you change a mapping (which often requires reindexing), you re-stream from the database and you are whole again. Never let data exist *only* in the search engine. Treat it as a cache that happens to be searchable.

> The single most useful sentence about operating search: if you deleted the entire cluster right now, could you rebuild every document from your database? If yes, you are holding it correctly. If no, you have put your only copy in the wrong place.

## "Why is the result order weird in tests?"

A subtler gotcha: because the index is distributed across pieces called **shards**, and BM25's statistics (like how rare a term is) are computed *per shard*, the exact `_score` for a document can vary slightly depending on which shard it landed on. With lots of data this evens out. With a tiny test dataset it can make rankings look unstable or surprising.

```text
small index, 2 shards:
  doc on shard A and doc on shard B can get different scores
  for the same terms, because each shard counts term rarity locally
```

*What just happened:* this is not a bug and not something to fix in production, it is just why a five-document test fixture can rank in an order that feels arbitrary. Knowing it exists saves an afternoon of chasing a ghost.

## The real question: did you need search at all?

Before you stand up a cluster, run the ladder. A search engine is real operational weight, another system to run, monitor, secure, keep in sync, and pay for. Often you do not need it:

- **A few exact filters on indexed columns?** Your database already does this well. Add the right index and move on. If a database query is slow, the fix is usually in the query plan, not a new system, see [/guides/why-is-my-query-slow](/guides/why-is-my-query-slow).
- **Modest data and the occasional `LIKE`?** It may be fine. "Slow on a million rows" is a real reason; "might be slow someday" is not yet one.
- **Postgres already in the stack and search needs are light-to-medium?** Its built-in full-text search (`tsvector`/`tsquery`) gives you a real inverted index without a second system. Reach for a dedicated engine when you outgrow it, not before.

You genuinely want Elasticsearch or OpenSearch when search *is* the feature: relevance ranking that must be good, typo tolerance, faceted navigation, autocomplete, searching huge text corpora, or aggregations over large volumes. Then the operational cost buys something your database cannot do.

```text
decision sketch:
  exact filters / joins / transactions      -> database
  light full-text, already on Postgres      -> Postgres FTS
  relevance, typos, facets, scale, big text -> Elasticsearch / OpenSearch
```

*What just happened:* the question was never "search engine or database." It is "which tool for which job," and the most senior move is the smallest one that works. Add search when the job is search; otherwise let the database you already run do what it is good at.

## In the wild

Mature systems almost always run both: the database as the transactional source of truth, the search engine fed asynchronously beside it for the searching, ranking, and faceting the database cannot do well. The reindex-from-scratch capability is treated as a first-class operation, not an emergency, because mapping changes and the occasional desync are normal life, not disasters. Build it that way from day one and search becomes boring, which is the goal.

```quiz
[
  {
    "q": "You index a document and immediately search for it, but get zero hits. What is happening?",
    "choices": [
      "The write failed silently and was never saved",
      "Search engines are near-real-time; a refresh delay means it is not searchable yet",
      "The document needs a manual commit like a SQL transaction",
      "The inverted index is corrupted"
    ],
    "answer": 1,
    "explain": "Writes are batched and made searchable on periodic refresh, so there is a short delay. The document is saved; it is just not visible to search yet."
  },
  {
    "q": "Why should the search engine never be your only copy of the data?",
    "choices": [
      "It is too slow to read from",
      "It cannot store JSON",
      "It is eventually consistent and not transactional; treat it as a rebuildable copy of an authoritative database",
      "It deletes documents after 30 days by default"
    ],
    "answer": 2,
    "explain": "A search engine is a derived, eventually-consistent copy with no transactions or joins. Keep the database authoritative so you can always rebuild the index."
  },
  {
    "q": "Your app already runs Postgres and needs light full-text search. What is the lazy correct move?",
    "choices": [
      "Stand up an Elasticsearch cluster immediately for future-proofing",
      "Use Postgres's built-in full-text search before adding a separate engine",
      "Switch the whole app to a search engine as the primary store",
      "Run LIKE '%term%' and add more CPU"
    ],
    "answer": 1,
    "explain": "Postgres full-text search gives a real inverted index with no second system. Add a dedicated engine only when you outgrow it, not preemptively."
  }
]
```
