# The ELK Stack

> Centralized logging with Elasticsearch, Logstash, and Kibana: ship logs from everywhere, index them, and search and visualize across your whole fleet.


---

# The ELK Stack

It's 2am, the alert fired, and the incident spans four services on a dozen machines. The old move is to SSH into a box, `tail` a file, guess which box, SSH into the next one, and `grep` until your eyes blur. That doesn't scale past a handful of servers, and it falls apart the moment a container dies and takes its logs with it. ELK fixes the shape of the problem: every log lands in one searchable place, and you ask questions across the whole fleet from a browser.

## How to read this

Three phases, in order. Phase 1 builds the mental model - the four pieces, what each one does, and why centralizing logs beats logging into boxes. Phase 2 is the everyday core: shipping logs with Beats, structuring them, and searching in Kibana. Phase 3 is production reality - the cost of indexing everything, index lifecycle and retention, and the failure modes that page you. Read 1 even if you're impatient; the model makes the rest obvious.

## The phases

1. [What ELK actually is](01-what-elk-actually-is.md) - the four pieces and why centralized logs win.
2. [Shipping, structuring, and searching](02-shipping-structuring-searching.md) - Beats, parsing, index patterns, and Kibana queries.
3. [Cost, retention, and production reality](03-cost-retention-production.md) - index lifecycle, the price of indexing, and what breaks.


---

# What ELK actually is

You already know how to read one log file. The skill that doesn't transfer is reading a *hundred* log files, scattered across machines that come and go, while a customer is yelling. That's the problem ELK was built for, and once you see the shape of the problem, the four pieces fall into place on their own.

## The pain ELK removes

Here's the workflow centralized logging kills:

```console
$ ssh web-03
$ tail -f /var/log/app/server.log | grep "order_id=88213"
# nothing here. wrong box.
$ exit
$ ssh web-07
$ tail -f /var/log/app/server.log | grep "order_id=88213"
# found it - but the request also touched the payment service. which box was that?
```

*What just happened:* you spent five minutes guessing which machine held the line you needed, and the trail went cold the moment the request crossed a service boundary. Multiply that by every box and every incident. SSH-and-grep is fine for one server; it does not survive a fleet.

The fix is to stop going to the logs and make the logs come to you. Every line from every machine streams into one store, gets indexed for fast search, and you query all of it from a single screen. That's the whole idea. ELK is one popular way to build that.

## The letters: E, L, K (and B)

ELK is three tools, plus a fourth that joined later and quietly took over the front door.

```text
  [your apps]          write logs to files / stdout
       │
   ┌───▼────┐
   │ Beats  │  lightweight shippers on each host - read & forward
   └───┬────┘
       │   (optionally through)
   ┌───▼──────┐
   │ Logstash │  parse, transform, enrich, route
   └───┬──────┘
       │
 ┌─────▼─────────┐
 │ Elasticsearch │  index & store - the searchable database
 └─────┬─────────┘
       │
   ┌───▼────┐
   │ Kibana │  search & visualize in the browser
   └────────┘
```

*What just happened:* logs flow left to right. **Beats** ship, **Logstash** parses, **Elasticsearch** indexes and stores, **Kibana** is the window you look through. The name "ELK" predates Beats; you'll also hear "Elastic Stack," which is the same thing with Beats included. Many setups skip Logstash entirely and let Beats send straight to Elasticsearch - more on that in phase 2.

Take each one on its own terms:

- **Elasticsearch** is the heart. It's a distributed search engine that stores your logs as JSON *documents* and builds an *inverted index* over them - the same trick a search engine uses, so "find every log mentioning `timeout` in the last hour" comes back in milliseconds instead of a full-file scan.
- **Logstash** is the pipeline. It takes messy input (a raw log line), pulls structure out of it (timestamp, level, request ID), and hands clean JSON to Elasticsearch. It's powerful and heavy; it runs on the JVM and likes a lot of memory.
- **Kibana** is the face. Search bar, tables, charts, dashboards, saved queries. It's where humans actually live during an incident.
- **Beats** are the couriers. Small, single-purpose agents you install on each host. **Filebeat** tails log files; **Metricbeat** collects metrics; there are others. They're deliberately dumb and cheap so you can run one on every box without thinking about it.

> A useful frame: Elasticsearch is a *search engine that happens to be great at logs*, not a logging tool that happens to search. Everything good and everything expensive about ELK traces back to that.

## Why "search engine" is the key word

A normal log file is a flat stream of text. To find something you read it top to bottom. Elasticsearch instead breaks every document into terms and builds an index from term back to document - exactly like the index at the back of a book.

```text
Document 1: "GET /orders 200 ok in 14ms"
Document 2: "GET /orders 500 timeout"

inverted index (term → which docs contain it):
  "GET"     → [1, 2]
  "orders"  → [1, 2]
  "200"     → [1]
  "timeout" → [2]
```

*What just happened:* asking "which logs say `timeout`?" is now a single lookup, not a scan of every line. This is why Elasticsearch search feels instant - and it's also the seed of the cost story in phase 3, because building and holding that index for *every field of every log* is not free.

## What this is, and what it isn't

ELK gives you centralized, searchable **logs**. That's one of the three pillars of observability - logs, metrics, and traces. ELK can stretch toward metrics (Metricbeat) and there are tracing add-ons, but its center of gravity is log search. If you want the full picture of how logs sit next to metrics and traces, see /guides/observability-logs-metrics-traces. And for the underlying skill of actually reading what you find - not drowning in volume - see /guides/reading-logs-without-drowning.

## For builders

You don't have to run ELK yourself to get the model. Cloud providers and Elastic itself offer managed Elasticsearch, and the mental model is identical: something ships logs, Elasticsearch indexes them, a UI searches them. The pieces you'll spend real time on are the *edges* - getting clean, structured logs in (phase 2) and keeping storage from eating you alive (phase 3). The middle mostly takes care of itself.

```quiz
[
  {
    "q": "What is the core job of Elasticsearch in the stack?",
    "choices": ["Tailing log files on each host", "Indexing and storing logs as searchable JSON documents", "Drawing dashboards in the browser", "Parsing raw log lines into fields"],
    "answer": 1,
    "explain": "Elasticsearch is the distributed search engine that indexes and stores documents; Beats tail, Logstash parses, Kibana visualizes."
  },
  {
    "q": "Why does searching for a term in Elasticsearch feel instant?",
    "choices": ["It scans every file in parallel", "It uses an inverted index mapping terms to documents", "It caches the last query", "It only searches the most recent log"],
    "answer": 1,
    "explain": "The inverted index maps each term back to the documents containing it, so a search is a lookup rather than a full scan."
  },
  {
    "q": "Where do Beats fit in the pipeline?",
    "choices": ["They store and index documents", "They are the browser-based search UI", "They are lightweight shippers on each host that read and forward logs", "They replace Elasticsearch"],
    "answer": 2,
    "explain": "Beats (like Filebeat) are small agents on each host that forward logs, often straight to Elasticsearch or via Logstash."
  }
]
```


---

# Shipping, structuring, and searching

The model from phase 1 only pays off when real logs are flowing. This phase is the daily loop: get logs off the boxes, give them structure so they're worth searching, and ask questions in Kibana. The single biggest lever here isn't any tool - it's structured logging. Get that right and everything downstream gets easier.

## Step one: ship the logs with Filebeat

Filebeat is the workhorse. You install it on a host, point it at some files, tell it where to send them, and it tails forever - surviving restarts by remembering its position.

```yaml
# filebeat.yml
filebeat.inputs:
  - type: filestream
    id: app-logs
    paths:
      - /var/log/app/*.log

output.elasticsearch:
  hosts: ["http://es01:9200"]
  index: "app-logs-%{+yyyy.MM.dd}"
```

*What just happened:* Filebeat now tails every `.log` file under `/var/log/app/` and ships each line to Elasticsearch, into a fresh index per day (`app-logs-2026.06.30`, and so on). The daily rollover isn't cosmetic - it's what makes deleting old data cheap later, which phase 3 leans on hard. You can point `output` at Logstash instead when you need heavier parsing.

Start it and confirm data is arriving:

```console
$ systemctl start filebeat
$ curl -s "http://es01:9200/_cat/indices/app-logs-*?v"
health status index               docs.count  store.size
green  open   app-logs-2026.06.30      12483       8.4mb
```

*What just happened:* the `_cat/indices` endpoint is your "is anything happening?" smoke test. A growing `docs.count` means logs are landing. If this stays empty, the problem is upstream - Filebeat config, network, or permissions on the log files - not Elasticsearch.

## Step two: structure beats parsing

You can ship raw text and parse it later, but the better move is to log structured data at the source. Compare these two log lines for the same event:

```text
# unstructured - a human sentence
2026-06-30 14:22:01 ERROR could not charge card for order 88213, gateway timed out

# structured - JSON
{"ts":"2026-06-30T14:22:01Z","level":"ERROR","msg":"charge failed","order_id":88213,"reason":"gateway_timeout","service":"payments"}
```

*What just happened:* the second line costs you nothing extra to write but means Elasticsearch stores `order_id`, `level`, and `service` as real fields. Now "show me every ERROR in `payments` for order 88213" is a precise query. With the first line you'd be writing fragile regex to claw those values back out. **Structure at the source is the cheapest win in the whole stack.**

If you can't change the app - third-party software, legacy code - that's exactly when Logstash earns its keep. Its `grok` filter pattern-matches raw lines into fields:

```text
# logstash pipeline filter
filter {
  grok {
    match => { "message" => "%{TIMESTAMP_ISO8601:ts} %{LOGLEVEL:level} %{GREEDYDATA:msg}" }
  }
}
```

*What just happened:* grok carved `ts`, `level`, and `msg` out of an otherwise opaque line. It's powerful, but every grok pattern is a small parser you now own and must keep matching as the log format drifts. Prefer structured logging; reach for grok when you have no other choice.

## Step three: search in Kibana

Open Kibana, create a **data view** (older versions call it an *index pattern*) that matches `app-logs-*`, and you get one searchable view across every daily index. The bar at the top speaks **KQL** (Kibana Query Language):

```text
level: "ERROR" and service: "payments"
order_id: 88213
status >= 500 and not url: "/healthcheck"
message: *timeout*
```

*What just happened:* each line is a real filter. The first narrows to payment errors; the second pulls one order's whole story across services; the third finds server errors while ignoring noisy health checks; the last does a wildcard text match. Because these run against the inverted index, they return fast even over millions of documents - and because your logs are *structured*, the field names actually exist to filter on.

> The `*timeout*` wildcard query is handy but slow at scale - leading wildcards can't use the index efficiently. For anything you search often, make it a real field (`reason: gateway_timeout`) instead of fishing in free text.

A data view also unlocks the rest of Kibana: time-series charts of error rate, a table of top failing endpoints, a dashboard you pin to a wall during an incident. The search bar is where you live; dashboards are how you spot trouble before the alert fires.

## The daily loop, end to end

```text
app logs ──▶ Filebeat tails ──▶ Elasticsearch indexes
                                      │
                                      ▼
                       Kibana: search, chart, dashboard
```

*What just happened:* that's the whole everyday cycle. Note Logstash isn't in it - plenty of healthy setups run Beats-straight-to-Elasticsearch and only add Logstash when parsing or routing demands it. Don't stand up Logstash because a diagram told you to; add it when you have a job for it.

## In the wild

The teams that get the most out of ELK share one habit: they agreed on a **log schema** early. Same field names everywhere - `service`, `level`, `request_id`, `ts` - so a query written for one service works for all of them. A consistent `request_id` threaded through every service is what turns a pile of logs into a traceable story across a request's whole journey. That discipline costs a meeting; the lack of it costs you every incident. For the human side of actually making sense of what you find, see /guides/reading-logs-without-drowning.

```quiz
[
  {
    "q": "Why is structured (JSON) logging preferred over parsing raw text later?",
    "choices": ["It makes log files smaller on disk", "Fields like order_id become real, queryable fields with no fragile regex", "It is required by Filebeat", "It removes the need for Elasticsearch"],
    "answer": 1,
    "explain": "Logging structured data at the source gives Elasticsearch real fields to index, avoiding brittle grok/regex parsing downstream."
  },
  {
    "q": "When does adding Logstash to the pipeline actually make sense?",
    "choices": ["Always - Beats cannot send to Elasticsearch", "When you need heavier parsing/transforming logs you can't change at the source", "Only for drawing dashboards", "Never - it has been removed from the stack"],
    "answer": 1,
    "explain": "Beats can ship straight to Elasticsearch; Logstash earns its place when you need to parse, enrich, or route logs you can't restructure at the source."
  },
  {
    "q": "What does a KQL query like `level: \"ERROR\" and service: \"payments\"` do in Kibana?",
    "choices": ["Deletes matching logs", "Filters to documents where both fields match", "Creates a new index", "Restarts Filebeat"],
    "answer": 1,
    "explain": "KQL filters the data view to documents matching the field conditions, fast because it uses the inverted index."
  }
]
```


---

# Cost, retention, and production reality

The first month of ELK is a honeymoon: logs flow, search is instant, dashboards look great. Then the bill arrives - in disk, in memory, in the 3am page where the cluster is red and nothing's indexing. Almost every ELK horror story has the same root cause: indexing everything forever because nobody decided not to. This phase is how you avoid that.

## The thing nobody warns you about: indexing isn't free

Remember the inverted index from phase 1 - the magic that makes search instant. The catch is that Elasticsearch builds and stores that index for *every field of every document*, and the index plus the original document often takes **more** disk than the raw log did.

```text
raw log line:           ~200 bytes
stored in Elasticsearch: ~200 bytes (the source) + index structures
                         → frequently MORE than the original on disk
```

*What just happened:* the speed you love has a storage tax, and you pay it on every log whether or not you ever search that field. A fleet emitting a few hundred gigabytes of logs a day will fill terabytes fast. The lever you control is **what you index** - drop fields you'll never query, and don't ship debug-level noise to a system that charges you to index it.

> The trap is treating ELK like an infinite bucket. It's a search engine with a storage bill. Every field you index is a recurring cost, not a one-time write.

## Retention: decide when logs die, automatically

Logs have a shelf life. Last hour's logs are gold during an incident; last quarter's are dead weight you're paying to keep searchable. The tool for this is **Index Lifecycle Management (ILM)** - a policy that ages indices through phases and eventually deletes them, with no human in the loop.

```text
hot   →  actively written & searched   (fast storage)   days 0–2
warm  →  searched, not written          (cheaper)        days 2–14
cold  →  rarely searched                 (cheapest)       days 14–30
delete → gone                                             day 30
```

*What just happened:* this is why phase 2 shipped to a *daily* index (`app-logs-2026.06.30`). ILM can roll over to a new index and drop whole old ones cheaply - deleting one day's index is one fast operation, where deleting individual old documents would be slow and painful. A retention policy is the single most important production decision you'll make. Pick a number, write the policy, walk away.

Sketch of an ILM policy attached via the API:

```console
$ curl -X PUT "http://es01:9200/_ilm/policy/app-logs-policy" -H 'Content-Type: application/json' -d'
{ "policy": { "phases": {
    "hot":    { "actions": { "rollover": { "max_age": "1d", "max_primary_shard_size": "50gb" } } },
    "delete": { "min_age": "30d", "actions": { "delete": {} } }
}}}'
{"acknowledged":true}
```

*What just happened:* you told Elasticsearch to roll over to a fresh index daily (or sooner if a shard hits ~50GB) and delete anything older than 30 days. From now on, retention runs itself. The `max_primary_shard_size` guard matters because oversized shards are a classic source of slow, unstable clusters.

## The failure modes that page you

A handful of problems cause most ELK outages. Knowing the shape of each saves the panic.

**Disk fills and the cluster goes read-only.** When a node crosses a disk *watermark*, Elasticsearch protects itself by refusing new writes - so indexing silently stops and logs pile up upstream.

```console
$ curl -s "http://es01:9200/_cluster/health?pretty" | grep status
  "status" : "red",
$ curl -s "http://es01:9200/_cat/allocation?v"
shards disk.used disk.avail disk.percent
   412    471gb       12gb           97
```

*What just happened:* `red` plus a near-full disk is the textbook "logs stopped flowing" incident. The fix is to free space (delete old indices - ILM should have, which is why you set it up) or add capacity. The deeper fix is the retention policy that stops you getting here.

**Mapping explosion / field blowup.** Send logs with unpredictable field names - say, a field per user ID - and Elasticsearch tries to index each as a new field, the *mapping* balloons, and the cluster strains. Cap it and keep high-cardinality junk out of indexed fields.

**Too many shards.** Every index is split into shards, and each shard has fixed overhead. Thousands of tiny daily indices each with several shards adds up to a cluster spending all its energy on bookkeeping. Fewer, larger shards beat many tiny ones - another reason ILM rollover by size matters.

```text
green  → all good
yellow → replicas unassigned (often single-node dev) - degraded, not down
red    → some primary data unavailable - STOP, investigate now
```

*What just happened:* cluster color is your at-a-glance health. Yellow is common and survivable on small setups; red means real data is unreachable and is always worth dropping what you're doing for.

## A sane production checklist

- **Set a retention policy on day one.** ILM with a real delete phase. Not "we'll figure it out later" - later is a full disk.
- **Index what you'll search, drop what you won't.** Debug logs and one-off fields are pure cost.
- **Watch disk and shard count**, not vanity dashboards. Those are what actually take the cluster down.
- **Run replicas in production** so a lost node doesn't lose data - and remember replicas roughly double storage.
- **Don't put Logstash in the path until you need it.** It's another JVM service to feed, tune, and keep alive.

## In the wild

The teams that stay happy with ELK treat it as a cost center they actively manage, not a black box. They know their daily ingest in gigabytes, they have a retention number they can defend, and they prune what they index. The teams that get burned are the ones who set it up once, indexed everything, and met it again only when the disk hit 100% during an incident. ELK is genuinely excellent at what it does - searching your whole fleet's logs from one screen beats SSH-and-grep every single time - as long as you respect that the search magic has a running bill. For where logs sit in the bigger observability picture, see /guides/observability-logs-metrics-traces.

```quiz
[
  {
    "q": "What is the most common root cause of ELK storage and stability problems?",
    "choices": ["Using Kibana instead of Grafana", "Indexing everything forever with no retention policy", "Logging in JSON", "Running Filebeat on too few hosts"],
    "answer": 1,
    "explain": "The inverted index has a storage cost on every field; without retention (ILM), disks fill and clusters go read-only."
  },
  {
    "q": "Why does shipping to a daily index make retention cheap?",
    "choices": ["Daily indices compress better", "ILM can delete a whole old index in one fast operation instead of deleting documents one by one", "Kibana only reads daily indices", "It avoids needing Elasticsearch"],
    "answer": 1,
    "explain": "Dropping an entire day's index is a single cheap operation; deleting individual aged documents would be slow and costly."
  },
  {
    "q": "Your cluster health is `red` and disk is at 97%. What does this mean?",
    "choices": ["Everything is healthy", "Replicas are unassigned but data is safe", "Some primary data is unavailable and writes have likely stopped - investigate now", "Kibana needs restarting"],
    "answer": 2,
    "explain": "Red means primary shards are unavailable; near-full disk crosses a watermark and turns the cluster read-only, halting indexing."
  }
]
```
