# Grafana Loki

> Logs that act like Prometheus: Loki indexes only labels (not full text), making centralized logging cheap, and ties them to your metrics in Grafana.


---

# Grafana Loki

You want centralized logs, but the full-text logging stack you priced out costs more than the app it watches. Indexing every word of every log line is expensive, and most of those words you never search. Loki takes the opposite bet: index only a handful of labels, store the raw log lines compressed and cheap, and lean on brute-force scanning for the rest. This guide takes you from "logging is too expensive to centralize" to running queries in the same Grafana pane as your metrics.

## How to read this

Read the phases in order. Phase 1 builds the mental model: why Loki indexes labels instead of content, and how that one decision shapes everything else. Phase 2 is the everyday core - shipping logs with an agent and querying them with LogQL. Phase 3 is production reality: the cardinality trap, the Elasticsearch tradeoff, and what breaks at scale. Each phase ends with a short quiz so you can check yourself before moving on.

## The phases

1. [What Loki actually is](01-what-loki-actually-is.md) - the mental model: index the labels, not the log content.
2. [Shipping logs and querying with LogQL](02-shipping-and-querying-logql.md) - agents, label streams, and the query language.
3. [Cardinality, cost, and the Elasticsearch tradeoff](03-cardinality-cost-tradeoffs.md) - where Loki shines, where it bites, and how to size it.


---

# What Loki actually is

Here is the reality Loki was built for. You have logs scattered across dozens of containers that come and go. You want them in one place you can search. So you look at the standard full-text logging stack, and you flinch at the bill - because to make every word searchable, that kind of system builds an index over the entire content of every line. Indexes are not free; a full-text index can rival or exceed the size of the data itself. For something as high-volume and low-value-per-line as logs, you end up paying to index a haystack you mostly never read.

Loki's founders, the same people behind Grafana and the Prometheus ecosystem, asked a sharper question: *what do you actually search by?* In practice, you almost always start narrow - "show me the logs from this service, in this namespace, at error level" - and then you grep within that slice. Loki is built around exactly that habit.

## It indexes labels, not content

This is the one idea that explains everything else about Loki. Loki does **not** index the text of your log lines. It indexes a small set of **labels** attached to each stream of logs - things like `app`, `env`, `level`, `namespace`. The log content itself is stored compressed in cheap object storage and left un-indexed.

Picture a single log line and how Loki splits it:

```text
{app="checkout", env="prod", level="error"}   user 90431 payment declined: card expired
└──────────────── indexed labels ───────────┘  └──────── NOT indexed, just stored ───────┘
```

*What just happened:* the part in braces is the only part Loki builds an index over. The message after it - the actual content - is compressed and parked in object storage. Loki never builds a word-by-word index of "payment declined" or "card expired."

If that sounds like it would make text search impossible, it doesn't - it makes it *deferred*. When you search for a string, Loki first uses the label index to find the matching streams, then **brute-force scans** the raw content of only those streams. The label index does the cheap narrowing; a fast linear scan does the rest. You trade "index everything up front" for "store cheaply, scan a small slice on demand."

## A stream is a unique set of labels

In Loki, a **stream** is the unit of everything. A stream is one unique combination of label key-value pairs, and the log lines for that combination, in time order.

```text
Stream A:  {app="checkout", env="prod", level="info"}   → lines, lines, lines...
Stream B:  {app="checkout", env="prod", level="error"}  → lines, lines, lines...
Stream C:  {app="api",      env="prod", level="info"}    → lines, lines, lines...
```

*What just happened:* each distinct label set is its own stream with its own ordered log lines. Change one label value - `info` to `error` - and you have a different stream. This is identical to how Prometheus treats a time series, which is the whole point: if you know Prometheus, you already know Loki's data model.

That deliberate symmetry with Prometheus is why Loki fits the rest of the stack so naturally. The label model, the query feel, the agent-based collection - they mirror metrics on purpose. If the broader picture of logs, metrics, and traces is still fuzzy, the [observability-logs-metrics-traces](/guides/observability-logs-metrics-traces) guide maps how the three pillars relate.

## Why this makes logging cheap

The expensive part of a logging system is the index, and the index is what scales with your data volume in the nastiest way. By indexing only labels, Loki keeps its index tiny - it grows with the *number of unique label combinations*, not with the *number of bytes of log text*. The bulk of your data, the raw lines, lands in object storage like S3 or GCS, which is about the cheapest durable storage you can buy.

```text
Full-text approach:  big index over all content  +  content        → index dominates cost
Loki approach:       tiny index over labels       +  compressed content in object storage
```

*What just happened:* Loki moves the cost from an ever-growing index into cheap object storage. That's the lever that makes "centralize all our logs" affordable instead of a budget fight.

> The flip side, and the thing that bites newcomers: because the index is keyed on labels, the *number of unique label values* is what you must protect. A label like `level` has a few values and is perfect. A label like `user_id` or `request_id` has millions of values, would create millions of streams, and would blow up the very index Loki works hard to keep small. This is **cardinality**, and Phase 3 is largely about respecting it.

## It lives inside Grafana, next to your metrics

Loki's other reason to exist is correlation. Because Loki shares Prometheus's label model and plugs into Grafana as a first-class data source, you can put a metrics graph and the matching logs on the **same dashboard**, scoped by the **same labels**.

```text
Grafana dashboard
  ┌─ panel: error rate for {app="checkout"}  (Prometheus)  ▲ spike at 14:22
  └─ panel: logs for {app="checkout", level="error"} (Loki) ── the lines behind the spike
```

*What just happened:* the metric tells you *something* spiked; the Loki panel right below it, filtered by the same `app` label, shows you the exact lines from that spike. You pivot from "the graph went red" to "here is why" without leaving the page or re-typing a query into a different tool.

**In the wild:** teams already running Prometheus and Grafana reach for Loki precisely because it's the path of least resistance - same labels, same UI, same mental model, and a fraction of the storage cost of a full-text stack. If you haven't met the metrics side yet, [prometheus-and-grafana](/guides/prometheus-and-grafana) is the natural companion to this guide.

```quiz
[
  {
    "q": "What does Loki build its index over?",
    "choices": [
      "Every word in every log line",
      "Only the labels attached to each log stream",
      "The timestamps only",
      "Nothing; it never indexes anything"
    ],
    "answer": 1,
    "explain": "Loki indexes only the labels (like app, env, level). The log content itself is stored compressed and un-indexed, which is what keeps the index tiny and storage cheap."
  },
  {
    "q": "How does Loki find a text string inside log content if it doesn't index that content?",
    "choices": [
      "It can't - text search is impossible in Loki",
      "It uses the label index to narrow to matching streams, then brute-force scans only those",
      "It rebuilds a full-text index on every query",
      "It searches a separate Elasticsearch cluster"
    ],
    "answer": 1,
    "explain": "Labels do the cheap narrowing; Loki then linearly scans the raw content of only the matching streams. Index up front is traded for scan on demand."
  },
  {
    "q": "What is a 'stream' in Loki?",
    "choices": [
      "A single log line",
      "One unique combination of label key-value pairs and its ordered log lines",
      "A network connection to the Loki server",
      "A Grafana dashboard panel"
    ],
    "answer": 1,
    "explain": "A stream is one unique label set plus its time-ordered lines - the same idea as a Prometheus time series, which is why the data models match."
  }
]
```


---

# Shipping logs and querying with LogQL

You have the mental model: index labels, store content cheap. Now you need to actually get logs into Loki and ask questions of them. There are two halves to the day-to-day: an **agent** that tails your logs and ships them with the right labels, and **LogQL**, the query language you use to read them back. Both will feel familiar if you've touched Prometheus, because both were built to rhyme with it.

## An agent tails the files and attaches labels

Loki doesn't pull logs; something pushes them to it. That something is a collection agent running next to your workloads. The classic one is **Promtail**; the current recommended agent is **Grafana Alloy**, which does the same job and more. Either way the agent's responsibilities are the same: find the log sources, attach labels, and push the lines to Loki.

Here's a stripped-down Promtail config for tailing a file:

```yaml
clients:
  - url: http://loki:3100/loki/api/v1/push   # where to send logs

scrape_configs:
  - job_name: checkout
    static_configs:
      - targets: [localhost]
        labels:
          app: checkout        # ← becomes an indexed label
          env: prod            # ← indexed label
          __path__: /var/log/checkout/*.log   # which files to tail
```

*What just happened:* the agent tails every file matching `__path__`, and stamps each line it ships with `app="checkout"` and `env="prod"`. Those labels are exactly what Loki will index, so they're the dimensions you'll be able to query and correlate on later. The `__path__` is a directive telling the agent what to read; it does not become a label.

The single most important discipline here lives in this config block: **choose labels that have few possible values.** `app`, `env`, `level`, `namespace` - good. Anything per-request or per-user - bad, for reasons Phase 3 makes painful. Promtail can extract fields from log content into labels with pipeline stages, and that power is exactly where people accidentally create cardinality disasters. When in doubt, attach fewer labels.

> In Kubernetes, you don't hand-write paths. The agent runs as a DaemonSet, auto-discovers pods, and turns Kubernetes metadata - namespace, pod, container, app - into labels for you. The model is the same; the label values come from the cluster instead of a static file.

## LogQL: pick the streams, then filter the lines

LogQL has a deliberate two-step shape that mirrors how Loki works internally. First you select streams by their labels - this is the cheap, indexed part. Then you filter the content of those streams - this is the scan.

```logql
{app="checkout", env="prod"} |= "payment declined"
└──── stream selector ──────┘ └── line filter ──┘
```

*What just happened:* the part in braces is the **stream selector** - Loki uses the label index to find only the checkout/prod streams, fast. Then `|=` is a **line filter** meaning "keep lines containing this string," applied by scanning just those streams. This is the brute-force scan from Phase 1, made visible in the query.

The stream selector is mandatory - every LogQL query must start with one. You cannot ask Loki "find this string everywhere," because there is no global content index to answer that. You always narrow by label first. The line filter operators are worth memorizing:

```text
|=  "text"      line contains the string
!=  "text"      line does NOT contain the string
|~  "regex"     line matches the regular expression
!~  "regex"     line does NOT match the regex
```

*What just happened:* these four operators are most of what you'll ever use for reading logs. Chain them - `{app="checkout"} |= "error" != "timeout"` - and each one further narrows the scan, left to right.

## Parse fields out of the line at query time

Because Loki stores raw content, you often want to pull structured fields out of a line *when you query*, not when you ingest. LogQL parsers do this. If your app logs JSON, the `json` parser turns fields into temporary labels you can filter on for that query only:

```logql
{app="checkout"} | json | status_code >= 500
```

*What just happened:* `| json` parses each line's JSON body into fields, then `status_code >= 500` filters on a parsed field. Crucially, `status_code` is **not** an indexed label - it's extracted at query time, so it costs you nothing in cardinality. This is the Loki way to get rich filtering without paying the index price: keep indexed labels tiny, parse the detail on demand.

## Turn logs into metrics with a range query

LogQL has a second mode. Wrap a log query in a **range aggregation** and Loki computes a number over time - letting a stream of log lines behave like a Prometheus metric.

```logql
sum by (status_code) (
  rate({app="checkout"} | json | __error__="" [5m])
)
```

*What just happened:* `rate(...[5m])` counts matching lines per second over a 5-minute window, and `sum by (status_code)` groups the result. You've turned raw logs into a graphable time series - error rate straight from log volume, no separate metric needed. This is why Loki panels sit so comfortably next to Prometheus panels in Grafana: the query language and the output shape match. (The [prometheus-and-grafana](/guides/prometheus-and-grafana) guide covers the `rate` and `sum by` machinery in depth.)

**For builders:** start every dashboard query with the narrowest stream selector that's correct, and add a small time range. The selector and the range together decide how much data Loki has to scan - and an unscoped query over a wide window is the number-one way to make Loki feel slow when it shouldn't.

```quiz
[
  {
    "q": "What must every LogQL query begin with?",
    "choices": [
      "A regular expression",
      "A stream selector (label matchers in braces)",
      "A time range",
      "The word SELECT"
    ],
    "answer": 1,
    "explain": "Every LogQL query starts with a stream selector. There's no global content index, so you always narrow by labels first, then filter the content of those streams."
  },
  {
    "q": "In `{app=\"checkout\"} |= \"payment declined\"`, what does the `|=` part do?",
    "choices": [
      "Selects which streams to read by label",
      "Keeps only lines that contain the string 'payment declined', by scanning the selected streams",
      "Assigns a new label to each line",
      "Deletes matching lines from Loki"
    ],
    "answer": 1,
    "explain": "`|=` is a line filter: after the label index narrows to the checkout streams, Loki scans them and keeps lines containing the string. That scan is the brute-force step."
  },
  {
    "q": "Why is filtering on a field via `| json | status_code >= 500` safe for cardinality?",
    "choices": [
      "Because JSON fields are automatically indexed",
      "Because the field is parsed at query time and is never an indexed label",
      "Because Loki ignores numeric fields",
      "It isn't safe; it creates a new stream per status code"
    ],
    "answer": 1,
    "explain": "Parsers extract fields at query time only. The parsed field is not an indexed label, so it adds no streams and no cardinality cost - that's how you get rich filtering cheaply."
  }
]
```


---

# Cardinality, cost, and the Elasticsearch tradeoff

Loki's cheapness is real, but it isn't free of rules. The same design that makes it cheap - index the labels, scan the content - has two sharp edges. The first is **cardinality**: get your labels wrong and the tiny index you were promised explodes. The second is the real tradeoff against full-text engines like Elasticsearch: Loki is cheaper and label-scoped, but it is not a free-text search engine. Knowing both edges is the difference between Loki that hums and Loki that pages you.

## The cardinality trap

Recall from Phase 1 that Loki's index grows with the number of unique label combinations, not with log volume. Every distinct combination of label values is a separate **stream**, and each stream carries overhead. The math is multiplicative: total streams is roughly the product of how many values each label can take.

```text
labels: app(20 values) × env(3) × level(5)        =     300 streams   ← fine
add a high-cardinality label:
labels: app(20) × env(3) × level(5) × user_id(1M)  = 300,000,000 streams  ← disaster
```

*What just happened:* adding one label with a million possible values multiplied the stream count into the hundreds of millions. Each stream needs index entries and its own write path; this floods the index Loki works to keep small, balloons memory, and grinds ingestion and queries to a crawl. The label that felt convenient quietly broke the system.

The rule is blunt and worth tattooing on your config: **never put unbounded or high-cardinality values in labels.** User IDs, request IDs, trace IDs, email addresses, raw URLs with IDs in them, timestamps, full IP addresses - none of these belong in labels. They belong **in the log line**, where Loki stores them cheaply and you filter them at query time with a line filter or a parser.

```text
WRONG:  {app="checkout", user_id="90431", request_id="a1b2c3"}   → a new stream per request
RIGHT:  {app="checkout", level="info"}  user_id=90431 request_id=a1b2c3 payment ok
                                        └──── these live in the content, not the index ────┘
```

*What just happened:* the right version keeps the high-cardinality detail in the message body. You can still find a specific user - `{app="checkout"} |= "user_id=90431"` - but you do it by scanning content, not by creating a stream for every user. Same answer, none of the index blowup.

> A useful gut check before you add a label: "how many distinct values can this take, ever?" If the answer is bounded and small - tens, maybe low hundreds - it's a candidate. If it's "one per user" or "one per request" or "I'm not sure," it goes in the line, not the label. When unsure, prefer fewer labels; you can always grep the content.

## The Elasticsearch tradeoff, stated plainly

This is the comparison everyone asks about, so here it is without spin. A full-text engine like the one in the [elk-elasticsearch-stack](/guides/elk-elasticsearch-stack) indexes the *content* of every log line. That makes arbitrary free-text search across everything fast - and makes the index large and the cost high. Loki indexes only labels, so it's far cheaper to run, but a content search must scan, and that scan is bounded by how well your stream selector narrowed things first.

```text
Elasticsearch:  index all content   → fast free-text search anywhere   → expensive, heavier to run
Loki:           index labels only    → cheap storage, label-scoped       → content search = scan within the slice
```

*What just happened:* the two tools optimize for different questions. Loki bets that you almost always know the labels to scope by before you search the text - and within a tight label scope, scanning is fast. Elasticsearch bets you need to search raw text broadly and will pay for the privilege.

So the clear decision rule: if your searches naturally start with "which service / namespace / level" and then grep - Loki is a great fit and will save you a lot of money. If you genuinely need ad-hoc full-text search across all logs with no obvious label to scope by, or rich text analytics and relevance ranking, that's Elasticsearch's home turf and Loki will fight you. Many teams run both: Loki for the high-volume operational firehose, a full-text engine for the smaller slice that needs deep text search.

## Retention and storage are a separate dial

Because raw content sits in object storage, retention is mostly a question of how long you keep objects, and it's configured independently of the index. Loki lets you set retention globally or **per stream** via label-matched rules - so you can keep noisy `level="debug"` logs for days and important `level="error"` logs for far longer.

```yaml
limits_config:
  retention_period: 168h          # global default: 7 days

# keep errors longer than the global default
retention_stream:
  - selector: '{level="error"}'
    period: 720h                  # 30 days for errors
```

*What just happened:* most logs age out after 7 days, but anything tagged `level="error"` is kept for 30. You spend your storage budget where it matters instead of paying to keep every debug line for a month. This per-stream control is another payoff of the label model.

## Failure modes that page you

A few real-world ways Loki goes wrong, all traceable to the design you now understand:

```text
Symptom                          Usual cause
queries crawl / OOM              high-cardinality labels → millions of streams
"too many outstanding requests"  query scans too wide a time range, no tight selector
logs missing                     agent stopped, wrong __path__, or label dropped at ingest
ingestion rejected               per-stream or per-tenant rate limits hit
```

*What just happened:* the top two - and most Loki pain in general - trace straight back to cardinality and unscoped queries. Get labels right and keep query scopes tight, and the rest is ordinary operations. Loki's failure modes are mostly self-inflicted, which is good news: they're the ones you control.

**In the wild:** the teams happiest with Loki are the ones who treat label design as the main decision, not an afterthought. They keep a deliberately short, boring list of labels, push everything else into the line, and lean on LogQL parsers at query time. That discipline is the whole game - it's what keeps Loki cheap and fast at the scale where a full-text stack would have been a budget conversation.

```quiz
[
  {
    "q": "Why is putting `user_id` in a Loki label a serious mistake?",
    "choices": [
      "User IDs are sensitive and can't be stored",
      "It creates a separate stream per user, exploding the index Loki tries to keep small",
      "Loki rejects numeric label values",
      "It makes the log lines larger"
    ],
    "answer": 1,
    "explain": "Each unique label combination is a stream. A high-cardinality label like user_id multiplies stream count into the millions, flooding the index and crushing performance."
  },
  {
    "q": "How should you handle high-cardinality data like a request ID in Loki?",
    "choices": [
      "Make it a label so you can filter on it fast",
      "Drop it entirely; Loki can't handle it",
      "Keep it in the log line content and filter with a line filter or parser at query time",
      "Store it in a separate Elasticsearch index automatically"
    ],
    "answer": 2,
    "explain": "High-cardinality values belong in the content, where storage is cheap. You still find them by scanning within a tight label scope - no stream explosion."
  },
  {
    "q": "Compared to Elasticsearch, what is Loki's core tradeoff?",
    "choices": [
      "Loki is faster at every kind of search",
      "Loki indexes only labels - cheaper to run, but content search means scanning within a label-scoped slice rather than free-text-anywhere",
      "Loki stores more data per dollar but can't query labels",
      "There is no tradeoff; Loki is strictly better"
    ],
    "answer": 1,
    "explain": "Elasticsearch indexes all content for fast broad free-text search at higher cost. Loki indexes labels only - cheaper, but text search is a scan bounded by how well labels narrowed it first."
  }
]
```
