# Datadog, From Zero

> The all-in-one observability platform: the agent, metrics and dashboards, APM traces, log management, and monitors - plus the bill that surprises teams.


---

# Datadog, From Zero

You have metrics in one tool, logs in another, and traces in a third - and when something breaks at 2am you're flipping between three tabs trying to line up timestamps by hand. Datadog's whole pitch is that those three are one story, told through one agent, sliced by one tag. This guide takes you from "we pay for Datadog but only I know how to use it" to a real mental model: how the agent collects everything, why tagging is the lever that makes the data useful, how to read a trace and wire up a monitor - and the plain part nobody puts on the marketing page, which is how the bill sneaks up on you.

## How to read this

Read the phases in order. Phase 1 builds the mental model: the three telemetry types, the one agent that ships them, and the tagging system that ties them together. Phase 2 is the daily loop: metrics on a dashboard, a distributed trace in APM, logs you can actually query, and a monitor that pages the right person. Phase 3 is production reality - the cost model, where the surprise charges come from (custom metrics cardinality, log volume, per-host pricing), and how to keep the platform useful without setting money on fire. Each phase ends with a short quiz so you can check yourself.

## The phases

1. [What Datadog actually is](01-what-datadog-actually-is.md) - one agent, three signals, and tags as the connective tissue.
2. [The everyday loop](02-the-everyday-loop.md) - dashboards, APM traces, log queries, and monitors that page the right human.
3. [The bill, and how it sneaks up](03-the-bill-and-how-it-sneaks-up.md) - custom-metric cardinality, log volume, host pricing, and keeping the cost in check.


---

# What Datadog actually is

Here's the reality most teams start from: the metrics live in one place, the logs in another, the traces somewhere else, and a "dashboard" is three browser tabs and a lot of squinting at timestamps. When a request is slow, you can see the latency spike on a graph, but the graph can't tell you *which* request or *why* - for that you go hunting in logs, and now you're correlating by eyeball.

Datadog's core idea is simple to say and the whole reason it exists: those three kinds of data are one story. It collects all of them through one agent, stamps them with the same tags, and lets you pivot from a spiking graph straight to the exact trace and the exact log lines behind it. The product surface is enormous, but the spine is small. Learn the spine and the rest is menus.

## The three signals

Almost everything in Datadog is one of three telemetry types. If you've read [Observability: Logs, Metrics, Traces](/guides/observability-logs-metrics-traces) this will be familiar - Datadog is one vendor's take on exactly those three.

- **Metrics** - numbers over time. CPU at 73%, 1,200 requests per second, 14 items in a queue. Cheap to store, great for trends and alerts, but they're aggregates: a metric tells you *that* latency rose, never *which* request.
- **Traces (APM)** - the story of one request as it moves through your services. A trace is made of spans, each span a unit of work (a DB query, an HTTP call), nested to show what called what and how long each took. This is the *why* behind a latency metric.
- **Logs** - the timestamped lines your code already writes. Datadog ingests, parses, and indexes them so you can search across every service at once instead of SSH-ing into boxes.

```text
METRIC   api.request.latency  p95 = 840ms   ← something is slow (the "what")
   │
TRACE    request abc123: web → auth → db    ← which request, where the time went (the "why")
   │       └─ db query span took 790ms
   │
LOG      "slow query: SELECT * FROM orders…" ← the exact line, with context
```

*What just happened:* one incident, read top to bottom - the metric raises the alarm, the trace localizes it to the DB span, the log shows the offending query. The value isn't any single signal; it's being able to walk between them without changing tools.

## The agent: one collector to ship them all

You don't send data to Datadog from a hundred places. You run **the Datadog Agent** - a small process on each host (or as a DaemonSet pod on each Kubernetes node) - and it does the collecting. It scrapes system metrics, receives traces from your instrumented apps, tails log files, runs integration checks against things like Postgres or nginx, and forwards all of it to Datadog over HTTPS.

```yaml
# /etc/datadog-agent/datadog.yaml - the agent's main config
api_key: "<YOUR_API_KEY>"
site: datadoghq.com        # which Datadog region/site to report to

logs_enabled: true         # off by default; logs are a separate paid product
apm_config:
  enabled: true            # turn on the trace receiver (listens on :8126)

tags:                      # host-level tags applied to everything this agent sends
  - env:prod
  - service:checkout
  - team:payments
```

*What just happened:* one config file turns on the three pipelines. Note `logs_enabled` is `false` until you set it - Datadog won't quietly start billing you for log ingestion. The `tags` block is the most important part of the file, and that's the next section.

> **Heads up:** the agent runs as a privileged process and ships data off your network. Treat the `api_key` like a credential - environment variable or a secrets manager, never committed to git.

## Tags: the one idea that makes it all useful

This is the concept that separates people who *have* Datadog from people who *use* it. A **tag** is a `key:value` label attached to your telemetry - `env:prod`, `service:checkout`, `region:us-east-1`, `version:1.4.2`. Tags are applied at the agent (host tags), in your app (service/span tags), and by integrations automatically.

Why they matter: tags are how you *slice* the data. Untagged, a latency metric is a single line - the average across everything, which hides every interesting problem. Tagged, the same metric becomes "p95 latency, grouped by `version`," and suddenly you can see that only `version:1.4.2` is slow, which means the last deploy did it.

```text
# Reading the same metric two ways

avg:api.request.latency                          → one line, tells you almost nothing

avg:api.request.latency{env:prod} by {version}   → one line per version
   version:1.4.1  → 120ms
   version:1.4.2  → 840ms   ← the deploy that broke it, found in one query
```

*What just happened:* the `{filter} by {grouping}` syntax is the heart of every Datadog query - metrics, traces, and logs all use the same tag-based filtering. Good tagging is what lets one query answer "is it this environment? this version? this customer?" Bad tagging makes Datadog an expensive line chart.

The flip side, which phase 3 returns to: every distinct combination of tag values on a metric is a separate thing Datadog stores and bills as a *custom metric*. Tags are the power and the cost in the same breath.

> **In the wild:** the teams who love Datadog enforce a small, consistent tag vocabulary - `env`, `service`, `version`, `team` on everything - usually through a shared agent config or a deploy template. The teams who fight it let every service invent its own tag names, and nothing lines up across dashboards.

## So what is it, really?

Datadog is one agent that collects metrics, traces, and logs, a tagging system that ties them together so you can pivot between them, and a pile of UI (dashboards, monitors, SLOs) built on top of that data. The "all-in-one" promise is real - and so is the bill that comes with sending it everything. Phase 2 is the daily loop; phase 3 is the money.

```quiz
[
  {
    "q": "What is the Datadog Agent's job?",
    "choices": [
      "It's the web dashboard you log into",
      "A process on each host that collects metrics, traces, and logs and forwards them to Datadog",
      "A billing console for tracking spend",
      "A database where your telemetry is stored long-term"
    ],
    "answer": 1,
    "explain": "The agent runs on your hosts (or as a DaemonSet in Kubernetes), gathers the three signals plus integration checks, and ships them to Datadog over HTTPS."
  },
  {
    "q": "Why are tags described as the key to slicing data?",
    "choices": [
      "They compress the data to reduce storage cost",
      "They let you filter and group telemetry, e.g. latency by version, to isolate exactly what's affected",
      "They are required for the agent to start",
      "They encrypt the data in transit"
    ],
    "answer": 1,
    "explain": "Tags turn a flat metric into something you can pivot on - by env, service, version, customer - which is how you go from 'it's slow' to 'only this version is slow.'"
  },
  {
    "q": "Which signal best answers 'WHY was this specific request slow,' as opposed to 'THAT latency rose'?",
    "choices": [
      "Metrics, because they're numbers over time",
      "A trace (APM), because it shows the spans of one request and where the time went",
      "Logs, because they're timestamped",
      "Tags, because they group data"
    ],
    "answer": 1,
    "explain": "A metric is an aggregate that flags the 'what.' A trace breaks one request into spans, showing which call (e.g. a DB query) consumed the time - the 'why.'"
  }
]
```


---

# The everyday loop

Phase 1 was the map. This is the territory: the four things you'll actually do in Datadog on a normal week. You'll graph a metric on a dashboard, follow a slow request through a trace, search logs across every service at once, and set up a monitor so the platform pages you instead of a customer telling you it's down. Each of these leans on the same tag-based query syntax, so once one clicks the others come quickly.

## Dashboards: metrics you can read at a glance

A dashboard is a saved set of widgets - graphs, query values, heatmaps - each backed by a metric query. The query language is the `{filter} by {grouping}` shape from phase 1, wrapped in an aggregation.

```text
# A timeseries widget query, read left to right:

sum:web.requests{env:prod, service:checkout} by {status_code}.as_rate()
 │   │            │                            │              │
 │   │            │                            │              └─ per-second rate, not raw count
 │   │            │                            └─ one line per HTTP status
 │   │            └─ scope: only prod checkout traffic
 │   └─ the metric name
 └─ space aggregation across matching hosts/tags
```

*What just happened:* one line of query language produces a graph of checkout request rate, split by status code, in prod only. Watch the `5xx` line and you have a live error-rate panel. Datadog ships pre-built dashboards for its integrations (Postgres, Kubernetes, nginx), so you rarely start from a blank page - you clone one and adjust the scope.

A useful habit: **template variables**. Define `$env` and `$service` at the top of a dashboard, reference them in every query (`{env:$env}`), and one dropdown re-scopes the whole board. That's how a single dashboard serves staging and prod without duplication.

## APM: reading a trace

When a metric says latency is up, APM tells you where the time went. The **service map** shows your services and the calls between them; the **flame graph** breaks one request into spans.

```text
Trace 7f3a… total: 840ms
─────────────────────────────────────────────────────────
web          ████████████████████████████████████  840ms
 └ auth      ██                                      40ms
 └ db.query  ████████████████████████████████       790ms   ← the culprit
      "SELECT * FROM orders WHERE user_id = ?"
 └ render    █                                       10ms
```

*What just happened:* the flame graph makes the bottleneck obvious - 790 of 840ms is a single database span, and it shows you the query. Without the trace you'd know the request was slow; with it you know to add an index on `orders.user_id`. Traces carry the same tags as metrics, so you can filter to `version:1.4.2` and confirm only the new deploy is slow.

Two terms that trip people up: **sampling** and **trace retention**. Datadog doesn't necessarily keep every trace forever - high-throughput services sample (keep a representative fraction), and retention filters decide which traces are indexed for search. The defaults usually keep error traces and slow traces, which are the ones you want. This matters for cost, which phase 3 covers.

## Logs: search across everything at once

Once `logs_enabled: true` is set, the agent tails your log files and ships them. In the Log Explorer you query with the same facet/tag idea, not grep.

```text
# Log Explorer query
service:checkout status:error @http.status_code:500 env:prod

# meaning: checkout service, error level, HTTP 500, in prod
```

*What just happened:* one query searched every checkout host's logs at once and returned only prod 500s. The `@http.status_code` syntax (with the `@`) queries a *parsed attribute* - a field Datadog extracted from structured (JSON) logs. This is the payoff of logging in JSON: you get queryable fields instead of full-text search across a wall of strings.

The strongest move here is **trace-log correlation**. If your app injects the `trace_id` into its logs, Datadog links them - from a slow trace you jump to that exact request's log lines, and from a log line you jump to its trace. That's the "one story" promise made literal.

> **For builders:** correlation is worth setting up early. Most Datadog tracing libraries can auto-inject `trace_id` and `span_id` into your log context. Do it once and every future incident gets shorter, because "show me the logs for *this* request" becomes a click instead of a timestamp hunt.

## Monitors: getting paged at the right time

A **monitor** watches a query and changes state (OK → Warn → Alert) when it crosses a threshold, then notifies a channel - Slack, PagerDuty, email. This is what turns passive graphs into something that wakes the right person.

```text
# A metric monitor, in plain terms
Query:    avg(last_5m): avg:web.request.latency{env:prod} by {service} > 500
Warn:     > 400 ms
Alert:    > 500 ms
Notify:   @slack-payments-oncall  @pagerduty-payments
Message:  "p95 latency on {{service.name}} is {{value}}ms in prod. Runbook: …"
```

*What just happened:* this monitor evaluates per-service latency every few minutes and pages the payments on-call only when prod latency crosses 500ms. The `{{service.name}}` and `{{value}}` are template variables filled in at alert time, so the message names the actual broken service. A good monitor message includes a runbook link - future-you, half-asleep, will thank present-you.

Two refinements that keep monitors from becoming noise:

- **Notify on the right grain.** A monitor grouped `by {service}` alerts per service, so one bad service doesn't silence alerts for the others. Without grouping, the first thing to break masks everything after it.
- **SLOs.** A Service Level Objective tracks a target over a window - "99.9% of requests under 300ms over 30 days" - and burns down an *error budget*. Alert on the budget burn rate, not every blip, and you page for trends that matter instead of every transient spike. This is the difference between actionable alerts and alert fatigue. [Prometheus and Grafana](/guides/prometheus-and-grafana) frames the same SLO/error-budget idea in the open-source stack if you want a second angle.

## The loop, end to end

Here's a normal incident, using all four: a monitor pages that prod checkout latency crossed 500ms → you open the linked dashboard and see `5xx` and latency both spiking on `version:1.4.2` → you click into APM, filter to that version, and the flame graph shows a 790ms DB span → you jump from the trace to its correlated logs and read the slow query → you ship an index and watch the monitor recover. Four tools, one tab-free path, because the tags line up. Phase 3 is what that convenience costs.

```quiz
[
  {
    "q": "In a flame graph for one request, what does a single very wide span usually tell you?",
    "choices": [
      "That span is where most of the request's time was spent - the bottleneck",
      "That span errored",
      "That span was sampled out",
      "The request was retried that many times"
    ],
    "answer": 0,
    "explain": "Span width is duration. A span that fills most of the trace's total time is where the latency lives - e.g. a slow DB query - which is exactly what you go fix."
  },
  {
    "q": "What does the `@` prefix mean in a log query like `@http.status_code:500`?",
    "choices": [
      "It mentions a user",
      "It queries a parsed attribute (a structured field) rather than full text",
      "It's a comment",
      "It escapes a reserved word"
    ],
    "answer": 1,
    "explain": "@ targets an extracted/parsed attribute, which you get from structured (JSON) logs. It's far more precise than full-text search and is a big reason to log in JSON."
  },
  {
    "q": "Why alert on an SLO's error-budget burn rate instead of on every metric spike?",
    "choices": [
      "Burn-rate alerts are cheaper to evaluate",
      "It pages for sustained problems that threaten your target, cutting noise from transient blips",
      "SLOs disable all other monitors",
      "Error budgets never run out, so you never get paged"
    ],
    "answer": 1,
    "explain": "An SLO tracks a target over a window with an error budget. Alerting on burn rate fires for trends that actually endanger the objective, instead of every momentary spike - far less alert fatigue."
  }
]
```


---

# The bill, and how it sneaks up

Everything so far is the part Datadog wants you to love, and you will. This phase is the conversation that doesn't happen until the invoice does. Datadog is genuinely excellent and genuinely expensive, and the worst version is the surprise: a bill that doubles month over month because of a config change nobody flagged as a spending decision. The pricing isn't a trick - the things that make Datadog powerful (send everything, tag everything) are the same things that make it costly, and the meter runs quietly. Knowing where the money goes is the difference between a tool you control and a tool that controls your budget.

## The shape of the bill

Datadog bills each product separately, mostly on volume. The exact numbers change and depend on your contract, so this guide stays qualitative - the *categories* are what you need to internalize:

- **Infrastructure** - billed **per host per month**. A "host" is a machine reporting metrics. This is the floor, and it scales with your fleet.
- **APM** - billed **per host** (often a higher per-host rate than infra), sometimes plus ingested/indexed span volume.
- **Logs** - billed on two separate axes: **ingestion** (bytes received) *and* **indexing/retention** (how many you make searchable, and for how long). You can ingest a log without indexing it, which matters a lot below.
- **Custom metrics** - billed per **custom metric**, where the unit is not what you'd guess. This is the single most common source of a shock bill, so it gets its own section.

The trap is that each is individually reasonable and they're additive across a growing fleet. Three products times more hosts times more volume compounds fast.

## The custom-metrics cardinality trap

This is the one that gets everyone. A **custom metric** in Datadog's billing is not "a metric name." It's a **unique combination of metric name plus tag values** - each distinct combination is a separate billed metric. The number of those combinations is the metric's **cardinality**.

```text
Metric: orders.processed
Tags:   env (2 values) × region (4) × service (3)
        = 2 × 4 × 3 = 24 custom metrics.  Fine.

Now someone adds a tag: user_id (say 50,000 active users)
        = 2 × 4 × 3 × 50,000 = 1,200,000 custom metrics.
        From ONE metric name. From ONE line of code.
```

*What just happened:* multiplying tag values multiplies billed metrics. Adding a high-cardinality tag - `user_id`, `request_id`, `email`, a raw URL with IDs in it, a container ID - explodes the count, and nothing in your code looks wrong. The graph still works; the invoice quietly grew by a million metrics. **The rule: never put unbounded or per-request identifiers in metric tags.** Those belong on logs and traces (where they're searchable and not multiplied), never on metrics.

> **Heads up:** the dangerous part is that the explosion is invisible at write time. Your code emits one `orders.processed` either way - the cost is decided by how many *distinct tag combinations* show up over the month. A single new tag can be a five-figure decision that looks like a one-line change.

Datadog gives you a **Metrics Summary** page and a **Metrics without Limits** feature to see per-metric cardinality and to drop tags you don't query on. Auditing cardinality is the highest-leverage cost work you can do.

## Log volume: ingest cheap, index deliberately

Logs are the second classic surprise, because **ingestion and indexing are billed separately** and teams forget the second axis. The expensive part is usually *indexing* - making logs searchable and retaining them.

The fix is built into the product: **ingest everything, index selectively.** Datadog lets you ingest all your logs (cheaper, and you keep them via archive to S3) but apply **exclusion filters** in the pipeline so only the logs worth searching get indexed.

```text
# Log indexing pipeline (conceptual)
INGEST  all logs ─────────────────────────────►  archive to your S3 (cheap, rehydrate later)
            │
            ├─ index:  status:error, status:warn        ← keep searchable
            └─ exclude: status:info on /healthz 200      ← ingest but DON'T index (90% of volume)
```

*What just happened:* the health-check and routine 200-OK lines - often the overwhelming majority of log volume - get ingested and archived but not indexed, so you stop paying to make noise searchable. If you ever need them, you **rehydrate** from the archive on demand. The art is excluding the chatter while keeping every error and warning indexed.

## Hosts and APM: watch the fleet definition

Per-host pricing sounds simple until autoscaling and short-lived containers get involved. Datadog typically bills the **high-water mark** or a percentile of concurrent hosts over the billing period, not a flat snapshot - so a fleet that bursts to 200 hosts under load costs more than its steady-state 50 suggests. Ephemeral containers can each register as billable depending on how they're counted, so a chatty autoscaler or a CI pipeline that spins up agent-enabled hosts can quietly inflate the count.

The practical guardrails: don't run the full agent (especially APM) on machines that don't need observability, scope APM to the services that matter rather than everything, and check whether short-lived CI/build hosts are reporting as billable infra.

## Keeping it in check

You don't fix Datadog cost once; you keep it in check with a few habits:

- **Set a tag budget.** Agree on a small set of allowed metric tags (`env`, `service`, `version`, `region`, `team`) and forbid per-user/per-request IDs on metrics. Most cardinality blowups are one rogue tag.
- **Audit cardinality monthly.** The Metrics Summary page sorts by volume - the top few metrics are usually most of the bill.
- **Index logs deliberately.** Exclusion filters for health checks and routine successes; archive the rest to your own storage.
- **Treat config changes as spend changes.** "Add a tag," "turn on logs for this service," "enable APM everywhere" are budget decisions. Route them through whoever owns the bill.
- **Use Datadog's own usage dashboards.** It bills you using data it'll happily graph - build a usage/cost dashboard and put a monitor on *that*. Page yourself when spend spikes, the same way you'd page on latency.

## The real mental model, completed

Datadog earns its reputation: one agent, three correlated signals, dashboards and monitors that genuinely shorten incidents. The catch is that its pricing rewards exactly the behavior the product encourages - send everything, tag everything, keep everything - and the meter runs in categories (per-host, per-custom-metric-cardinality, per-log-byte) that don't show up in your code review. Use it fully, but treat cardinality and log indexing as first-class engineering concerns, not afterthoughts. The teams who are happy with Datadog aren't the ones who use it less - they're the ones who decided, on purpose, what to send.

> **In the wild:** the recurring horror story is identical every time - a tag with `user_id` or a raw request path lands in a metric, the next invoice arrives with an extra zero, and someone spends a frantic afternoon in the Metrics Summary page hunting the one tag. Set the tag budget before that happens, not after.

```quiz
[
  {
    "q": "In Datadog billing, what counts as one 'custom metric'?",
    "choices": [
      "One metric name, regardless of tags",
      "One unique combination of metric name plus tag values (its cardinality)",
      "One host reporting the metric",
      "One dashboard widget using the metric"
    ],
    "answer": 1,
    "explain": "Each distinct combination of name + tag values is a separate billed custom metric. That's why adding a high-cardinality tag like user_id can explode the count from one line of code."
  },
  {
    "q": "Why is adding `user_id` as a tag on a metric dangerous, while putting it on a log or trace is fine?",
    "choices": [
      "user_id is personally identifiable and banned on metrics",
      "Metrics bill per tag-value combination, so a high-cardinality tag multiplies billed metrics; logs/traces don't multiply that way",
      "Logs can't store user_id",
      "It makes the metric query slower but costs the same"
    ],
    "answer": 1,
    "explain": "Tag-value combinations on metrics are the billing unit, so an unbounded id multiplies cost. On logs and traces that same id is a searchable field, not a multiplier - which is where per-request identifiers belong."
  },
  {
    "q": "What's the standard way to cut log cost without losing the ability to investigate later?",
    "choices": [
      "Turn off logging entirely",
      "Index every log to be safe",
      "Ingest everything but apply exclusion filters so only useful logs (errors, warnings) are indexed; archive the rest",
      "Lower the agent's CPU limit"
    ],
    "answer": 2,
    "explain": "Ingestion and indexing are billed separately, and indexing is the costly part. Ingest all, index selectively with exclusion filters, archive the rest to your own storage, and rehydrate on demand."
  }
]
```
