# kubectl, Day to Day

> The Kubernetes commands you actually use: get/describe/logs/exec to see what's happening, apply to change it, and the debugging loop when a pod won't start.


---

# kubectl, Day to Day

You don't need to understand all of Kubernetes to be useful with it. You need maybe ten commands, the discipline to read what they tell you, and a debugging loop you can run half-asleep. Most days you are not architecting clusters; you are asking "what is this thing doing right now?" and "why won't it start?" - and answering those is a small, learnable skill.

This guide is the daily driver, not the theory. If a pod is stuck and Slack is getting loud, the commands here are what you reach for.

## How to read this

Read it in order the first time; after that, treat it as a lookup. Phase 1 builds the mental model: kubectl is a typewriter to the cluster's API, and almost everything is a variation on one verb-plus-resource shape. Phase 2 is the handful of commands you'll run hourly. Phase 3 is the debugging loop - the exact sequence for a pod that won't come up. Type the commands as you go; reading them is not the same as feeling them.

If you've never met the moving parts (pods, deployments, namespaces), skim [/guides/kubernetes-without-the-hype](/guides/kubernetes-without-the-hype) first so the nouns aren't a mystery.

## The phases

1. [The mental model: kubectl talks to one API](01-the-mental-model.md)
2. [The commands you actually run](02-the-everyday-commands.md)
3. [When a pod won't start: the debugging loop](03-when-pods-wont-start.md)


---

# The mental model: kubectl talks to one API

The first time you watch someone fly through kubectl, it looks like a hundred memorized incantations. It isn't. There's one shape under almost everything, and once you see it, the "hundred commands" collapse into a small grammar you can recombine.

## kubectl is a typewriter, not the cluster

Here is the single most useful thing to internalize: **kubectl does almost nothing itself.** It takes your command, turns it into an HTTP request, and sends it to the Kubernetes API server. The API server is the brain. kubectl is the keyboard.

```text
you ──► kubectl ──► (HTTPS) ──► API server ──► etcd (the cluster's memory)
                                     │
                                     ▼
                              controllers + kubelets
                              make reality match the request
```

*What just happened:* every command you run is a read or a write against one API. `kubectl get pods` is a GET. `kubectl apply` is roughly a PATCH. This is why the commands feel so regular - they're all CRUD against the same set of resources.

Why does this matter on a normal Tuesday? Because it tells you where to look when things are weird. If a command hangs, the question is "can I reach the API server?" not "is kubectl broken?" If a change "didn't take," the question is "did the write reach the API, and did a controller act on it?" You stop blaming the tool and start reading the system.

## The grammar: verb, resource, name

Most kubectl commands fit one template:

```text
kubectl <verb> <resource-type> <name> [flags]
```

```bash
kubectl get pods                  # verb=get, resource=pods, no name → list them all
kubectl get pod web-7c9f          # one specific pod
kubectl describe pod web-7c9f     # verb=describe → the full story of one object
kubectl delete pod web-7c9f       # verb=delete
kubectl logs web-7c9f             # logs is special: it targets a pod by name directly
```

*What just happened:* the same four pieces rearrange into different commands. Learn the verbs (`get`, `describe`, `logs`, `exec`, `apply`, `delete`) and the common resources (`pods`, `deployments`, `services`, `nodes`, `events`) and you can already express most of what you need.

A few resources have short aliases you'll see constantly - `po` for pods, `deploy` for deployments, `svc` for services, `ns` for namespaces. They're typing savers, nothing more:

```bash
kubectl get po          # same as kubectl get pods
kubectl get deploy      # deployments
kubectl get svc         # services
```

*What just happened:* `po` and `pods` hit the exact same API endpoint. Use whichever your fingers prefer; the cluster can't tell the difference.

## Context and namespace: where am I, and where am I looking?

Two settings silently shape every command, and forgetting them is the number-one source of "it works on my machine but not in the demo."

**Context** = which cluster (and which user/credentials) you're talking to. Your `kubectl` config can hold many clusters - laptop minikube, staging, production - and exactly one is "current."

```bash
kubectl config current-context        # which cluster am I pointed at RIGHT NOW?
kubectl config get-contexts           # list all of them; the * marks current
kubectl config use-context staging    # switch
```

*What just happened:* `current-context` answers the scariest question in Kubernetes - "wait, is this prod?" Run it before any command that changes things. The asterisk in `get-contexts` is your you-are-here marker.

**Namespace** = a folder inside one cluster. Pods, services, and deployments live in a namespace. By default kubectl looks only at the `default` namespace, which is why `get pods` can come back empty even though the cluster is busy - your workload is in `payments` or `kube-system`, not `default`.

```bash
kubectl get pods                       # only the 'default' namespace
kubectl get pods -n payments           # look in the 'payments' namespace
kubectl get pods -A                     # ALL namespaces (the full, unfiltered picture)
```

*What just happened:* `-A` (short for `--all-namespaces`) is the command that ends the "but there's nothing running!" confusion. When a cluster looks empty, run `get pods -A` and the truth appears.

> **Two-question habit:** before any command that matters, ask *which context* and *which namespace*. Ninety percent of "kubectl is lying to me" moments are actually "I was looking in the wrong place." A wrong context can also mean you're about to change the wrong cluster - that one isn't only confusing, it's dangerous.

If typing `-n payments` on every command gets old, you can pin the default namespace for your current context:

```bash
kubectl config set-context --current --namespace=payments
```

*What just happened:* from now on, in this context, bare commands target `payments`. It's a per-context setting, so switching clusters with `use-context` resets you to that cluster's default - which is exactly what you want.

## Why this mental model pays off

Everything in the next two phases is built from these pieces. "Read what's happening" is `get` and `describe` and `logs`. "Change it" is `apply`. "Get inside" is `exec` and `port-forward`. The debugging loop is nothing more than running the read verbs in a deliberate order. You're not memorizing a phrasebook - you're learning a grammar, and grammar generalizes.

In the wild, the engineers who look fastest with kubectl aren't the ones who memorized the most. They're the ones who always know their context and namespace, and who read the output instead of skimming it.

```quiz
[
  {
    "q": "What does kubectl actually do when you run a command?",
    "choices": [
      "Runs the workload directly on your laptop",
      "Turns the command into an HTTP request to the Kubernetes API server",
      "Edits etcd files on disk over SSH",
      "Restarts the affected pods itself"
    ],
    "answer": 1,
    "explain": "kubectl is a client. It builds an HTTP request to the API server, which is the real brain; controllers and kubelets do the work."
  },
  {
    "q": "You run `kubectl get pods` and get nothing, but you know the cluster is busy. What's the most likely cause?",
    "choices": [
      "The cluster is down",
      "kubectl needs reinstalling",
      "Your workload is in a different namespace than 'default'",
      "Pods don't show up in 'get' until they crash"
    ],
    "answer": 2,
    "explain": "Bare `get pods` looks only at the 'default' namespace. Use `-n <namespace>` or `-A` to see all namespaces."
  },
  {
    "q": "Which command answers 'am I about to run this against production?'",
    "choices": [
      "kubectl get pods -A",
      "kubectl config current-context",
      "kubectl describe cluster",
      "kubectl version"
    ],
    "answer": 1,
    "explain": "`config current-context` shows which cluster and credentials kubectl is pointed at right now. Run it before any change."
  }
]
```


---

# The commands you actually run

This is the working set - the commands you'll type dozens of times a day. There aren't many. Split them into two jobs: **seeing** what's happening (get, describe, logs, exec) and **changing** it (apply, port-forward). Get fluent in these and you can handle the large majority of real work without looking anything up.

## See the shape: `get`

`get` is your overview. It lists objects and their high-level state. Run it first, always, to orient yourself.

```bash
kubectl get pods
```

```text
NAME                    READY   STATUS    RESTARTS   AGE
web-7c9f5d8b6-2xk4p     1/1     Running   0          3d
web-7c9f5d8b6-9mlpq     1/1     Running   0          3d
worker-5f6c8d9b-tq2vn   0/1     Pending   0          12s
```

*What just happened:* one line per pod. Read the columns: `READY` is `running-containers / desired-containers` (so `0/1` means it's not up yet), `STATUS` is the headline, `RESTARTS` is a smell - a climbing number means something keeps dying, and `AGE` tells you if this is fresh or has been limping for days.

Two flags turn `get` from a snapshot into a tool:

```bash
kubectl get pods -o wide          # adds node + pod IP columns
kubectl get pods -w               # WATCH: stream changes live, don't re-run
```

*What just happened:* `-o wide` answers "which node is this on, and what's its IP?" `-w` keeps the command open and prints a new line every time a pod's state changes - perfect for watching a rollout settle without spamming the up-arrow key. Ctrl-C to stop.

You can `get` any resource the same way: `kubectl get deploy`, `kubectl get svc`, `kubectl get nodes`. Same grammar, different noun.

## Get the full story: `describe`

`get` gives you the headline; `describe` gives you the article. When a pod looks wrong, `describe` is where the *why* lives - especially the **Events** section at the bottom.

```bash
kubectl describe pod worker-5f6c8d9b-tq2vn
```

```text
Name:         worker-5f6c8d9b-tq2vn
Namespace:    payments
Status:       Pending
Containers:
  worker:
    Image:    registry.example.com/worker:1.4.2
    State:    Waiting
      Reason: ImagePullBackOff
...
Events:
  Type     Reason     Age              From     Message
  ----     ------     ----             ----     -------
  Warning  Failed     20s (x3 over 1m) kubelet  Failed to pull image "...worker:1.4.2": not found
```

*What just happened:* the Events section narrated the failure in plain language - Kubernetes tried to pull an image that doesn't exist. `get` only showed you `Pending`; `describe` told you exactly why. **When something is stuck, the Events at the bottom of `describe` are usually the answer.** Train your eyes to scroll straight there.

## Read the program's own voice: `logs`

`describe` tells you what Kubernetes thinks about your pod from the outside. `logs` shows you what your application wrote to stdout/stderr from the inside. Both matter, and they answer different questions.

```bash
kubectl logs web-7c9f5d8b-2xk4p          # the container's stdout/stderr
kubectl logs web-7c9f5d8b-2xk4p -f       # follow (tail -f style), stream live
kubectl logs web-7c9f5d8b-2xk4p --tail=50    # last 50 lines only
```

*What just happened:* `logs` printed whatever your app logs. `-f` follows it live; `--tail=50` spares you scrolling through a day of output to see the last few lines.

One flag earns its keep over and over. When a pod has restarted, the *current* container's logs are often empty or boring - the interesting crash is in the container that died:

```bash
kubectl logs web-7c9f5d8b-2xk4p --previous
```

*What just happened:* `--previous` (or `-p`) shows the logs of the *prior* container instance - the one that crashed and got replaced. For anything in a restart loop, this is where the real error message hides.

If a pod runs more than one container, `logs` needs to know which:

```bash
kubectl logs web-7c9f5d8b-2xk4p -c sidecar    # pick a container by name
```

*What just happened:* `-c` selects a container inside a multi-container pod. Without it, `logs` defaults to the first container, which may not be the one you care about.

## Step inside: `exec`

Sometimes you need to be *in* the container - check a file, hit localhost, see what an env var actually resolved to. `exec` runs a command inside a running container; with `-it` it gives you an interactive shell.

```bash
kubectl exec -it web-7c9f5d8b-2xk4p -- /bin/sh
```

```text
/app # ls
config.yaml  server  static
/app # echo $DATABASE_URL
postgres://db.internal:5432/app
/app # exit
```

*What just happened:* `-it` gave you a terminal inside the container; everything after `--` is the command to run there (here, a shell). The `--` matters - it tells kubectl "stop reading flags, the rest is the container's command." Use a slim image's `/bin/sh` if `/bin/bash` isn't present.

You don't need a full shell for a one-off check:

```bash
kubectl exec web-7c9f5d8b-2xk4p -- env        # dump environment variables
kubectl exec web-7c9f5d8b-2xk4p -- cat /etc/config/app.yaml
```

*What just happened:* a single command runs in the container and its output comes back to you, no interactive session needed. Great for quick "is the config what I think it is?" checks.

## Reach the service: `port-forward`

A service inside the cluster usually isn't reachable from your laptop. `port-forward` tunnels a local port straight to a pod or service so you can poke it with a browser or curl, without exposing anything publicly.

```bash
kubectl port-forward svc/web 8080:80
```

```text
Forwarding from 127.0.0.1:8080 -> 80
Forwarding from [::1]:8080 -> 80
```

*What just happened:* traffic to `localhost:8080` on your machine now flows to port 80 of the `web` service in the cluster. Open `http://localhost:8080` and you're hitting the in-cluster app. The tunnel lives only as long as the command runs - Ctrl-C closes it. The format is `LOCAL:REMOTE`, so `8080:80` means "my 8080 → its 80."

## Change it: `apply`

So far everything has been read-only. `apply` is how you make changes the right way. You hand it a YAML file describing the desired state, and Kubernetes makes reality match it.

```bash
kubectl apply -f deployment.yaml
```

```text
deployment.apps/web configured
```

*What just happened:* Kubernetes compared your file to what's running and made the difference real. The output verb tells you what it did - `created` (new), `configured` (changed), or `unchanged` (already matched). `apply` is declarative: the file is the source of truth, and you can run it repeatedly with the same result.

> **apply vs edit:** `kubectl edit deploy web` opens the live object in your editor for a quick in-place change. It's handy for a hotfix at 3am - but the change exists only in the cluster, not in your YAML or git, so it vanishes the next time someone runs `apply`. Treat `edit` as a temporary probe; treat `apply -f` (from version-controlled files) as how real changes ship. If you `edit` something to recover, port the fix back into the YAML before you forget.

A few change commands you'll use alongside `apply`:

```bash
kubectl rollout status deploy/web        # watch a deployment finish rolling out
kubectl rollout restart deploy/web       # restart all pods (e.g. to reload config)
kubectl rollout undo deploy/web          # roll back to the previous version
```

*What just happened:* `rollout status` blocks until the new version is fully up (or fails), so you know when a deploy is actually done. `rollout restart` cycles the pods without changing the spec. `rollout undo` is your panic button - it reverts to the last known-good revision.

## The 90-percent set, in one place

If you remember nothing else, remember this list. It covers most days:

```text
kubectl get pods [-A] [-o wide] [-w]     # what exists, what state
kubectl describe pod <name>              # why it's in that state (read Events!)
kubectl logs <name> [-f] [--previous]    # what the app itself said
kubectl exec -it <name> -- /bin/sh       # get inside
kubectl port-forward svc/<name> 8080:80  # reach it from localhost
kubectl apply -f <file>.yaml             # change it, declaratively
```

*What just happened:* six commands, each with a flag or two. This is the actual working vocabulary of most Kubernetes users. The next phase shows how to chain the read commands into a debugging loop when a pod refuses to start.

```quiz
[
  {
    "q": "A pod keeps restarting and `kubectl logs <pod>` shows nothing useful. What should you try?",
    "choices": [
      "kubectl logs <pod> --previous",
      "kubectl delete <pod>",
      "kubectl get pods -w",
      "kubectl apply -f again"
    ],
    "answer": 0,
    "explain": "`--previous` shows logs from the container instance that just crashed - where the real error usually is - instead of the fresh, empty one."
  },
  {
    "q": "What's the key risk of fixing something with `kubectl edit` instead of `kubectl apply -f`?",
    "choices": [
      "edit is slower than apply",
      "edit can only change one field at a time",
      "The change lives only in the cluster and is lost on the next apply",
      "edit requires cluster-admin permissions"
    ],
    "answer": 2,
    "explain": "`edit` mutates the live object but not your YAML or git, so the next `apply` from version control overwrites it. Use edit for temporary probes, apply for real changes."
  },
  {
    "q": "In `kubectl port-forward svc/web 8080:80`, what does 8080:80 mean?",
    "choices": [
      "Cluster port 8080 maps to your local port 80",
      "Your local port 8080 maps to the service's port 80",
      "It forwards both ports 8080 and 80 simultaneously",
      "It's a timeout in seconds"
    ],
    "answer": 1,
    "explain": "The format is LOCAL:REMOTE. Traffic to localhost:8080 on your machine is tunneled to port 80 of the in-cluster service."
  }
]
```


---

# When a pod won't start: the debugging loop

This is the phase that earns its keep. A pod is stuck, the deploy is red, and you need to find out why - calmly, in order, without flailing. The good news: there's a fixed loop that handles the large majority of "pod won't start" problems, and the status string itself usually tells you which branch to take.

## The loop, in three moves

Whatever the symptom, run these in order. Each step narrows it down.

```text
1. kubectl get pods          → read the STATUS column
2. kubectl describe pod X    → scroll to Events (the why)
3. kubectl logs X [-p]       → read what the app said before it died
```

*What just happened:* `get` tells you the *category* of failure (the status), `describe` tells you what Kubernetes tried and what went wrong (the events), and `logs` tells you what your own code said on the way down. Status is the diagnosis; events and logs are the detail. Resist the urge to skip to logs - the status often tells you logs won't even exist yet.

```text
get pods → STATUS?
   │
   ├─ ImagePullBackOff ─► describe → fix the image name / registry auth
   ├─ Pending ─────────► describe → no room to schedule (resources / affinity)
   ├─ CrashLoopBackOff ► logs --previous → the app is crashing on startup
   └─ Running but 0/1 ─► describe → readiness/liveness probe failing
```

*What just happened:* the status routes you to the right next step. Three statuses cover most real outages. Let's take them one at a time.

## ImagePullBackOff: Kubernetes can't get the image

The node tried to pull your container image and failed. Nothing about your code is wrong yet - the cluster can't even get the image to run.

```bash
kubectl describe pod worker-5f6c8d9b-tq2vn
```

```text
Events:
  Warning  Failed  30s  kubelet  Failed to pull image "registry.example.com/worker:1.4.2":
                                 manifest unknown: manifest unknown
```

*What just happened:* the event names the exact image string and the exact failure. The usual culprits, in rough order: a typo in the image name or tag, a tag that was never pushed, or the cluster lacking credentials for a private registry (you'd see `unauthorized` or `pull access denied` instead). `logs` is useless here - there's no container to log. Fix the image reference (or the registry secret) and re-`apply`.

> **Why "BackOff"?** Kubernetes doesn't retry instantly forever - it backs off, waiting longer between attempts. So the status flickers between `ErrImagePull` (the latest attempt failed) and `ImagePullBackOff` (waiting before the next try). Same root cause; the BackOff version means it's been failing for a bit.

## Pending: nowhere to put it

`Pending` means the pod has been accepted but the scheduler can't place it on any node. The pod isn't broken - there's no room, or no node that matches its requirements.

```bash
kubectl describe pod worker-5f6c8d9b-tq2vn
```

```text
Events:
  Warning  FailedScheduling  45s  default-scheduler
    0/3 nodes are available: 3 Insufficient memory.
```

*What just happened:* the scheduler told you, per node, why it said no. `Insufficient memory` (or cpu) means your pod requested more than any node has free - shrink the request or add capacity. Other common reasons: `node(s) had untolerated taint` (the pod isn't allowed on the available nodes), or an unbound `PersistentVolumeClaim` (it's waiting for storage). The fix is almost never in the pod's code; it's in resources, scheduling rules, or storage.

## CrashLoopBackOff: it starts, then dies, on repeat

This is the one everybody meets. The container *does* start - then exits or crashes within seconds, Kubernetes restarts it, it crashes again, and Kubernetes backs off between restarts (hence the name). Your `RESTARTS` count climbs steadily.

Here `describe` tells you it's crashing, but the *reason* is in the application logs - and crucially, in the **previous** container's logs, because the current one may have already died.

```bash
kubectl logs worker-5f6c8d9b-tq2vn --previous
```

```text
Traceback (most recent call last):
  File "/app/server.py", line 11, in <module>
    DATABASE_URL = os.environ["DATABASE_URL"]
KeyError: 'DATABASE_URL'
```

*What just happened:* the real cause appeared - the app crashed on boot because a required environment variable was missing. Without `--previous`, you might have seen an empty log (the freshest container hadn't even gotten far enough to print). CrashLoopBackOff is almost always an application-level problem: a missing env var or config, a bad migration, a failed dependency connection, or an unhandled exception at startup. Read the stack trace, fix the cause, re-`apply`.

It's also worth glancing at the exit code in `describe`, under the container's `Last State`:

```text
    Last State:     Terminated
      Reason:       Error
      Exit Code:    1
```

*What just happened:* exit code 1 is a generic application error (read the logs). One code worth recognizing: **137** means the container was killed - usually the OOMKilled signal, meaning it used more memory than its limit allowed. If you see 137, the fix is a memory limit or a memory leak, not a logic bug.

## Running but not Ready: the probe is failing

A subtle one. `get pods` shows `STATUS: Running` but `READY: 0/1`. The container is up, but Kubernetes won't send it traffic because its **readiness probe** is failing - the app hasn't reported itself healthy.

```bash
kubectl describe pod web-7c9f5d8b-2xk4p
```

```text
Events:
  Warning  Unhealthy  10s (x6 over 1m)  kubelet
    Readiness probe failed: HTTP probe failed with statuscode: 503
```

*What just happened:* the app is running but its health endpoint returns 503, so Kubernetes keeps it out of the load balancer. Either the app genuinely isn't ready (still connecting to its database, warming a cache) or the probe is misconfigured (wrong path, wrong port, too short a timeout). Check the app logs to see which, then fix the app or the probe definition.

## The discipline that makes it fast

The loop only works if you actually *read* the output instead of skimming it. Three habits separate the people who fix it in two minutes from the people who flail for twenty:

- **Trust the status first.** `ImagePullBackOff` means don't bother with logs. `CrashLoopBackOff` means go straight to `logs --previous`. The status is a signpost; follow it.
- **Scroll to Events.** In `describe`, the Events section at the very bottom is where Kubernetes confesses what it tried and why it failed. It's the highest-value output in the whole tool.
- **Use `--previous` reflexively for restart loops.** The crash you want is in the container that already died, not the fresh one.

If the noun "readiness probe" or "PersistentVolumeClaim" is still fuzzy, [/guides/kubernetes-without-the-hype](/guides/kubernetes-without-the-hype) explains the moving parts; and if the image itself is the suspect, building and tagging it correctly is covered in [/guides/docker-without-the-magic](/guides/docker-without-the-magic).

In the wild, on-call engineers don't have a secret sense for this. They run the same three commands in the same order every single time, and they let the output tell them where to go next. That's the whole trick - a loop you can run half-asleep, which is often exactly when you'll need it.

```quiz
[
  {
    "q": "A pod is in CrashLoopBackOff. What's the right first move?",
    "choices": [
      "kubectl delete the pod and hope it reschedules cleanly",
      "kubectl logs <pod> --previous to read why the crashed container died",
      "kubectl describe the node it's running on",
      "Increase the pod's CPU request"
    ],
    "answer": 1,
    "explain": "CrashLoopBackOff means the app started then died. The cause is in the application logs - and in the PREVIOUS container's logs, since the current one may already be gone."
  },
  {
    "q": "A pod is stuck in Pending and describe shows '0/3 nodes are available: 3 Insufficient memory.' What's the problem?",
    "choices": [
      "The container image can't be pulled",
      "The application is crashing on startup",
      "The scheduler can't place the pod because no node has enough free memory",
      "The readiness probe is failing"
    ],
    "answer": 2,
    "explain": "Pending means the scheduler can't place the pod. 'Insufficient memory' means the pod requests more than any node has free - reduce the request or add capacity."
  },
  {
    "q": "In a container's Last State, you see Exit Code 137. What does that usually mean?",
    "choices": [
      "A generic application error - read the logs",
      "The image was not found in the registry",
      "The container was killed for exceeding its memory limit (OOMKilled)",
      "The readiness probe returned 503"
    ],
    "answer": 2,
    "explain": "Exit code 137 typically means the container was OOMKilled - it used more memory than its limit allowed. The fix is a higher limit or a memory leak, not a logic bug."
  }
]
```
