# AI in the Terminal (CLIs)

> Coding agents that live in your terminal can read files, run commands, and edit code on your say-so. The workflow, the habits, and how the main CLIs compare.


---

# AI in the Terminal (CLIs)

Most people meet AI through a chat box. You type, it talks back, you copy what it gives you and paste it somewhere useful. That works, but there's a whole other shape of AI tool that a lot of developers have quietly moved to: an agent that runs in your terminal, inside your actual project, and can read your files, run commands, and change your code - when you let it.

This guide is about that shape. It's written for someone who's comfortable opening a terminal and running a command, but who isn't necessarily a hardcore engineer. You don't need to know how these tools work under the hood. You need to know what they do, what the day-to-day rhythm of using one feels like, and how to stay in the driver's seat instead of handing over the keys.

We'll go in three steps. First, why putting an agent in your terminal is a real shift and not a gimmick - the difference between a chatbot that suggests and an agent that acts. Then the everyday loop: how a normal task actually goes, from describing what you want to reviewing what changed and approving or pushing back. Finally, a clear-eyed look at the three main terminal coding agents - Claude Code, OpenAI's Codex CLI, and Google's Gemini CLI - what they have in common, where they differ, and how to pick one without overthinking it. By the end you should be able to install one, run a small task on a real project, and trust yourself to know when to say yes and when to say no.


---

# Why an Agent in Your Terminal

Picture the old way. You're stuck on something - a script that won't run, a config you don't understand. You open a chat window, describe the problem as best you can, maybe paste in a chunk of code. The AI gives you a fix. Now you copy it, switch back to your editor, find the right spot, paste it in, fix the indentation, run it, find out it referenced a file that doesn't exist in your project, switch back to the chat, explain that, get a new version, and repeat.

The chatbot is smart but blind. It can't see your project. It doesn't know your file names, your folder layout, what version of anything you have installed, or what error your terminal actually printed. You are the messenger, carrying context back and forth by hand. Most of the friction isn't the thinking - it's the ferrying.

A terminal agent removes the ferrying. You run it from inside your project folder, and it can look around for itself.

## What "it can act" really means

When people say a terminal agent can "act," they mean three concrete things:

- **It can read your files.** It opens the files in your project and reads them the way you would, so its suggestions fit your actual code instead of a generic template.
- **It can run commands.** It can run your tests, install a package, search the codebase, check what's installed, or reproduce the error you described - and then read the output and react to it.
- **It can edit code.** It can change files directly, not hand you a snippet to paste. You see exactly what it changed before anything is final.

That third one is the leap. The agent doesn't only describe a fix; it makes the change, runs the thing to check it worked, sees the real result, and adjusts. That's a loop a chatbot can't close on its own, because the chatbot never gets to see what happened next.

Here's the same stuck-script scenario with an agent:

```text
You:   This script crashes when I run it. Can you fix it?
Agent: Reading run.py... I see it reads from data/input.csv.
       Running it to reproduce the error...
       It's failing because the file path is relative and you're
       running from a different folder. I'll make the path absolute.
       [shows you the one-line change]
       Re-running... it works now.
```

No copy-paste. No re-explaining your folder layout. It found the real cause by running the thing, not by guessing.

## Why the terminal, specifically

You might wonder why the terminal and not a nice graphical app. A few reasons, and they're practical rather than ideological.

The terminal is where your project already lives. Your files, your version control, your test commands, your package manager - they all run there. An agent that lives in the same place doesn't need a bridge to reach any of it.

It's also upfront about what's happening. Every command the agent wants to run shows up as text you can read before it runs. There's no hidden machinery. You can watch it work the same way you'd watch over a colleague's shoulder.

And it's composable. The terminal is built for tools that chain together and run in scripts. Once you're comfortable, you can hand the agent a task and have it run as part of a larger process. (Plenty of these agents also offer editor extensions and graphical versions - but the terminal is the common denominator, and it's where the model is clearest about its actions.)

## The mental model to carry forward

Stop thinking of it as a smarter search box. Start thinking of it as a fast, eager junior colleague who happens to be sitting at your machine.

A junior colleague is genuinely useful. They can do a lot of legwork quickly. But you wouldn't let one push changes to your project without looking. They sometimes misunderstand the goal. They sometimes do the wrong thing confidently. They occasionally break something while trying to fix something else. None of that makes them useless - it only means you stay in the loop, you review their work, and you keep the authority to say "no, not like that."

That's the whole posture for working with a terminal agent: delegate the legwork, keep the judgment. It can read, run, and edit - but it does those things on your say-so, and you're the one who decides what's good enough to keep.

The next phase walks through exactly how that back-and-forth goes in a normal session: how you describe a task, how you watch it work, how you review what changed, and how the permission settings let you tune how much rope it gets before it has to stop and ask.


---

# The Everyday Loop

Once you've got a terminal agent running, almost everything you do follows the same four-beat rhythm. Learn the rhythm and the specific tool barely matters.

1. **Describe the goal.** Tell it what you want, in plain language.
2. **Let it work.** It reads files, runs commands, and proposes changes.
3. **Review the diff.** You look at exactly what it changed.
4. **Approve or correct.** You keep it, tweak it, or tell it what's wrong and let it try again.

Then you loop. Most real tasks take two or three trips around this circle, not one perfect shot. That's normal and it's fine - the loop is the feature, not a sign it's failing.

## Beat one: describe the goal

The single biggest lever you have is how you describe the task. Vague in, vague out.

"Make the login better" gives the agent nothing to aim at. Compare:

```text
The login form lets you submit with an empty password field and
then shows a confusing server error. Make it block empty passwords
on the client side and show "Password is required" under the field.
```

That second version tells it the symptom, the desired behavior, and even the wording. You don't have to be this thorough every time, but when a result comes back wrong, the fix is usually upstream - you under-specified.

A good habit: state the goal, name any constraints ("don't touch the database schema," "keep the existing tests passing"), and mention where to look if you know ("it's in the auth folder"). You're not writing a spec. You're pointing.

## Beat two: let it work

Now you watch. The agent narrates what it's doing - opening files, running a search, running your tests. This part is genuinely useful to read, not noise to skip. It's where you catch a wrong turn early.

If you see it heading somewhere you didn't intend - editing the wrong file, about to install something you don't want - you can stop it and redirect. You don't have to wait for it to finish to course-correct.

## Beat three: review the diff

This is the beat you must not skip.

When the agent edits code, it shows you a **diff** - a before-and-after view of every line it changed. Removed lines are usually marked in red, added lines in green. Reading the diff is the moment you actually exercise judgment.

```text
  function login(user, pass) {
-   submit(user, pass);
+   if (!pass) {
+     showError("Password is required");
+     return;
+   }
+   submit(user, pass);
  }
```

You don't need to understand every character. You're checking three things: Did it do what you asked? Did it change anything you didn't ask it to? Does anything look wrong or risky? If the diff is huge and sprawling for a small request, that itself is a warning sign - ask why before you accept.

Treat every diff like a pull request from that eager junior colleague. The agent is confident even when it's wrong, so confidence in the explanation is not evidence the code is right. The diff is the evidence.

## Beat four: approve or correct

If it's good, approve it. If it's close, you can edit it yourself or describe the adjustment ("good, but also trim whitespace before checking"). If it's wrong, say what's wrong - specifically - and let it try again with that feedback. Specific correction ("you broke the case where the password is just spaces") works far better than "that's not right."

## Permission modes: how much rope

Here's the part that determines how safe and how fast the whole thing feels. Every serious terminal agent asks permission before it does something consequential - but you control how often it asks. The exact names differ per tool, but they land on roughly the same ladder:

| Mode | What it does | When to use it |
|------|-------------|----------------|
| Ask every time | Pauses for your approval before each file edit or command | New project, unfamiliar agent, anything important |
| Auto-approve reads | Lets it read files freely, asks before editing or running | The comfortable default for most work |
| Auto-approve edits | Edits files without asking, still asks before risky commands | A trusted, well-scoped task you're watching |
| Full auto / "yolo" | Does almost everything without asking | Throwaway code, sandboxes - not your real project |

The trade-off is plain: more autonomy is faster but gives you fewer chances to catch a mistake before it happens. The reads that the agent does are harmless; it's the writes and the commands that need a gate.

A sane starting policy: let it read freely, make it ask before it edits or runs anything, and only loosen that once you've watched a particular agent enough to trust its judgment on a particular kind of task. The danger isn't the agent reading your code - it's a command that deletes files, force-pushes, or installs something, run without you looking. Keep those gated.

One more guardrail worth the small effort: use version control. If your project is in git, you have an undo button for everything the agent does. Commit before you start a task, let the agent work, and if it makes a mess you roll back to the commit and lose nothing. That safety net is what lets you give the agent more rope without anxiety - the worst case is a `git reset`, not a ruined afternoon.

That's the loop. Describe, watch, review, decide - with permission modes setting how often you're pulled in and version control catching anything that slips through. Next we'll compare the actual tools you'd run this loop in.


---

# Claude Code, Codex, and Gemini CLI

Three terminal coding agents get most of the attention right now: **Claude Code** from Anthropic, the **Codex CLI** from OpenAI, and the **Gemini CLI** from Google. There are others, and the landscape moves fast enough that specifics here may drift - but these three are the ones a normal person is most likely to reach for, and understanding them tells you how to read any newcomer.

The plain headline first: they are more alike than different. All three follow the loop from the last phase - describe, work, review, approve. All three can read your files, run commands, and edit code. All three ask permission before consequential actions and let you tune how often. If you learn one, you can use the others within an afternoon. So don't agonize over the choice; the differences are real but mostly at the margins.

## What they all share

- **The same basic workflow.** Run it in your project folder, talk to it in plain language, review diffs, approve.
- **Permission gating.** Each defaults to asking before edits and risky commands, with modes to loosen that as you trust it.
- **Project memory files.** Each reads a plain-text file you can drop in your project to give it standing instructions - your conventions, what to avoid, how to run your tests. (They use different file names, but the idea is identical.)
- **Tool connections (MCP).** All three support a shared standard for plugging in extra tools and data sources, so an agent can reach a database, an issue tracker, or your docs. You don't need this on day one, but it's there when you grow into it.

That shared core is why "which one" matters less than people expect. The skill you build transfers.

## Where they differ

The differences cluster in a few areas: which AI model is behind them, how they're priced, and the feel of the tool.

| | Claude Code | Codex CLI | Gemini CLI |
|---|---|---|---|
| Maker | Anthropic | OpenAI | Google |
| Model family | Claude | GPT / o-series | Gemini |
| Open source | No | Yes | Yes |
| Typical pull | Strong at multi-step coding tasks and staying on track | Tight integration with the OpenAI ecosystem | Large free allowance; ties into Google's tools |
| Usual access | Paid subscription or API | Subscription or API | Generous free tier, then paid |

A few notes on that table, kept accurate:

- **The model matters most, and it changes constantly.** The quality of any of these comes mostly from the model underneath, and all three makers ship new models often. Any claim that one is flatly "the best coder" has a short shelf life. By the time you read this, the rankings may have shuffled.
- **Open source isn't only about cost.** Codex CLI and Gemini CLI being open source means you can read exactly what they do and the community can extend them. Claude Code is closed but widely used and heavily polished. None of these positions is wrong - they're different bets.
- **Pricing is the most likely thing to be out of date here.** Gemini CLI has stood out for a large free tier, which makes it a low-risk place to start. Claude Code and Codex are typically reached through a paid plan or pay-as-you-go API. Check current pricing yourself before committing - this is exactly the kind of detail that moves.

## How to pick

Don't overthink it. Here's a decision that holds up:

- **Want to try this for free first?** Start with the Gemini CLI's free tier. You'll learn the loop at no cost and that skill carries to the others.
- **Already paying for ChatGPT or living in OpenAI's tools?** The Codex CLI fits that world and is the natural pick.
- **Already paying for Claude, or you've heard good things about it for coding?** Claude Code is a strong default and many developers reach for it for exactly the multi-step "go fix this across a few files" work this guide describes.

And the real answer underneath all three: pick one, run a small task on a real project, and judge it yourself. A guide can tell you the shape of the field, but the feel of a tool on your kind of work is something you only learn by using it. Spend twenty minutes, not twenty hours of research. Because they share a workflow, switching later costs you almost nothing - your habits transfer even when the tool doesn't.

## The takeaway

You came in thinking the choice was the hard part. It isn't. The hard part - and the valuable part - is the habit: describe precisely, watch the work, read every diff, keep the risky actions gated, and lean on version control as your undo. That habit makes any of these three a genuine force multiplier and keeps all of them from making a confident mess of your project. The tool is interchangeable. The discipline is what you keep.
