# Running AI Locally (Ollama)

> Run a capable model on your own machine: private, free per use, and offline. What Ollama does, how to start, and the real tradeoffs versus the big cloud models.


---

# Running AI Locally (Ollama)

Most AI you use lives in someone else's data center. You type, your words travel to a company's servers, a model answers, and the reply comes back. That works well, but it means every prompt leaves your machine, every request can cost money, and nothing works on a plane with no Wi-Fi. Running a model locally flips all three. The model sits on your own computer, answers from there, and never phones home.

Ollama is the tool that made this approachable. It packages an open model, downloads it with one command, and gives you a chat prompt in your terminal - no accounts, no API keys, no setup ritual. This guide is for normal smart people who are curious about local AI: founders weighing privacy, writers who want a draft buddy offline, anyone who has wondered whether they can skip the monthly bill. You do not need to be an engineer, and there is no math or model-training here.

Across three phases you will get the why, the how, and the real catch. Phase 1 lays out the real reasons to run a model yourself - and who should not bother. Phase 2 walks you from nothing to a working chat in a few commands, plus the rough hardware it takes. Phase 3 is the part marketing pages skip: where a local model falls short of a frontier cloud model, and how to tell which job needs which tool. By the end you will know whether local AI fits your work, and how to try it in an afternoon.


---

# Why Run a Model Yourself

You already have ChatGPT, Claude, Gemini - fast, smart, always available. So why would you run a model on your own laptop, where it is slower and not as sharp? Four reasons, and they are real ones. But they do not apply to everyone, so let's be clear about who this is for.

## Privacy: nothing leaves the machine

When you use a cloud model, your prompt travels to a company's servers. Reputable providers have privacy policies, enterprise tiers that promise not to train on your data, and security teams. For most chats, that is fine.

But some text you would rather not send anywhere. A draft of an unannounced product. A contract with a client's name in it. Notes from a therapy session. Medical history. A founder's cap table. With a local model, none of that crosses the network. The text goes from your keyboard to a model running a few inches away and back. No server logs, no policy to trust, no breach that could expose it. For regulated work - health, legal, finance - that difference can be the whole reason local AI exists.

This is the strongest argument, and it is concrete. "Private" here is not a marketing word; it is a fact about where the bytes go.

## No per-use cost

Cloud models charge by usage. For light use, the bill is small. But if you run thousands of requests - summarizing a folder of documents, tagging a backlog, generating draft after draft - the meter adds up, and you watch it the whole time.

A local model has no meter. You pay once, in hardware and electricity, and then every request is free. If your work involves grinding through volume - and you do not need top-tier quality on each one - running it yourself can turn a recurring bill into a fixed cost you already paid for.

## Offline: works with no internet

A local model needs no connection. On a flight, in a cabin, on bad hotel Wi-Fi, during an outage - it still answers. The model lives on your disk. If your laptop boots, the model runs.

This matters more than it sounds. Cloud AI quietly assumes you are always online. The moment you are not, it is a dead box. A local model is the one that still works when nothing else does.

## Tinkering: it's yours to poke at

The last reason is curiosity. A local model is something you control end to end. You can swap models, try a smaller or larger one, point your own scripts at it, build a little tool over a weekend, and learn how these things behave when you can see all the dials. For the technically curious, that is its own payoff - you understand AI better when one is sitting on your machine instead of behind a paywall.

## Who it's for - and who it isn't

Local AI is a good fit if you:

- Handle sensitive text you would rather not send to a third party.
- Run high volume and want to stop paying per request.
- Work offline often, or need AI that does not depend on a connection.
- Like to experiment and own your stack.

It is a poor fit if you:

- Want the single best answer every time. Frontier cloud models are still sharper than anything you can run at home. (Phase 3 covers this in detail.)
- Have a modest laptop. Local models need real memory and ideally a decent graphics chip; on weak hardware they crawl. (Phase 2 covers the requirements.)
- Want it to work with zero setup. Cloud AI is a website and a login. Local AI is a download and a terminal - not hard, but not nothing.

Here is the plain summary: for most people, most of the time, cloud AI is the right default. It is sharper, faster, and needs no setup. Local AI earns its place in specific situations - privacy, cost at scale, offline, and tinkering. If one of those describes you, the rest of this guide gets you running. If none does, it is still worth knowing the option exists, because the situation that calls for it tends to arrive without warning.

A useful way to hold it: cloud AI is renting a sports car, local AI is owning a reliable sedan. The rental is faster and you never maintain it, but it is not in your garage, the company knows every trip you take, and you pay each time you drive. The sedan is yours, it is paid for, and it starts even when the rental office is closed.


---

# Getting Started with Ollama

Going from "I'm curious" to "I'm chatting with a model on my own machine" takes about ten minutes and four commands. Let's walk it.

## Step 1: Install Ollama

Go to ollama.com and download the installer for your system - macOS, Windows, or Linux. It installs like any normal app. On macOS and Windows you get a small background program plus a command you can run in the terminal. On Linux it is a one-line install script the site gives you.

Once it is installed, open a terminal and confirm it is there:

```bash
ollama --version
```

If that prints a version number, you are ready.

## Step 2: Pull a model

A "model" is the AI itself - a single (large) file Ollama downloads and stores. You pick one by name. A sensible first choice in 2026 is one of the small, current open models; Llama, Gemma, Qwen, and Mistral all publish versions sized to run at home. Start small:

```bash
ollama pull llama3.2
```

This downloads the model. Expect a wait - these files run from a couple of gigabytes for the smallest models up to tens of gigabytes for larger ones. The number in a model's name (like "3B" or "8B") is roughly how big it is: bigger usually means smarter but slower and hungrier for memory. For a first run, smaller is friendlier.

You only pull a model once. After that it lives on your disk.

## Step 3: Chat with it

Now talk to it:

```bash
ollama run llama3.2
```

You will get a prompt. Type a question, press enter, and the model answers right there in your terminal - generated on your machine, no internet needed. Type `/bye` to leave the chat.

That is the whole loop. Install, pull, run, chat.

```text
>>> Write me a two-line note thanking a coworker for covering my shift.
Thanks so much for covering my shift - you really saved me, and
I owe you one. I'll happily return the favor whenever you need it.
>>> /bye
```

## Step 4 (optional): see what you've got

A couple of commands worth knowing:

```bash
ollama list      # shows every model you've downloaded
ollama rm llama3.2   # deletes one to free up disk space
```

Models take real disk space, so `ollama rm` is how you clean up the ones you tried and did not keep.

## The hardware you actually need

This is where local AI gets real. The model has to fit in your computer's memory while it runs, and how fast it responds depends mostly on your hardware. Here is a rough guide - not a spec sheet, only the shape of it:

| Your machine | What runs well | What to expect |
|---|---|---|
| 8 GB RAM, no dedicated graphics | Small models (≈3B) only | Works, but slow; fine for short tasks |
| 16 GB RAM, integrated graphics | Small to mid models (3B–8B) | Comfortable for everyday chat |
| 16–32 GB RAM + a decent GPU | Mid models (8B–14B) | Fast, genuinely usable daily |
| Apple Silicon Mac (M-series), 16 GB+ | Surprisingly large models | Strong; Apple's unified memory helps a lot |
| 32 GB+ RAM + a strong GPU | Large models (30B+) | Closest you'll get to cloud quality at home |

Two things drive the experience:

**Memory (RAM, or video memory on a graphics card).** The model must fit. If it does not, it either refuses to load or spills over and grinds to a crawl. This is the hard limit - pick a model that fits your memory, not the other way around.

**The graphics chip (GPU).** A dedicated GPU does the heavy lifting and makes responses come fast. Without one, the model runs on your main processor (CPU), which works but is much slower - you will watch words appear one at a time. Apple's M-series Macs are a happy exception: their shared memory design runs local models well without a separate graphics card.

If you are not sure what you have, try the smallest model first. If it feels quick enough, step up to a bigger one and see where your machine taps out. That trial-and-error takes minutes and teaches you your ceiling faster than any chart.

A realistic expectation: on a typical modern laptop, a small model is fine for drafting, summarizing, and quick questions. It will not feel as instant or as sharp as a cloud model - but it is yours, it is private, and it is free to run. That trade is the whole point, and the next phase is about when it is worth making.


---

# Limits vs the Cloud

A local model is a real win in the right spot. But it is not a free lunch, and pretending otherwise sets you up to be disappointed. The plain picture: the model running on your laptop is smaller, slower, and more forgetful than the frontier models the big labs serve from their data centers. Here is each gap, why it exists, and when it actually matters.

## Quality: smaller models, smaller skill

The biggest cloud models are enormous - far larger than anything that fits on a personal computer. Size is not everything, but in these models it tracks closely with skill: bigger models reason better, follow instructions more reliably, make fewer mistakes, and handle weird edge cases with more grace.

The model you run at home is a fraction of that size, because it has to fit in your memory. So it is genuinely less capable. On a casual question - "rewrite this email to sound warmer" - you may not notice. On something hard - a tricky bit of code, a nuanced legal question, multi-step reasoning, anything where being wrong is expensive - the gap shows. A local model is more likely to be confidently wrong, miss a subtlety, or wander off the instruction.

This is not a knock on local models; it is physics and economics. You cannot fit a data center's model in a laptop. Set your expectations to "a capable junior who is fast and private" rather than "the sharpest expert available."

## Speed: depends entirely on your machine

Cloud providers run racks of specialized hardware, so responses stream back quickly and consistently. Your local speed depends on what is under your desk. With a good GPU, a small model feels snappy. Without one, you watch words trickle out, and a long answer can take a genuinely uncomfortable while.

There is also no scaling. A cloud service can answer ten of your requests at once. Your machine does one thing at a time and gets warm doing it. For interactive back-and-forth that is fine; for "process these 500 documents right now," it will be a long, patient afternoon.

## Context: how much it can read at once

Every model has a limit on how much text it can hold in mind at one time - the "context window." Cloud models keep pushing this higher; some can read a whole book, a large codebase, or hundreds of pages of documents in a single go.

Local models have context windows too, but running with a large one eats a lot of memory - the same memory the model already needs to exist. So in practice you often run a local model with a smaller window than its cloud cousins. That means feeding it a giant contract or an entire repository may not fit, or may force you to chop the input into pieces. For short prompts this never comes up. For "read all of this and reason across it," it is a real constraint.

## A quick comparison

| | Local model (Ollama) | Frontier cloud model |
|---|---|---|
| Quality on hard tasks | Good, not best | Best available |
| Speed | Depends on your hardware | Fast and consistent |
| Context (how much it reads) | Often smaller | Very large |
| Privacy | Total - nothing leaves | Depends on the provider |
| Cost per use | Free after hardware | Per request |
| Works offline | Yes | No |

## So which do you reach for?

A clean rule:

**Use a local model when the deciding factor is privacy, cost at scale, or being offline** - and the task is within reach of a smaller model. Drafting, summarizing, rewriting, tagging, answering routine questions, working through sensitive text, grinding high-volume jobs you do not want metered. For these, local is not a compromise; it is the better tool.

**Reach for a frontier cloud model when the deciding factor is getting it right** - the hard problem, the long document, the answer you will act on without double-checking, the multi-step reasoning where a subtle mistake costs you. When quality is the whole point and the text is not sensitive, pay for the best.

Plenty of people run both, and that is the mature setup. Local for the private, the routine, and the offline; cloud for the hard and the high-stakes. You are not picking a side - you are matching the tool to the job. Start a sensitive draft locally, and if you hit a wall the small model cannot climb, move the non-sensitive parts to the cloud for the final pass.

The thing to carry away: local AI is not a worse version of cloud AI. It is a different deal - you trade some quality, speed, and reach for privacy, zero per-use cost, and independence from the network. Knowing which of those you need on a given task is the whole skill. Now you have both tools, and you know when each one earns its keep.
