# Context Engineering

> The model only knows what is in front of it. Context engineering is the craft of controlling that window: what to include, what to cut, and what to pull in on demand.


---

# Context Engineering

Here is the single fact that explains most of what goes right and wrong with AI tools: the model only knows what is in front of it. Not what you meant. Not what you told it yesterday in a different chat. Not the document sitting in your other tab. Whatever is inside the current window of text - your message, the files you attached, the earlier turns of this conversation - that is the model's entire world for this answer. Everything outside it might as well not exist.

Context engineering is the practice of deciding what goes into that window. It is less about writing a clever prompt and more about being a good editor: choosing what the model needs to see, leaving out what would distract it, and pulling in the right reference material at the right moment. People obsess over phrasing, but a well-fed model with a plain prompt beats a starved model with a beautiful one almost every time.

This guide is for anyone who works with AI and wants better, more reliable results - founders, operators, writers, support leads, anyone who has watched a chatbot confidently get something wrong and wondered why. You do not need to know how models are built. The arc runs across three phases. First, the core idea: the context window as the model's working memory, and why what you include matters more than how you word it. Second, the practical moves: choosing what to include, summarizing long material, pulling in documents on demand (the thing people call retrieval or RAG), and giving a tool memory that survives across sessions. Third, the failure mode nobody warns you about: long conversations that slowly fill with noise until the model loses the plot - and the handful of fixes that keep the signal high. By the end you will stop blaming the model for forgetting and start managing what it can see.


---

# The Model Only Knows What It Sees

Picture a brilliant contractor who shows up to your house, does excellent work, and then gets their memory wiped the second they walk out the door. Every visit, they arrive knowing nothing about your house except what you hand them on a sheet of paper at the front door. That sheet is the only thing they read. If the answer to "where's the water shutoff?" is on the sheet, they nail it. If it's not, they guess - and a confident guess from a skilled contractor sounds exactly like knowledge.

That is the model, and that sheet of paper is the context window.

## The window is the whole world

The context window is the block of text the model reads to produce its answer. It holds your current message, any files or images you attached, the system instructions the tool sets behind the scenes, and the back-and-forth so far in this conversation. That is the complete set of things the model can use. It has training knowledge baked in - general facts about the world - but anything specific to *you*, *your* company, *today*, has to be in the window or it doesn't exist for this answer.

This explains a pile of behavior that otherwise looks like the model being broken:

- You ask about a document "you sent earlier" in a different chat, and it has no idea. Different chat, different window. Nothing carried over.
- You correct it, it agrees, and ten messages later it makes the same mistake. The correction scrolled far back; in a long enough conversation, older material gets crowded out.
- It cites a policy that sounds plausible but is wrong. The real policy wasn't in the window, so it filled the gap with something shaped like an answer. People call this a hallucination. Often it's a context gap.

Once you internalize "the window is the whole world," you stop asking *why did it forget* and start asking *was it ever in front of the model in the first place*.

## Working memory, not a hard drive

Think of the window as working memory - like the few things you can hold in your head at once - not as long-term storage. It is finite. Every model has a limit on how much text fits, measured in tokens (roughly, chunks of words; a token is about three-quarters of a word). Modern tools hold a lot - many thousands of words, sometimes the equivalent of a small book - but it is still a ceiling, and you share that ceiling with everything else in the conversation.

Two consequences follow.

First, more is not always better. If you paste a 40-page contract to ask one question about the cancellation clause, the relevant clause is now floating in a sea of mostly irrelevant text. The model can find it, but its attention is split, and the odds of a clean answer drop. Giving it the two relevant pages beats giving it everything.

Second, what falls out of the window is gone. As a chat grows past the limit, tools quietly drop or compress the oldest turns to make room. That early instruction you gave - "always write in British English" - can silently age out. The model isn't disobeying. It can no longer see the rule.

## Why content beats wording

There is a whole industry of advice about magic phrases: say "you are an expert," promise it a tip, threaten it, ask it to "think step by step." Some phrasing tweaks help at the margins. But they are a rounding error next to *what information is in the window*.

Run the comparison in your head. A perfectly worded prompt asking the model to summarize a meeting it cannot see will produce a confident, fictional summary. A clumsy, one-line prompt - "summarize this" - pasted above the actual transcript will produce a real summary. The second one wins, and it isn't close. The transcript did the work, not the wording.

So the first question for any AI task is not "how do I phrase this?" It is "does the model have what it needs to answer?" If you want feedback on your pricing page, paste the pricing page. If you want it to match your brand voice, give it two examples of your writing. If it keeps getting your product's name wrong, put the correct name and a one-line description right there in the message. You are not training the model. You are setting the table.

A useful gut check before you hit send: *if a sharp new hire read only what I typed - nothing else - could they do this?* If the answer is no, the model can't either. Add the missing piece. That instinct - noticing the gap and filling it - is the entire skill. The next phase turns it into a set of repeatable moves: what to include, how to compress long material, how to pull in documents on demand, and how to give a tool memory that outlives a single chat.


---

# Managing the Window

If phase one was the diagnosis - the model only knows what it sees - this is the treatment. There are four moves you reach for, roughly in order of effort: choose what to include, compress what's too long, pull in material on demand, and give the tool a memory that outlives the chat. You will use the first two constantly, the third when your material won't fit, and the fourth when you find yourself re-explaining the same things every session.

## Move 1: Choose what to include

The default mistake is dumping everything; the second mistake is dumping nothing and hoping. The target is in between: the relevant material, plus enough surrounding context to make sense of it, and not much more.

A workable habit before any non-trivial request - ask yourself what the model needs to see:

- **The goal.** What does a good answer look like? "Draft a refund email" is vague; "draft a refund email, apologetic but not groveling, offering store credit, under 120 words" gives the model a target.
- **The specifics.** The actual customer message, the actual numbers, the actual clause. Not your paraphrase - the source.
- **The constraints.** Tone, length, format, things to avoid. "Don't promise a refund date" prevents a whole class of bad output.
- **An example, when you can.** One sample of the output you want - a past email, a paragraph in your voice - teaches more than a paragraph of description. Showing beats telling.

What to leave out is half the craft. The legal footer on every email, the boilerplate, the eleven attachments when one is relevant - cut them. Each irrelevant block dilutes the model's attention and eats room you may need later.

## Move 2: Compress what's too long

Sometimes the material genuinely matters but won't fit, or fitting it would crowd out everything else. The answer is to compress before you feed.

The reliable pattern is summarize-then-work. Ask the model to distill the long thing first - "summarize this 30-page report into the ten decisions it asks me to make" - then work from the summary. You can even do this across a long chat: when a conversation gets unwieldy, ask "summarize what we've decided so far in a tight bullet list," then start fresh with that summary as your opening message. You've hand-rolled a clean window out of a messy one.

Be clear about the tradeoff: summarizing loses detail. If exact wording matters - a contract clause, a quoted figure - keep the original text for that part and summarize the rest. Compression is for the background, not the load-bearing details.

## Move 3: Pull in material on demand (retrieval / RAG)

When your knowledge is too big to ever fit - a 500-page handbook, three years of support tickets, a whole documentation site - you can't paste it. Instead, the tool fetches only the relevant slices at the moment you ask, and drops them into the window for that one answer. This is what people mean by **retrieval**, or **RAG** (retrieval-augmented generation). The name sounds technical; the idea is a librarian who, instead of handing you the whole library, pulls the three pages that answer your question.

You meet retrieval more often than you might think:

- A "chat with your documents" feature, where you upload a folder and ask questions across it.
- A customer-support bot that answers from your help center.
- An enterprise assistant wired into your company wiki or drive.

The practical thing to know as a user: retrieval is only as good as what it pulls. If the bot gives a wrong answer, the failure is usually that it pulled the wrong slice - or the right answer wasn't in the source at all - not that the model is dumb. That's why retrieval systems that show their sources are worth more than ones that don't: when you can see *which* pages it used, you can tell at a glance whether it grabbed the right ones. If you're choosing or configuring such a tool, that visibility is the feature to insist on.

## Move 4: Give it memory across sessions

Everything so far lives and dies inside one conversation. Memory is the layer that survives across them. Many tools now offer it: a place where the system stores durable facts about you - your role, your projects, your preferences - and quietly slips them into the window at the start of new chats so you stop repeating yourself.

It is worth understanding what's actually happening. The model still has no real long-term memory of its own. "Memory" is a notes file the tool keeps on the side and re-injects into the context window each time. It is the same mechanism - text in the window - with the tool doing the pasting for you.

That framing tells you how to use it well:

- **Put stable facts in memory, not transient ones.** "I manage a five-person support team" is durable. "I'm debugging a bug today" is not - that belongs in the chat.
- **Review what's stored.** Most tools let you see and edit the memory. A stale or wrong fact gets re-injected into *every* future chat, quietly poisoning answers. Prune it.
- **Mind the privacy line.** Whatever lands in memory rides along into future sessions. Don't store secrets or sensitive client details you wouldn't want resurfacing later.

These four moves cover the deliberate side of context: what you choose to put in. The next phase covers what creeps in on its own - the slow accumulation of clutter in a long session, and how to clear it before it drags the model down.


---

# Context Rot and Fixes

You've all had the long session that started sharp and went soft. The first hour, the model was nailing every request. By hour three it's contradicting earlier decisions, reintroducing a bug you fixed together, ignoring an instruction you gave at the top. Nothing changed about the model. What changed is the window - it filled up with hours of back-and-forth, and the useful signal got buried in the noise.

That slow decline has a name people are starting to use: **context rot**. It is not an official term - the vocabulary here is still settling, and you'll see it called context pollution, context decay, or "the model getting confused in long chats." The label matters less than recognizing the pattern, because once you see it you can fix it in about thirty seconds.

## What actually goes wrong

A few distinct problems hide under the same symptom, and they compound:

- **Crowding.** The window is finite. As the chat grows, the oldest turns get dropped or compressed to make room. Your careful opening instructions can age out entirely. The model isn't ignoring the rule - it can no longer see it.
- **Distraction.** Even when everything still fits, more text means the model's attention is spread thinner. Three abandoned tangents, two pasted error messages you've moved past, a draft you rejected - all still sitting in the window, all still competing for attention against the thing you actually care about now.
- **Stale state.** Long sessions accumulate decisions that got reversed. "Let's use approach A... actually, B." Both are in the window. The model can pick up the dead one and run with it, because to the model, text that's still present is still live.

The through-line: the window doesn't clean itself. Everything you and the model said is still in there, weighted roughly the same, whether it's the current goal or a wrong turn from forty messages ago.

## The fixes

The good news is that the cure is cheap. The instinct to fight against is "let me explain it *again*, more forcefully." Re-explaining adds *more* text to an already-crowded window - you're treating the disease with more of the disease. Reach for these instead.

**Start fresh - the most underused move there is.** When a chat has gone sideways, a new conversation is a blank window with none of the accumulated junk. The fear is losing what you built up, so bridge it: ask the current chat to "summarize everything we've decided and the current state in a tight brief," copy that summary, and paste it as the opening message of a new chat. You keep the signal, drop the noise, and the model is sharp again. Make this a reflex, not a last resort - a fresh start every time the topic shifts cleanly costs you nothing.

**Compaction.** Some tools do a version of the fresh-start automatically: when the window fills, they summarize the older part of the conversation in place and continue from the summary. This is called compaction. It buys room, but it's lossy - details get smoothed away in the summary. If you notice the model getting vaguer about specifics after a long run, compaction may have quietly eaten the detail. The fix is the same: if a precise fact matters, restate it explicitly rather than trusting it survived.

**Prune as you go.** You don't have to wait for a full reset. If you've pasted a long block - a log, a document, a draft - and you're done with it, say so: "we're finished with that error log, ignore it from here." It stays in the window, but you've told the model what's live. Better still, in tools that let you edit or delete earlier messages, remove the dead weight outright. One clean correction beats five increasingly frustrated repeats.

**Keep the goal in front.** In a genuinely long working session, periodically restate where you are: "current goal: finish the onboarding email sequence; tone decided: warm and brief; still open: the subject line for email three." It costs a few seconds and re-anchors the model's attention on what matters now, near the front of the window where it carries the most weight.

## The mindset that ties it together

Step back and the whole guide collapses to one habit: **treat the context window as something you actively curate, not a bucket you keep tossing things into.** A good answer comes from a clean, relevant window. A bad answer usually comes from a window that's missing what it needed, or drowning in what it didn't.

So the moves all rhyme. Include what's needed; cut what isn't. Compress what's too long. Pull in references on demand. Store durable facts as memory, transient ones in the chat. And when a session rots, start fresh with a clean summary rather than shouting into the clutter.

You will never have perfect control over what the model can see - tools do things behind the scenes, windows have limits, summaries lose detail. But you have far more control than most people use. The difference between someone who finds AI unreliable and someone who gets steady, useful work out of it is rarely the model. It's whether they manage the window or let it manage them.
