# AI Privacy and What Not to Paste

> Before you paste that document into a chatbot, know where it goes and what could come back to bite you. A practical privacy guide for everyday AI use.


---

# AI Privacy and What Not to Paste

You're mid-task. You have a messy email thread, a spreadsheet, a contract draft, maybe a customer's complaint with their name and account number in it. The chatbot is right there, and it would clean all of this up in ten seconds. Your finger hovers over paste.

Stop for a moment. That paste is the whole privacy question in one gesture. Whatever you put in that box leaves your machine, travels to a company's servers, gets processed, and may be stored for a while - sometimes a long while. Most of the time nothing bad happens. But "most of the time" is not a plan, and the times it goes wrong tend to involve exactly the data you'd least want loose: a password, a customer's medical detail, an unannounced acquisition, the API key that runs your production system.

This guide is for normal people doing real work - founders, ops folks, writers, support agents, anyone who reaches for AI to get through the day. You don't need to understand how these models are built. You need three things, and that's the arc of the four phases here. First, a clear picture of where your words actually go after you hit send: who keeps them, for how long, and the real gap between a free consumer account and a paid business one. Second, a concrete list of what should never go into a general chatbot - secrets, credentials, customer personal data, regulated information, and your company's unreleased plans - with examples so you can recognize them in the wild. Third, the habits that let you keep using AI for almost everything anyway: redacting before you paste, sticking to tools your organization has actually approved, knowing when a model running on your own machine is the right call, and finding out what your own AI policy says before someone in legal finds out for you.

None of this is about fear, and none of it means you should stop using these tools. It's about drawing one line in the right place so you can move fast everywhere else without lying awake about it.


---

# Where Your Words Go

When you type into a chatbot and hit send, your text doesn't stay on your laptop. It travels over the internet to the AI company's servers, gets processed by the model, and the reply comes back. That round trip is unavoidable for any cloud chatbot - the model is too big to run on your device. So the real questions aren't "does my data leave?" (it does) but "what do they do with it once it arrives, and how long do they hang on to it?"

There are three things happening to your words, and people mix them up constantly. Keep them separate in your head.

## Three different things, not one

**Processing.** To answer you, the system has to read what you sent. This is true everywhere and isn't optional. Nobody can help you write the email without reading the email.

**Retention.** After answering, most services keep a copy of the conversation for some period - to show you your history, to debug problems, to catch abuse. This is storage, and storage is where data sits waiting for a breach, a subpoena, or an employee who shouldn't be looking.

**Training.** Separately, the company may use your conversations to improve their future models. This is the one people fear most, and it's the one most often turned off by default on paid and business tiers. Training is not the same as retention - a service can keep your chats for thirty days without ever using them to train anything.

The headline you want: training and retention are two different switches. Check both.

## Does it train on my data?

Depends entirely on which account you're using.

- **Free and personal consumer accounts** often use your conversations to train future models by default. There's usually a setting to turn this off, buried somewhere in privacy or data controls. Most people never touch it.
- **Paid consumer plans** vary. Some still train unless you opt out; some don't. Read the specific product's page - don't assume paying means private.
- **Business, team, and enterprise plans** from the major providers generally do *not* train on your inputs by default, and say so contractually. This is the single biggest reason to push your company toward a proper business agreement rather than everyone using personal logins.

The practical move: in any account you use for real work, go find the "improve the model for everyone" or "data controls" toggle and turn it off. It takes two minutes and removes the scariest possibility.

One plain caveat: turning off training does not mean your text vanishes. It still gets sent, still gets processed, and is usually still retained for a window. Off-training is necessary, not sufficient.

## How long do they keep it?

Retention windows differ by provider and by plan, and they change, so treat any specific number as "check it yourself." The shape is usually:

| Tier | Typical retention behavior |
|---|---|
| Free / consumer | Kept until you delete; deleted chats often purged within ~30 days |
| Business / enterprise | Configurable; admins may set or shorten retention |
| "Temporary" / incognito chat modes | Short window (often ~30 days) for abuse review, then gone, not used for training |

Two things worth knowing. First, even when you delete a conversation, there's frequently a lag - a backend copy may persist for weeks for safety and legal reasons before it's actually erased. Second, legal holds override everything: if a company is required to preserve data for a lawsuit, your "deleted" chats can be kept regardless of the normal schedule. That has already happened in real litigation involving major AI providers, where a court ordered chat logs preserved that would otherwise have been deleted.

## Consumer vs business: why it's a different animal

The gap is bigger than a nicer interface. A consumer account is a deal between you and the AI company under their standard terms. A business or enterprise account is a contract your organization signs, and that contract typically includes: no training on your data, an admin who controls retention and access, the ability to delete data on demand, and security commitments (encryption, access logs, sometimes compliance certifications like SOC 2). Some offer "zero data retention" arrangements where nothing is stored after the response is generated.

```text
Consumer login:  you  ->  AI company's standard terms  ->  maybe trained, kept a while
Business account: your org  ->  signed contract  ->  no training, admin-controlled retention
```

This matters for one reason above all: if you're doing work that touches anyone else's data - customers, patients, employees - a personal account quietly makes promises on your employer's behalf that your employer never agreed to. The business tier exists precisely so those promises are written down and enforceable.

The takeaway for this phase: every chatbot keeps and may learn from what you type, the defaults differ wildly by account type, and the fix is mostly about *which account* you're typing into, not how careful you are with any single message. Next, we get specific about the messages themselves - what should never go into a general chatbot no matter how the account is configured.


---

# What Not to Paste

Here's the rule of thumb that does most of the work: before you paste, ask "if this exact text showed up in a news article, or in a stranger's inbox, would I be in trouble?" If the answer is yes, don't paste it into a general chatbot. Everything below is a category where the answer is usually yes.

A useful mental test alongside it: would you put this on a whiteboard in a coffee shop and walk away? A pasted prompt is closer to that than to a private notebook.

## Secrets and credentials

These are the worst offenders because the damage is immediate and total. A leaked password isn't embarrassing - it's a working key to a real lock.

- **Passwords and PINs** - yours or anyone else's.
- **API keys, access tokens, secret keys** - the long random strings that let software talk to other software. Pasting your AWS or payment-processor key into a chatbot to "help debug this error" is one of the most common real-world leaks.
- **Private keys and certificates** - SSH keys, signing keys, anything starting with `-----BEGIN PRIVATE KEY-----`.
- **Connection strings** - database URLs often have the username and password baked right in: `postgres://admin:hunter2@db.internal:5432`. That one line is full credentials.
- **Two-factor codes and recovery codes.**

Why it bites: credentials are designed to be used by whoever holds them. The moment a key sits in a chat log on a server you don't control, you have to assume it could be used. The fix isn't redaction here - it's *rotate it*. If you've already pasted a key, go invalidate it and generate a new one. Today.

## Customer and personal information (PII)

PII is any data that identifies a specific living person. It's the category most people leak without noticing, because it's sitting inside the very documents they want help with - the support ticket, the spreadsheet, the email thread.

- Full names paired with contact details, addresses, dates of birth.
- Government IDs: Social Security numbers, passport numbers, national insurance numbers, driver's licenses.
- Financial details: credit card numbers, bank account and routing numbers.
- Anything tied to an identifiable person that they'd consider private.

Concrete example: you paste a customer's angry email so the AI can draft a calm reply. That email has their name, their account number, their phone, and a line about their billing dispute. You've now sent a real person's data to a third party they never consented to. Strip the identifying bits first - the AI can write an equally good reply to "[Customer]" about "[the billing issue]."

## Regulated data

Some data isn't only sensitive - it's governed by law, and the penalties are real money. You may be personally fine pasting it and still put your employer in legal jeopardy.

- **Health information (HIPAA in the US, and similar elsewhere):** diagnoses, treatments, medical records, anything linking a person to a health condition.
- **Children's data (COPPA):** information about kids under 13 carries special rules.
- **EU/UK personal data (GDPR):** strict consent and transfer rules; sending an EU resident's personal data to a US AI service can itself be a violation.
- **Payment card data (PCI DSS):** full card numbers have their own handling standard.
- **Financial and legal records** under sector-specific rules (banking, insurance, securities).

Why it bites: with regulated data, "nothing bad happened" doesn't save you. The act of mishandling it is the violation, fine or no fine, breach or no breach. If your work touches any of these, this is exactly the conversation to have with whoever owns compliance before you build AI into your routine.

## Unreleased and confidential company information

This one feels lower-stakes because it's "just internal." It isn't.

- **Unannounced products, features, launch dates.**
- **Financials before they're public** - revenue, fundraising terms, an acquisition in progress. (For public companies, leaking these can be a securities-law problem, not a privacy one.)
- **Source code and proprietary algorithms**, especially anything that's a competitive advantage.
- **Internal strategy, legal matters, HR cases, layoffs in planning.**
- **Anything under an NDA** - yours or a partner's. If you signed a contract promising to keep it confidential, a chatbot is a third party.

Concrete example: pasting your entire codebase or a strategy deck into a consumer chatbot to "summarize it for the team." On a free tier that might train future models, you've handed your edge to a system that millions of people query.

## The quick gut-check table

| You're about to paste... | General chatbot? |
|---|---|
| An API key or password | No - and rotate it if you already did |
| A customer's email with their details | No - redact the personal bits first |
| Patient health info, kids' data, EU personal data | No - regulated, check compliance first |
| An unreleased product plan or financials | No - confidential, use an approved tool |
| A public blog post you want shortened | Yes - it's already public |
| Generic code with no secrets or business logic | Usually fine - strip keys first |

The pattern across every category is the same: the chatbot is a third party, and pasting is publishing to it. Public or anonymous, paste freely. Identifying, secret, regulated, or confidential - don't, or strip it down until it isn't. Next we turn that "strip it down" instinct into a handful of habits that let you keep almost everything on the green list.


---

# Safer Habits and Org Rules

The previous phases drew the line. This one is about living comfortably on the right side of it without turning every task into a chore. The good news: a few small habits cover the vast majority of cases, and once they're muscle memory you stop thinking about them.

## Redact before you paste

Most documents you want help with are 95% harmless and 5% sensitive. You don't have to throw away the whole thing - strip the 5% and paste the rest. The model almost never needs the real names and numbers to do the job.

The move is to replace specifics with placeholders:

```text
Before: "Draft a refund reply to Maria Gonzalez, account 4471-8832,
         who was double-charged $89.99 on her Visa ending 4012."

After:  "Draft a refund reply to [CUSTOMER], account [ACCT],
         who was double-charged [AMOUNT] on her card."
```

The reply you get back is identical in quality - you fill the real details back in yourself afterward. A few practical notes:

- **Watch the corners.** People redact the obvious name and forget the phone number three lines down, or the email address in the signature, or the case number in the subject. Read the whole thing before pasting.
- **Beware "anonymized" that isn't.** Removing a name doesn't help if you leave "the only left-handed VP in our Tokyo office." Combinations of small details can re-identify someone. When in doubt, generalize harder.
- **Don't redact secrets - kill them.** A redacted password is still a leaked password if you fat-finger it. Credentials get rotated, not masked.
- **Files are sneaky.** Uploading a document or spreadsheet sends everything in it, including hidden columns, tracked changes, comments, and metadata you forgot were there. Uploading is pasting the whole file.

## Use the tools your org actually approved

If your company has chosen a sanctioned AI tool - an enterprise ChatGPT, a Microsoft Copilot tenant, a Claude for Work account, an internal wrapper - use that one, even if your personal account feels faster. The approved tool usually sits behind a business agreement that says they won't train on your data, an admin who controls retention, and logging that protects both you and the company.

The opposite of this is "shadow AI": people quietly using personal accounts and random browser extensions for work because IT was slow or said no. It's understandable and it's a real problem - it's how sensitive data ends up in places nobody is tracking. If the sanctioned tool is missing a feature you need, that's a conversation to have with IT, not a reason to route company data through your personal login.

Be especially wary of free AI browser extensions, "summarize this page" plugins, and no-name apps. Many send whatever they touch to servers with terms nobody has read. A flashy free tool is often paying its bills with your data.

## When a local model helps

For the most sensitive work, there's an option that sidesteps the whole "where does it go" question: run the model on your own machine. Tools like Ollama or LM Studio let you download an open-weights model (such as Llama or Mistral) and run it entirely offline. Nothing leaves your computer, so there's no third party to worry about.

This is genuinely useful for confidential drafts, sensitive analysis, or regulated data where you can't risk a cloud service. But be clear about the trade-offs:

- **Quality is lower.** A model that fits on your laptop is smaller and less capable than the big cloud ones. Fine for many tasks, frustrating for hard ones.
- **It needs decent hardware.** A capable machine with plenty of memory; modest laptops struggle.
- **"Local" only counts if it's actually local.** Plenty of tools say "private" while still calling a cloud API. Verify it runs offline - pull your network connection and see if it still answers.
- **You own the security now.** No vendor is encrypting and patching for you. The data's safety is your machine's safety.

Local models are a specialized tool, not the default. For everyday work, an approved cloud business account is the right balance of safety and capability. Reach for local when the data is too sensitive to leave the building at all.

## Read your org's AI policy

This is the step everyone skips, and it's the one that keeps you out of trouble. Most organizations of any size now have an AI policy - and if yours doesn't yet, assume it will, possibly retroactively. Find it before you need it. It's usually in the employee handbook, the IT or security wiki, or a one-pager from legal.

When you read it, look for the answers to these specific questions:

- **Which tools are approved**, and for what kinds of data?
- **What's explicitly banned** - customer data, code, financials?
- **Do you need to disclose** when AI helped produce work?
- **Who owns AI output** and who's accountable if it's wrong?
- **Who do you ask** when you're unsure?

If there is no policy, that's not a green light - it's a gap, and you'll be the example if something goes wrong. Ask. A two-line email to IT or your manager ("Is it okay to use [tool] for [task]? Anything I should keep out of it?") protects you and pushes the org to think it through.

## The whole guide in one breath

Everything you type into a cloud chatbot leaves your machine, may be stored, and on the wrong account may train a model. So: turn off training where you can, never paste secrets or credentials or other people's personal or regulated data, redact the sensitive bits out of everything else, stick to the tools your org blessed, keep a local model in your back pocket for the truly sensitive stuff, and read the policy before someone reads it to you. Do that, and you get almost all of AI's speed with almost none of the regret.
