# gitignore, LFS, and Submodules

> Keep junk, secrets, and giant files out of your repo, and tame the submodule: the settings that stop you committing node_modules or a 2GB video.


---

# gitignore, LFS, and Submodules

Your repo should hold your work and nothing else. But somehow it keeps filling up with junk: a `node_modules/` folder the size of a small planet, build output, an `.env` file with a live API key in it, a 2GB demo video that makes every clone crawl. And then there's the submodule - that nested repo someone added that now shows up as a confusing single line in your diff and dumps you into "detached HEAD" the moment you `cd` into it.

This guide is about the three tools that draw the line between *your work* and *everything else*: the ignore file (what Git pretends not to see), Git LFS (how to version huge binaries without bloating history), and submodules (a repo inside a repo, with all the pain that implies). By the end you'll know why a file you ignored *still* shows up, how to get a leaked secret out, and when a submodule is the right call versus a trap.

## How to read this

Read the phases in order - each builds on the last. Phase 1 gives you the mental model for all three (what Git is actually tracking, and why "ignore" is narrower than you think). Phase 2 is the everyday core: writing ignore patterns that work, untracking files, and reaching for LFS when a binary is too big. Phase 3 is where it bites: leaked secrets, submodule detached-HEAD pain, and knowing when to walk away from a submodule entirely.

If you're new to how Git tracks files at all, skim [/guides/git-from-zero](/guides/git-from-zero) first - this guide assumes you've committed before.

## The phases

1. [What Git tracks (and what "ignore" really means)](01-what-git-tracks.md)
2. [Ignoring, untracking, and LFS for big files](02-ignoring-untracking-lfs.md)
3. [Leaked secrets and the submodule trap](03-secrets-and-submodules.md)


---

# What Git tracks (and what "ignore" really means)

Here's the thing that trips almost everyone up at least once: you add a file to `.gitignore`, you commit, and the file *still shows up* in `git status`. You double-check the spelling. It's right. The pattern is right. And Git keeps tracking it anyway. You start to wonder if `.gitignore` is broken.

It isn't. You've hit the single most important fact about ignoring files, and it's the one nobody tells you up front. So let's get the mental model right before we touch a single pattern, because once this clicks, three different mysteries stop being mysteries.

## Git sorts every file into one of three buckets

At any moment, every file in your working directory is in exactly one of these states from Git's point of view:

```text
tracked     → Git knows about it, watches it for changes
untracked   → Git sees it but isn't watching it (shows in `git status`)
ignored     → Git is told to not even mention it
```

*What just happened:* Notice that **ignored** and **untracked** are different states, and - this is the key - `.gitignore` only ever affects **untracked** files. It tells Git "don't bother nagging me about this file you aren't tracking yet." It says *nothing* about files Git is *already tracking*.

That one sentence explains the mystery. If a file was committed before you ignored it, it's **tracked**. `.gitignore` doesn't apply to tracked files at all. Git keeps watching it, keeps showing its changes, keeps committing it - exactly as if the ignore rule weren't there.

> The fix, which we'll do properly in Phase 2, is to *untrack* the file (remove it from the index) so it falls back to being merely ignored. The ignore rule only gets a chance to work once the file is no longer tracked.

## The index: the part of Git people forget exists

To really get this, you need one more piece: the **index** (also called the *staging area*). It sits between your working directory and your commits.

```text
working directory  →  index (staging area)  →  commit history
   (your files)        (what's tracked,         (the snapshots
                        staged for next          you've saved)
                        commit)
```

*What just happened:* "Tracked" really means "present in the index." When you `git add` a file, you put it in the index. From then on Git follows it. `.gitignore` is a filter on what's *allowed into* the index on its own - it can't evict something that's already there. That's why ignoring a tracked file does nothing: the file is already past the gate.

So the rule in one line: **`.gitignore` keeps files OUT of the index; it cannot kick files OUT of it.** Kicking a file out is a separate, deliberate act (Phase 2).

## What belongs out of the repo

Before patterns, the *why*. Three kinds of things should almost never live in version control:

- **Generated output** - `node_modules/`, `dist/`, `build/`, `target/`, compiled binaries, `__pycache__/`. These are reproducible from your source. Committing them bloats the repo and creates pointless merge conflicts. Anyone can regenerate them with one command.
- **Local and personal files** - editor settings, `.env` files, OS cruft like `.DS_Store`. These are about *your machine*, not the project.
- **Secrets** - API keys, passwords, tokens, private certificates. These should never enter history, and as Phase 3 will show, getting them *out* after the fact is genuinely hard.

```bash
# A starter .gitignore for a typical Node project
node_modules/
dist/
.env
.DS_Store
*.log
```

*What just happened:* Each line is a pattern. `node_modules/` ignores that whole folder; `*.log` ignores every file ending in `.log`; `.env` ignores that one file. Drop this in the repo root *before* your first commit and those files never get tracked in the first place - no untracking dance needed later. Prevention is far cheaper than cleanup.

## Big binaries are a different problem

Generated junk you keep out entirely. But sometimes you genuinely need a large file *in* the project - a design source, a sample dataset, a demo video. You can't ignore it; the project needs it. But Git was built for text, and it stores every version of every file forever.

```text
edit a 2GB video 5 times  →  Git stores ~10GB in history
                              (every version, full size, forever)
```

*What just happened:* Git diffs text beautifully but treats binaries as opaque blobs - every change stores a fresh full copy. A handful of edits to one big file can balloon your `.git` folder past the size of the actual project. Every clone then drags all of it down. This is the exact problem **Git LFS** (Large File Storage) exists to solve, and we'll set it up in Phase 2. The short version: LFS stores a tiny text *pointer* in your repo and stashes the real bytes elsewhere.

## For builders

When you scaffold a new project, the very first commit should already include a `.gitignore`. Most frameworks generate a sensible one for you (`create-react-app`, `cargo new`, `django-admin startproject` all do). GitHub also publishes a maintained collection at github.com/github/gitignore covering most languages - copy the one for your stack rather than hand-rolling it. Getting this right on commit one means you never fight a tracked-junk-file later.

```quiz
[
  {
    "q": "You add `config.local.json` to `.gitignore`, but `git status` still lists it as modified. Why?",
    "choices": [
      "The pattern syntax is wrong",
      "The file was already tracked before you ignored it; .gitignore only affects untracked files",
      ".gitignore needs to be committed before it works",
      "Git ignores JSON files by default"
    ],
    "answer": 1,
    "explain": ".gitignore only filters untracked files. A file already in the index stays tracked until you explicitly untrack it."
  },
  {
    "q": "What does it mean for a file to be 'tracked' in Git?",
    "choices": [
      "It appears somewhere in the working directory",
      "It is present in the index (staging area)",
      "It has been pushed to a remote",
      "It is not listed in .gitignore"
    ],
    "answer": 1,
    "explain": "Tracked means present in the index. .gitignore controls what is allowed into the index, not what is already there."
  },
  {
    "q": "Why is committing a frequently-edited 2GB video directly into Git a problem?",
    "choices": [
      "Git refuses to commit files over 1GB",
      "Git stores a full copy of every version, so history balloons and every clone is huge",
      "Binary files corrupt the index",
      "Git automatically deletes large files after 30 days"
    ],
    "answer": 1,
    "explain": "Git keeps every version of every file forever, and stores binaries as full blobs. Large files inflate history and slow every clone, which is what Git LFS addresses."
  }
]
```


---

# Ignoring, untracking, and LFS for big files

Now the hands-on part. You know *why* an ignored-but-tracked file misbehaves. This phase is the muscle memory: how to write patterns that match what you mean, how to untrack a file without deleting it, and how to wire up LFS so a giant binary stops bloating your history. These are the moves you'll reach for weekly.

## Ignore patterns, decoded

A `.gitignore` is a list of patterns, one per line, matched against file paths relative to the file's location. Most of the syntax is intuitive once you've seen each piece once:

```text
node_modules/      # trailing slash → match directories only
*.log              # * matches anything except a slash
build/             # ignores a folder and everything in it
/secret.txt        # leading slash → only at repo root, not nested
**/temp            # ** matches across directories (any depth)
!keep.log          # leading ! → un-ignore (exception to a rule above)
# this is a comment
```

*What just happened:* Each pattern is a rule. The two that surprise people: a leading `/` anchors to the root (so `/secret.txt` ignores the root one but not `docs/secret.txt`), and a leading `!` *re-includes* a file an earlier pattern ignored. Order matters for `!` - the exception must come *after* the rule it overrides.

A common real pattern: ignore a whole folder but keep one file in it.

```text
logs/
!logs/.gitkeep
```

*What just happened:* `logs/` ignores everything in the folder, then `!logs/.gitkeep` carves out an exception so the empty-ish folder still exists in the repo. (Git won't commit empty folders, so people add a placeholder file like `.gitkeep` to force the folder to exist for everyone who clones.)

> One gotcha: `!` cannot re-include a file if its *parent directory* is ignored. `logs/` then `!logs/important.log` works, but if you ignore `logs/` you can't selectively un-ignore a file inside a *sub*-folder that's also blanket-ignored. Re-include the path step by step if you hit this.

## Untracking a file you already committed

This is the fix for the Phase 1 mystery. You committed `config.local.json`, *then* realized it shouldn't be tracked. Adding it to `.gitignore` did nothing because it's already in the index. You need to remove it from the index - but keep it on disk.

```bash
# Stop tracking the file, but DO NOT delete it from your folder
git rm --cached config.local.json

# Then add it to .gitignore so it stays out
echo "config.local.json" >> .gitignore

git commit -m "Stop tracking config.local.json"
```

*What just happened:* `git rm --cached` removes the file from the index (untracks it) while leaving the actual file untouched on your disk - that's what `--cached` means. Now the file is *untracked*, so the `.gitignore` rule finally applies and Git stops nagging you. Drop `--cached` and `git rm` would delete the file from disk too, which is usually not what you want here.

For a whole folder you committed by mistake (the classic `node_modules/`):

```bash
git rm -r --cached node_modules/
echo "node_modules/" >> .gitignore
git commit -m "Stop tracking node_modules"
```

*What just happened:* `-r` recurses into the folder. After this commit, the folder still sits on your disk (your app still runs), but it's gone from Git's tracking and won't come back. Teammates who pull this commit will have it untracked too.

> Important nuance: this stops *future* tracking, but the file still lives in past commits and in history. For build junk that's harmless. For a *secret*, it is not enough - the secret is still recoverable from history. That's the whole of Phase 3.

## When to reach for LFS

If a binary is large *and* you genuinely need it versioned (not ignored), use Git LFS. The mental model from Phase 1: LFS swaps the real bytes for a small text pointer in your repo, and stores the actual file on an LFS server (GitHub, GitLab, etc. provide this).

```text
What's in your commit:        What LFS stores elsewhere:
┌──────────────────────┐      ┌──────────────────────┐
│ video.mp4 (pointer)  │ ───→ │ video.mp4 (the real  │
│ ~130 bytes of text   │      │ 2GB of bytes)        │
└──────────────────────┘      └──────────────────────┘
```

*What just happened:* Your repo and its history stay small because they only ever hold tiny pointers. When someone checks out the commit, LFS fetches the real file behind the scenes. Clones are fast because they don't drag every version of every big file along.

Setting it up is three steps:

```bash
# 1. Install LFS for your user (once per machine)
git lfs install

# 2. Tell LFS which files to handle (writes to .gitattributes)
git lfs track "*.psd"
git lfs track "*.mp4"

# 3. Commit the .gitattributes file so teammates get the same rules
git add .gitattributes
git commit -m "Track PSD and MP4 files with LFS"
```

*What just happened:* `git lfs install` sets up the LFS hooks. `git lfs track` records a pattern in a file called `.gitattributes` - from now on any `*.psd` or `*.mp4` you add gets stored as an LFS pointer automatically. Committing `.gitattributes` is what makes it work for everyone, not just you. After this, you `git add` and `git commit` big files exactly as normal; LFS handles the swap invisibly.

You can confirm what LFS is managing:

```bash
git lfs ls-files
```

*What just happened:* This lists every file currently stored via LFS, with a short hash and the filename. If a file you expected isn't here, its pattern probably isn't in `.gitattributes`, or you tracked it *after* it was already committed normally (LFS tracking, like ignore, only catches files going forward).

## For builders

Decide your LFS rules at project start, the same as `.gitignore`. If your repo will hold design files, datasets, or media, run `git lfs track` and commit `.gitattributes` in the first commit - retrofitting LFS onto files already in history means rewriting history (the same painful operation as scrubbing a secret, covered next). One more rule of thumb: LFS is for *large binaries you need to version*, not a dumping ground. If a file is reproducible build output, ignore it instead - LFS storage isn't free and has quotas.

```quiz
[
  {
    "q": "You accidentally committed node_modules/. What command stops tracking it without deleting it from your disk?",
    "choices": [
      "git rm -r node_modules/",
      "git rm -r --cached node_modules/",
      "git ignore node_modules/",
      "rm -rf node_modules/"
    ],
    "answer": 1,
    "explain": "git rm -r --cached removes the folder from the index (untracks it) but leaves it on disk. Dropping --cached would delete the actual files too."
  },
  {
    "q": "In a .gitignore, what does a line beginning with `!` do?",
    "choices": [
      "Marks a high-priority ignore rule",
      "Re-includes (un-ignores) a file matched by an earlier rule",
      "Ignores the file only on the current branch",
      "Comments out the line"
    ],
    "answer": 1,
    "explain": "A leading ! creates an exception, re-including a file an earlier pattern ignored. It must come after the rule it overrides, and can't rescue a file inside a blanket-ignored parent directory."
  },
  {
    "q": "How does Git LFS keep your repository small when versioning a 2GB video?",
    "choices": [
      "It compresses the video on each commit",
      "It stores a small text pointer in the repo and keeps the real bytes on an LFS server",
      "It only commits the video once and ignores later changes",
      "It splits the video into smaller tracked chunks"
    ],
    "answer": 1,
    "explain": "LFS replaces the file content in your commit with a tiny pointer and stores the actual bytes elsewhere, so history and clones stay small."
  }
]
```


---

# Leaked secrets and the submodule trap

This is where the stakes get real. The two topics here are the ones that turn into incidents: a secret committed to history that you *thought* you removed, and a submodule that drops you into "detached HEAD" and quietly commits the wrong version of a nested repo. Both punish the casual fix. Let's handle them with eyes open.

## A committed secret is a leaked secret

Say you committed `.env` with a live database password, noticed, and ran `git rm --cached .env` from Phase 2. The file's gone from your *current* commit. Problem solved?

No. It's still sitting in history.

```bash
# The secret is still right there in the old commit
git log --all --oneline -- .env
git show <that-commit>:.env   # prints the password in full
```

*What just happened:* `git rm --cached` only stops *future* tracking. Every commit that ever included the file still contains it, byte for byte. Anyone with the repo - or anyone who already cloned it - can recover the secret. If the repo was ever pushed anywhere public, treat the secret as compromised the instant it landed.

So the first and most important step is the one that isn't a Git command at all:

> **Rotate the secret. Immediately.** Revoke the leaked key or change the password at its source. Scrubbing history is damage control; rotation is the actual fix, because you can never be sure who already grabbed the old value. Do this *first*, every time.

Only then is it worth scrubbing history, so the secret isn't lying around in old commits:

```bash
# git-filter-repo is the modern, recommended tool (install separately)
git filter-repo --invert-paths --path .env

# force-push the rewritten history (coordinate with your team first!)
git push --force
```

*What just happened:* `git filter-repo` rewrites every commit to remove `.env` from history entirely, as if it had never been committed. This **changes the hash of every affected commit**, which is why it needs a force-push and why it wrecks everyone else's clones - they'll have to re-clone or reset. It's a heavy, disruptive operation. That's exactly why keeping secrets out from the start (Phase 1) matters so much: getting them out is genuinely painful.

> Tools like `git secrets`, `gitleaks`, or a pre-commit hook can block a secret *before* it's ever committed. For anything beyond a solo hobby repo, this is worth setting up - see [/guides/git-with-other-people](/guides/git-with-other-people) for hooks and team workflow.

## Submodules: a repo inside a repo

A **submodule** lets you embed one Git repository inside another. The outer (parent) repo doesn't store the inner repo's files - it stores a *pointer to one specific commit* of the inner repo.

```bash
# Add a library as a submodule
git submodule add https://github.com/some/library.git vendor/library
```

*What just happened:* This clones the library into `vendor/library` and creates two things in the parent repo: a `.gitmodules` file (recording the submodule's URL and path) and a special entry that pins the submodule to the *exact commit* you added. From then on, the parent doesn't track the library's files - it tracks "the library, at commit abc123."

In a diff, that pointer is all you see:

```text
-Subproject commit a1b2c3d4...
+Subproject commit e5f6a7b8...
```

*What just happened:* When you update a submodule, the parent's diff is just this one cryptic line - the old pinned commit replaced by the new one. It looks like nothing changed, but you've actually moved the parent to depend on a different version of the nested repo. This terseness is a big part of why submodules confuse people.

## The detached-HEAD pain

Here's the trap that catches everyone. When Git checks out a submodule, it puts you at that *specific pinned commit* - not on a branch.

```bash
cd vendor/library
git status
# HEAD detached at a1b2c3d
```

*What just happened:* "Detached HEAD" means you're sitting on a commit directly, with no branch checked out. If you make changes and commit here, the commit isn't on any branch - and a later `git checkout` inside the submodule can leave it stranded and easy to lose. This is the classic way people "lose" submodule work: they edit inside a submodule, commit on a detached HEAD, and the parent never gets updated to point at it.

The safe workflow when you *do* need to change a submodule:

```bash
cd vendor/library
git checkout main          # get onto an actual branch first
# ...make your changes, commit, push the submodule...

cd ../..                   # back to the parent
git add vendor/library     # stage the new pointer
git commit -m "Bump library to include fix"
```

*What just happened:* You checked out a real branch *inside* the submodule before changing anything, so your commit lands somewhere findable. Then back in the parent, `git add vendor/library` records the new pinned commit, and committing that is what actually moves the parent's dependency forward. Skip that last step and your submodule change is invisible to teammates - they'll still get the old pinned commit.

And the move people forget after cloning a repo *with* submodules:

```bash
git clone <repo>
git submodule update --init --recursive
```

*What just happened:* A plain `git clone` brings down the parent and the empty submodule folders, but not the submodules' contents. `git submodule update --init --recursive` fetches each submodule at its pinned commit (and any submodules-of-submodules). Forget this and you'll find empty folders where the library should be - a very common "it works on my machine but not the build server" cause.

## When to use one - and when to run

Submodules are the right tool in a narrow set of cases:

- You need to pin a dependency to an *exact* commit and control upgrades deliberately.
- The nested repo is genuinely separate (different team, different release cycle) and you want to keep its history out of yours.

But for most situations, reach for something simpler first:

- **A package manager** (npm, Cargo, pip, Go modules) handles versioned dependencies far better - real version ranges, lockfiles, no detached-HEAD surprises. If your dependency is published as a package, use the package manager, not a submodule.
- **A monorepo** (one repo holding multiple projects) sidesteps the whole pointer-juggling problem when the projects are developed together.
- **Git subtree** is an alternative that vendors the code directly into your repo, trading the pinning model for simplicity.

> Rule of thumb: if you *can* express the relationship with a package manager, do that. Submodules earn their complexity only when you truly need commit-level pinning of a separate repository. Many teams adopt them, get bitten by detached HEAD and forgotten `submodule update` steps, and migrate away.

## For builders

If you inherit a project with submodules, put `git submodule update --init --recursive` in your setup script and your CI checkout step - it's the single most common reason a submodule project "won't build" on a fresh machine. And before adding a *new* submodule, ask the YAGNI question: would a package dependency do the job? Most of the time it would, with far less pain for everyone who clones after you.

```quiz
[
  {
    "q": "You committed a live API key, then ran `git rm --cached .env`. What is the FIRST thing you should do?",
    "choices": [
      "Force-push to overwrite the remote",
      "Rotate the secret - revoke or change the key at its source",
      "Add .env to .gitignore",
      "Run git filter-repo to scrub history"
    ],
    "answer": 1,
    "explain": "The key is still in history and may already have been copied. Rotation is the real fix; scrubbing history is only damage control and must come after rotating."
  },
  {
    "q": "What does the parent repository actually store for a submodule?",
    "choices": [
      "A full copy of the submodule's files",
      "A pointer to one specific commit of the submodule",
      "The submodule's entire history merged into its own",
      "Just the submodule's URL, nothing else"
    ],
    "answer": 1,
    "explain": "The parent pins the submodule to an exact commit. Its diff shows only the old and new 'Subproject commit' hashes, not the files."
  },
  {
    "q": "When is a submodule usually the WRONG choice?",
    "choices": [
      "When the dependency is published as a package you could install with a package manager",
      "When you need to pin a separate repo to an exact commit",
      "When the nested repo has a different release cycle",
      "When you want the nested repo's history kept out of yours"
    ],
    "answer": 0,
    "explain": "If a package manager can express the dependency, it does so far better (version ranges, lockfiles, no detached-HEAD pain). Submodules earn their complexity only for true commit-level pinning of a separate repo."
  }
]
```
